system

The system addresses the lack of personalized features in music playback by integrating playlist creation, song recommendation, karaoke, and voice training, enhancing user satisfaction through tailored musical experiences.

JP7808655B2Active Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024161810
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-09-19
Filing Date
2024-09-19
Publication Date
2026-01-29
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

Conventional music playback systems lack integration of functions such as creating playlists tailored to user mood and preferences, mixing favorite songs to create new songs, displaying karaoke scores, and voice training, making it difficult for users to enjoy a personalized music experience.

Method used

A system that integrates playlist creation based on user input, song recommendation considering preferred melody and language, mixing favorite songs to create new songs, displaying karaoke scores, and voice training functions, enabling users to enjoy a musical experience tailored to their preferences.

Benefits of technology

The system enhances user satisfaction by automatically creating playlists and integrating karaoke and voice training functions, allowing users to enjoy a personalized and diverse musical experience in a single system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007808655000001
    Figure 0007808655000001
  • Figure 0007808655000002
    Figure 0007808655000002
  • Figure 0007808655000003
    Figure 0007808655000003
Patent Text Reader

Abstract

To provide a system.SOLUTION: A system includes: means for generating a playlist based on a mood and preference of a user; means for recommending music taking into account user's preferred melody and language; means for generating new music by combining preferred musics; means for displaying scores of Karaoke; means for providing voice training; means for detecting emotion of the user by using an emotion engine; and means for adjusting score display of Karaoke based on detected emotion of the user by using a generative AI model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional music playback systems do not integrate functions such as creating playlists tailored to the user's mood and preferences, mixing favorite songs to create new songs, displaying karaoke scores, or voice training functions, making it difficult for users to enjoy a music experience tailored to their preferences. [Means for solving the problem]

[0005] The present invention provides a system that integrates a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred melody, language, etc., a means for mixing favorite songs to create new songs, a means for displaying scores using a karaoke function, and a voice training function, thereby enabling users to enjoy a musical experience that suits their preferences. [Brief explanation of the drawings]

[0006] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 2 is a sequence diagram showing a flow of processing in the data processing system according to the first embodiment of the first form example. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Embodiment 1. [Figure 13]FIG. 10 is a sequence diagram showing a processing flow of a data processing system in a second embodiment of the second form example. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Embodiment Example 2. [Figure 15] FIG. 10 is a sequence diagram showing the flow of processing in a data processing system according to a third embodiment of the third embodiment. [Figure 16] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Embodiment 3. [Figure 17] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the first embodiment of the first form example when an emotion engine is combined. [Figure 18] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the second embodiment of the second form example when an emotion engine is combined. [Figure 20] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the third embodiment of the third form example when an emotion engine is combined. [Figure 22] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0007] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0008] First, the terms used in the following description will be explained.

[0009] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Also, the processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), Examples include a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)).

[0010] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0011] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0012] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0014] [First embodiment]

[0015] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0016] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0018] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0019] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0022] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0023] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0025] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0026] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[0027] "Example 1"

[0028] In one embodiment of the present invention, a system is provided that includes an interface that allows a user to input their mood, preferred musical style, language, etc. This system automatically creates a playlist based on the user's input. For example, if a user inputs "I want to relax," the system creates a playlist by selecting songs that match that mood. Furthermore, if the user indicates a preference for a particular musical style or language, the system recommends songs based on that preference.

[0029] "Example 2"

[0030] Furthermore, as an embodiment of the present invention, a function is provided that allows a user to select their favorite songs and mix them to create a new song. This function combines the beats and melodies of the selected songs to generate a new song. For example, if a user selects "Song A" and "Song B," the features of those songs are combined to create a new song called "Song C."

[0031] "Example 3"

[0032] In addition, the present invention provides a karaoke function and a voice training function. The karaoke function displays the user's singing results as a score. The voice training function provides training for the user to improve their singing ability. For example, when a user selects a specific song and sings along with the song, their singing ability is evaluated and a score is displayed. Furthermore, if a user wants to improve their singing ability, they can use the voice training function to practice specific singing techniques.

[0033] The processing flow of each embodiment will be described below.

[0034] "Example 1"

[0035] Step 1: The user inputs their mood, preferred melody, language, etc. through the system interface.

[0036] Step 2: The system automatically creates a playlist based on user input. For example, if a user inputs "I want to relax," the system will create a playlist by selecting songs that fit that mood.

[0037] Step 3: The system allows users to indicate preferences for specific musical styles and languages, and recommends songs based on those preferences.

[0038] "Example 2"

[0039] Step 1: The user selects his favorite song through the system interface.

[0040] Step 2: The system combines the beats and melodies of the selected songs to create a new song. For example, if a user selects "Song A" and "Song B," the system combines the features of those songs to create a new song called "Song C."

[0041] "Example 3"

[0042] Step 1: A user utilizes the system's karaoke function to select a particular song and sing along to it.

[0043] Step 2: The system evaluates the user's singing ability and displays a score.

[0044] Step 3: If the user wants to improve their singing ability, they can use the system's voice training feature to practice specific singing techniques.

[0045] Example 1

[0046] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0047] Conventional music playback systems lack the functionality to automatically create playlists based on the user's mood and preferences, resulting in the user having to manually select songs. Furthermore, since music recommendations cannot take into account the user's past song selection history or the user's mood that day, user satisfaction may decrease. Furthermore, few systems offer integrated karaoke and voice training functions, forcing users to use multiple applications.

[0048] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0049] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit for creating new songs by combining favorite songs, a karaoke score display unit, a voice training function, a user interface provision unit, a data analysis unit, a song search unit using the API of a music streaming service, and a unit for providing the generated playlist. This enables automatic creation of playlists tailored to the user's mood and preferences, thereby improving user satisfaction. Furthermore, by integrating the karaoke and voice training functions, users can utilize multiple functions in a single system.

[0050] The "means for creating a playlist based on user input" is a function that automatically creates a playlist based on the mood and preference information entered by the user.

[0051] The "means for recommending music that takes into consideration the user's preferred music style, language, etc." is a function that recommends appropriate music based on the user's preferred music style and language.

[0052] The "means of creating new music by combining favorite songs" is a function that allows a user to mix multiple songs selected by the user to create a new song.

[0053] The "means for displaying scores in the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[0054] The "voice training function" is a training function that allows the user to improve their singing ability.

[0055] The "means for providing a user interface" is a function that provides an interface for the user to input information about their moods and preferences.

[0056] The "means for receiving user input" is a function for receiving information input by a user through an interface.

[0057] "Means for analyzing data" is a function that analyzes information entered by the user and extracts their moods and preferences.

[0058] "Means for searching for songs using the API of a music streaming service" refers to a function that uses the API of a music streaming service to search for songs that match the user's mood and preferences.

[0059] The "means for providing the generated playlist" is a function for providing the generated playlist to the user.

[0060] MODE FOR CARRYING OUT THE INVENTION

[0061] The present invention is a music playback system that automatically creates a playlist based on the user's mood and preferences, and further integrates karaoke and voice training functions. Specific embodiments of this system are described below.

[0062] Providing a user interface

[0063] An interface is provided for users to input their mood, preferred melody, language, etc. The device displays this interface through a web browser or mobile application. For example, an interface that runs on a web browser can be built using HTML5 and JavaScript. Users can input their mood and preferences using text boxes and drop-down menus.

[0064] Receiving User Input

[0065] The server receives the information the user enters into the interface. If the user enters "I want to relax," that information is sent to the server via an HTTP request. The server temporarily stores the received data and proceeds to the next processing step.

[0066] Data analysis and processing

[0067] The server analyzes the received user input data. Using Python's natural language processing libraries (NLTK and spaCy), it analyzes the user's input text and extracts their mood and preferences. For example, from the input "I want to relax," it extracts the mood of "relaxation." Based on the results of this analysis, the server selects music in the next step.

[0068] Playlist Generation

[0069] The server selects songs that match the user's mood and preferences based on the analysis results. It uses the APIs of music streaming services such as Spotify® API and Apple Music® API. For example, use the Spotify API to search for songs that match "relaxation." The server generates a playlist from the search results and formats the information in JSON format.

[0070] Providing playlists

[0071] The generated playlist is provided to the user. The server sends the generated playlist information to the user's terminal. The terminal displays the received playlist information on an interface. The user can play the generated playlist. For example, each song in the playlist is displayed in list format, and a song can be played by clicking the play button.

[0072] Specific examples

[0073] Example 1: When a user enters "I want to relax"

[0074] 1. The user types "I want to relax" into the interface.

[0075] 2. The server receives this input and uses natural language processing to extract the mood "relaxed."

[0076] 3. The server uses the Spotify API to search for songs that match "relaxation."

[0077] 4. Generate a playlist from the search results and format it in JSON format.

[0078] 5. The server sends the generated playlist to the user's device, and the device displays the playlist on the interface.

[0079] Example 2: When a user enters "I want to listen to up-tempo English songs"

[0080] 1. The user types into the interface, "I want to listen to some up-tempo English songs."

[0081] 2. The server receives this input and uses natural language processing to extract the preferences of "uptempo" and "English."

[0082] 3. The server uses the Spotify API to search for "uptempo" and "English" songs.

[0083] 4. Generate a playlist from the search results and format it in JSON format.

[0084] 5. The server sends the generated playlist to the user's device, and the device displays the playlist on the interface.

[0085] Prompt Sentence Examples

[0086] Example 1: When you want to relax

[0087] If the user enters "I want to relax," generate a playlist by selecting songs that are suitable for relaxation.

[0088] Example 2: If you want to listen to up-tempo English songs

[0089] If a user types "I want to listen to up-tempo English songs," generate a playlist by selecting up-tempo English songs.

[0090] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0091] Step 1:

[0092] The user accesses the interface and inputs their mood, preferred melody, language, etc.

[0093] Input: A user enters text into the interface, such as "I want to relax" or "I want to listen to some upbeat English music."

[0094] Output: The user's input data is generated.

[0095] Specific behavior: The device displays an interface through a web browser or mobile application, and the user enters information using text boxes and drop-down menus.

[0096] Step 2:

[0097] The server receives the information entered by the user.

[0098] Input: Text data entered by the user into the interface.

[0099] Output: User input data sent to the server.

[0100] What happens: When the user completes the input and clicks the submit button, the data is sent to the server via an HTTP request, which the server then temporarily stores.

[0101] Step 3:

[0102] The server parses the received user input data.

[0103] Input: User input data stored on the server.

[0104] Output: Parsed mood and preference information.

[0105] How it works: The server uses Python's natural language processing libraries (NLTK and spaCy) to parse the user's input text and extract moods and preferences such as "relaxed" or "uptempo."

[0106] Step 4:

[0107] Based on the analysis results, the server selects music that matches the user's mood and preferences.

[0108] Input: Parsed mood and preference information.

[0109] Output: A list of selected songs.

[0110] Specific operation: The server uses the API of music streaming services such as Spotify API and Apple Music API to search for songs that match the user's mood and preferences. For example, it uses the Spotify API to search for songs that match "relaxation."

[0111] Step 5:

[0112] The server provides the generated playlist to the user.

[0113] Input: A list of selected songs.

[0114] Output: Playlist information sent to the user's device.

[0115] Specific operation: The server generates a playlist from the search results and formats the information in JSON format. The generated playlist is sent to the user's device, which displays the playlist on its interface. The user can then play the generated playlist.

[0116] (Application example 1)

[0117] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0118] Conventional music streaming services require users to manually create playlists that match their moods and preferences, which is a time-consuming process. Furthermore, music recommendations based on users' moods and preferences are often insufficient, resulting in low user satisfaction. Furthermore, it is difficult to automatically generate playlists that match users' moods and preferences, leaving a need for an improved user experience.

[0119] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0120] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for creating new songs by combining favorite songs, and a means for generating prompt sentences using a generative AI model and automatically generating a playlist based on the user's mood and preferences, thereby enabling the automatic generation of a playlist that matches the user's mood and preferences.

[0121] The "means for creating a playlist based on user input" is a function that automatically creates a playlist by selecting appropriate songs based on information entered by the user.

[0122] "Means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that recommends appropriate music based on the user's preferences and past selection history.

[0123] The "means for creating new music by combining favorite songs" is a function for creating new music by combining multiple songs selected by the user.

[0124] The "means for displaying scores in the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[0125] The "voice training function" is a training function for improving the user's singing ability.

[0126] "Means for generating prompt sentences using a generative AI model and automatically generating a playlist based on the user's mood and preferences" refers to a function that uses a generative AI model to convert user input information into prompt sentences, and automatically generates a playlist that matches the user's mood and preferences based on those prompt sentences.

[0127] A system for carrying out this invention includes a plurality of means for automatically creating a playlist based on user input, specifically a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred musical style, language, etc., a means for creating a new song by combining favorite songs, a means for displaying scores in a karaoke function, a voice training function, and a means for generating prompt sentences using a generative AI model to automatically generate a playlist based on the user's mood and preferences.

[0128] Program processing explanation

[0129] The server receives information such as mood, preferred melody, and language input from the user's smartphone or other device. Based on this input information, a generative AI model is used to generate a prompt. The generated prompt may have the following format, for example:

[0130] Example prompt sentence:

[0131] The user's mood is Relax, their favorite music style is Jazz, and their language is English. Create a playlist based on this.

[0132] Based on the generated prompt, a generative AI model (e.g., GPT-3®) is used to automatically generate a playlist that matches the user's mood and preferences. This playlist is then displayed on the user's device, allowing the user to play the playlist.

[0133] Hardware and software used

[0134] Hardware: Smartphones, servers

[0135] Software: Python, OpenAI® API

[0136] Specific examples

[0137] For example, if a user inputs "I want to relax," "Jazz," and "English," the server receives this information and uses the generative AI model to generate a prompt. Based on this prompt, the generative AI model automatically generates a playlist that matches the user's mood and preferences, and displays it on the user's smartphone.

[0138] In this way, users can easily create playlists that suit their moods and preferences, improving their experience using music streaming services.

[0139] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0140] Step 1:

[0141] The user inputs information such as mood, preferred music style, language, etc. from a device such as a smartphone.

[0142] Input: User's mood, preferred tune, language

[0143] Output: Sending input information

[0144] Specific actions: The user opens the application on their smartphone, enters information such as "I want to relax," "Jazz," and "English" into the text boxes, and presses the send button.

[0145] Step 2:

[0146] The terminal sends the input information to the server.

[0147] Input: Information entered by the user

[0148] Output: Send data to the server

[0149] Specific operation: The smartphone application makes an API request to send the information entered by the user to the server.

[0150] Step 3:

[0151] The server generates a prompt based on the input information it receives.

[0152] Input: User's mood, preferred tune, language

[0153] Output: prompt statement

[0154] Specific operation: The server analyzes the received information and generates a prompt sentence of the form "The user's mood is to relax, their favorite music style is jazz, and their language is English. Please create a playlist based on this."

[0155] Step 4:

[0156] The server sends the generated prompts to the generative AI model to generate a playlist.

[0157] Input: prompt statement

[0158] Output: Playlist

[0159] Specific operation: The server sends the generated prompt sentence to the OpenAI API and generates a playlist using a generative AI model (e.g., GPT-3).

[0160] Step 5:

[0161] The server transmits the generated playlist to the terminal.

[0162] Input: Playlist

[0163] Output: Send playlist

[0164] Specific operation: The server makes an API response to send the generated playlist to the smartphone application.

[0165] Step 6:

[0166] The terminal displays the received playlist to the user.

[0167] Input: Playlist

[0168] Output: Playlist display

[0169] Specific operation: The smartphone application displays the received playlist on the screen and allows the user to play it.

[0170] Example 2

[0171] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0172] Conventional music playback systems make it difficult for users to find music that suits their tastes and lack the ability to create new music by combining favorite songs. Furthermore, the means for providing the created music to users is insufficient, preventing user satisfaction. Furthermore, the lack of integrated karaoke and voice training functions makes it difficult for users to enjoy a diverse musical experience in a single system.

[0173] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0174] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a unit that analyzes the beat and melody of selected songs and generates new songs using a generative AI model, a unit that provides the generated songs to the user, a unit that displays a score using a karaoke function, and a voice training function. This allows users to easily find songs that suit their preferences and create new songs by combining their favorite songs. Furthermore, the rapid provision of generated songs increases user satisfaction. Furthermore, by integrating the karaoke and voice training functions, users can enjoy a diverse musical experience on a single system.

[0175] The "means for creating a playlist based on user input" is a function that automatically generates a list of songs that meets specific conditions or preferences based on information entered by the user.

[0176] "Means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that analyzes the user's past music selection history and input information and recommends music that suits the user's preferences.

[0177] "Means for creating new music by mixing favorite songs" is a function that combines multiple songs selected by the user to create new music.

[0178] "Means of analyzing the beat and melody of a selected song and generating a new song using a generative AI model" refers to a function that analyzes the beat and melody of a song selected by the user and generates a new song using a generative AI model based on that information.

[0179] The "means for providing the generated music to the user" is a function for providing the generated new music to the user in a format that can be listened to.

[0180] The "means for displaying scores in the karaoke function" is a function that provides a karaoke function that displays scores for songs sung by the user.

[0181] The "voice training function" is a function that provides training to improve the user's singing ability.

[0182] This invention is a system that allows users to easily find songs that suit their tastes and combine their favorite songs to create new songs. Furthermore, it provides users with a diverse musical experience by quickly providing the created songs and integrating karaoke and voice training functions.

[0183] System configuration

[0184] Subject: User

[0185] First, the user selects their favorite songs through the system interface. For example, they select "Song A" and "Song B." Next, the user performs an operation to mix the selected songs. Specifically, they click the MIX button.

[0186] Subject: Terminal

[0187] The device sends data about the user-selected "song A" and "song B" to the server. This data includes information about the beat and melody of the songs. The device receives the data about the user-selected songs and uses software (e.g., music editing software) to analyze them.

[0188] Subject: Server

[0189] The server receives the data for "Song A" and "Song B" sent from the device. It analyzes the received data and extracts beat and melody features. A generative AI model (e.g., OpenAI's GPT-3 or Google's Magenta) is used for the analysis. The server inputs the following prompt sentence into the generative AI model:

[0190] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[0191] The generative AI model generates a new song, "Song C," based on this prompt. The generated song is saved on the server.

[0192] Subject: Terminal

[0193] The device receives the newly generated song "Song C" from the server, and the received song is converted into a format that can be played on the device's media player.

[0194] Subject: User

[0195] The user listens to the new song "Song C" through the device. The user can save the created song or share it on social media. For example, the user can save the song in MP3 format and send it to a friend.

[0196] Specific examples

[0197] If the user selects "Song A" and "Song B," the device sends the song data to the server, which uses the generative AI model to input the following prompt:

[0198] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[0199] The generative AI model generates a new song, "Song C," based on this prompt. The generated song, "Song C," is then sent from the server to the device, where it becomes available for the user to listen to.

[0200] In this way, users can easily find songs that suit their tastes and combine their favorite songs to create new songs. Furthermore, by quickly providing the created songs, user satisfaction can be increased. Furthermore, by integrating karaoke and voice training functions, users can enjoy a diverse musical experience in one system.

[0201] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0202] Step 1:

[0203] The user selects their favorite songs through the system interface. For example, they select "Song A" and "Song B." The input is the information of the songs selected by the user, and the output is a list of the selected songs. Specifically, the user enters "Song A" and "Song B" in the search bar and selects from the displayed list.

[0204] Step 2:

[0205] The user performs an operation to mix the selected songs. Specifically, the user clicks the MIX button. The input is a list of the selected songs, and the output is instructions for the MIX operation. As a specific operation, the user clicks the MIX button on the interface.

[0206] Step 3:

[0207] The device sends data on "song A" and "song B" selected by the user to the server. The input is a list of selected songs, and the output is the song data sent to the server. Specifically, the device sends data including beat and melody information of the selected songs to the server.

[0208] Step 4:

[0209] The server receives the data for "Song A" and "Song B" sent from the device. The input is the song data sent from the device, and the output is the received song data. Specifically, the server analyzes the received data and extracts the beat and melody characteristics.

[0210] Step 5:

[0211] The server uses a generative AI model to generate new music. The input is the analyzed beat and melody characteristics, and the output is the generated new music. Specifically, the server inputs the following prompt sentence into the generative AI model:

[0212] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[0213] The generative AI model generates a new song, "Song C," based on this prompt.

[0214] Step 6:

[0215] The server sends the newly generated song "Song C" to the terminal. The input is the generated song, and the output is the song data to be sent to the terminal. In concrete terms, the server sends the generated song to the terminal.

[0216] Step 7:

[0217] The device receives a new song, "Song C," generated by the server. The input is the song data sent from the server, and the output is the received song data. Specifically, the device converts the received song into a format that can be played on a media player.

[0218] Step 8:

[0219] The user listens to the new song "Song C" through the device. The input is the received song data, and the output is a song that can be listened to. Specifically, the user plays the song using the device's media player. The user can save the created song or share it on social media.

[0220] (Application example 2)

[0221] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0222] Conventional music creation systems offer limited functionality for users to mix songs to their own tastes and lack the ability to preview, save, and share created songs. Furthermore, there is no easy way for users to share their created songs with other users, limiting how users can enjoy music. Furthermore, the lack of integrated karaoke and voice training features results in an inconsistent user experience.

[0223] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0224] In this invention, the server includes a playlist creation unit based on user input, a music recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a unit that previews, saves, and shares the created songs, a unit that displays a score using a karaoke function, and a voice training function. This allows users to create songs that suit their preferences and preview, save, and share them. Furthermore, integrating the karaoke and voice training functions provides a consistent and enriched music experience.

[0225] The "means for creating a playlist based on user input" is a function that automatically generates a playlist based on the songs selected by the user and their mood that day.

[0226] "Means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music style, language, etc., and recommends appropriate music based on that.

[0227] "Means to mix your favorite songs to create new music" is a function that generates new music by combining the beats and melodies of multiple songs selected by the user.

[0228] "Means for previewing, saving, and sharing the generated music" refers to a function that allows users to listen to the new music they have generated, save it if they like it, and share it with other users via social media, messaging apps, etc.

[0229] The "means for displaying a score using the karaoke function" is a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays a score.

[0230] The "voice training function" is a function that assists users in practicing to improve their singing ability, and provides feedback on pitch and rhythm.

[0231] A system for implementing this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred style, language, etc., a means for creating a new song by mixing favorite songs, a means for previewing, saving, and sharing the created song, a means for displaying scores in a karaoke function, and a voice training function.

[0232] System Program

[0233] Hardware and software used

[0234] Hardware: Smartphone, Head-Mounted Display (HMD)

[0235] Software: Python, librosa, pydub

[0236] Data processing and calculation

[0237] The server loads user-selected songs and uses the librosa library to analyze the beats and melodies. It then combines the waveform data from multiple selected songs to generate a new song, which the user can preview and, if they like it, save and share.

[0238] Specific examples

[0239] When a user selects "Song A" and "Song B" and mixes them to create a new "Song C," the following steps are performed:

[0240] 1. The user selects "Song A" and "Song B" within the app.

[0241] 2. The server loads the selected song using the librosa library and analyzes the beat and melody.

[0242] 3. Based on the analysis results, the server combines the waveform data of the two songs to generate a new "Song C."

[0243] 4. The user previews the generated "Song C" and saves it if they like it.

[0244] 5. Share the saved "Song C" on social media or messaging apps.

[0245] Prompt Sentence Examples

[0246] Generate a new song, "Song C," by mixing the user-selected "Song A" and "Song B." The generated song must be a combination of beats and melodies. Also, provide the ability to preview, save, and share the generated song.

[0247] In this way, users can create songs to their liking, preview, save and share them, and the integration of karaoke and voice training features makes the music experience more consistent and fulfilling.

[0248] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0249] Step 1:

[0250] The user selects "Song A" and "Song B" within the app.

[0251] Input: User selected music files (song A, song B)

[0252] Output: Path of the selected music file

[0253] Specific operation: The user selects "Song A" and "Song B" from the library through the application interface. The paths of the selected music files are sent to the server.

[0254] Step 2:

[0255] The server loads the selected song using the librosa library and analyzes the beat and melody.

[0256] Input: Path of selected music file (song A, song B)

[0257] Output: Analysis data of the beat and melody of the song

[0258] Specific operation: The server uses the librosa library to load the selected music file, analyzes the beat and melody of the loaded music, and generates analysis data.

[0259] Step 3:

[0260] Based on the analysis results, the server combines the waveform data of the two songs to generate a new "Song C."

[0261] Input: Analysis data of the beat and melody of the songs (song A, song B)

[0262] Output: New song file (song C)

[0263] Specific operation: Based on the analysis data, the server combines the waveform data of the two songs by averaging them or other methods to generate a new song, "Song C."

[0264] Step 4:

[0265] The user previews the generated "Song C" and saves it if they like it.

[0266] Input: New song file (song C)

[0267] Output: Path of saved song file (song C)

[0268] Specific behavior: The user previews the new song "Song C" through the application interface. If they like it, they press the save button to save the song.

[0269] Step 5:

[0270] Share the saved "Song C" on social media or messaging apps.

[0271] Input: Path of saved music file (song C)

[0272] Output: Link or file of the shared song

[0273] Specific operation: The user uses the application's sharing function to share the saved song "Song C" with other users via social media or messaging apps.

[0274] Example 3

[0275] Next, a description will be given of a third embodiment of the third embodiment. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0276] Conventional karaoke systems and voice training systems have limited functions for evaluating and training users' singing ability, making it difficult to effectively support users in improving their singing ability. Furthermore, they lack the ability to recommend songs and create playlists tailored to users' preferences, making it difficult to increase user satisfaction. This has resulted in a lack of motivation for users to continue using the system.

[0277] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[0278] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit for mixing favorite songs to create a new song, a unit for displaying a score using the karaoke function, a unit for playing a song selected by the user and displaying the score after singing, and a unit for providing training for the user's selected singing technique and displaying feedback after practice. This makes it possible to effectively evaluate and improve the user's singing ability. Furthermore, song recommendations and playlist creation based on the user's preferences can be made, increasing user satisfaction and encouraging continued use.

[0279] The "means for creating a playlist based on user input" is a function that automatically creates a list of songs based on information entered by the user.

[0280] "Means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music genre, language, etc., and recommends music based on that.

[0281] "Means for creating new music by mixing favorite songs" is a function that combines multiple songs selected by the user to create new music.

[0282] The "means for displaying a score in a karaoke function" is a function that analyzes the results of the user's singing and displays the singing ability as a score.

[0283] The "means for playing a song selected by the user and displaying a score after singing" is a function for playing a song selected by the user and displaying the result as a score after singing.

[0284] "Means for providing training for a singing technique selected by the user and displaying feedback after practice" is a function that provides training for a specific singing technique selected by the user and displays the results as feedback after practice.

[0285] The present invention is a system that provides a karaoke function and a voice training function for evaluating and improving a user's singing ability. Specific embodiments of this system will be described below.

[0286] Karaoke function

[0287] 1. Select a song

[0288] The user selects the song he wants to sing through the terminal interface, for example, the user selects "Song 1."

[0289] 2. Sending song data

[0290] The server sends the audio data and lyrics data of the selected song to the terminal. The hardware used is the server and the terminal, and the software used is a music database and a communication protocol.

[0291] 3. Play a song

[0292] The device plays the received audio data and displays the lyrics on the screen. The hardware used is the device (smartphone, tablet, PC, etc.), and the software used is a media player and lyric display software.

[0293] 4. Singing

[0294] The user sings along to the song through the device's microphone.

[0295] 5. Audio Analysis

[0296] The device analyzes the user's singing voice in real time and evaluates pitch and rhythm using software such as voice analysis algorithms and pitch detection software.

[0297] 6. Calculation of points

[0298] The server receives the analysis results and calculates the score. The software used is the score calculation algorithm.

[0299] 7. Display of score

[0300] The terminal displays the score received from the server on the screen. For example, after the user finishes singing "Song 1," a score of 85 is displayed.

[0301] Voice training function

[0302] 1. Select a training menu

[0303] The user selects the singing technique they want to practice through the device interface, for example, by selecting "practice vibrato."

[0304] 2. Sending training data

[0305] The server transmits audio data and instructional videos corresponding to the selected training menu to the terminal. The hardware used is the server and the terminal, and the software used is the training database and communication protocol.

[0306] 3. Training playback

[0307] The device plays the received audio data and instructional video. The hardware used is the device (smartphone, tablet, PC, etc.), and the software used is a media player and video playback software.

[0308] 4. Practice

[0309] The user practices singing techniques by following instructional videos.

[0310] 5. Audio Analysis

[0311] The device analyzes the user's practice voice in real time and provides feedback on areas for improvement. The software used is a voice analysis algorithm and feedback generation software.

[0312] 6. Track your progress

[0313] The server receives the analysis results and records the user's progress. The software used is a progress management system.

[0314] 7. Viewing Feedback

[0315] The device displays the feedback it receives from the server on the screen. For example, after the user has finished practicing vibrato, the device displays feedback such as, "Your vibrato duration is too short, so try to hold it for a little longer."

[0316] Examples and prompts

[0317] Specific examples

[0318] Karaoke function:

[0319] The user selects "Song 1" and after finishing singing, a score of 85 is displayed.

[0320] Voice training features:

[0321] The user selects "Practice vibrato" and after practicing, feedback is displayed saying, "The vibrato duration is short, try to hold it for a little longer."

[0322] Prompt Sentence Examples

[0323] Karaoke function:

[0324] "Generate a program that plays a song selected by the user and displays the score after singing."

[0325] Voice training features:

[0326] "Please create a program that provides training for a user-selected singing technique and displays feedback after practice." The flow of the specific process in the third embodiment will be described with reference to FIG.

[0327] Karaoke function

[0328] Step 1: Select a song

[0329] The user selects the song they want to sing through the terminal interface.

[0330] Input: Information about the song selected by the user (e.g. "Song 1")

[0331] Output: The information of the selected songs will be saved on your device.

[0332] Specific action: The user taps "Song 1" on the device screen.

[0333] Step 2: Send the song data

[0334] The server transmits the audio data and lyrics data of the selected song to the terminal.

[0335] Input: Information about the song selected by the user

[0336] Output: Audio data and lyrics data are sent to the device.

[0337] Specific operation: The server retrieves the audio file and lyrics file for "Song 1" from the music database and sends them to the device.

[0338] Step 3: Play a song

[0339] The device plays the received audio data and displays the lyrics on the screen.

[0340] Input: Audio data and lyrics data sent from the server

[0341] Output: The audio is played and the lyrics are displayed on the screen.

[0342] Specific operation: The device's media player plays "Song 1" and the lyrics display software scrolls the lyrics.

[0343] Step 4: Singing

[0344] The user sings along to the song through the device's microphone.

[0345] Input: Audio to be played and lyrics to be displayed

[0346] Output: User's singing voice is input through a microphone

[0347] Specific action: The user sings "Song 1" into the microphone.

[0348] Step 5: Analyze the audio

[0349] The device analyzes the user's singing voice in real time and evaluates pitch and rhythm.

[0350] Input: User singing

[0351] Output: Pitch and rhythm analysis data

[0352] How it works: The device's voice analysis algorithm analyzes the user's singing voice using pitch detection software to generate pitch and rhythm data.

[0353] Step 6: Calculate your score

[0354] The server receives the analysis results and calculates the score.

[0355] Input: Pitch and rhythm analysis data sent from the device

[0356] Output: Calculated score

[0357] Specific operation: Based on the pitch and rhythm data received by the server, the score calculation algorithm calculates 85 points.

[0358] Step 7: View your scores

[0359] The terminal displays the score received from the server on the screen.

[0360] Input: Score sent from the server

[0361] Output: The score displayed on the screen

[0362] Specific action: "85 points" will be displayed on the device screen.

[0363] Voice training function

[0364] Step 1: Select a training menu

[0365] The user selects the singing technique they wish to practice through the terminal interface.

[0366] Input: User-selected training menu (e.g., "Vibrato practice")

[0367] Output: The selected training menu is saved to the device.

[0368] Specific action: The user taps "Practice Vibrato" on the device screen.

[0369] Step 2: Submitting training data

[0370] The server transmits audio data and instructional videos corresponding to the selected training menu to the terminal.

[0371] Input: User selected training menu

[0372] Output: Audio data and instructional videos are sent to the device.

[0373] Specific operation: The server retrieves the audio file and instructional video for "vibrato practice" from the training database and sends them to the terminal.

[0374] Step 3: Playback the training

[0375] The device plays the received audio data and instructional video.

[0376] Input: Audio data and instruction video sent from the server

[0377] Output: Audio is played and instructional video is displayed on the screen

[0378] Specific operation: The device's media player plays the audio for "Vibrato Practice," and the video playback software displays the instructional video.

[0379] Step 4: Practice

[0380] The user practices singing techniques by following instructional videos.

[0381] Input: Audio to be played and instructional video

[0382] Output: User's practice voice is input through microphone

[0383] Specific actions: The user practices vibrato while watching an instructional video.

[0384] Step 5: Analyze the audio

[0385] The device analyzes the user's practice audio in real time and provides feedback on areas for improvement.

[0386] Input: User's practice voice

[0387] Output: Analysis results and feedback data

[0388] Specific operation: The device's audio analysis algorithm analyzes the user's practice audio, and the feedback generation software generates feedback such as "the vibrato duration is short."

[0389] Step 6: Record your progress

[0390] The server receives the analysis results and records the user's progress.

[0391] Input: Analysis results sent from the device

[0392] Output: Recorded progress data

[0393] Specific operation: The analysis results received by the server are saved in the progress management system.

[0394] Step 7: View your feedback

[0395] The device displays the feedback received from the server on the screen.

[0396] Input: Feedback data sent from the server

[0397] Output: On-screen feedback

[0398] What it does: The device displays feedback on the screen saying, "The vibrato duration is too short, try to hold it for a little longer."

[0399] (Application example 3)

[0400] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0401] Conventional karaoke systems and voice training systems have limited functionality for evaluating a user's singing ability, making it difficult for users to objectively evaluate their own singing ability and receive specific feedback to improve it. In addition, there has been a lack of systems that allow users to effectively train to improve their singing ability.

[0402] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[0403] In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score using a karaoke function, a system including a voice training function, a means for reading voice data and extracting features, a means for evaluating singing ability based on the extracted features, and a means for displaying the evaluation results, thereby enabling users to objectively evaluate their own singing ability and receive specific feedback.

[0404] The "means for creating a playlist based on user input" is a means for automatically creating a list of songs that match the user's preferences based on information entered by the user.

[0405] "Means for recommending music that takes into consideration the user's preferred music style, language, etc." refers to a means for recommending music that is suitable for a user based on information such as the user's past music selection history, preferred music style, language, etc.

[0406] "Means for creating new music by mixing favorite songs" refers to a means for combining multiple songs selected by the user to create new music.

[0407] The "means for displaying a score in the karaoke function" is a means for evaluating the results of the user's singing and displaying the evaluation as a score.

[0408] The "voice training function" is a function that provides training to help users improve their singing ability.

[0409] The "means for reading voice data and extracting features" refers to a means for reading voice data sung by a user and extracting features from the voice data.

[0410] The "means for evaluating singing ability based on extracted features" is a means for evaluating the singing ability of a user based on the extracted feature amounts of voice data.

[0411] The "means for displaying the evaluation result" is a means for visually displaying the evaluation result of singing ability to the user.

[0412] The following system configuration will be described as an embodiment of the present invention.

[0413] System Configuration

[0414] This system includes a means for creating a playlist based on user input, a means for recommending songs taking into consideration the user's preferred style, language, etc., a means for creating new songs by mixing favorite songs, a means for displaying scores using a karaoke function, a voice training function, a means for reading voice data and extracting features, a means for evaluating singing ability based on the extracted features, and a means for displaying the evaluation results.

[0415] Program processing

[0416] The server creates a playlist based on information entered by the user using their smartphone. The server takes into consideration the user's mood and past music selection history when creating the playlist. It then recommends songs based on the user's preferred style and language. It is also possible to create new songs by mixing the songs selected by the user.

[0417] When a user uses the karaoke function, their smartphone records their singing and sends the audio data to a server. The server uses Librosa to extract features from the audio data and evaluates their singing ability using a generative AI model trained with TENSORFLOW (registered trademark). The evaluation results are displayed to the user as a score.

[0418] The voice training function provides training programs for users to practice specific singing techniques. The user's singing voice data is sent back to the server, where features are extracted and evaluated in the same way. Based on the evaluation results, specific feedback is provided to the user.

[0419] Hardware and software used

[0420] Hardware: Smartphone

[0421] Software: Python, Librosa, TensorFlow

[0422] Specific examples

[0423] For example, when a user sings "Song 1," the audio data recorded on the smartphone is sent to the server. The server uses Librosa to extract MFCCs (Mel Frequency Cepstrum Coefficients) from the audio data and evaluates the user's singing ability using a generative AI model trained with TensorFlow. The evaluation result is displayed to the user as "Your vocal performance score is: 85.75."

[0424] Prompt Sentence Examples

[0425] An example of a prompt to input to a generative AI model is as follows:

[0426] Load the audio file sung by the user, extract the MFCCs, input them into the singing ability evaluation model, and calculate the score.

[0427] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[0428] Step 1:

[0429] A user starts a karaoke application using a smartphone and selects a song they want to sing.

[0430] Input: User song selection

[0431] Output: Selected song information

[0432] Specific operation: The user operates the app interface to select a song, and the selection information is saved within the app.

[0433] Step 2:

[0434] The user begins singing along with the selected song, and the smartphone records the voice.

[0435] Input: User's singing voice

[0436] Output: Recorded audio data

[0437] Specific operation: The user's singing voice is recorded using the smartphone's microphone and saved as audio data.

[0438] Step 3:

[0439] The recorded audio data is sent to the server.

[0440] Input: Recorded audio data

[0441] Output: Audio data sent to the server

[0442] Specific operation: The smartphone uploads the voice data to the server via the Internet.

[0443] Step 4:

[0444] The server uses Librosa to extract MFCCs (Mel-Frequency Cepstral Coefficients) from the audio data.

[0445] Input: Audio data sent to the server

[0446] Output: Extracted MFCC features

[0447] Specific operation: The server analyzes the audio data using the Librosa library and calculates MFCC features.

[0448] Step 5:

[0449] The server uses a generative AI model trained with TensorFlow to evaluate singing ability based on the extracted MFCC features.

[0450] Input: Extracted MFCC features

[0451] Output: Singing ability evaluation score

[0452] Specific operation: The server inputs MFCC features into the generative AI model, and the model outputs a singing ability evaluation score.

[0453] Step 6:

[0454] The server sends the evaluation results to the smartphone.

[0455] Input: Singing ability evaluation score

[0456] Output: Evaluation results sent to your smartphone

[0457] Specific operation: The server sends the evaluation score to the smartphone and provides feedback to the user.

[0458] Step 7:

[0459] The smartphone displays the evaluation results to the user.

[0460] Input: Evaluation results sent to your smartphone

[0461] Output: The evaluation results displayed to the user

[0462] Specific behavior: The evaluation score is displayed on the smartphone screen so that the user can check the results.

[0463] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0464] "Example 1"

[0465] In one embodiment of the present invention, a system incorporating an emotion engine is provided. The system recognizes a user's emotion and creates a playlist based on that emotion. Specifically, if a user indicates to the system that they are feeling "sad," the emotion engine receives that information and creates a playlist by selecting songs that fit the sad mood.

[0466] "Example 2"

[0467] The emotion engine also recommends songs based on the user's emotions. For example, if the user says "

[0468] If a user expresses the emotion "fun," the emotion engine will receive that information and recommend songs that fit their happy mood. This recommendation can also take into account the user's past song selection history, preferred music style, language, etc.

[0469] "Example 3"

[0470] Furthermore, the emotion engine also displays karaoke scores and provides vocal training according to the user's emotions. For example, if the user expresses the emotion "nervous," the emotion engine receives that information and provides vocal training to help them relax. Also, if the user expresses the emotion "confident," the emotion engine receives that information and provides the optimal musical experience according to the user's emotions, such as by setting the karaoke score display more strictly.

[0471] The processing flow of each embodiment will be described below.

[0472] "Example 1"

[0473] Step 1: The user indicates an emotion to the system, for example, "sad."

[0474] Step 2: The emotion engine receives input from the user.

[0475] Step 3: Based on the emotion received by the emotion engine, select a song that matches the sad mood.

[0476] Step 4: Create a playlist based on the selected songs and provide it to the user.

[0477] "Example 2"

[0478] Step 1: The user expresses an emotion to the system. For example, the user expresses the emotion "fun."

[0479] Step 2: The emotion engine receives input from the user.

[0480] Step 3: Based on the emotion received by the emotion engine, select a song that matches the happy mood.

[0481] Step 4: Recommend the selected song to the user. This recommendation can take into account the user's past song selection history, preferred style, language, etc.

[0482] "Example 3"

[0483] Step 1: The user indicates their feelings to the system. For example, they indicate that they are "nervous."

[0484] Step 2: The emotion engine receives input from the user.

[0485] Step 3: Based on the emotion received by the emotion engine, provide voice training for relaxation.

[0486] Step 4: If the user expresses the emotion "confident," the emotion engine receives that information and provides the optimal music experience according to the user's emotions, such as by setting the karaoke score display more strictly.

[0487] Example 1

[0488] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0489] Conventional music recommendation systems have difficulty automatically generating playlists based on the user's mood and emotions, making it difficult to provide songs that perfectly match the user's preferences. Furthermore, there is a lack of technology to accurately analyze user input and recommend appropriate songs based on the results. Furthermore, there is also a lack of means to quickly and efficiently provide the generated playlists to the user's device.

[0490] The specification process by the specification processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means. In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into account the user's preferred music genre, language, etc., an emotion engine means for recognizing the user's emotions and creating a playlist based on those emotions, a natural language processing means for analyzing the user's input, and a means for transmitting the created playlist to the user's terminal. This makes it possible to automatically create a playlist based on the user's mood and emotions, and to quickly and efficiently provide songs that match the user's preferences.

[0491] The "means for creating a playlist based on user input" is a means for automatically generating a playlist by selecting appropriate songs based on the mood and preference information input by the user through the interface.

[0492] "Means for recommending music that take into consideration the user's preferred music genre, language, etc." refers to a means for recommending appropriate music based on the user's specified preferences for music genre and language.

[0493] The "emotion engine means for recognizing the user's emotions and creating a playlist based on those emotions" is a means for analyzing the user's emotions, selecting songs that match those emotions, and creating a playlist.

[0494] The "natural language processing means for analyzing user input" is a means for analyzing text information entered by a user and using natural language processing technology to understand the meaning.

[0495] The "means for transmitting the generated playlist to the user's terminal" refers to a means for transmitting the playlist generated by the server to the user's terminal.

[0496] MODE FOR CARRYING OUT THE INVENTION

[0497] This invention is a system equipped with an interface for inputting a user's mood, preferred music genre, language, etc., and automatically creates a playlist based on the user's input. Specific embodiments of this system are described below.

[0498] Providing a user interface

[0499] It provides an interface for users to input their mood, preferred music genre, language, etc. This interface is often implemented as a web application or mobile application. For example, it could be a web application using HTML5 and JavaScript (registered trademark) or an iOS application using Swift. Users enter information using text boxes and drop-down menus.

[0500] Receiving and Parsing User Input

[0501] The server receives information entered by the user through the interface. The entered information includes the user's mood (e.g., "I want to relax") and their preferred music genre and language (e.g., "Jazz" or "Japanese"). The server analyzes this information and processes it as data to generate an appropriate playlist. Natural language processing (NLP) techniques are used for the analysis. For example, Python's NLTK library or the Google Cloud Natural Language API could be used.

[0502] Automatic playlist generation

[0503] The server automatically creates a playlist based on the user's input. This process uses an emotion engine and a music database. The emotion engine recognizes the user's emotion and selects music that matches that emotion. For example, if the user enters "sad," the emotion engine analyzes that information and selects music that matches the sad mood. The music database often uses the API of a music streaming service (e.g., Spotify API, Apple Music API).

[0504] Providing playlists

[0505] The generated playlist is sent to the user's device. Specifically, the playlist information is sent as an HTTP response. The user can then play the provided playlist using a music streaming service application such as Spotify or Apple Music.

[0506] Specific examples

[0507] Example 1: When you want to relax

[0508] The user enters "I want to relax" into the interface. The server receives this information and extracts the keyword "relax" using natural language processing technology. The emotion engine selects songs that fit the relaxing mood based on the keyword "relax." For example, jazz or classical music is often selected. The generated playlist is sent to the user's device as an HTTP response. The user then opens the Spotify app and plays the playlist.

[0509] Example 2: Preferences for specific music genres or languages

[0510] The user inputs that they like "jazz" and "Japanese." The server receives this information and uses natural language processing technology to extract the keywords "jazz" and "Japanese." The emotion engine then selects jazz songs with Japanese lyrics based on these keywords. The generated playlist is sent to the user's device as an HTTP response. The user then opens the Apple Music app and plays the playlist.

[0511] Prompt Sentence Examples

[0512] Prompt 1: If you want to relax

[0513] If a user types "I want to relax," create a playlist with songs that fit that relaxing mood.

[0514] Prompt 2: Preferences for specific music genres or languages

[0515] If a user inputs that they like "jazz" and "Japanese," create a playlist by selecting jazz songs with Japanese lyrics.

[0516] In this way, a system is provided that automatically generates a playlist according to the user's mood and preferences.

[0517] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0518] Step 1: Provide a user interface

[0519] It provides an interface for users to input their mood, preferred music genre, language, etc. Specifically, it is implemented as a web application or mobile application. For example, it could be a web application using HTML5 and JavaScript, or an iOS application using Swift. Users enter information using text boxes and drop-down menus. Inputs include the user's mood (e.g., "I want to relax") and preferred music genre and language (e.g., "Jazz," "Japanese"). The output is the user's input information.

[0520] Step 2: Receiving User Input

[0521] The server receives information entered by the user through the interface. Specifically, the information entered by the user is sent to the server as an HTTP request. The server analyzes the received request and extracts the user's mood and preferences. The input includes the user's input information. The output is the analyzed information on the user's mood and preferences.

[0522] Step 3: Parsing User Input

[0523] The server analyzes the received information. Specifically, it uses natural language processing (NLP) technology to analyze the user's input. For example, it could use Python's NLTK library or the Google Cloud Natural Language API. If the user inputs "I want to relax," the server analyzes the information and extracts the keyword "relax." The input includes information about the user's mood and preferences. The output is the analyzed keyword.

[0524] Step 4: Automatically generate a playlist

[0525] The server automatically creates a playlist based on user input. Specifically, it uses an emotion engine and a music database. The emotion engine recognizes the user's emotion and selects music that matches that emotion. For example, if the user enters "sad," the emotion engine analyzes that information and selects music that matches the sad mood. The music database often uses the API of a music streaming service (e.g., Spotify API, Apple Music API). The input includes the analyzed keywords. The output is the generated playlist.

[0526] Step 5: Serve the playlist

[0527] The generated playlist is sent to the user's device. Specifically, the playlist information is sent as an HTTP response. The user can then play the provided playlist. To play, they use a music streaming service application such as Spotify or Apple Music. The input includes the generated playlist. The output is the playlist sent to the user's device.

[0528] (Application example 1)

[0529] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0530] Conventional music recommendation systems simply recommend songs based on the user's past music selection history and preferences, without considering the user's mood or emotions. This makes it difficult for the user to find songs that match the user's current mood or emotions, resulting in low satisfaction. In addition, it is inconvenient because the user has to input their mood and emotions.

[0531] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score in a karaoke function, a voice training function, an emotion recognition means for recognizing the user's emotion, and a means for automatically generating a playlist based on the recognized emotion. This makes it possible to automatically recommend songs that match the user's current mood and emotion, thereby improving user satisfaction.

[0532] The "means for creating a playlist based on user input" is a function that selects appropriate songs and creates a playlist based on information entered by the user.

[0533] The "means for recommending music in consideration of the user's preferred music style, language, etc." is a function that recommends the most suitable music in consideration of information such as the user's preferred music style and language.

[0534] "Means for creating new music by mixing favorite songs" is a function that allows users to combine multiple songs selected by the user to create new music.

[0535] The "means for displaying scores in the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[0536] The "voice training function" is a training function for improving the user's singing ability.

[0537] The "emotion recognition means for recognizing the user's emotions" is a function for analyzing and recognizing emotions from the user's facial expressions, voice, etc.

[0538] The "means for automatically generating a playlist based on the recognized emotion" is a function for automatically generating a playlist by selecting appropriate songs based on the recognized emotion of the user.

[0539] A system for carrying out this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred musical style, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying scores in a karaoke function, a voice training function, an emotion recognition means for recognizing the user's emotions, and a means for automatically generating a playlist based on the recognized emotions.

[0540] Hardware and software used

[0541] Hardware: Smartphone (camera, microphone)

[0542] software:

[0543] Emotion recognition engine (e.g., Microsoft® Azure® Emotion API)

[0544] Music recommendation systems (e.g., Spotify API)

[0545] User interface (e.g., React Native (registered trademark))

[0546] System processing overview

[0547] The server receives the mood, preferred music style, and language input by the user through their smartphone and creates a playlist based on that.The user interface is built using React Native, providing an interface for users to input their mood, preferred music style, and language.

[0548] The emotion recognition method uses the Microsoft Azure Emotion API to analyze the user's facial expressions and voice using the smartphone's camera and microphone. The recognized emotion information is then used to automatically generate playlists.

[0549] As a music recommendation method, the Spotify API is used to recommend the best songs based on user input and recognized emotions. This makes it possible to automatically recommend songs that match the user's current mood and emotions, improving user satisfaction.

[0550] Specific examples

[0551] For example, if a user inputs "I want to relax," a playlist is created by selecting songs that match that mood. The emotion recognition means recognizes the emotion "I want to relax" from the user's facial expressions and voice, and automatically generates a playlist based on that.

[0552] Prompt Sentence Examples

[0553] Build an application that, when a user types "I want to relax," selects songs that match that mood and creates a playlist. Use the Microsoft Azure Emotion API for emotion recognition and the Spotify API for playlist generation. Build the user interface using React Native.

[0554] In this way, a system can be realized that recommends optimal music based on the user's mood and emotions.

[0555] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0556] Step 1:

[0557] The user inputs their mood, preferred music style, and language through the smartphone's user interface. The input data is sent to the server. Examples of input data include "I want to relax," "pop music," and "English music."

[0558] Step 2:

[0559] The server receives the input data and activates the emotion recognition means. It uses the smartphone's camera and microphone to capture the user's facial expressions and voice, and sends them to an emotion recognition engine (e.g., Microsoft Azure Emotion API). The emotion recognition engine analyzes the captured data and recognizes the user's emotion. For example, it may recognize the emotion "I want to relax."

[0560] Step 3:

[0561] The server combines the recognized emotion information with the user's input data to prepare data for generating a playlist. Specifically, it creates a query to select appropriate songs based on the user's mood, preferred musical style, and language.

[0562] Step 4:

[0563] The server sends a query to a music recommendation system (e.g., Spotify API) to recommend songs that match the user's mood and preferences. The Spotify API searches for songs based on the query and returns a list of recommended songs to the server. For example, pop English songs that match the user's mood of "wanting to relax" are recommended.

[0564] Step 5:

[0565] The server receives the recommended song list and generates a playlist, which is then sent to the user's smartphone, where the user can play the playlist through an application on the smartphone.

[0566] Step 6:

[0567] When a user plays a playlist, they can use the karaoke and voice training features. The karaoke feature displays a score for each song the user sings. The voice training feature provides training to improve the user's singing ability.

[0568] In this way, a system is realized that recommends optimal music based on the user's mood and emotions, thereby improving user satisfaction.

[0569] Example 2

[0570] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0571] Conventional music playback systems lack the ability to recommend songs based on the user's emotions and preferences, or to create new songs by mixing multiple songs. Furthermore, since they are unable to recommend songs based on the user's emotions, it is difficult to increase user satisfaction. Furthermore, since karaoke and voice training functions are not integrated, users are unable to enjoy a diverse musical experience with a single system.

[0572] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0573] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for creating new songs by mixing favorite songs, a means for recommending songs based on the user's emotions, a means for displaying a score in a karaoke function, and a voice training function. This makes it possible to recommend songs according to the user's emotions and preferences, and to create new songs. Furthermore, by integrating the karaoke and voice training functions, the user can enjoy a variety of musical experiences in one system.

[0574] The "means for creating a playlist based on user input" is a function that automatically creates a list of songs based on information entered by the user.

[0575] "Means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music genre, language, etc., and recommends appropriate music based on that.

[0576] "Means for creating new music by mixing favorite songs" is a function that combines multiple songs selected by the user to create new music.

[0577] The "means for recommending music based on the user's emotions" is a function that recommends music that suits the emotions based on the emotional information input by the user.

[0578] The "means for displaying a score in the karaoke function" is a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays the result as a score.

[0579] The "voice training function" is a training function aimed at improving the user's singing ability, and is a function that allows for vocal practice and pitch checking.

[0580] MODE FOR CARRYING OUT THE INVENTION

[0581] This invention is a system including a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred musical style, language, etc., a means for creating a new song by mixing favorite songs, a means for recommending songs based on the user's emotions, a means for displaying scores in a karaoke function, and a voice training function.

[0582] Hardware and software used

[0583] Hardware

[0584] Devices: smartphones, PCs, tablets, etc.

[0585] Server: Cloud server, on-premise server

[0586] software

[0587] Music processing library: LibROSA

[0588] Music generation AI model: OpenAI's Jukedeck

[0589] Sentiment Analysis Engine: IBM Watson® Sentiment Analysis API

[0590] Recommendation Engine: Spotify's Recommendation API

[0591] Program processing explanation

[0592] Music Mix function

[0593] 1. The user selects a song

[0594] A user logs in to the system using a terminal and selects the songs they want to mix. Specifically, the user selects "Song A" and "Song B" on the system interface.

[0595] 2. The server receives the data

[0596] The server receives the data of the user's selected "Song A" and "Song B." Specifically, the server retrieves the song metadata and audio files via an HTTP request.

[0597] 3. The server mixes the music

[0598] The server analyzes the beats and melodies of the selected songs and combines them to create a new song. Specifically, the server uses the LibROSA library to extract the beats of "Song A" and combines them with the melody of "Song B" using OpenAI's Jukedeck.

[0599] 4. The server provides new music

[0600] The server then sends the newly generated song, "Song C," to the user's device. Specifically, the server returns the generated audio file as an HTTP response, which the user can download or stream.

[0601] Music recommendation function using an emotional engine

[0602] 1. User inputs emotion

[0603] The user inputs their emotions into the system using a terminal. Specifically, the user selects the emotion "fun" on the system interface.

[0604] 2. The server receives the emotion data

[0605] The server receives the user's emotion data. Specifically, the server obtains the emotion data through an HTTP request.

[0606] 3. The server recommends songs

[0607] The server recommends appropriate songs based on the user's emotional data, past song selection history, preferred music style, language, etc. Specifically, the server analyzes the emotional data using IBM Watson's sentiment analysis API and recommends songs using Spotify's recommendation API.

[0608] 4. The server provides recommended songs

[0609] The server sends the recommended songs to the user's device. Specifically, the server returns a list of recommended songs as an HTTP response, and the user can play them.

[0610] Examples of concrete examples and prompts

[0611] Specific examples

[0612] The user selects "Song A" and "Song B" and generates a new "Song C."

[0613] The user selects "Song A" and "Song B" on the terminal.

[0614] The server receives the selected song data, extracts the beat using LibROSA, and combines the melody using Jukedeck.

[0615] The server provides the generated "Song C" to the user.

[0616] The user inputs the emotion "fun," and the server recommends "happy songs."

[0617] The user inputs the emotion "fun" into the terminal.

[0618] The server receives the emotion data and analyzes it using IBM Watson.

[0619] The server uses Spotify's API to recommend "happy songs" and provide them to the user.

[0620] Prompt Sentence Examples

[0621] "Mix song A and song B selected by the user to create a new song."

[0622] "Recommend suitable music when the user expresses joy."

[0623] In this way, the system performs specific processing to generate and recommend music based on the user's selection and emotions.

[0624] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0625] Music Mix function

[0626] Step 1: User selects a song

[0627] Input: The user uses the terminal to select the songs they want to mix.

[0628] Specific operation: The user selects "Song A" and "Song B" on the system interface.

[0629] Output: The information of the selected song (metadata and audio file path) is sent to the server.

[0630] Step 2: The server receives the data

[0631] Input: Information about "Song A" and "Song B" selected by the user.

[0632] What happens: The server retrieves the song's metadata and audio file via an HTTP request.

[0633] Output: The captured audio files and metadata are saved on the server.

[0634] Step 3: The server mixes the music

[0635] Input: Saved audio files and metadata for "Song A" and "Song B".

[0636] How it works: The server uses the LibROSA library to extract the beat of "Song A" and combines it with the melody of "Song B" using OpenAI's Jukedeck.

[0637] Output: An audio file of the newly generated song "Song C".

[0638] Step 4: The server serves up a new song

[0639] Input: The generated audio file for "Song C".

[0640] Specific operation: The server returns the generated audio file as an HTTP response.

[0641] Output: The audio file of "Song C" is sent to the user's device, where they can download or stream it.

[0642] Music recommendation function using an emotional engine

[0643] Step 1: User enters emotion

[0644] Input: The user uses a terminal to input their emotions into the system.

[0645] Specific action: The user selects the emotion "fun" on the system interface.

[0646] Output: The input emotion data is sent to the server.

[0647] Step 2: The server receives the emotion data

[0648] Input: Emotion data entered by the user.

[0649] Specific operation: The server obtains emotion data through an HTTP request.

[0650] Output: The acquired emotion data is stored in the server.

[0651] Step 3: The server recommends songs

[0652] Input: Stored emotional data, the user's past music selection history, preferred music style, language, and other data.

[0653] Specific operation: The server analyzes the sentiment data using IBM Watson's sentiment analysis API and recommends songs using Spotify's recommendation API.

[0654] Output: A list of recommended songs.

[0655] Step 4: The server provides song recommendations

[0656] Input: A list of recommended songs.

[0657] Specific operation: The server returns a list of recommended songs as an HTTP response.

[0658] Output: A list of recommended songs is sent to the user's device, where the user can play them.

[0659] (Application example 2)

[0660] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0661] Conventional music recommendation systems and playlist creation systems do not fully consider the user's emotions and moods when recommending music. Furthermore, the functionality for creating new music by mixing user-selected songs is limited, making it difficult to meet the diverse needs of users. Furthermore, the lack of integrated karaoke and voice training functions means that it is difficult to provide a comprehensive music experience.

[0662] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0663] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for creating new songs by mixing favorite songs, an emotion engine means that recommends songs based on the user's emotions, a means for displaying scores in the karaoke function, and a voice training function. This makes it possible to recommend songs that take into account the user's emotions and mood, to create new songs by mixing selected songs, and to provide a comprehensive music experience that integrates the karaoke function and voice training function.

[0664] The "means for creating a playlist based on user input" is a function that automatically generates a playlist based on the songs and conditions selected by the user.

[0665] "Means for recommending music that takes into consideration the user's preferred music style, language, etc." is a function that recommends appropriate music based on the user's past music selection history, preferred music genre, language used, etc.

[0666] "Means to mix your favorite songs to create new music" is a function that generates new music by combining the beats and melodies of multiple songs selected by the user.

[0667] The "emotion engine means for recommending music based on the user's emotions" is a function that analyzes the user's current emotional state and recommends music that matches that emotion.

[0668] The "means for displaying a score using the karaoke function" is a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays a score.

[0669] The "voice training function" is a training function aimed at improving the user's singing ability, and is a function that allows for vocal practice and pitch correction.

[0670] A system for carrying out this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, an emotion engine means for recommending songs based on the user's emotions, a means for displaying scores in a karaoke function, and a voice training function.

[0671] System Program

[0672] The system is implemented using the following hardware and software:

[0673] Hardware

[0674] Smartphone

[0675] head-mounted display

[0676] software

[0677] Python

[0678] Music MIX library (e.g. music_mixer)

[0679] Emotion engine library (e.g. emotion_engine)

[0680] Processing Description

[0681] Playlist creation method

[0682] The server automatically generates a playlist based on the songs selected by the user and other conditions. For example, if a user inputs "I want to relax," a playlist of songs suitable for relaxation will be generated.

[0683] Music recommendation method

[0684] The server recommends appropriate songs based on the user's past music selection history, preferred music genre, language used, etc. For example, if the user has listened to many pop songs in the past, pop songs will be mainly recommended.

[0685] Music Mixing Method

[0686] The server generates a new song by combining the beats and melodies of multiple songs selected by the user. For example, if a user selects "Song A" and "Song B," the server will generate "Song C" by combining the features of those songs.

[0687] Emotion Engine Means

[0688] The server analyzes the user's current emotional state and recommends music that matches that emotion. For example, if the user expresses the emotion "happy," music that matches that happy mood will be recommended.

[0689] Karaoke function

[0690] The server evaluates the pitch and rhythm of the song sung by the user and displays a score. For example, after a user sings karaoke, a score of 90 is displayed.

[0691] Voice training function

[0692] The server provides training functions aimed at improving the user's singing ability, such as vocal practice and pitch correction.

[0693] Specific examples

[0694] When a user selects and mixes "songA" and "songB," a new song "MixedSong" is generated.

[0695] If the user expresses the emotion "happy," "RecommendedSong1" and "RecommendedSong2" are recommended taking into account the user's past history and preferences.

[0696] Prompt Sentence Examples

[0697] Create an application that mixes user-selected songs to generate new songs and recommends songs based on the user's emotions, taking into account the user's past song selection history and preferred style and language.

[0698] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0699] Step 1:

[0700] The user selects songs using the terminal. The user selects multiple songs through the application interface and specifies the songs to be mixed. The input is a list of songs selected by the user, and the output is a list of selected songs.

[0701] Step 2:

[0702] The server receives the selected song list and generates a new song using a music mix library. Specifically, it uses the music_mixer library to combine the beats and melodies of the selected songs. The input is the selected song list, and the output is the generated new song.

[0703] Step 3:

[0704] The user inputs their current emotion using the terminal. The user selects an emotion such as "happy" or "sad" through the application interface. The input is the emotion selected by the user, and the output is the selected emotion.

[0705] Step 4:

[0706] The server receives the selected emotion and recommends songs using the emotion engine library. Specifically, the emotion_engine library is used to recommend songs taking into account the user's emotion, past music selection history, and preferred music genres. The input is the selected emotion and the user's past music selection history and preference data, and the output is a list of recommended songs.

[0707] Step 5:

[0708] A user uses the karaoke function on a terminal. The user selects the karaoke mode through the application interface and starts singing. The input is the user's singing data, and the output is the score displayed by the karaoke function.

[0709] Step 6:

[0710] The server analyzes the user's singing data and evaluates pitch, rhythm, etc. Specifically, it uses a voice analysis algorithm to evaluate the user's singing data and calculate a score. The input is the user's singing data, and the output is the calculated score.

[0711] Step 7:

[0712] A user uses the voice training function on a terminal. The user selects the voice training mode through the application interface and starts training. The input is the user's singing data, and the output is the voice training feedback.

[0713] Step 8:

[0714] The server analyzes the user's singing data and provides feedback such as vocal practice and pitch correction. Specifically, it uses a voice analysis algorithm to evaluate the user's singing data and provides feedback on areas for improvement. The input is the user's singing data, and the output is the feedback content.

[0715] Example 3

[0716] Next, a description will be given of a third embodiment of the third embodiment. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0717] Conventional karaoke and voice training systems provide uniform evaluations and training without considering the user's emotional state, making it difficult to provide an optimal musical experience that suits the user's psychological state. Furthermore, they lacked the functionality to adjust karaoke score displays and voice training content based on the user's emotions, making it difficult to increase user satisfaction.

[0718] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[0719] In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred style, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score in a karaoke function, and a voice training function, thereby enabling a system including an emotion engine that analyzes the user's emotions and adjusts the karaoke score display and the content of the voice training.

[0720] The "playlist creation means" is a function that generates a list of songs based on user input.

[0721] The "music recommendation means" is a function that suggests appropriate music by taking into consideration the user's preferred musical style, language, etc.

[0722] The "music mix means" is a function that allows a user to combine their favorite songs to create new music.

[0723] The "karaoke function" is a function that allows the user to sing along with a song selected by the user and displays the singing result in the form of a score.

[0724] The "voice training function" is a function that provides training to help users improve their singing ability.

[0725] The "emotion engine" is a function that analyzes the user's emotions and adjusts the karaoke score display and voice training content based on the analysis results.

[0726] This invention is a system including a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying scores in a karaoke function, a voice training function, and an emotion engine that analyzes the user's emotions and adjusts the karaoke score display and the content of voice training.

[0727] Hardware and software used

[0728] This system uses the following hardware and software:

[0729] Terminal: The device operated by the user (smartphone, tablet, PC, etc.)

[0730] Server: A remote server for data analysis and evaluation.

[0731] Microphone: An input device for recording the user's singing voice

[0732] Speaker: An output device for playing music and training audio.

[0733] Emotion Engine: A software module for analyzing user emotions

[0734] Generative AI model: an algorithm for analyzing and evaluating the user's singing voice

[0735] Program processing

[0736] Karaoke function

[0737] 1. The user selects the karaoke function.

[0738] 2. The device plays the song selected by the user.

[0739] 3. The user sings along to the song.

[0740] 4. The device records the user's singing voice through the microphone.

[0741] 5. The server analyzes the recorded singing voice.

[0742] 6. The server calculates the evaluation results as a score.

[0743] 7. The device will display the score on the screen.

[0744] Voice training function

[0745] 1. The user selects the voice training function.

[0746] 2. The device selects the specific singing technique the user wants to practice.

[0747] 3. The device plays a training program based on the selected technique.

[0748] 4. The user sings according to the training program.

[0749] 5. The device records the user's singing voice through the microphone.

[0750] 6. The server analyzes the recorded singing voice.

[0751] 7. The server calculates the evaluation results as a score.

[0752] 8. The device will display the score on the screen.

[0753] Emotion Engine

[0754] 1. The user inputs an emotion while using the karaoke function or voice training function.

[0755] 2. The device sends the user's emotion data to the emotion engine.

[0756] 3. The server's emotion engine analyzes the user's emotions.

[0757] 4. The server adjusts the karaoke score display and voice training content based on the analysis results.

[0758] 5. The device provides the adjusted content to the user.

[0759] Specific examples

[0760] Karaoke function example

[0761] The user selects "Song 1" and begins singing.

[0762] The device plays the song and records the user's singing.

[0763] The server analyzes the recorded singing voice and evaluates pitch and rhythm.

[0764] The server calculates the evaluation result as a score and sends it to the terminal.

[0765] The device displays "85 points" on the screen.

[0766] Examples of voice training functions

[0767] The user selects "Practice Vibrato."

[0768] The device plays a vibrato training program.

[0769] The user sings according to the program.

[0770] The device records the singing voice and sends it to the server.

[0771] The server analyzes the recorded singing voice and evaluates the degree of vibrato mastery.

[0772] The server calculates the evaluation result as a score and sends it to the terminal.

[0773] The device will display "70 points" on the screen.

[0774] Examples of emotion engines

[0775] A user inputs "I'm nervous" while using the karaoke function.

[0776] The terminal transmits the emotion data to the emotion engine.

[0777] The server's emotion engine analyzes the emotion "tense."

[0778] The server provides voice training for relaxation.

[0779] The device plays relaxing voice training.

[0780] Prompt Sentence Examples

[0781] "Describe a program that uses a karaoke function to play a song selected by the user and display a score for the singing result."

[0782] "Describe a program that uses voice training features to help users practice specific singing techniques and evaluate the results."

[0783] "Please explain a program that uses an emotion engine to display karaoke scores and provide voice training according to the user's emotions." The flow of the specific processing in the third embodiment will be explained with reference to FIG.

[0784] Karaoke function processing steps

[0785] Step 1:

[0786] The user selects the karaoke function.

[0787] Input: The user taps the "Karaoke" button on the device screen.

[0788] Output: Karaoke function is activated.

[0789] Step 2:

[0790] The device plays the song selected by the user.

[0791] Input: The user selects a song.

[0792] Output: The device retrieves the song data from the server and plays it through the speaker.

[0793] Step 3:

[0794] The user sings along to the song.

[0795] Input: User sings into a microphone.

[0796] Output: The user's singing voice is input to the device through a microphone.

[0797] Step 4:

[0798] The device records the user's singing voice through a microphone.

[0799] Input: User's singing voice.

[0800] Output: The recorded vocal data is saved on the device.

[0801] Step 5:

[0802] The server analyzes the recorded singing voice.

[0803] Input: Recording data sent from the device.

[0804] Output: Analysis results such as pitch, rhythm, and vocal strength.

[0805] Step 6:

[0806] The server calculates the evaluation results as a score.

[0807] Input: Analysis results.

[0808] Output: Overall evaluation score.

[0809] Step 7:

[0810] The device will display the score on the screen.

[0811] Input: The score sent by the server.

[0812] Output: The score is displayed on the terminal screen.

[0813] Voice Training Function Processing Steps

[0814] Step 1:

[0815] The user selects the voice training function.

[0816] Input: The user taps the "Voice Training" button on the device screen.

[0817] Output: The voice training function is activated.

[0818] Step 2:

[0819] The device selects the particular singing technique that the user wants to practice.

[0820] Input: The user selects a particular technique from an on-screen menu.

[0821] Output: A training program based on the selected technique is determined.

[0822] Step 3:

[0823] The terminal plays a training program based on the selected technique.

[0824] Input: Selected training program.

[0825] Output: Training audio and video are played through speakers and a screen.

[0826] Step 4:

[0827] The user sings according to the training program.

[0828] Input: Training program instructions.

[0829] Output: The user's singing voice is input to the device through a microphone.

[0830] Step 5:

[0831] The device records the user's singing voice through a microphone.

[0832] Input: User's singing voice.

[0833] Output: The recorded vocal data is saved on the device.

[0834] Step 6:

[0835] The server analyzes the recorded singing voice.

[0836] Input: Recording data sent from the device.

[0837] Output: Analysis results assessing the mastery of specific techniques.

[0838] Step 7:

[0839] The server calculates the evaluation results as a score.

[0840] Input: Analysis results.

[0841] Output: Calculate the mastery of a particular technique as a score.

[0842] Step 8:

[0843] The device will display the score on the screen.

[0844] Input: The score sent by the server.

[0845] Output: The score is displayed on the terminal screen.

[0846] Emotion Engine Processing Steps

[0847] Step 1:

[0848] The user inputs emotions while using the karaoke function or the voice training function.

[0849] Input: The user taps the "Emotion Input" button on the device screen and selects an emotion.

[0850] Output: Emotion data is input to the terminal.

[0851] Step 2:

[0852] The terminal transmits the user's emotion data to the emotion engine.

[0853] Input: User emotion data.

[0854] Output: The emotion data is sent to the emotion engine on the server.

[0855] Step 3:

[0856] The server's emotion engine analyzes the user's emotions.

[0857] Input: Emotion data.

[0858] Output: Analysis results showing the user's emotional state.

[0859] Step 4:

[0860] Based on the analysis results, the server adjusts the karaoke score display and voice training content.

[0861] Input: Sentiment analysis results.

[0862] Output: Adjusted karaoke score display and voice training content.

[0863] Step 5:

[0864] The terminal provides the adjusted content to the user.

[0865] Input: The adjustment sent by the server.

[0866] Output: The adjusted content is displayed on the device screen.

[0867] (Application example 3)

[0868] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0869] Conventional karaoke systems and voice training systems provide uniform evaluations and training without considering the user's emotional state, making it difficult to provide an optimal musical experience that corresponds to the user's psychological state. Furthermore, even when used in physical stores, it is not possible to provide services that correspond to the user's emotions, making it difficult to improve user satisfaction.

[0870] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[0871] In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred musical style, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score in a karaoke function, a means including a voice training function, an emotion engine means for detecting the user's emotion, a means for adjusting the karaoke score display and voice training based on the detected emotion, and a smartphone application for use in a physical store. This provides an optimal music experience according to the user's emotional state, and makes it possible to improve user satisfaction even when using the system in a physical store.

[0872] The "playlist creation means" is a function that generates a list of songs based on user input.

[0873] The "music recommendation means" is a function that suggests appropriate music by taking into consideration the user's preferred musical style, language, etc.

[0874] The "music mix means" is a function that allows a user to combine their favorite songs to create new music.

[0875] The "karaoke function" is a function that displays the results of the user's singing in the form of a score.

[0876] The "voice training function" is a function that provides training to improve the user's singing ability.

[0877] The "emotion engine means" is a function that detects the user's emotions and adjusts the system's operation based on that information.

[0878] The "adjustment means based on emotion" is a function that changes the karaoke score display and voice training content according to the detected emotion of the user.

[0879] A "smartphone application used in a physical store" is an application that a user uses on a smartphone in a physical store.

[0880] As an embodiment of the present invention, a method for realizing an emotion-responsive karaoke and voice training system as a smartphone application for use in a physical store will be described.

[0881] System Program

[0882] The system implements a program that includes the following main functions:

[0883] 1. Playlist creation method: Generates a list of songs based on user input.

[0884] 2. Music recommendation method: Suggests appropriate music based on the user's preferred style and language.

[0885] 3. Music Mixing: Create new music by combining the user's favorite songs.

[0886] 4. Karaoke function: The user's singing results are displayed as a score.

[0887] 5. Voice training function: Provides training to improve the user's singing ability.

[0888] 6. Emotion engine means: Detects the user's emotions and adjusts the system's behavior based on that information.

[0889] 7. Emotion-based adjustment: The karaoke score display and voice training content are changed according to the detected user emotions.

[0890] Hardware and Software

[0891] This system is implemented using the following hardware and software:

[0892] Hardware: Smartphone

[0893] Software: Python, emotion engine (EmotionEngine class), karaoke system (KaraokeSystem class)

[0894] Data processing and calculation

[0895] The server receives the user's input data and generates a list of songs using a playlist creation means. It then uses a song recommendation means to suggest songs that take the user's preferences into consideration. It then uses a song mix means to create new songs based on the songs selected by the user.

[0896] When a user sings using the karaoke function, the results are displayed as a score. In addition, a voice training function is provided to help users improve their singing ability.

[0897] The emotion engine means detects the user's emotions and adjusts the karaoke score display and voice training content based on that information. For example, if the user is nervous, it provides training to help them relax, and if the user is confident, it gives a stricter evaluation.

[0898] Specific examples

[0899] As a concrete example, consider a scenario after a user sings in a karaoke booth. The user launches a smartphone application and uses the karaoke function to sing. The emotion engine detects the user's emotions and adjusts the score of the singing result. If the user is nervous, it provides training to help them relax, and if the user is confident, it gives a stricter evaluation.

[0900] Prompt Sentence Examples

[0901] "After a user sings in a karaoke booth, the emotion engine detects the user's emotions and adjusts the score of the singing result. If the user is nervous, it will provide training to help them relax, and if they are confident, it will give a stricter evaluation."

[0902] In this way, it is possible to provide the optimal music experience according to the user's emotional state, and to improve user satisfaction even when using the service in a physical store.

[0903] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[0904] Step 1:

[0905] The user starts the smartphone application and selects the karaoke function.

[0906] Input: User actions

[0907] Output: Karaoke function activation

[0908] Specific operation: The user selects the karaoke function from the application menu, and the karaoke screen is displayed.

[0909] Step 2:

[0910] The server creates a playlist based on the user's input.

[0911] Input: User song selection

[0912] Output: Playlist

[0913] Specific operation: The user selects their favorite songs, and the server receives that information and generates a playlist.

[0914] Step 3:

[0915] The server recommends songs taking into account the user's preferred style and language.

[0916] Input: User's past song selection history

[0917] Output: Recommended song list

[0918] Specific operation: The server analyzes the user's past music selection data and recommends songs that match their preferences.

[0919] Step 4:

[0920] The server mixes the user's favorite songs to create new music.

[0921] Input: User's favorite songs

[0922] Output: New Mixed Songs

[0923] Specific operation: The server combines the selected songs to generate a new song.

[0924] Step 5:

[0925] The user sings using the karaoke function.

[0926] Input: User's singing voice

[0927] Output: Singing result score

[0928] Specific operation: The user sings along with the selected song, and the system records the singing and displays the score.

[0929] Step 6:

[0930] The server detects the user's emotions using an emotion engine.

[0931] Input: User's voice data and facial expression data

[0932] Output: Detected emotion information

[0933] Specific operation: The server performs voice analysis and facial expression analysis to identify the user's emotions.

[0934] Step 7:

[0935] The server adjusts the karaoke score display and voice training content based on the detected emotions.

[0936] Input: Detected emotion information

[0937] Output: Adjusted score display and training content

[0938] Specific operation: The server adjusts the score display based on emotional information and provides appropriate voice training.

[0939] Step 8:

[0940] The user can check the adjusted score display and voice training.

[0941] Input: Adjusted score display and training content

[0942] Output: User feedback

[0943] Specific operation: The user checks the displayed score and training content and selects the next action.

[0944] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0945] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt including an instruction, voice data indicating a voice, a text message, and a text message.

[0946] Inference data, such as text data representing text and image data representing images, is input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with instructions indicated by prompts, and outputs the inference results in the form of data such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0947] Another example of generative AI is Gemini (registered trademark) (Internet search engine). <url: https: gemini.google.com ?hl="ja">) are mentioned.

[0948] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0949] [Second embodiment]

[0950] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0951] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0952] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0953] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0954] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0955] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0956] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0957] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0958] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0959] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0960] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0961] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[0962] "Example 1"

[0963] In one embodiment of the present invention, a system is provided that includes an interface that allows a user to input their mood, preferred musical style, language, etc. This system automatically creates a playlist based on the user's input. For example, if a user inputs "I want to relax," the system creates a playlist by selecting songs that match that mood. Furthermore, if the user indicates a preference for a particular musical style or language, the system recommends songs based on that preference.

[0964] "Example 2"

[0965] Furthermore, as an embodiment of the present invention, a function is provided that allows a user to select their favorite songs and mix them to create a new song. This function combines the beats and melodies of the selected songs to generate a new song. For example, if a user selects "Song A" and "Song B," the features of those songs are combined to create a new song called "Song C."

[0966] "Example 3"

[0967] In addition, the present invention provides a karaoke function and a voice training function. The karaoke function displays the user's singing results as a score. The voice training function provides training for the user to improve their singing ability. For example, when a user selects a specific song and sings along with the song, their singing ability is evaluated and a score is displayed. Furthermore, if a user wants to improve their singing ability, they can use the voice training function to practice specific singing techniques.

[0968] The processing flow of each embodiment will be described below.

[0969] "Example 1"

[0970] Step 1: The user inputs their mood, preferred melody, language, etc. through the system interface.

[0971] Step 2: The system automatically creates a playlist based on user input. For example, if a user inputs "I want to relax," the system will create a playlist by selecting songs that fit that mood.

[0972] Step 3: The system allows users to indicate preferences for specific musical styles and languages, and recommends songs based on those preferences.

[0973] "Example 2"

[0974] Step 1: The user selects his favorite song through the system interface.

[0975] Step 2: The system combines the beats and melodies of the selected songs to create a new song. For example, if a user selects "Song A" and "Song B," the system combines the features of those songs to create a new song called "Song C."

[0976] "Example 3"

[0977] Step 1: A user utilizes the system's karaoke function to select a particular song and sing along to it.

[0978] Step 2: The system evaluates the user's singing ability and displays a score.

[0979] Step 3: If the user wants to improve their singing ability, they can use the system's voice training feature to practice specific singing techniques.

[0980] Example 1

[0981] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0982] Conventional music playback systems lack the functionality to automatically create playlists based on the user's mood and preferences, resulting in the user having to manually select songs. Furthermore, since music recommendations cannot take into account the user's past song selection history or the user's mood that day, user satisfaction may decrease. Furthermore, few systems offer integrated karaoke and voice training functions, forcing users to use multiple applications.

[0983] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0984] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit for creating new songs by combining favorite songs, a karaoke score display unit, a voice training function, a user interface provision unit, a data analysis unit, a song search unit using the API of a music streaming service, and a unit for providing the generated playlist. This enables automatic creation of playlists tailored to the user's mood and preferences, thereby improving user satisfaction. Furthermore, by integrating the karaoke and voice training functions, users can utilize multiple functions in a single system.

[0985] The "means for creating a playlist based on user input" is a function that automatically creates a playlist based on the mood and preference information entered by the user.

[0986] The "means for recommending music that takes into consideration the user's preferred music style, language, etc." is a function that recommends appropriate music based on the user's preferred music style and language.

[0987] The "means of creating new music by combining favorite songs" is a function that allows a user to mix multiple songs selected by the user to create a new song.

[0988] The "means for displaying scores in the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[0989] The "voice training function" is a training function that allows the user to improve their singing ability.

[0990] The "means for providing a user interface" is a function that provides an interface for the user to input information about their moods and preferences.

[0991] The "means for receiving user input" is a function for receiving information input by a user through an interface.

[0992] "Means for analyzing data" is a function that analyzes information entered by the user and extracts their moods and preferences.

[0993] "Means for searching for songs using the API of a music streaming service" refers to a function that uses the API of a music streaming service to search for songs that match the user's mood and preferences.

[0994] The "means for providing the generated playlist" is a function for providing the generated playlist to the user.

[0995] MODE FOR CARRYING OUT THE INVENTION

[0996] The present invention is a music playback system that automatically creates a playlist based on the user's mood and preferences, and further integrates karaoke and voice training functions. Specific embodiments of this system are described below.

[0997] Providing a user interface

[0998] An interface is provided for users to input their mood, preferred melody, language, etc. The device displays this interface through a web browser or mobile application. For example, an interface that runs on a web browser can be built using HTML5 and JavaScript. Users can input their mood and preferences using text boxes and drop-down menus.

[0999] Receiving User Input

[1000] The server receives the information the user enters into the interface. If the user enters "I want to relax," that information is sent to the server via an HTTP request. The server temporarily stores the received data and proceeds to the next processing step.

[1001] Data analysis and processing

[1002] The server analyzes the received user input data. Using Python's natural language processing libraries (NLTK and spaCy), it analyzes the user's input text and extracts their mood and preferences. For example, from the input "I want to relax," it extracts the mood of "relaxation." Based on the results of this analysis, the server selects music in the next step.

[1003] Playlist Generation

[1004] The server selects songs that match the user's mood and preferences based on the analysis results. It uses the APIs of music streaming services such as Spotify API and Apple Music API. For example, you can use the Spotify API to search for songs that match "relaxation." The server generates a playlist from the search results and formats the information in JSON format.

[1005] Providing playlists

[1006] The generated playlist is provided to the user. The server sends the generated playlist information to the user's terminal. The terminal displays the received playlist information on an interface. The user can play the generated playlist. For example, each song in the playlist is displayed in list format, and a song can be played by clicking the play button.

[1007] Specific examples

[1008] Example 1: When a user enters "I want to relax"

[1009] 1. The user types "I want to relax" into the interface.

[1010] 2. The server receives this input and uses natural language processing to extract the mood "relaxed."

[1011] 3. The server uses the Spotify API to search for songs that match "relaxation."

[1012] 4. Generate a playlist from the search results and format it in JSON format.

[1013] 5. The server sends the generated playlist to the user's device, and the device displays the playlist on the interface.

[1014] Example 2: When a user enters "I want to listen to up-tempo English songs"

[1015] 1. The user types into the interface, "I want to listen to some up-tempo English songs."

[1016] 2. The server receives this input and uses natural language processing to extract the preferences of "uptempo" and "English."

[1017] 3. The server uses the Spotify API to search for "uptempo" and "English" songs.

[1018] 4. Generate a playlist from the search results and format it in JSON format.

[1019] 5. The server sends the generated playlist to the user's device, and the device displays the playlist on the interface.

[1020] Prompt Sentence Examples

[1021] Example 1: When you want to relax

[1022] If the user enters "I want to relax," generate a playlist by selecting songs that are suitable for relaxation.

[1023] Example 2: If you want to listen to up-tempo English songs

[1024] If a user types "I want to listen to up-tempo English songs," generate a playlist by selecting up-tempo English songs.

[1025] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1026] Step 1:

[1027] The user accesses the interface and inputs their mood, preferred melody, language, etc.

[1028] Input: A user enters text into the interface, such as "I want to relax" or "I want to listen to some upbeat English music."

[1029] Output: The user's input data is generated.

[1030] Specific behavior: The device displays an interface through a web browser or mobile application, and the user enters information using text boxes and drop-down menus.

[1031] Step 2:

[1032] The server receives the information entered by the user.

[1033] Input: Text data entered by the user into the interface.

[1034] Output: User input data sent to the server.

[1035] What happens: When the user completes the input and clicks the submit button, the data is sent to the server via an HTTP request, which the server then temporarily stores.

[1036] Step 3:

[1037] The server parses the received user input data.

[1038] Input: User input data stored on the server.

[1039] Output: Parsed mood and preference information.

[1040] How it works: The server uses Python's natural language processing libraries (NLTK and spaCy) to parse the user's input text and extract moods and preferences such as "relaxed" or "uptempo."

[1041] Step 4:

[1042] Based on the analysis results, the server selects music that matches the user's mood and preferences.

[1043] Input: Parsed mood and preference information.

[1044] Output: A list of selected songs.

[1045] Specific operation: The server uses the API of music streaming services such as Spotify API and Apple Music API to search for songs that match the user's mood and preferences. For example, it uses the Spotify API to search for songs that match "relaxation."

[1046] Step 5:

[1047] The server provides the generated playlist to the user.

[1048] Input: A list of selected songs.

[1049] Output: Playlist information sent to the user's device.

[1050] Specific operation: The server generates a playlist from the search results and formats the information in JSON format. The generated playlist is sent to the user's device, which displays the playlist on its interface. The user can then play the generated playlist.

[1051] (Application example 1)

[1052] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1053] Conventional music streaming services require users to manually create playlists that match their moods and preferences, which is a time-consuming process. Furthermore, music recommendations based on users' moods and preferences are often insufficient, resulting in low user satisfaction. Furthermore, it is difficult to automatically generate playlists that match users' moods and preferences, leaving a need for an improved user experience.

[1054] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1055] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for creating new songs by combining favorite songs, and a means for generating prompt sentences using a generative AI model and automatically generating a playlist based on the user's mood and preferences, thereby enabling the automatic generation of a playlist that matches the user's mood and preferences.

[1056] The "means for creating a playlist based on user input" is a function that automatically creates a playlist by selecting appropriate songs based on information entered by the user.

[1057] "Means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that recommends appropriate music based on the user's preferences and past selection history.

[1058] The "means for creating new music by combining favorite songs" is a function for creating new music by combining multiple songs selected by the user.

[1059] The "means for displaying scores in the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[1060] The "voice training function" is a training function for improving the user's singing ability.

[1061] "Means for generating prompt sentences using a generative AI model and automatically generating a playlist based on the user's mood and preferences" refers to a function that uses a generative AI model to convert user input information into prompt sentences, and automatically generates a playlist that matches the user's mood and preferences based on those prompt sentences.

[1062] A system for carrying out this invention includes a plurality of means for automatically creating a playlist based on user input, specifically a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred musical style, language, etc., a means for creating a new song by combining favorite songs, a means for displaying scores in a karaoke function, a voice training function, and a means for generating prompt sentences using a generative AI model to automatically generate a playlist based on the user's mood and preferences.

[1063] Program processing explanation

[1064] The server receives information such as mood, preferred melody, and language input from the user's smartphone or other device. Based on this input information, a generative AI model is used to generate a prompt. The generated prompt may have the following format, for example:

[1065] Example prompt sentence:

[1066] The user's mood is Relax, their favorite music style is Jazz, and their language is English. Create a playlist based on this.

[1067] Based on the generated prompt, a generative AI model (e.g., GPT-3) is used to automatically generate a playlist that matches the user's mood and preferences. This playlist is then displayed on the user's device, allowing the user to play the playlist.

[1068] Hardware and software used

[1069] Hardware: Smartphones, servers

[1070] Software: Python, OpenAI API

[1071] Specific examples

[1072] For example, if a user inputs "I want to relax," "Jazz," and "English," the server receives this information and uses the generative AI model to generate a prompt. Based on this prompt, the generative AI model automatically generates a playlist that matches the user's mood and preferences, and displays it on the user's smartphone.

[1073] In this way, users can easily create playlists that suit their moods and preferences, improving their experience using music streaming services.

[1074] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1075] Step 1:

[1076] The user inputs information such as mood, preferred music style, language, etc. from a device such as a smartphone.

[1077] Input: User's mood, preferred tune, language

[1078] Output: Sending input information

[1079] Specific actions: The user opens the application on their smartphone, enters information such as "I want to relax," "Jazz," and "English" into the text boxes, and presses the send button.

[1080] Step 2:

[1081] The terminal sends the input information to the server.

[1082] Input: Information entered by the user

[1083] Output: Send data to the server

[1084] Specific operation: The smartphone application makes an API request to send the information entered by the user to the server.

[1085] Step 3:

[1086] The server generates a prompt based on the input information it receives.

[1087] Input: User's mood, preferred tune, language

[1088] Output: prompt statement

[1089] Specific operation: The server analyzes the received information and generates a prompt sentence of the form "The user's mood is to relax, their favorite music style is jazz, and their language is English. Please create a playlist based on this."

[1090] Step 4:

[1091] The server sends the generated prompts to the generative AI model to generate a playlist.

[1092] Input: prompt statement

[1093] Output: Playlist

[1094] Specific operation: The server sends the generated prompt sentence to the OpenAI API and generates a playlist using a generative AI model (e.g., GPT-3).

[1095] Step 5:

[1096] The server transmits the generated playlist to the terminal.

[1097] Input: Playlist

[1098] Output: Send playlist

[1099] Specific operation: The server makes an API response to send the generated playlist to the smartphone application.

[1100] Step 6:

[1101] The terminal displays the received playlist to the user.

[1102] Input: Playlist

[1103] Output: Playlist display

[1104] Specific operation: The smartphone application displays the received playlist on the screen and allows the user to play it.

[1105] Example 2

[1106] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1107] Conventional music playback systems make it difficult for users to find music that suits their tastes and lack the ability to create new music by combining favorite songs. Furthermore, the means for providing the created music to users is insufficient, preventing user satisfaction. Furthermore, the lack of integrated karaoke and voice training functions makes it difficult for users to enjoy a diverse musical experience in a single system.

[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1109] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a unit that analyzes the beat and melody of selected songs and generates new songs using a generative AI model, a unit that provides the generated songs to the user, a unit that displays a score using a karaoke function, and a voice training function. This allows users to easily find songs that suit their preferences and create new songs by combining their favorite songs. Furthermore, the rapid provision of generated songs increases user satisfaction. Furthermore, by integrating the karaoke and voice training functions, users can enjoy a diverse musical experience on a single system.

[1110] The "means for creating a playlist based on user input" is a function that automatically generates a list of songs that meets specific conditions or preferences based on information entered by the user.

[1111] "Means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that analyzes the user's past music selection history and input information and recommends music that suits the user's preferences.

[1112] "Means for creating new music by mixing favorite songs" is a function that combines multiple songs selected by the user to create new music.

[1113] "Means of analyzing the beat and melody of a selected song and generating a new song using a generative AI model" refers to a function that analyzes the beat and melody of a song selected by the user and generates a new song using a generative AI model based on that information.

[1114] The "means for providing the generated music to the user" is a function for providing the generated new music to the user in a format that can be listened to.

[1115] The "means for displaying scores in the karaoke function" is a function that provides a karaoke function that displays scores for songs sung by the user.

[1116] The "voice training function" is a function that provides training to improve the user's singing ability.

[1117] This invention is a system that allows users to easily find songs that suit their tastes and combine their favorite songs to create new songs. Furthermore, it provides users with a diverse musical experience by quickly providing the created songs and integrating karaoke and voice training functions.

[1118] System configuration

[1119] Subject: User

[1120] First, the user selects their favorite songs through the system interface. For example, they select "Song A" and "Song B." Next, the user performs an operation to mix the selected songs. Specifically, they click the MIX button.

[1121] Subject: Terminal

[1122] The device sends data about the user-selected "song A" and "song B" to the server. This data includes information about the beat and melody of the songs. The device receives the data about the user-selected songs and uses software (e.g., music editing software) to analyze them.

[1123] Subject: Server

[1124] The server receives the data for "Song A" and "Song B" sent from the device. It analyzes the received data and extracts beat and melody features. A generative AI model (such as OpenAI's GPT-3 or Google's Magenta) is used for the analysis. The server inputs the following prompt sentence into the generative AI model:

[1125] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[1126] The generative AI model generates a new song, "Song C," based on this prompt. The generated song is saved on the server.

[1127] Subject: Terminal

[1128] The device receives the newly generated song "Song C" from the server, and the received song is converted into a format that can be played on the device's media player.

[1129] Subject: User

[1130] The user listens to the new song "Song C" through the device. The user can save the created song or share it on social media. For example, the user can save the song in MP3 format and send it to a friend.

[1131] Specific examples

[1132] If the user selects "Song A" and "Song B," the device sends the song data to the server, which uses the generative AI model to input the following prompt:

[1133] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[1134] The generative AI model generates a new song, "Song C," based on this prompt. The generated song, "Song C," is then sent from the server to the device, where it becomes available for the user to listen to.

[1135] In this way, users can easily find songs that suit their tastes and combine their favorite songs to create new songs. Furthermore, by quickly providing the created songs, user satisfaction can be increased. Furthermore, by integrating karaoke and voice training functions, users can enjoy a diverse musical experience in one system.

[1136] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1137] Step 1:

[1138] The user selects their favorite songs through the system interface. For example, they select "Song A" and "Song B." The input is the information of the songs selected by the user, and the output is a list of the selected songs. Specifically, the user enters "Song A" and "Song B" in the search bar and selects from the displayed list.

[1139] Step 2:

[1140] The user performs an operation to mix the selected songs. Specifically, the user clicks the MIX button. The input is a list of the selected songs, and the output is instructions for the MIX operation. As a specific operation, the user clicks the MIX button on the interface.

[1141] Step 3:

[1142] The device sends data on "song A" and "song B" selected by the user to the server. The input is a list of selected songs, and the output is the song data sent to the server. Specifically, the device sends data including beat and melody information of the selected songs to the server.

[1143] Step 4:

[1144] The server receives the data for "Song A" and "Song B" sent from the device. The input is the song data sent from the device, and the output is the received song data. Specifically, the server analyzes the received data and extracts the beat and melody characteristics.

[1145] Step 5:

[1146] The server uses a generative AI model to generate new music. The input is the analyzed beat and melody characteristics, and the output is the generated new music. Specifically, the server inputs the following prompt sentence into the generative AI model:

[1147] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[1148] The generative AI model generates a new song, "Song C," based on this prompt.

[1149] Step 6:

[1150] The server sends the newly generated song "Song C" to the terminal. The input is the generated song, and the output is the song data to be sent to the terminal. In concrete terms, the server sends the generated song to the terminal.

[1151] Step 7:

[1152] The device receives a new song, "Song C," generated by the server. The input is the song data sent from the server, and the output is the received song data. Specifically, the device converts the received song into a format that can be played on a media player.

[1153] Step 8:

[1154] The user listens to the new song "Song C" through the device. The input is the received song data, and the output is a song that can be listened to. Specifically, the user plays the song using the device's media player. The user can save the created song or share it on social media.

[1155] (Application example 2)

[1156] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1157] Conventional music creation systems offer limited functionality for users to mix songs to their own tastes and lack the ability to preview, save, and share created songs. Furthermore, there is no easy way for users to share their created songs with other users, limiting how users can enjoy music. Furthermore, the lack of integrated karaoke and voice training features results in an inconsistent user experience.

[1158] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1159] In this invention, the server includes a playlist creation unit based on user input, a music recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a unit that previews, saves, and shares the created songs, a unit that displays a score using a karaoke function, and a voice training function. This allows users to create songs that suit their preferences and preview, save, and share them. Furthermore, integrating the karaoke and voice training functions provides a consistent and enriched music experience.

[1160] The "means for creating a playlist based on user input" is a function that automatically generates a playlist based on the songs selected by the user and their mood that day.

[1161] "Means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music style, language, etc., and recommends appropriate music based on that.

[1162] "Means to mix your favorite songs to create new music" is a function that generates new music by combining the beats and melodies of multiple songs selected by the user.

[1163] "Means for previewing, saving, and sharing the generated music" refers to a function that allows users to listen to the new music they have generated, save it if they like it, and share it with other users via social media, messaging apps, etc.

[1164] The "means for displaying a score using the karaoke function" is a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays a score.

[1165] The "voice training function" is a function that assists users in practicing to improve their singing ability, and provides feedback on pitch and rhythm.

[1166] A system for implementing this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred style, language, etc., a means for creating a new song by mixing favorite songs, a means for previewing, saving, and sharing the created song, a means for displaying scores in a karaoke function, and a voice training function.

[1167] System Program

[1168] Hardware and software used

[1169] Hardware: Smartphone, Head-Mounted Display (HMD)

[1170] Software: Python, librosa, pydub

[1171] Data processing and calculation

[1172] The server loads user-selected songs and uses the librosa library to analyze the beats and melodies. It then combines the waveform data from multiple selected songs to generate a new song, which the user can preview and, if they like it, save and share.

[1173] Specific examples

[1174] When a user selects "Song A" and "Song B" and mixes them to create a new "Song C," the following steps are performed:

[1175] 1. The user selects "Song A" and "Song B" within the app.

[1176] 2. The server loads the selected song using the librosa library and analyzes the beat and melody.

[1177] 3. Based on the analysis results, the server combines the waveform data of the two songs to generate a new "Song C."

[1178] 4. The user previews the generated "Song C" and saves it if they like it.

[1179] 5. Share the saved "Song C" on social media or messaging apps.

[1180] Prompt Sentence Examples

[1181] Generate a new song, "Song C," by mixing the user-selected "Song A" and "Song B." The generated song must be a combination of beats and melodies. Also, provide the ability to preview, save, and share the generated song.

[1182] In this way, users can create songs to their liking, preview, save and share them, and the integration of karaoke and voice training features makes the music experience more consistent and fulfilling.

[1183] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1184] Step 1:

[1185] The user selects "Song A" and "Song B" within the app.

[1186] Input: User selected music files (song A, song B)

[1187] Output: Path of the selected music file

[1188] Specific operation: The user selects "Song A" and "Song B" from the library through the application interface. The paths of the selected music files are sent to the server.

[1189] Step 2:

[1190] The server loads the selected song using the librosa library and analyzes the beat and melody.

[1191] Input: Path of selected music file (song A, song B)

[1192] Output: Analysis data of the beat and melody of the song

[1193] Specific operation: The server uses the librosa library to load the selected music file, analyzes the beat and melody of the loaded music, and generates analysis data.

[1194] Step 3:

[1195] Based on the analysis results, the server combines the waveform data of the two songs to generate a new "Song C."

[1196] Input: Analysis data of the beat and melody of the songs (song A, song B)

[1197] Output: New song file (song C)

[1198] Specific operation: Based on the analysis data, the server combines the waveform data of the two songs by averaging them or other methods to generate a new song, "Song C."

[1199] Step 4:

[1200] The user previews the generated "Song C" and saves it if they like it.

[1201] Input: New song file (song C)

[1202] Output: Path of saved song file (song C)

[1203] Specific behavior: The user previews the new song "Song C" through the application interface. If they like it, they press the save button to save the song.

[1204] Step 5:

[1205] Share the saved "Song C" on social media or messaging apps.

[1206] Input: Path of saved music file (song C)

[1207] Output: Link or file of the shared song

[1208] Specific operation: The user uses the application's sharing function to share the saved song "Song C" with other users via social media or messaging apps.

[1209] Example 3

[1210] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1211] Conventional karaoke systems and voice training systems have limited functions for evaluating and training users' singing ability, making it difficult to effectively support users in improving their singing ability. Furthermore, they lack the ability to recommend songs and create playlists tailored to users' preferences, making it difficult to increase user satisfaction. This has resulted in a lack of motivation for users to continue using the system.

[1212] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1213] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit for mixing favorite songs to create a new song, a unit for displaying a score using the karaoke function, a unit for playing a song selected by the user and displaying the score after singing, and a unit for providing training for the user's selected singing technique and displaying feedback after practice. This makes it possible to effectively evaluate and improve the user's singing ability. Furthermore, song recommendations and playlist creation based on the user's preferences can be made, increasing user satisfaction and encouraging continued use.

[1214] The "means for creating a playlist based on user input" is a function that automatically creates a list of songs based on information entered by the user.

[1215] "Means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music genre, language, etc., and recommends music based on that.

[1216] "Means for creating new music by mixing favorite songs" is a function that combines multiple songs selected by the user to create new music.

[1217] The "means for displaying a score in a karaoke function" is a function that analyzes the results of the user's singing and displays the singing ability as a score.

[1218] The "means for playing a song selected by the user and displaying a score after singing" is a function for playing a song selected by the user and displaying the result as a score after singing.

[1219] "Means for providing training for a singing technique selected by the user and displaying feedback after practice" is a function that provides training for a specific singing technique selected by the user and displays the results as feedback after practice.

[1220] The present invention is a system that provides a karaoke function and a voice training function for evaluating and improving a user's singing ability. Specific embodiments of this system will be described below.

[1221] Karaoke function

[1222] 1. Select a song

[1223] The user selects the song he wants to sing through the terminal interface, for example, the user selects "Song 1."

[1224] 2. Sending song data

[1225] The server sends the audio data and lyrics data of the selected song to the terminal. The hardware used is the server and the terminal, and the software used is a music database and a communication protocol.

[1226] 3. Play a song

[1227] The device plays the received audio data and displays the lyrics on the screen. The hardware used is the device (smartphone, tablet, PC, etc.), and the software used is a media player and lyric display software.

[1228] 4. Singing

[1229] The user sings along to the song through the device's microphone.

[1230] 5. Audio Analysis

[1231] The device analyzes the user's singing voice in real time and evaluates pitch and rhythm using software such as voice analysis algorithms and pitch detection software.

[1232] 6. Calculation of points

[1233] The server receives the analysis results and calculates the score. The software used is the score calculation algorithm.

[1234] 7. Display of score

[1235] The terminal displays the score received from the server on the screen. For example, after the user finishes singing "Song 1," a score of 85 is displayed.

[1236] Voice training function

[1237] 1. Select a training menu

[1238] The user selects the singing technique they want to practice through the device interface, for example, by selecting "practice vibrato."

[1239] 2. Sending training data

[1240] The server transmits audio data and instructional videos corresponding to the selected training menu to the terminal. The hardware used is the server and the terminal, and the software used is the training database and communication protocol.

[1241] 3. Training playback

[1242] The device plays the received audio data and instructional video. The hardware used is the device (smartphone, tablet, PC, etc.), and the software used is a media player and video playback software.

[1243] 4. Practice

[1244] The user practices singing techniques by following instructional videos.

[1245] 5. Audio Analysis

[1246] The device analyzes the user's practice voice in real time and provides feedback on areas for improvement. The software used is a voice analysis algorithm and feedback generation software.

[1247] 6. Track your progress

[1248] The server receives the analysis results and records the user's progress. The software used is a progress management system.

[1249] 7. Viewing Feedback

[1250] The device displays the feedback it receives from the server on the screen. For example, after the user has finished practicing vibrato, the device displays feedback such as, "Your vibrato duration is too short, so try to hold it for a little longer."

[1251] Examples and prompts

[1252] Specific examples

[1253] Karaoke function:

[1254] The user selects "Song 1" and after finishing singing, a score of 85 is displayed.

[1255] Voice training features:

[1256] The user selects "Practice vibrato" and after practicing, feedback is displayed saying, "The vibrato duration is short, try to hold it for a little longer."

[1257] Prompt Sentence Examples

[1258] Karaoke function:

[1259] "Generate a program that plays a song selected by the user and displays the score after singing."

[1260] Voice training features:

[1261] "Please create a program that provides training for a user-selected singing technique and displays feedback after practice." The flow of the specific process in the third embodiment will be described with reference to FIG.

[1262] Karaoke function

[1263] Step 1: Select a song

[1264] The user selects the song they want to sing through the terminal interface.

[1265] Input: Information about the song selected by the user (e.g. "Song 1")

[1266] Output: The information of the selected songs will be saved on your device.

[1267] Specific action: The user taps "Song 1" on the device screen.

[1268] Step 2: Send the song data

[1269] The server transmits the audio data and lyrics data of the selected song to the terminal.

[1270] Input: Information about the song selected by the user

[1271] Output: Audio data and lyrics data are sent to the device.

[1272] Specific operation: The server retrieves the audio file and lyrics file for "Song 1" from the music database and sends them to the device.

[1273] Step 3: Play a song

[1274] The device plays the received audio data and displays the lyrics on the screen.

[1275] Input: Audio data and lyrics data sent from the server

[1276] Output: The audio is played and the lyrics are displayed on the screen.

[1277] Specific operation: The device's media player plays "Song 1" and the lyrics display software scrolls the lyrics.

[1278] Step 4: Singing

[1279] The user sings along to the song through the device's microphone.

[1280] Input: Audio to be played and lyrics to be displayed

[1281] Output: User's singing voice is input through a microphone

[1282] Specific action: The user sings "Song 1" into the microphone.

[1283] Step 5: Analyze the audio

[1284] The device analyzes the user's singing voice in real time and evaluates pitch and rhythm.

[1285] Input: User singing

[1286] Output: Pitch and rhythm analysis data

[1287] How it works: The device's voice analysis algorithm analyzes the user's singing voice using pitch detection software to generate pitch and rhythm data.

[1288] Step 6: Calculate your score

[1289] The server receives the analysis results and calculates the score.

[1290] Input: Pitch and rhythm analysis data sent from the device

[1291] Output: Calculated score

[1292] Specific operation: Based on the pitch and rhythm data received by the server, the score calculation algorithm calculates 85 points.

[1293] Step 7: View your scores

[1294] The terminal displays the score received from the server on the screen.

[1295] Input: Score sent from the server

[1296] Output: The score displayed on the screen

[1297] Specific action: "85 points" will be displayed on the device screen.

[1298] Voice training function

[1299] Step 1: Select a training menu

[1300] The user selects the singing technique they wish to practice through the terminal interface.

[1301] Input: User-selected training menu (e.g., "Vibrato practice")

[1302] Output: The selected training menu is saved to the device.

[1303] Specific action: The user taps "Practice Vibrato" on the device screen.

[1304] Step 2: Submitting training data

[1305] The server transmits audio data and instructional videos corresponding to the selected training menu to the terminal.

[1306] Input: User selected training menu

[1307] Output: Audio data and instructional videos are sent to the device.

[1308] Specific operation: The server retrieves the audio file and instructional video for "vibrato practice" from the training database and sends them to the terminal.

[1309] Step 3: Playback the training

[1310] The device plays the received audio data and instructional video.

[1311] Input: Audio data and instruction video sent from the server

[1312] Output: Audio is played and instructional video is displayed on the screen

[1313] Specific operation: The device's media player plays the audio for "Vibrato Practice," and the video playback software displays the instructional video.

[1314] Step 4: Practice

[1315] The user practices singing techniques by following instructional videos.

[1316] Input: Audio to be played and instructional video

[1317] Output: User's practice voice is input through microphone

[1318] Specific actions: The user practices vibrato while watching an instructional video.

[1319] Step 5: Analyze the audio

[1320] The device analyzes the user's practice audio in real time and provides feedback on areas for improvement.

[1321] Input: User's practice voice

[1322] Output: Analysis results and feedback data

[1323] Specific operation: The device's audio analysis algorithm analyzes the user's practice audio, and the feedback generation software generates feedback such as "the vibrato duration is short."

[1324] Step 6: Record your progress

[1325] The server receives the analysis results and records the user's progress.

[1326] Input: Analysis results sent from the device

[1327] Output: Recorded progress data

[1328] Specific operation: The analysis results received by the server are saved in the progress management system.

[1329] Step 7: View your feedback

[1330] The device displays the feedback received from the server on the screen.

[1331] Input: Feedback data sent from the server

[1332] Output: On-screen feedback

[1333] What it does: The device displays feedback on the screen saying, "The vibrato duration is too short, try to hold it for a little longer."

[1334] (Application example 3)

[1335] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1336] Conventional karaoke systems and voice training systems have limited functionality for evaluating a user's singing ability, making it difficult for users to objectively evaluate their own singing ability and receive specific feedback to improve it. In addition, there has been a lack of systems that allow users to effectively train to improve their singing ability.

[1337] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[1338] In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score using a karaoke function, a system including a voice training function, a means for reading voice data and extracting features, a means for evaluating singing ability based on the extracted features, and a means for displaying the evaluation results, thereby enabling users to objectively evaluate their own singing ability and receive specific feedback.

[1339] The "means for creating a playlist based on user input" is a means for automatically creating a list of songs that match the user's preferences based on information entered by the user.

[1340] "Means for recommending music that takes into consideration the user's preferred music style, language, etc." refers to a means for recommending music that is suitable for a user based on information such as the user's past music selection history, preferred music style, language, etc.

[1341] "Means for creating new music by mixing favorite songs" refers to a means for combining multiple songs selected by the user to create new music.

[1342] The "means for displaying a score in the karaoke function" is a means for evaluating the results of the user's singing and displaying the evaluation as a score.

[1343] The "voice training function" is a function that provides training to help users improve their singing ability.

[1344] The "means for reading voice data and extracting features" refers to a means for reading voice data sung by a user and extracting features from the voice data.

[1345] The "means for evaluating singing ability based on extracted features" is a means for evaluating the singing ability of a user based on the extracted feature amounts of voice data.

[1346] The "means for displaying the evaluation result" is a means for visually displaying the evaluation result of singing ability to the user.

[1347] The following system configuration will be described as an embodiment of the present invention.

[1348] System Configuration

[1349] This system includes a means for creating a playlist based on user input, a means for recommending songs taking into consideration the user's preferred style, language, etc., a means for creating new songs by mixing favorite songs, a means for displaying scores using a karaoke function, a voice training function, a means for reading voice data and extracting features, a means for evaluating singing ability based on the extracted features, and a means for displaying the evaluation results.

[1350] Program processing

[1351] The server creates a playlist based on information entered by the user using their smartphone. The server takes into consideration the user's mood and past music selection history when creating the playlist. It then recommends songs based on the user's preferred style and language. It is also possible to create new songs by mixing the songs selected by the user.

[1352] When a user uses the karaoke function, their smartphone records their singing and sends the audio data to a server. The server uses Librosa to extract features from the audio data and evaluates their singing ability using a generative AI model trained with TensorFlow. The evaluation results are displayed to the user as a score.

[1353] The voice training function provides training programs for users to practice specific singing techniques. The user's singing voice data is sent back to the server, where features are extracted and evaluated in the same way. Based on the evaluation results, specific feedback is provided to the user.

[1354] Hardware and software used

[1355] Hardware: Smartphone

[1356] Software: Python, Librosa, TensorFlow

[1357] Specific examples

[1358] For example, when a user sings "Song 1," the audio data recorded on the smartphone is sent to the server. The server uses Librosa to extract MFCCs (Mel Frequency Cepstrum Coefficients) from the audio data and evaluates the singing ability using a generative AI model trained with TensorFlow. The evaluation result is displayed to the user as "Your vocal performance score is: 85.75."

[1359] Prompt Sentence Examples

[1360] An example of a prompt to input to a generative AI model is as follows:

[1361] Load the audio file sung by the user, extract the MFCCs, input them into the singing ability evaluation model, and calculate the score.

[1362] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1363] Step 1:

[1364] A user starts a karaoke application using a smartphone and selects a song they want to sing.

[1365] Input: User song selection

[1366] Output: Selected song information

[1367] Specific operation: The user operates the app interface to select a song, and the selection information is saved within the app.

[1368] Step 2:

[1369] The user begins singing along with the selected song, and the smartphone records the voice.

[1370] Input: User's singing voice

[1371] Output: Recorded audio data

[1372] Specific operation: The user's singing voice is recorded using the smartphone's microphone and saved as audio data.

[1373] Step 3:

[1374] The recorded audio data is sent to the server.

[1375] Input: Recorded audio data

[1376] Output: Audio data sent to the server

[1377] Specific operation: The smartphone uploads the voice data to the server via the Internet.

[1378] Step 4:

[1379] The server uses Librosa to extract MFCCs (Mel-Frequency Cepstral Coefficients) from the audio data.

[1380] Input: Audio data sent to the server

[1381] Output: Extracted MFCC features

[1382] Specific operation: The server analyzes the audio data using the Librosa library and calculates MFCC features.

[1383] Step 5:

[1384] The server uses a generative AI model trained with TensorFlow to evaluate singing ability based on the extracted MFCC features.

[1385] Input: Extracted MFCC features

[1386] Output: Singing ability evaluation score

[1387] Specific operation: The server inputs MFCC features into the generative AI model, and the model outputs a singing ability evaluation score.

[1388] Step 6:

[1389] The server sends the evaluation results to the smartphone.

[1390] Input: Singing ability evaluation score

[1391] Output: Evaluation results sent to your smartphone

[1392] Specific operation: The server sends the evaluation score to the smartphone and provides feedback to the user.

[1393] Step 7:

[1394] The smartphone displays the evaluation results to the user.

[1395] Input: Evaluation results sent to your smartphone

[1396] Output: The evaluation results displayed to the user

[1397] Specific behavior: The evaluation score is displayed on the smartphone screen so that the user can check the results.

[1398] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1399] "Example 1"

[1400] In one embodiment of the present invention, a system incorporating an emotion engine is provided. The system recognizes a user's emotion and creates a playlist based on that emotion. Specifically, if a user indicates to the system that they are feeling "sad," the emotion engine receives that information and creates a playlist by selecting songs that fit the sad mood.

[1401] "Example 2"

[1402] The emotion engine also recommends songs based on the user's emotions. For example, if the user expresses the emotion "fun," the emotion engine will receive that information and recommend songs that match that happy mood. This recommendation can also take into account the user's past song selection history, preferred musical style, language, etc.

[1403] "Example 3"

[1404] Furthermore, the emotion engine also displays karaoke scores and provides vocal training according to the user's emotions. For example, if the user expresses the emotion "nervous," the emotion engine receives that information and provides vocal training to help them relax. Also, if the user expresses the emotion "confident," the emotion engine receives that information and provides the optimal musical experience according to the user's emotions, such as by setting the karaoke score display more strictly.

[1405] The processing flow of each embodiment will be described below.

[1406] "Example 1"

[1407] Step 1: The user indicates an emotion to the system, for example, "sad."

[1408] Step 2: The emotion engine receives input from the user.

[1409] Step 3: Based on the emotion received by the emotion engine, select a song that matches the sad mood.

[1410] Step 4: Create a playlist based on the selected songs and provide it to the user.

[1411] "Example 2"

[1412] Step 1: The user expresses an emotion to the system. For example, the user expresses the emotion "fun."

[1413] Step 2: The emotion engine receives input from the user.

[1414] Step 3: Based on the emotion received by the emotion engine, select a song that matches the happy mood.

[1415] Step 4: Recommend the selected song to the user. This recommendation can take into account the user's past song selection history, preferred style, language, etc.

[1416] "Example 3"

[1417] Step 1: The user indicates their feelings to the system. For example, they indicate that they are "nervous."

[1418] Step 2: The emotion engine receives input from the user.

[1419] Step 3: Based on the emotion received by the emotion engine, provide voice training for relaxation.

[1420] Step 4: If the user expresses the emotion "confident," the emotion engine receives that information and provides the optimal music experience according to the user's emotions, such as by setting the karaoke score display more strictly.

[1421] Example 1

[1422] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1423] Conventional music recommendation systems have difficulty automatically generating playlists based on the user's mood and emotions, making it difficult to provide songs that perfectly match the user's preferences. Furthermore, there is a lack of technology to accurately analyze user input and recommend appropriate songs based on the results. Furthermore, there is also a lack of means to quickly and efficiently provide the generated playlists to the user's device.

[1424] The specification process by the specification processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means. In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into account the user's preferred music genre, language, etc., an emotion engine means for recognizing the user's emotions and creating a playlist based on those emotions, a natural language processing means for analyzing the user's input, and a means for transmitting the created playlist to the user's terminal. This makes it possible to automatically create a playlist based on the user's mood and emotions, and to quickly and efficiently provide songs that match the user's preferences.

[1425] The "means for creating a playlist based on user input" is a means for automatically generating a playlist by selecting appropriate songs based on the mood and preference information input by the user through the interface.

[1426] "Means for recommending music that take into consideration the user's preferred music genre, language, etc." refers to a means for recommending appropriate music based on the user's specified preferences for music genre and language.

[1427] The "emotion engine means for recognizing a user's emotions and creating a playlist based on those emotions" is a means for analyzing a user's emotions, selecting songs that match those emotions, and creating a playlist.

[1428] "Natural language processing means for analyzing user input" refers to a means for analyzing text information entered by a user and using natural language processing technology to understand the meaning.

[1429] The "means for transmitting the generated playlist to the user's terminal" refers to a means for transmitting the playlist generated by the server to the user's terminal.

[1430] MODE FOR CARRYING OUT THE INVENTION

[1431] This invention is a system equipped with an interface for inputting a user's mood, preferred music genre, language, etc., and automatically creates a playlist based on the user's input. Specific embodiments of this system are described below.

[1432] Providing a user interface

[1433] It provides an interface for users to input their mood, preferred music genre, language, etc. This interface is often implemented as a web or mobile application. For example, it could be a web application using HTML5 and JavaScript, or an iOS application using Swift. Users enter information using text boxes and drop-down menus.

[1434] Receiving and Parsing User Input

[1435] The server receives information entered by the user through the interface. The entered information includes the user's mood (e.g., "I want to relax") and their preferred music genre and language (e.g., "Jazz" or "Japanese"). The server analyzes this information and processes it as data to generate an appropriate playlist. Natural language processing (NLP) techniques are used for the analysis. For example, Python's NLTK library or the Google Cloud Natural Language API could be used.

[1436] Automatic playlist generation

[1437] The server automatically creates a playlist based on the user's input. This process uses an emotion engine and a music database. The emotion engine recognizes the user's emotion and selects music that matches that emotion. For example, if the user enters "sad," the emotion engine analyzes that information and selects music that matches the sad mood. The music database often uses the API of a music streaming service (e.g., Spotify API, Apple Music API).

[1438] Providing playlists

[1439] The generated playlist is sent to the user's device. Specifically, the playlist information is sent as an HTTP response. The user can then play the provided playlist using a music streaming service application such as Spotify or Apple Music.

[1440] Specific examples

[1441] Example 1: When you want to relax

[1442] The user enters "I want to relax" into the interface. The server receives this information and extracts the keyword "relax" using natural language processing technology. The emotion engine selects songs that fit the relaxing mood based on the keyword "relax." For example, jazz or classical music is often selected. The generated playlist is sent to the user's device as an HTTP response. The user then opens the Spotify app and plays the playlist.

[1443] Example 2: Preferences for specific music genres or languages

[1444] The user inputs that they like "jazz" and "Japanese." The server receives this information and uses natural language processing technology to extract the keywords "jazz" and "Japanese." The emotion engine then selects jazz songs with Japanese lyrics based on these keywords. The generated playlist is sent to the user's device as an HTTP response. The user then opens the Apple Music app and plays the playlist.

[1445] Prompt Sentence Examples

[1446] Prompt 1: If you want to relax

[1447] If a user types "I want to relax," create a playlist with songs that fit that relaxing mood.

[1448] Prompt 2: Preferences for specific music genres or languages

[1449] If a user enters that they like "jazz" and "Japanese," create a playlist by selecting jazz songs with Japanese lyrics.

[1450] In this way, a system is provided that automatically generates a playlist according to the user's mood and preferences.

[1451] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1452] Step 1: Provide a user interface

[1453] It provides an interface for users to input their mood, preferred music genre, language, etc. Specifically, it is implemented as a web application or mobile application. For example, it could be a web application using HTML5 and JavaScript, or an iOS application using Swift. Users enter information using text boxes and drop-down menus. Inputs include the user's mood (e.g., "I want to relax") and preferred music genre and language (e.g., "Jazz," "Japanese"). The output is the user's input information.

[1454] Step 2: Receiving User Input

[1455] The server receives information entered by the user through the interface. Specifically, the information entered by the user is sent to the server as an HTTP request. The server analyzes the received request and extracts the user's mood and preferences. The input includes the user's input information. The output is the analyzed information on the user's mood and preferences.

[1456] Step 3: Parsing User Input

[1457] The server analyzes the received information. Specifically, it uses natural language processing (NLP) technology to analyze the user's input. For example, it could use Python's NLTK library or the Google Cloud Natural Language API. If the user inputs "I want to relax," the server analyzes the information and extracts the keyword "relax." The input includes information about the user's mood and preferences. The output is the analyzed keyword.

[1458] Step 4: Automatically generate a playlist

[1459] The server automatically creates a playlist based on user input. Specifically, it uses an emotion engine and a music database. The emotion engine recognizes the user's emotion and selects music that matches that emotion. For example, if a user enters "sad," the emotion engine analyzes that information and selects music that matches the sad mood. The music database often uses the API of a music streaming service (e.g., Spotify API, Apple Music API). The input includes the analyzed keywords. The output is the generated playlist.

[1460] Step 5: Serve the playlist

[1461] The generated playlist is sent to the user's device. Specifically, the playlist information is sent as an HTTP response. The user can play the provided playlist. To play, they use a music streaming service application such as Spotify or Apple Music. The input includes the generated playlist. The output is the playlist sent to the user's device.

[1462] (Application example 1)

[1463] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1464] Conventional music recommendation systems simply recommend songs based on past music selection history and preferences, without considering the user's mood or emotions. This makes it difficult for users to find songs that match their current mood or emotions, resulting in low satisfaction. In addition, it is inconvenient because users have to take the time to input their mood and emotions.

[1465] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means. In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into account the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score in a karaoke function, a voice training function, an emotion recognition means for recognizing the user's emotion, and a means for automatically generating a playlist based on the recognized emotion. This makes it possible to automatically recommend songs that match the user's current mood and emotion, thereby improving user satisfaction.

[1466] The "playlist creation means based on user input" is a function that selects appropriate songs and generates a playlist based on information entered by the user.

[1467] The "means for recommending music taking into consideration the user's preferred music style, language, etc." is a function that recommends the most suitable music by taking into consideration information such as the user's preferred music style and language.

[1468] "A way to mix your favorite songs to create new songs" is a function that allows users to combine multiple songs selected by the user to generate new songs.

[1469] The "means for displaying scores using the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[1470] The "voice training function" is a training function for improving the user's singing ability.

[1471] "Emotion recognition means for recognizing user emotions" is a function that analyzes and recognizes emotions from the user's facial expressions, voice, etc.

[1472] The "means for automatically generating a playlist based on recognized emotions" is a function that selects appropriate songs and automatically generates a playlist based on the recognized emotions of the user.

[1473] A system for implementing this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred style, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying scores in a karaoke function, a voice training function, an emotion recognition means for recognizing the user's emotions, and a means for automatically generating a playlist based on the recognized emotions.

[1474] Hardware and software used

[1475] Hardware: Smartphone (camera, microphone)

[1476] software:

[1477] Emotion recognition engine (e.g. Microsoft Azure Emotion API)

[1478] Music recommendation systems (e.g., Spotify API)

[1479] User interface (e.g., React Native)

[1480] System processing overview

[1481] The server receives the mood, preferred music style, and language input by the user through their smartphone and creates a playlist based on that.The user interface is built using React Native, providing an interface for users to input their mood, preferred music style, and language.

[1482] The app uses the Microsoft Azure Emotion API to recognize emotions by analyzing the user's facial expressions and voice via the smartphone camera and microphone. The recognized emotional information is then used to automatically generate playlists.

[1483] As a music recommendation method, it uses the Spotify API to recommend the best songs based on user input and recognized emotions, which allows it to automatically recommend songs that match the user's current mood and emotions, improving user satisfaction.

[1484] Specific examples

[1485] For example, if a user inputs "I want to relax," the system will select songs that match that mood and create a playlist. The emotion recognition means recognizes the emotion "I want to relax" from the user's facial expressions and voice, and automatically generates a playlist based on that.

[1486] Prompt Sentence Examples

[1487] Build an application that, when a user types "I want to relax," selects songs that fit that mood and creates a playlist. Use the Microsoft Azure Emotion API for emotion recognition and the Spotify API for playlist generation. Build the user interface using React Native.

[1488] In this way, a system can be realized that recommends optimal music based on the user's mood and emotions.

[1489] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1490] Step 1:

[1491] The user inputs their mood, preferred music style, and language through the smartphone's user interface. The input data is sent to the server. Examples of input data include "I want to relax," "pop music," and "English music."

[1492] Step 2:

[1493] The server receives the input data and activates the emotion recognition means. It uses the smartphone's camera and microphone to capture the user's facial expressions and voice, and sends them to an emotion recognition engine (e.g., Microsoft Azure Emotion API). The emotion recognition engine analyzes the captured data and recognizes the user's emotion. For example, it may recognize the emotion "I want to relax."

[1494] Step 3:

[1495] The server combines the recognized emotion information with the user's input data to prepare data for generating a playlist. Specifically, it creates a query to select appropriate songs based on the user's mood, preferred musical style, and language.

[1496] Step 4:

[1497] The server sends a query to a music recommendation system (e.g., Spotify API) to recommend songs that match the user's mood and preferences. The Spotify API searches for songs based on the query and returns a list of recommended songs to the server. For example, pop English songs that match the user's mood of "wanting to relax" are recommended.

[1498] Step 5:

[1499] The server receives the recommended song list and generates a playlist, which is then sent to the user's smartphone, where the user can play the playlist through an application on the smartphone.

[1500] Step 6:

[1501] When a user plays a playlist, they can use the karaoke and voice training features. The karaoke feature displays a score for each song the user sings. The voice training feature provides training to improve the user's singing ability.

[1502] In this way, a system is realized that recommends optimal music based on the user's mood and emotions, thereby improving user satisfaction.

[1503] Example 2

[1504] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1505] Conventional music playback systems lack the ability to recommend songs based on the user's emotions and preferences, or to create new songs by mixing multiple songs. Furthermore, since they are unable to recommend songs based on the user's emotions, it is difficult to increase user satisfaction. Furthermore, since karaoke and voice training functions are not integrated, users are unable to enjoy a diverse music experience with a single system.

[1506] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1507] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for mixing favorite songs to create new songs, a means for recommending songs based on the user's emotions, a means for displaying scores in a karaoke function, and a voice training function. This makes it possible to recommend songs according to the user's emotions and preferences, and to create new songs. Furthermore, by integrating the karaoke and voice training functions, the user can enjoy a diverse musical experience in one system.

[1508] The "means for creating a playlist based on user input" is a function that automatically generates a list of songs based on information entered by the user.

[1509] "A means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music genre, language, etc., and recommends appropriate music based on that.

[1510] "A way to mix your favorite songs to create new songs" is a function that allows users to combine multiple songs selected by the user to create new songs.

[1511] "Means for recommending music based on the user's emotions" is a function that recommends music that suits the emotions based on the emotional information entered by the user.

[1512] "Means for displaying scores using the karaoke function" refers to a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays the results as a score.

[1513] The "voice training function" is a training function aimed at improving the user's singing ability, and allows for vocal practice and pitch checking.

[1514] MODE FOR CARRYING OUT THE INVENTION

[1515] This invention is a system including a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred style, language, etc., a means for creating a new song by mixing favorite songs, a means for recommending songs based on the user's emotions, a means for displaying scores in a karaoke function, and a voice training function.

[1516] Hardware and software used

[1517] Hardware

[1518] Devices: smartphones, PCs, tablets, etc.

[1519] Server: Cloud server, on-premise server

[1520] software

[1521] Music processing library: LibROSA

[1522] Music generation AI model: OpenAI's Jukedeck

[1523] Sentiment Analysis Engine: IBM Watson Sentiment Analysis API

[1524] Recommendation Engine: Spotify's Recommendation API

[1525] Program processing explanation

[1526] Music Mix function

[1527] 1. The user selects a song

[1528] The user logs in to the system using a terminal and selects the songs they want to mix. Specifically, the user selects "Song A" and "Song B" on the system interface.

[1529] 2. The server receives the data

[1530] The server receives the user's selected data for "Song A" and "Song B." Specifically, the server retrieves the song metadata and audio files via an HTTP request.

[1531] 3. The server mixes the music

[1532] The server analyzes the beats and melodies of the selected songs and combines them to create a new song. Specifically, the server uses the LibROSA library to extract the beats of "Song A" and combines them with the melody of "Song B" using OpenAI's Jukedeck.

[1533] 4. The server provides new music

[1534] The server then sends the generated new song, "Song C," to the user's device. Specifically, the server returns the generated audio file as an HTTP response, which the user can download or stream.

[1535] Music recommendation function using an emotional engine

[1536] 1. Users input their emotions

[1537] The user inputs their emotions into the system using a terminal. Specifically, the user selects the emotion "fun" on the system interface.

[1538] 2. The server receives the emotion data

[1539] The server receives the user's emotional data. Specifically, the server obtains the emotional data through an HTTP request.

[1540] 3. The server recommends songs

[1541] The server recommends appropriate songs based on the user's emotional data, past song selection history, preferred music style, language, etc. Specifically, the server analyzes emotional data using IBM Watson's sentiment analysis API and recommends songs using Spotify's recommendation API.

[1542] 4. The server provides recommended songs

[1543] The server sends the recommended songs to the user's device. Specifically, the server returns a list of recommended songs as an HTTP response, which the user can play.

[1544] Examples of concrete examples and prompts

[1545] Specific examples

[1546] The user selects "Song A" and "Song B" and generates a new "Song C."

[1547] The user selects "Song A" and "Song B" on the device.

[1548] The server receives the selected song data, extracts the beat using LibROSA, and combines the melody using Jukedeck.

[1549] The server provides the generated "song C" to the user.

[1550] The user inputs the emotion "fun," and the server recommends "happy songs."

[1551] The user inputs the emotion "fun" into the device.

[1552] The server receives the emotion data and analyzes it using IBM Watson.

[1553] The server uses Spotify's API to recommend "happy songs" and provide them to users.

[1554] Prompt Sentence Examples

[1555] "Mix song A and song B selected by the user to create a new song."

[1556] "Recommend suitable songs when the user expresses joy."

[1557] In this way, the system performs specific processing to generate and recommend songs based on the user's selection and emotions.

[1558] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1559] Music Mix function

[1560] Step 1: User selects a song

[1561] Input: The user uses the device to select the songs they want to mix.

[1562] Specific operation: The user selects "Song A" and "Song B" on the system interface.

[1563] Output: The information of the selected song (metadata and audio file path) is sent to the server.

[1564] Step 2: The server receives the data

[1565] Input: User selected "Song A" and "Song B" information.

[1566] What happens: The server retrieves the song's metadata and audio file via an HTTP request.

[1567] Output: The captured audio files and metadata are saved on the server.

[1568] Step 3: The server mixes the music

[1569] Input: Saved audio files and metadata for "Song A" and "Song B".

[1570] How it works: The server uses the LibROSA library to extract the beat of "Song A" and combines it with the melody of "Song B" using OpenAI's Jukedeck.

[1571] Output: An audio file of the newly generated song "Song C".

[1572] Step 4: The server serves up a new song

[1573] Input: The generated audio file for "Song C".

[1574] Specific operation: The server returns the generated audio file as an HTTP response.

[1575] Output: The audio file for "Song C" is sent to the user's device, where they can download or stream it.

[1576] Music recommendation function using an emotional engine

[1577] Step 1: User enters emotion

[1578] Input: The user uses a terminal to input their emotions into the system.

[1579] Specific action: The user selects the emotion "fun" on the system interface.

[1580] Output: The input emotion data is sent to the server.

[1581] Step 2: The server receives the emotion data

[1582] Input: Emotion data entered by the user.

[1583] Specific operation: The server obtains emotion data through an HTTP request.

[1584] Output: The acquired emotion data is stored in the server.

[1585] Step 3: The server recommends songs

[1586] Input: Stored emotional data, user's past music selection history, preferred music style, language, etc.

[1587] Specific operation: The server analyzes the sentiment data using IBM Watson's sentiment analysis API and recommends songs using Spotify's recommendation API.

[1588] Output: A list of recommended songs.

[1589] Step 4: The server provides song recommendations

[1590] Input: A list of recommended songs.

[1591] Specific operation: The server returns a list of recommended songs as an HTTP response.

[1592] Output: A list of recommended songs is sent to the user's device, where the user can play them.

[1593] (Application example 2)

[1594] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1595] Conventional music recommendation systems and playlist creation systems do not fully consider the user's emotions and moods when recommending music. Furthermore, the ability to create new music by mixing user-selected songs is limited, making it difficult to meet the diverse needs of users. Furthermore, the lack of integrated karaoke and voice training functions makes it difficult to provide a comprehensive music experience.

[1596] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1597] In this invention, the server includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating new songs by mixing favorite songs, an emotion engine means for recommending songs based on the user's emotions, a means for displaying scores in the karaoke function, and a voice training function. This makes it possible to recommend songs taking into consideration the user's emotions and mood, to create new songs by mixing selected songs, and to provide a comprehensive music experience that integrates the karaoke function and voice training function.

[1598] The "means for creating a playlist based on user input" is a function that automatically generates a playlist based on the songs and conditions selected by the user.

[1599] "A means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that recommends appropriate music based on the user's past music selection history, preferred music genre, language used, etc.

[1600] "A way to mix your favorite songs to create new songs" is a function that generates new songs by combining the beats and melodies of multiple songs selected by the user.

[1601] The "emotion engine means for recommending music based on the user's emotions" is a function that analyzes the user's current emotional state and recommends music that matches those emotions.

[1602] The "means for displaying scores using the karaoke function" is a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays a score.

[1603] The "voice training function" is a training function aimed at improving the user's singing ability, and includes vocal practice and pitch correction.

[1604] A system for carrying out this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, an emotion engine means for recommending songs based on the user's emotions, a means for displaying scores in a karaoke function, and a voice training function.

[1605] System Program

[1606] The system is implemented using the following hardware and software:

[1607] Hardware

[1608] Smartphone

[1609] head-mounted display

[1610] software

[1611] Python

[1612] Music MIX library (e.g. music_mixer)

[1613] Emotion engine library (e.g. emotion_engine)

[1614] Processing Description

[1615] Playlist creation method

[1616] The server automatically generates a playlist based on the songs and conditions selected by the user. For example, if a user inputs "I want to relax," a playlist of songs suitable for relaxation will be generated.

[1617] Music recommendation method

[1618] The server recommends appropriate songs based on the user's past music selection history, preferred music genre, language used, etc. For example, if the user has listened to a lot of pop songs in the past, the server will recommend mainly pop songs.

[1619] Music Mixing Method

[1620] The server generates a new song by combining the beats and melodies of multiple songs selected by the user. For example, if a user selects "Song A" and "Song B," the server will generate "Song C" by combining the characteristics of those songs.

[1621] Emotion Engine Means

[1622] The server analyzes the user's current emotional state and recommends music that matches that emotion. For example, if the user expresses the emotion "happy," music that matches that happy mood will be recommended.

[1623] Karaoke function

[1624] The server evaluates the pitch and rhythm of the song sung by the user and displays a score. For example, after a user sings karaoke, a score of 90 is displayed.

[1625] Voice training function

[1626] The server provides training functions aimed at improving the user's singing ability, such as vocal practice and pitch correction.

[1627] Specific examples

[1628] When a user selects and mixes "songA" and "songB," a new song, "MixedSong," is generated.

[1629] If the user expresses the emotion "happy," "RecommendedSong1" and "RecommendedSong2" will be recommended taking into account their past history and preferences.

[1630] Prompt Sentence Examples

[1631] Create an application that mixes user-selected songs to generate new songs and recommends songs based on the user's emotions, taking into account the user's past song selection history and preferred style and language.

[1632] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1633] Step 1:

[1634] The user selects songs using the terminal. The user selects multiple songs through the application interface and specifies the songs to be mixed. The input is a list of songs selected by the user, and the output is a list of selected songs.

[1635] Step 2:

[1636] The server receives the selected song list and generates a new song using a music mix library. Specifically, it uses the music_mixer library to combine the beats and melodies of the selected songs. The input is the selected song list, and the output is the generated new song.

[1637] Step 3:

[1638] The user inputs their current emotion using the terminal. The user selects an emotion such as "happy" or "sad" through the application interface. The input is the emotion selected by the user, and the output is the selected emotion.

[1639] Step 4:

[1640] The server receives the selected emotion and recommends songs using the emotion engine library. Specifically, the emotion_engine library is used to recommend songs taking into account the user's emotion, past music selection history, and preferred music genres. The input is the selected emotion and the user's past music selection history and preference data, and the output is a list of recommended songs.

[1641] Step 5:

[1642] A user uses the karaoke function on a terminal. The user selects the karaoke mode through the application interface and starts singing. The input is the user's singing data, and the output is the score displayed by the karaoke function.

[1643] Step 6:

[1644] The server analyzes the user's singing data and evaluates pitch, rhythm, etc. Specifically, it uses a voice analysis algorithm to evaluate the user's singing data and calculate a score. The input is the user's singing data, and the output is the calculated score.

[1645] Step 7:

[1646] A user uses the voice training function on a terminal. The user selects the voice training mode through the application interface and starts training. The input is the user's singing data, and the output is the voice training feedback.

[1647] Step 8:

[1648] The server analyzes the user's singing data and provides feedback such as vocal practice and pitch correction. Specifically, it uses a voice analysis algorithm to evaluate the user's singing data and provides feedback on areas for improvement. The input is the user's singing data, and the output is the feedback content.

[1649] Example 3

[1650] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1651] Conventional karaoke and voice training systems provide uniform evaluations and training without considering the user's emotional state, making it difficult to provide an optimal musical experience that reflects the user's psychological state. Furthermore, they lacked the functionality to adjust karaoke score displays and voice training content based on the user's emotions, making it difficult to increase user satisfaction.

[1652] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[1653] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score in a karaoke function, and a voice training function. This enables a system that includes an emotion engine that analyzes the user's emotions and adjusts the karaoke score display and the content of voice training.

[1654] The "playlist creation means" is a function that generates a list of songs based on user input.

[1655] The "music recommendation means" is a function that suggests appropriate music by taking into consideration the user's preferred musical style and language.

[1656] "Music Mixing" is a function that allows users to combine their favorite songs to create new songs.

[1657] The "karaoke function" allows the user to sing along with a song of their choice and displays the results of their singing in the form of a score.

[1658] The "voice training function" is a function that provides training to help users improve their singing ability.

[1659] The "Emotion Engine" is a function that analyzes the user's emotions and adjusts the karaoke score display and voice training content based on the analysis results.

[1660] This invention is a system including a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred melody, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying scores in a karaoke function, a voice training function, and an emotion engine that analyzes the user's emotions and adjusts the karaoke score display and the content of voice training.

[1661] Hardware and software used

[1662] This system uses the following hardware and software:

[1663] Device: The device operated by the user (smartphone, tablet, PC, etc.)

[1664] Server: A remote server for data analysis and evaluation.

[1665] Microphone: An input device for recording the user's singing voice

[1666] Speaker: An output device for playing music and training audio.

[1667] Emotion Engine: A software module for analyzing user emotions

[1668] Generative AI model: an algorithm for analyzing and evaluating the user's singing voice

[1669] Program processing

[1670] Karaoke function

[1671] 1. The user selects the karaoke function.

[1672] 2. The device plays the song selected by the user.

[1673] 3. The user sings along to the song.

[1674] 4. The device records the user's singing voice through the microphone.

[1675] 5. The server analyzes the recorded singing voice.

[1676] 6. The server calculates the evaluation results as a score.

[1677] 7. The device will display the score on the screen.

[1678] Voice training function

[1679] 1. The user selects the voice training function.

[1680] 2. The device selects the specific singing technique the user wants to practice.

[1681] 3. The device plays a training program based on the selected technique.

[1682] 4. The user sings according to the training program.

[1683] 5. The device records the user's singing voice through the microphone.

[1684] 6. The server analyzes the recorded singing voice.

[1685] 7. The server calculates the evaluation results as a score.

[1686] 8. The device will display the score on the screen.

[1687] Emotion Engine

[1688] 1. The user inputs an emotion while using the karaoke function or voice training function.

[1689] 2. The device sends the user's emotion data to the emotion engine.

[1690] 3. The server's emotion engine analyzes the user's emotions.

[1691] 4. The server adjusts the karaoke score display and voice training content based on the analysis results.

[1692] 5. The device provides the adjusted content to the user.

[1693] Specific examples

[1694] Karaoke function example

[1695] The user selects "Song 1" and begins singing.

[1696] The device plays the song and records the user's singing.

[1697] The server analyzes the recorded singing voice and evaluates pitch and rhythm.

[1698] The server calculates the evaluation result as a score and sends it to the terminal.

[1699] The device displays "85 points" on the screen.

[1700] Examples of voice training functions

[1701] The user selects "Practice Vibrato."

[1702] The device plays a vibrato training program.

[1703] The user sings according to the program.

[1704] The device records the singing voice and sends it to the server.

[1705] The server analyzes the recorded singing voice and evaluates the degree of vibrato mastery.

[1706] The server calculates the evaluation result as a score and sends it to the terminal.

[1707] The device will display "70 points" on the screen.

[1708] Examples of emotion engines

[1709] A user inputs "I'm nervous" while using the karaoke function.

[1710] The terminal transmits the emotion data to the emotion engine.

[1711] The server's emotion engine analyzes the emotion "tense."

[1712] The server provides voice training for relaxation.

[1713] The device plays relaxing voice training.

[1714] Prompt Sentence Examples

[1715] "Describe a program that uses a karaoke function to play a song selected by the user and display a score for the singing result."

[1716] "Describe a program that uses voice training features to help users practice specific singing techniques and evaluate the results."

[1717] "Please explain a program that uses an emotion engine to display karaoke scores and provide voice training according to the user's emotions." The flow of the specific processing in the third embodiment will be explained with reference to FIG.

[1718] Karaoke function processing steps

[1719] Step 1:

[1720] The user selects the karaoke function.

[1721] Input: The user taps the "Karaoke" button on the device screen.

[1722] Output: Karaoke function is activated.

[1723] Step 2:

[1724] The device plays the song selected by the user.

[1725] Input: The user selects a song.

[1726] Output: The device retrieves the song data from the server and plays it through the speaker.

[1727] Step 3:

[1728] The user sings along to the song.

[1729] Input: User sings into a microphone.

[1730] Output: The user's singing voice is input to the device through a microphone.

[1731] Step 4:

[1732] The device records the user's singing voice through a microphone.

[1733] Input: User's singing voice.

[1734] Output: The recorded vocal data is saved on the device.

[1735] Step 5:

[1736] The server analyzes the recorded singing voice.

[1737] Input: Recording data sent from the device.

[1738] Output: Analysis results such as pitch, rhythm, and vocal strength.

[1739] Step 6:

[1740] The server calculates the evaluation results as a score.

[1741] Input: Analysis results.

[1742] Output: Overall evaluation score.

[1743] Step 7:

[1744] The device will display the score on the screen.

[1745] Input: The score sent by the server.

[1746] Output: The score is displayed on the terminal screen.

[1747] Voice Training Function Processing Steps

[1748] Step 1:

[1749] The user selects the voice training function.

[1750] Input: The user taps the "Voice Training" button on the device screen.

[1751] Output: The voice training function is activated.

[1752] Step 2:

[1753] The device selects the particular singing technique that the user wants to practice.

[1754] Input: The user selects a particular technique from an on-screen menu.

[1755] Output: A training program based on the selected technique is determined.

[1756] Step 3:

[1757] The terminal plays a training program based on the selected technique.

[1758] Input: Selected training program.

[1759] Output: Training audio and video are played through speakers and a screen.

[1760] Step 4:

[1761] The user sings according to the training program.

[1762] Input: Training program instructions.

[1763] Output: The user's singing voice is input to the device through a microphone.

[1764] Step 5:

[1765] The device records the user's singing voice through a microphone.

[1766] Input: User's singing voice.

[1767] Output: The recorded vocal data is saved on the device.

[1768] Step 6:

[1769] The server analyzes the recorded singing voice.

[1770] Input: Recording data sent from the device.

[1771] Output: Analysis results assessing the mastery of specific techniques.

[1772] Step 7:

[1773] The server calculates the evaluation results as a score.

[1774] Input: Analysis results.

[1775] Output: Calculate the mastery of a particular technique as a score.

[1776] Step 8:

[1777] The device will display the score on the screen.

[1778] Input: The score sent by the server.

[1779] Output: The score is displayed on the terminal screen.

[1780] Emotion Engine Processing Steps

[1781] Step 1:

[1782] The user inputs emotions while using the karaoke function or the voice training function.

[1783] Input: The user taps the "Emotion Input" button on the device screen and selects an emotion.

[1784] Output: Emotion data is input to the terminal.

[1785] Step 2:

[1786] The terminal transmits the user's emotion data to the emotion engine.

[1787] Input: User emotion data.

[1788] Output: The emotion data is sent to the emotion engine on the server.

[1789] Step 3:

[1790] The server's emotion engine analyzes the user's emotions.

[1791] Input: Emotion data.

[1792] Output: Analysis results showing the user's emotional state.

[1793] Step 4:

[1794] Based on the analysis results, the server adjusts the karaoke score display and voice training content.

[1795] Input: Sentiment analysis results.

[1796] Output: Adjusted karaoke score display and voice training content.

[1797] Step 5:

[1798] The terminal provides the adjusted content to the user.

[1799] Input: The adjustment sent by the server.

[1800] Output: The adjusted content is displayed on the device screen.

[1801] (Application example 3)

[1802] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1803] Conventional karaoke systems and voice training systems provide uniform evaluations and training without considering the user's emotional state, making it difficult to provide an optimal musical experience that suits the user's psychological state. Furthermore, even when used in physical stores, it was not possible to provide services that match the user's emotions, making it difficult to improve user satisfaction.

[1804] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[1805] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for creating a new song by mixing favorite songs, a means for displaying a score in a karaoke function, a means including a voice training function, an emotion engine means for detecting the user's emotion, a means for adjusting the karaoke score display and voice training based on the detected emotion, and a smartphone application for use in a physical store. This provides an optimal music experience according to the user's emotional state, and makes it possible to improve user satisfaction even when using the system in a physical store.

[1806] The "playlist creation means" is a function that generates a list of songs based on user input.

[1807] The "music recommendation means" is a function that suggests appropriate music by taking into consideration the user's preferred musical style and language.

[1808] "Music Mixing" is a function that allows users to combine their favorite songs to create new songs.

[1809] The "karaoke function" is a function that displays the user's singing results in the form of a score.

[1810] The "voice training function" is a function that provides training to improve the user's singing ability.

[1811] The "emotion engine means" is a function that detects the user's emotions and adjusts the system's operation based on that information.

[1812] "Emotion-based adjustment means" is a function that changes the karaoke score display and voice training content according to the detected emotions of the user.

[1813] A "smartphone application used in a physical store" is an application that a user uses on a smartphone in a physical store.

[1814] As an embodiment of the present invention, a method for realizing an emotion-responsive karaoke and voice training system as a smartphone application for use in a physical store will be described.

[1815] System Program

[1816] The system implements a program that includes the following main functions:

[1817] 1. Playlist creation method: Generates a list of songs based on user input.

[1818] 2. Music recommendation method: Suggests appropriate music based on the user's preferred style and language.

[1819] 3. Music Mixing: Users can combine their favorite songs to create new songs.

[1820] 4. Karaoke function: Displays the user's singing results as a score.

[1821] 5. Voice training function: Provides training to improve users' singing ability.

[1822] 6. Emotion engine means: Detects user emotions and adjusts the system's behavior based on that information.

[1823] 7. Emotion-based adjustment measures: Change the karaoke score display and voice training content according to the detected user emotions.

[1824] Hardware and Software

[1825] This system is implemented using the following hardware and software:

[1826] Hardware: Smartphone

[1827] Software: Python, emotion engine (EmotionEngine class), karaoke system (KaraokeSystem class)

[1828] Data processing and calculation

[1829] The server receives the user's input data and generates a list of songs using a playlist creation means. It then uses a song recommendation means to suggest songs that take the user's preferences into consideration. It then uses a song mixing means to create new songs based on the songs selected by the user.

[1830] When users sing using the karaoke function, the results are displayed as a score, and the voice training function provides training to improve users' singing ability.

[1831] The emotion engine detects the user's emotions and adjusts the karaoke score display and voice training content based on that information. For example, if the user is nervous, it provides training to help them relax, and if the user is confident, it gives a stricter evaluation.

[1832] Specific examples

[1833] As a concrete example, consider a scenario after a user sings in a karaoke booth. The user launches a smartphone application and uses the karaoke function to sing. The emotion engine detects the user's emotions and adjusts the score of the singing result. If the user is nervous, it provides training to help them relax, and if the user is confident, it gives a stricter evaluation.

[1834] Prompt Sentence Examples

[1835] "After a user sings in a karaoke booth, the emotion engine detects the user's emotions and adjusts the score of the singing result. If the user is nervous, it will provide training to help them relax, and if they are confident, it will give a stricter evaluation."

[1836] In this way, it is possible to provide the optimal music experience according to the user's emotional state, and to improve user satisfaction even when using the system in a physical store.

[1837] The flow of the specific processing in Application Example 3 will be described with reference to FIG.

[1838] Step 1:

[1839] The user starts the smartphone application and selects the karaoke function.

[1840] Input: User actions

[1841] Output: Karaoke function activation

[1842] Specific operation: The user selects the karaoke function from the application menu, and the karaoke screen is displayed.

[1843] Step 2:

[1844] The server creates a playlist based on the user's input.

[1845] Input: User song selection

[1846] Output: Playlist

[1847] Specific operation: The user selects their favorite songs, and the server receives that information and generates a playlist.

[1848] Step 3:

[1849] The server recommends songs taking into account the user's preferred style and language.

[1850] Input: User's past song selection history

[1851] Output: Recommended song list

[1852] Specific operation: The server analyzes the user's past music selection data and recommends songs that match their preferences.

[1853] Step 4:

[1854] The server mixes the user's favorite songs to create new music.

[1855] Input: User's favorite songs

[1856] Output: New Mixed Songs

[1857] Specific operation: The server combines the selected songs to generate a new song.

[1858] Step 5:

[1859] The user sings using the karaoke function.

[1860] Input: User's singing voice

[1861] Output: Singing result score

[1862] Specific operation: The user sings along with the selected song, and the system records the singing and displays the score.

[1863] Step 6:

[1864] The server detects the user's emotions using an emotion engine.

[1865] Input: User's voice data and facial expression data

[1866] Output: Detected emotion information

[1867] Specific operation: The server performs voice analysis and facial expression analysis to identify the user's emotions.

[1868] Step 7:

[1869] The server adjusts the karaoke score display and voice training content based on the detected emotions.

[1870] Input: Detected emotion information

[1871] Output: Adjusted score display and training content

[1872] Specific operation: The server adjusts the score display based on emotional information and provides appropriate voice training.

[1873] Step 8:

[1874] The user can check the adjusted score display and voice training.

[1875] Input: Adjusted score display and training content

[1876] Output: User feedback

[1877] Specific operation: The user checks the displayed score and training content and selects the next action.

[1878] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1879] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1880] Another example of generative AI is Gemini (internet search engine). <url: https: gemini.google.com ?hl="ja">) are mentioned.

[1881] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1882] [Third embodiment]

[1883] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1884] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1885] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1886] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1887] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1888] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1889] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1890] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1891] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1892] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1893] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1894] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.

[1895] "Example 1"

[1896] In one embodiment of the present invention, a system is provided that includes an interface that allows a user to input their mood, preferred musical style, language, etc. This system automatically creates a playlist based on the user's input. For example, if a user inputs "I want to relax," the system creates a playlist by selecting songs that match that mood. In addition, if the user indicates a preference for a particular musical style or language, the system recommends songs based on that preference.

[1897] "Example 2"

[1898] Furthermore, as an embodiment of the present invention, a function is provided that allows users to select their favorite songs and mix them to create new songs. This function combines the beats and melodies of the selected songs to generate new songs. For example, if a user selects "Song A" and "Song B," the features of those songs are combined to create a new song, "Song C."

[1899] "Example 3"

[1900] In addition, the present invention provides a karaoke function and a voice training function. The karaoke function displays the user's singing results as a score. The voice training function provides training for the user to improve their singing ability. For example, when a user selects a specific song and sings along with it, their singing ability is evaluated and a score is displayed. Furthermore, if a user wants to improve their singing ability, they can use the voice training function to practice specific singing techniques.

[1901] The processing flow of each embodiment will be described below.

[1902] "Example 1"

[1903] Step 1: The user inputs their mood, preferred melody, language, etc. through the system interface.

[1904] Step 2: The system automatically creates a playlist based on user input. For example, if a user inputs "I want to relax," the system will create a playlist by selecting songs that fit that mood.

[1905] Step 3: The system allows users to indicate preferences for specific musical styles and languages, and recommends songs based on those preferences.

[1906] "Example 2"

[1907] Step 1: The user selects his favorite song through the system interface.

[1908] Step 2: The system combines the beats and melodies of the selected songs to create a new song. For example, if a user selects "Song A" and "Song B," the system will combine the features of those songs to create a new song called "Song C."

[1909] "Example 3"

[1910] Step 1: The user utilizes the system's karaoke function to select a specific song and sing along to it.

[1911] Step 2: The system evaluates the user's singing ability and displays a score.

[1912] Step 3: If the user wants to improve their singing ability, they can use the system's voice training feature to practice specific singing techniques.

[1913] Example 1

[1914] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1915] Conventional music playback systems lack the functionality to automatically create playlists based on the user's mood and preferences, resulting in the user having to manually select songs. Furthermore, they are unable to recommend songs that take into account the user's past song selection history or the user's mood that day, which can lead to reduced user satisfaction. Furthermore, few systems offer integrated karaoke and voice training functions, forcing users to use multiple applications.

[1916] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1917] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit for creating new songs by combining favorite songs, a karaoke score display unit, a voice training function, a user interface provision unit, a data analysis unit, a song search unit using the API of a music streaming service, and a unit for providing the generated playlist. This enables automatic creation of playlists based on the user's mood and preferences, thereby improving user satisfaction. Furthermore, by integrating the karaoke and voice training functions, users can utilize multiple functions in a single system.

[1918] The "means for creating a playlist based on user input" is a function that automatically generates a playlist based on the mood and preference information entered by the user.

[1919] The "means for recommending music taking into consideration the user's preferred music style, language, etc." is a function that recommends appropriate music based on the user's specified music style and language preferences.

[1920] "A way to create new songs by combining your favorite songs" is a function that allows users to mix multiple songs they have selected to create new songs.

[1921] The "means for displaying scores using the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[1922] The "voice training function" is a training function that allows users to improve their singing ability.

[1923] The "means for providing a user interface" is a function that provides an interface for the user to input information about their moods and preferences.

[1924] "Means for receiving user input" refers to the function of receiving information entered by the user through the interface.

[1925] "Means for analyzing data" refers to the function of analyzing the information entered by the user and extracting their moods and preferences.

[1926] "Means for searching for songs using the API of a music streaming service" refers to a function that uses the API of a music streaming service to search for songs that match the user's mood and preferences.

[1927] The "means for providing the generated playlist" is a function for providing the generated playlist to the user.

[1928] MODE FOR CARRYING OUT THE INVENTION

[1929] This invention is a music playback system that automatically creates a playlist based on the user's mood and preferences, and also integrates karaoke and voice training functions. Specific embodiments of this system are described below.

[1930] Providing a user interface

[1931] An interface is provided for users to input their mood, preferred melody, language, etc. The device displays this interface through a web browser or mobile application. For example, an interface that runs on a web browser can be built using HTML5 and JavaScript. Users can input their mood and preferences using text boxes and drop-down menus.

[1932] Receiving user input

[1933] The server receives the information the user enters into the interface. If the user enters "I want to relax," that information is sent to the server via an HTTP request. The server temporarily stores the received data and proceeds to the next processing step.

[1934] Data analysis and processing

[1935] The server analyzes the received user input data. Using Python's natural language processing libraries (NLTK and spaCy), it analyzes the user's input text and extracts their mood and preferences. For example, from the input "I want to relax," it extracts the mood of "relaxation." Based on the results of this analysis, the server selects music in the next step.

[1936] Playlist Generation

[1937] The server selects songs that match the user's mood and preferences based on the analysis results. It uses the APIs of music streaming services such as Spotify API and Apple Music API. For example, you can use the Spotify API to search for songs that match "relaxation." The server generates a playlist from the search results and formats the information in JSON format.

[1938] Providing playlists

[1939] The generated playlist is provided to the user. The server sends the generated playlist information to the user's terminal. The terminal displays the received playlist information on an interface. The user can play the generated playlist. For example, each song in the playlist is displayed in list format, and a song can be played by clicking the play button.

[1940] Specific examples

[1941] Example 1: When a user enters "I want to relax"

[1942] 1. The user types "I want to relax" into the interface.

[1943] 2. The server receives this input and uses natural language processing to extract the mood "relaxed."

[1944] 3. The server uses the Spotify API to search for songs that match "relaxation."

[1945] 4. Generate a playlist from the search results and format it in JSON format.

[1946] 5. The server sends the generated playlist to the user's device, and the device displays the playlist on the interface.

[1947] Example 2: When a user enters "I want to listen to up-tempo English songs"

[1948] 1. The user types into the interface, "I want to listen to some up-tempo English songs."

[1949] 2. The server receives this input and uses natural language processing to extract the preferences of "uptempo" and "English."

[1950] 3. The server uses the Spotify API to search for "uptempo" and "English" songs.

[1951] 4. Generate a playlist from the search results and format it in JSON format.

[1952] 5. The server sends the generated playlist to the user's device, and the device displays the playlist on the interface.

[1953] Prompt Sentence Examples

[1954] Example 1: When you want to relax

[1955] If a user enters "I want to relax," generate a playlist by selecting songs that are suitable for relaxation.

[1956] Example 2: If you want to listen to up-tempo English songs

[1957] If a user types "I want to listen to up-tempo English songs," generate a playlist by selecting up-tempo English songs.

[1958] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1959] Step 1:

[1960] The user accesses the interface and inputs their mood, preferred melody, language, etc.

[1961] Input: A user enters text into the interface, such as "I want to relax" or "I want to listen to some upbeat English music."

[1962] Output: The user's input data is generated.

[1963] Specific behavior: The device displays an interface through a web browser or mobile application, and the user enters information using text boxes and drop-down menus.

[1964] Step 2:

[1965] The server receives the information entered by the user.

[1966] Input: Text data entered by the user into the interface.

[1967] Output: User input data sent to the server.

[1968] What happens: When the user completes the input and clicks the submit button, the data is sent to the server via an HTTP request, which the server then temporarily stores.

[1969] Step 3:

[1970] The server parses the received user input data.

[1971] Input: User input data stored on the server.

[1972] Output: Parsed mood and preference information.

[1973] How it works: The server uses Python's natural language processing libraries (NLTK and spaCy) to parse the user's input text and extract moods and preferences such as "relaxed" or "uptempo."

[1974] Step 4:

[1975] Based on the analysis results, the server selects music that matches the user's mood and preferences.

[1976] Input: Parsed mood and preference information.

[1977] Output: A list of selected songs.

[1978] Specific operation: The server uses the API of music streaming services such as Spotify API and Apple Music API to search for songs that match the user's mood and preferences. For example, it uses the Spotify API to search for songs that match "relaxation."

[1979] Step 5:

[1980] The server provides the generated playlist to the user.

[1981] Input: A list of selected songs.

[1982] Output: Playlist information sent to the user's device.

[1983] Specific operation: The server generates a playlist from the search results and formats the information in JSON format. The generated playlist is sent to the user's device, which displays the playlist on its interface. The user can then play the generated playlist.

[1984] (Application example 1)

[1985] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1986] Conventional music streaming services require users to manually create playlists that match their moods and preferences, which is a time-consuming process. Furthermore, song recommendations based on users' moods and preferences are often insufficient, resulting in low user satisfaction. Furthermore, it is difficult to automatically generate playlists that match users' moods and preferences, leading to a demand for an improved user experience.

[1987] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1988] In this invention, the server includes a playlist creation means based on user input, a song recommendation means that takes into account the user's preferred musical style, language, etc., a means for creating new songs by combining favorite songs, and a means for generating prompt sentences using a generative AI model and automatically generating a playlist based on the user's mood and preferences, thereby enabling the automatic generation of a playlist that matches the user's mood and preferences.

[1989] "Playlist creation means based on user input" is a function that automatically creates a playlist by selecting appropriate songs based on information entered by the user.

[1990] "A means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that recommends appropriate music based on the user's preferences and past selection history.

[1991] "A means to create new songs by combining your favorite songs" is a function that allows users to combine multiple songs selected by the user to generate new songs.

[1992] The "means for displaying scores using the karaoke function" is a karaoke function that displays scores for songs sung by the user.

[1993] The "voice training function" is a training function for improving the user's singing ability.

[1994] "Means for generating prompt sentences using a generative AI model and automatically generating a playlist based on the user's mood and preferences" refers to a function that uses a generative AI model to convert user input information into prompt sentences, and automatically generates a playlist that matches the user's mood and preferences based on those prompt sentences.

[1995] A system for carrying out this invention includes multiple means for automatically creating a playlist based on user input, specifically, a playlist creation means based on user input, a song recommendation means taking into account the user's preferred style, language, etc., a means for creating a new song by combining favorite songs, a means for displaying scores in a karaoke function, a voice training function, and a means for generating prompt sentences using a generative AI model to automatically generate a playlist based on the user's mood and preferences.

[1996] Program processing explanation

[1997] The server receives information such as mood, preferred melody, and language input from the user's smartphone or other device. Based on this input information, a generative AI model is used to generate a prompt. The generated prompt may have the following format, for example:

[1998] Example prompt sentence:

[1999] The user's mood is relaxation, their preferred style is jazz, and their language is English. Create a playlist based on this.

[2000] Based on the generated prompt, a generative AI model (e.g., GPT-3) is used to automatically generate a playlist that matches the user's mood and preferences. This playlist is then displayed on the user's device, and the user can play it.

[2001] Hardware and software used

[2002] Hardware: Smartphones, servers

[2003] Software: Python, OpenAI API

[2004] Specific examples

[2005] For example, if a user inputs "I want to relax," "Jazz," and "English," the server receives this information and uses the generative AI model to generate a prompt. Based on this prompt, the generative AI model automatically generates a playlist that matches the user's mood and preferences and displays it on the user's smartphone.

[2006] In this way, users can easily create playlists that suit their moods and preferences, improving their experience using music streaming services.

[2007] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2008] Step 1:

[2009] The user inputs information such as mood, preferred music style, language, etc. from a device such as a smartphone.

[2010] Input: User's mood, preferred tune, language

[2011] Output: Sending input information

[2012] Specific actions: The user opens the application on their smartphone, enters information such as "I want to relax," "Jazz," and "English" into the text boxes, and presses the send button.

[2013] Step 2:

[2014] The terminal sends the input information to the server.

[2015] Input: Information entered by the user

[2016] Output: Send data to the server

[2017] Specific operation: The smartphone application makes an API request to send the information entered by the user to the server.

[2018] Step 3:

[2019] The server generates a prompt based on the input information it receives.

[2020] Input: User's mood, preferred tune, language

[2021] Output: prompt statement

[2022] Specific behavior: The server analyzes the received information and generates a prompt in the form of "The user's mood is to relax, their preferred music style is jazz, and their language is English. Please create a playlist based on this."

[2023] Step 4:

[2024] The server sends the generated prompts to the generative AI model to generate a playlist.

[2025] Input: prompt statement

[2026] Output: Playlist

[2027] Specific operation: The server sends the generated prompt sentence to the OpenAI API and generates a playlist using a generative AI model (e.g., GPT-3).

[2028] Step 5:

[2029] The server transmits the generated playlist to the terminal.

[2030] Input: Playlist

[2031] Output: Send playlist

[2032] Specific operation: The server makes an API response to send the generated playlist to the smartphone application.

[2033] Step 6:

[2034] The terminal displays the received playlist to the user.

[2035] Input: Playlist

[2036] Output: Playlist display

[2037] Specific operation: The smartphone application displays the received playlist on the screen and allows the user to play it.

[2038] Example 2

[2039] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2040] Conventional music playback systems make it difficult for users to find music that suits their tastes and lack the ability to create new music by combining favorite songs. Furthermore, the means to provide the created music to users are insufficient, preventing user satisfaction. Furthermore, the lack of integrated karaoke and voice training functions makes it difficult for users to enjoy a diverse music experience in a single system.

[2041] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2042] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a unit that analyzes the beat and melody of selected songs and generates new songs using a generative AI model, a unit that provides the generated songs to the user, a unit that displays a score using a karaoke function, and a voice training function. This allows users to easily find songs that suit their preferences and create new songs by combining their favorite songs. Furthermore, the rapid provision of generated songs increases user satisfaction. Furthermore, by integrating the karaoke and voice training functions, users can enjoy a diverse musical experience on a single system.

[2043] The "means for creating a playlist based on user input" is a function that automatically generates a list of songs that meets specific conditions or preferences based on information entered by the user.

[2044] "Means for recommending music that takes into consideration the user's preferred style, language, etc." is a function that analyzes the user's past music selection history and input information to recommend music that suits the user's preferences.

[2045] "A way to mix your favorite songs to create new songs" is a function that allows users to combine multiple songs selected by the user to create new songs.

[2046] "Means of analyzing the beat and melody of a selected song and generating a new song using a generative AI model" refers to a function that analyzes the beat and melody of a song selected by the user and generates a new song using a generative AI model based on that information.

[2047] The "means for providing the generated music to the user" is a function for providing the generated new music to the user in a format that can be listened to.

[2048] The "means for displaying scores in a karaoke function" is a function that provides a karaoke function that displays scores for songs sung by a user.

[2049] The "voice training function" is a function that provides training functions to improve the user's singing ability.

[2050] This invention is a system that allows users to easily find songs that suit their tastes and combine their favorite songs to create new songs. Furthermore, it provides users with a diverse musical experience by quickly providing the created songs and integrating karaoke and voice training functions.

[2051] System configuration

[2052] Subject: User

[2053] First, the user selects their favorite songs through the system interface. For example, they select "Song A" and "Song B." Next, the user performs an operation to mix the selected songs. Specifically, they click the MIX button.

[2054] Subject: Terminal

[2055] The device sends data about the user-selected "song A" and "song B" to the server. This data includes information about the beat and melody of the songs. The device receives the data about the user-selected songs and uses software (e.g., music editing software) to analyze them.

[2056] Subject: Server

[2057] The server receives the data for "Song A" and "Song B" sent from the device. It analyzes the received data and extracts beat and melody features. A generative AI model (such as OpenAI's GPT-3 or Google's Magenta) is used for the analysis. The server inputs the following prompt sentence into the generative AI model:

[2058] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[2059] The generative AI model generates a new song, "Song C," based on this prompt. The generated song is saved on the server.

[2060] Subject: Terminal

[2061] The device receives the newly generated song "Song C" from the server, and the received song is converted into a format that can be played on the device's media player.

[2062] Subject: User

[2063] The user listens to the new song "Song C" through the device. The user can save the created song or share it on social media. For example, the user can save the song in MP3 format and send it to a friend.

[2064] Specific examples

[2065] If the user selects "Song A" and "Song B," the device sends the song data to the server, which uses the generative AI model to input the following prompt:

[2066] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[2067] The generative AI model generates a new song, "Song C," based on this prompt. The generated song, "Song C," is then sent from the server to the device, where it becomes available for the user to listen to.

[2068] In this way, users can easily find songs that suit their tastes and combine their favorite songs to create new songs. Furthermore, by quickly providing the created songs, user satisfaction can be increased. Furthermore, by integrating karaoke and voice training functions, users can enjoy a diverse musical experience in one system.

[2069] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2070] Step 1:

[2071] The user selects their favorite songs through the system interface. For example, they select "Song A" and "Song B." The input is the information of the songs selected by the user, and the output is a list of the selected songs. Specifically, the user enters "Song A" and "Song B" in the search bar and selects from the displayed list.

[2072] Step 2:

[2073] The user performs an operation to mix the selected songs. Specifically, the user clicks the MIX button. The input is a list of the selected songs, and the output is instructions for the MIX operation. As a specific operation, the user clicks the MIX button on the interface.

[2074] Step 3:

[2075] The device sends data on "song A" and "song B" selected by the user to the server. The input is a list of selected songs, and the output is the song data sent to the server. Specifically, the device sends data including beat and melody information of the selected songs to the server.

[2076] Step 4:

[2077] The server receives the data for "Song A" and "Song B" sent from the device. The input is the song data sent from the device, and the output is the received song data. Specifically, the server analyzes the received data and extracts the beat and melody characteristics.

[2078] Step 5:

[2079] The server uses a generative AI model to generate new music. The input is the analyzed beat and melody characteristics, and the output is the generated new music. Specifically, the server inputs the following prompt sentence into the generative AI model:

[2080] Combine the beats and melodies of "Song A" and "Song B" to create a new song.

[2081] The generative AI model generates a new song, "Song C," based on this prompt.

[2082] Step 6:

[2083] The server sends the newly generated song "Song C" to the terminal. The input is the generated song, and the output is the song data to be sent to the terminal. In concrete terms, the server sends the generated song to the terminal.

[2084] Step 7:

[2085] The device receives a new song, "Song C," generated by the server. The input is the song data sent from the server, and the output is the received song data. Specifically, the device converts the received song into a format that can be played on a media player.

[2086] Step 8:

[2087] The user listens to the new song "Song C" through the device. The input is the received song data, and the output is a song that can be listened to. Specifically, the user plays the song using the device's media player. The user can save the created song or share it on social media.

[2088] (Application example 2)

[2089] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2090] Conventional music generation systems offer limited functionality for users to mix songs to their own tastes, and lack the ability to preview, save, or share the songs they create. Furthermore, there is no easy way for users to share their songs with other users, limiting how they can enjoy music. Furthermore, the lack of integrated karaoke and voice training features results in an inconsistent user experience.

[2091] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2092] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a unit that previews, saves, and shares the created songs, a unit that displays a score for the karaoke function, and a voice training function. This allows users to create songs that suit their preferences and preview, save, and share them. Furthermore, integrating the karaoke and voice training functions provides a consistent and enriched music experience.

[2093] The "means for creating a playlist based on user input" is a function that automatically generates a playlist based on the songs selected by the user and their mood that day.

[2094] "A means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music style, language, etc., and recommends appropriate music based on that.

[2095] "A way to mix your favorite songs to create new songs" is a function that generates new songs by combining the beats and melodies of multiple songs selected by the user.

[2096] "Means to preview, save, and share generated songs" refers to a function that allows users to listen to new songs they have generated, save them if they like them, and share them with other users via social media, messaging apps, etc.

[2097] "Means for displaying scores using the karaoke function" refers to a function that evaluates the pitch, rhythm, etc. of the song sung by the user and displays a score.

[2098] The "voice training function" is a function that helps users practice to improve their singing ability, providing feedback on pitch and rhythm.

[2099] A system for implementing this invention includes a playlist creation means based on user input, a song recommendation means taking into consideration the user's preferred style, language, etc., a means for creating a new song by mixing favorite songs, a means for previewing, saving, and sharing the created song, a means for displaying scores in a karaoke function, and a voice training function.

[2100] System Program

[2101] Hardware and software used

[2102] Hardware: Smartphone, Head-Mounted Display (HMD)

[2103] Software: Python, librosa, pydub

[2104] Data processing and calculation

[2105] The server loads user-selected songs and uses the librosa library to analyze the beats and melodies. It then combines the waveform data from multiple selected songs to generate a new song. This song can then be previewed by the user, and if they like it, saved and shared.

[2106] Specific examples

[2107] When a user selects "Song A" and "Song B" and mixes them to create a new "Song C," the following steps are taken:

[2108] 1. The user selects "Song A" and "Song B" within the app.

[2109] 2. The server loads the selected song using the librosa library and analyzes the beat and melody.

[2110] 3. Based on the analysis results, the server combines the waveform data of the two songs to generate a new "Song C."

[2111] 4. The user previews the generated "Song C" and saves it if they like it.

[2112] 5. Share the saved "Song C" on social media or messaging apps.

[2113] Prompt Sentence Examples

[2114] Mix the user-selected "Song A" and "Song B" to generate a new song, "Song C." The generated song must be a combination of beats and melodies. Also, provide the ability to preview, save, and share the generated song.

[2115] In this way, users can create songs to their liking, preview, save and share them, and the integration of karaoke and voice training functions makes the music experience more consistent and fulfilling.

[2116] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2117] Step 1:

[2118] The user selects "Song A" and "Song B" within the app.

[2119] Input: User selected music files (song A, song B)

[2120] Output: Path of the selected music file

[2121] Specific operation: The user selects "Song A" and "Song B" from the library through the application interface. The paths of the selected music files are sent to the server.

[2122] Step 2:

[2123] The server loads the selected song using the librosa library and analyzes the beat and melody.

[2124] Input: Path of selected music file (song A, song B)

[2125] Output: Analysis data of the beat and melody of the song

[2126] Specific operation: The server uses the librosa library to load the selected music file, analyzes the beat and melody of the loaded music, and generates analysis data.

[2127] Step 3:

[2128] Based on the analysis results, the server combines the waveform data of the two songs to generate a new "Song C."

[2129] Input: Analysis data of the beat and melody of the songs (song A, song B)

[2130] Output: New song file (song C)

[2131] Specific operation: Based on the analysis data, the server combines the waveform data of the two songs by averaging them or other methods to generate a new song, "Song C."

[2132] Step 4:

[2133] The user previews the generated "Song C" and saves it if they like it.

[2134] Input: New song file (song C)

[2135] Output: Path of saved song file (song C)

[2136] Specific behavior: The user previews the new song "Song C" through the application interface. If they like it, they press the save button to save the song.

[2137] Step 5:

[2138] Share the saved "Song C" on social media or messaging apps.

[2139] Input: Path of saved music file (song C)

[2140] Output: Link or file of the shared song

[2141] Specific behavior: The user uses the application's sharing function to share the saved song "Song C" with other users via social media or messaging apps.

[2142] Example 3

[2143] Next, a third embodiment of the third embodiment will be described. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2144] Conventional karaoke systems and voice training systems have limited functions for evaluating and training users' singing ability, making it difficult to effectively support users in improving their singing ability. Furthermore, they lack the ability to recommend songs and create playlists tailored to users' preferences, making it difficult to increase user satisfaction. This has resulted in a lack of motivation for users to continue using the system.

[2145] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.

[2146] In this invention, the server includes a playlist creation unit based on user input, a song recommendation unit that takes into account the user's preferred musical style, language, etc., a unit that mixes favorite songs to create new songs, a karaoke function score display unit, a unit that plays the user's selected song and displays the score after singing, and a unit that provides training for the user's selected singing technique and displays feedback after practice. This makes it possible to effectively evaluate and improve the user's singing ability. Furthermore, song recommendations and playlist creation based on the user's preferences can be made, increasing user satisfaction and encouraging continued use.

[2147] The "means for creating a playlist based on user input" is a function that automatically generates a list of songs based on information entered by the user.

[2148] "A means for recommending music that takes into account the user's preferred music style, language, etc." is a function that analyzes the user's past music selection history, preferred music genre, language, etc., and recommends music based on that.

[2149] "A way to mix your favorite songs to create new songs" is a function that allows users to combine multiple songs selected by the user to generate new songs.

[2150] The "means for displaying scores using the karaoke function" is a function that analyzes the results of a user's singing and displays their singing ability as a score.

[2151] The "means of playing a song selected by the user and displaying a score after singing" is a function of playing a song selected by the user and displaying the result as a score after singing.

[2152] "Means for providing training for a singing technique selected by the user and displaying feedback after practice" is a function that provides training for a specific singing technique selected by the user and displays the results as feedback after practice.

[2153] The present invention is a system that provides a karaoke function and a voice training function for evaluating and improving a user's singing ability. Specific embodiments of this system will be described below.

[2154] Karaoke function

[2155] 1. Select a song

[2156] The user selects the song he wants to sing through the terminal interface, for example, the user selects "Song 1."

[2157] 2. Sending song data

[2158] The server sends the audio data and lyrics data of the selected song to the terminal. The hardware used is the server and the terminal, and the software used is a music database and a communication protocol.

[2159] 3. Play a song

[2160] The device plays the received audio data and displays the lyrics on the screen. The hardware used is the device (smartphone, tablet, PC, etc.), and the software used is a media player and lyric display software.

[2161] 4. Singing

[2162] The user sings along to the song through the device's microphone.

[2163] 5. Audio Analysis

[2164] The device analyzes the user's singing voice in real time and evaluates pitch and rhythm using software such as voice analysis algorithms and pitch detection software.

[2165] 6. Calculation of points

[2166] The server receives the analysis results and calculates the score. The software used is the score calculation algorithm.

[2167] 7. Display of score

[2168] The terminal displays the score received from the server on the screen. For example, after the user finishes singing "Song 1," a score of 85 is displayed.

[2169] Voice training function

[2170] 1. Select a training menu

[2171] The user selects the singing technique they want to practice through the device interface, for example, by selecting "practice vibrato."

[2172] 2. Sending training data

[2173] The server transmits audio data and instructional videos corresponding to the selected training menu to the terminal. The hardware used is the server and the terminal, and the software used is the training database and communication protocol.

[2174] 3. Training playback

[2175] The device plays the received audio data and instructional video. The hardware used is the device (smartphone, tablet, PC, etc.), and the software used is a media player and video playback software.

[2176] 4. Practice

[2177] The user practices singing techniques by following instructional videos.

[2178] 5. Audio Analysis

[2179] The device analyzes the user's practice voice in real time and provides feedback on areas for improvement. The software used is a voice analysis algorithm and feedback generation software.

[2180] 6. Track your progress

[2181] The server receives the analysis results and records the user's progress. The software used is a progress management system.

[2182] 7. Viewing Feedback

[2183] The device displays the feedback it receives from the server on the screen. For example, after the user has finished practicing vibrato, the device displays feedback such as, "Your vibrato duration is too short, so try to hold it for a little longer."

[2184] Examples and prompts

[2185] Specific examples

[2186] Karaoke function:

[2187] The user selects "Song 1" and after finishing singing, a score of 85 is displayed.

[2188] Voice training features:

[2189] The user selects "Practice vibrato" and after practicing, feedback is displayed saying, "The vibrato duration is short, try to hold it for a little longer."

[2190] Prompt Sentence Examples

[2191] Karaoke function:

[2192] "Generate a program that plays a song selected by the user and displays the score after singing."

[2193] Voice training features:

[2194] "Please create a program that provides training for a user-selected singing technique and displays feedback after practice." The flow of the specific process in the third embodiment will be described with reference to FIG.

[2195] Karaoke function

[2196] Step 1: Select a song

[2197] The user selects the song they want to sing through the terminal interface.

[2198] Input: Information about the song selected by the user (e.g. "Song 1")

[2199] Output: The information of the selected songs will be saved on your device.

[2200] Specific action: The user taps "Song 1" on the device screen.

[2201] Step 2: Send the song data

[2202] The server transmits the audio data and lyrics data of the selected song to the terminal.

[2203] Input: Information about the song selected by the user

[2204] Output: Audio data and lyrics data are sent to the device.

[2205] Specific operation: The server retrieves the audio file and lyrics file for "Song 1" from the music database and sends them to the device.

[2206] Step 3: Play a song

[2207] The device plays the received audio data and displays the lyrics on the screen.

[2208] Input: Audio data and lyrics data sent from the server

[2209] Output: The audio is played and the lyrics are displayed on the screen.

[2210] Specific operation: The device's media player plays "Song 1" and the lyrics display software scrolls the lyrics.

[2211] Step 4: Singing

[2212] The user sings along to the song through the device's microphone.

[2213] Input: Audio to be played and lyrics to be displayed

[2214] Output: User's singing voice is input through a microphone

[2215] Specific action: The user sings "Song 1" into the microphone.

[2216] Step 5: Analyze the audio

[2217] The device analyzes the user's singing voice in real time and evaluates pitch and rhythm.

[2218] Input: User singing

[2219] Output: Pitch and rhythm analysis data

[2220] How it works: The device's voice analysis algorithm analyzes the user's singing voice using pitch detection software to generate pitch and rhythm data.

[2221] Step 6: Calculate your score

[2222] The server receives the analysis results and calculates the score.

[2223] Input: Pitch and rhythm analysis data sent from the device

[2224] Output: Calculated score

[2225] Specific operation: Based on the pitch and rhythm data received by the server, the score calculation algorithm calculates 85 points.

[2226] Step 7: View your scores

[2227] The terminal displays the score received from the server on the screen.

[2228] Input: Score sent from the server

[2229] Output: The score displayed on the screen

[2230] Specific action: "85 points" will be displayed on the device screen.

[2231] Voice training function

[2232] Step 1: Select a training menu

[2233] The user selects the singing technique they wish to practice through the terminal interface.

[2234] Input: User-selected training menu (e.g., "Vibrato practice")

[2235] Output: The selected training menu is saved to the device.

[2236] Specific action: The user taps "Practice Vibrato" on the device screen.

[2237] Step 2: Submitting training data

[2238] The server transmits audio data and instructional videos corresponding to the selected training menu to the terminal.

[2239] Input: User selected training menu

[2240] Output: Audio data and instructional videos are sent to the device.

[2241] Specific operation: The server retrieves the audio file and instructional video for "vibrato practice" from the training database and sends them to the terminal.

[2242] Step 3: Playback the training

[2243] The device plays the received audio data and instructional video.

[2244] Input: Audio data and instruction video sent from the server

[2245] Output: Audio is played and instructional video is displayed on the screen

[2246] Specific operation: The device's media player plays the audio for "Vibrato Practice," and the video playback software displays the instructional video.

[2247] Step 4: Practice

[2248] The user practices singing techniques by following instructional videos.

[2249] Input: Audio to be played and instructional video

[2250] Output: User's practice voice is input through microphone

[2251] Specific actions: The user practices vibrato while watching an instructional video.

[2252] Step 5: Analyze the audio

[2253] The device analyzes the user's practice audio in real time and provides feedback on areas for improvement.

[2254] Input: User's practice voice

[2255] Output: Analysis results and feedback data

[2256] Specific operation: The device's audio analysis algorithm analyzes the user's practice audio, and the feedback generation software generates feedback such as "the vibrato duration is short."

[2257] Step 6: Record your progress

[2258] The server receives the analysis results and records the user's progress.

[2259] Input: Analysis results sent from the device

[2260] Output: Recorded progress data

[2261] Specific operation: The analysis results received by the server are saved in the progress management system.

[2262] Step 7: View your feedback

[2263] The device displays the feedback received from the server on the screen.

[2264] Input: Feedback data sent from the server

[2265] Output: On-screen feedback

[2266] What it does: The device displays feedback on the screen saying, "The vibrato duration is too short, try to hold it for a little longer."

[2267] (Application example 3)

[2268] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2269] Conventional karaoke systems and voice training systems have limited functionality for evaluating a user's singing ability, making it difficult for users to objectively evaluate their own singing ability and receive specific feedback to improve it. In addition, there has been a lack of systems that allow users to effectively train to improve their singing ability.

[2270] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.

[2271] In this invention, the server includes a playlist creation means based on user input, a song recommendation means takin...

Claims

[Claim 1] means for generating a playlist based on a user's mood and preferences; a means for recommending music that takes into consideration the user's preferred musical style and language; A way to combine your favorite songs to create new ones, A means for displaying karaoke scores; a means of providing voice training; means for detecting the user's emotion using an emotion engine; a means for adjusting a karaoke score display based on the detected emotion of the user; a means for adjusting the content of voice training based on the detected emotion of the user; A system including:

Citation Information

Patent Citations

  • Karaoke medley music composition device

    JP1995295581A

  • Hierarchical playlist generator

    JP2007524955A

  • Karaoke system, server, karaoke terminal and music piece proposal method

    JP2009205114A

  • Karaoke device

    JP2017083792A

  • On-vehicle karaoke device

    JP2020003662A