System
A system that analyzes users' singing to determine the optimal key for karaoke, simplifying the process and improving the singing experience by automatically adjusting the pitch to match their vocal abilities.
Patent Information
- Application Number
- JP2024123839
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Finding the right key for singing in karaoke can be difficult, especially for beginners with little musical knowledge, making it challenging for them to enjoy singing effectively.
A system that allows users to select a song, sing a cappella, record their voice, analyze the audio using a music analysis API to identify the optimal key, and display the results for adjusting the pitch to a comfortable singing position.
Enables users to easily find and sing in the key that suits them best, enhancing their karaoke experience without requiring manual adjustments or specialized knowledge.
Smart Images

Figure 2026022322000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In karaoke, singing in a key that suits you doubles the enjoyment of singing, but finding the right key can be difficult for beginners. It is especially difficult for users with little musical knowledge to select the key that is best for them. A system that solves this problem and allows users to enjoy singing is needed. [Means for solving the problem]
[0005] In order to solve the above-mentioned problems, the present invention provides the following means. First, a selection means is provided that allows the user to select a song and sing a little a cappella along with it. Next, a recording means is provided that records the audio source of the user's singing. An analysis means is provided that analyzes the recorded audio source data. The analysis means analyzes the audio source data using a music analysis API and identifies the optimal key for the user's singing based on the results of the analysis. Finally, a display means is provided that displays the optimal key for the user's singing based on the results of the analysis by the analysis means. This allows the user to enjoy karaoke in a key that suits them.
[0006] The "selection means" refers to an interface and operation function that allows the user to select the song they want to sing in karaoke.
[0007] The "recording means" refers to a microphone and recording device for recording the user's voice when singing a cappella.
[0008] The "analysis means" refers to a processing function and device for analyzing the sound source data acquired by the recording means and identifying the optimum key for the user.
[0009] A "music analysis API" is an interface for external services or software that uses recorded audio data to analyze pitch and interval and derive the optimal key.
[0010] The "display means" refers to a display and a display device for visually presenting to the user the optimum key identified by the analysis means.
[0011] "Key" refers to the musical tone and scale of a song, and is a setting that allows users to adjust the pitch and tone to a comfortable singing position.
[0012] "Karaoke system" is a general term for equipment and software that includes a set of functions that allow users to select songs, sing, record and analyze their voice, and sing in the optimal key. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] To implement this invention, the following functions must be integrated into the karaoke system:
[0035] 1. Song selection method
[0036] The user selects the song he or she wants to sing by operating the interface on the terminal screen. This song selection means includes a function for displaying a song list and accepting user input.
[0037] 2. Recording Method
[0038] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[0039] 3. Analysis method
[0040] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[0041] 4. Display means
[0042] Based on the analysis results, the server derives the optimal key and sends the result to the terminal. The terminal's display means visually presents the analysis results to the user. For example, a notification such as "The optimal key is G" is displayed on the screen.
[0043] A natural language description of the program's operation
[0044] System Overview
[0045] The system provides functionality to allow users to sing in the optimal key for karaoke.
[0046] 1. Terminal Processing
[0047] The user selects the song they want to sing on the device.
[0048] The device displays a message inviting the user to sing a little a cappella.
[0049] When the user starts singing, the device records the voice through the microphone.
[0050] Once the recording is complete, the audio data is sent to the server.
[0051] 2. Server Processing
[0052] The server receives the voice data transmitted from the terminal.
[0053] To analyze the audio data, a request is sent to the music analysis API.
[0054] Receive analysis results containing relevant key information.
[0055] The analysis results are sent to the device.
[0056] 3. Terminal Processing (cont.)
[0057] The terminal displays the analysis results received from the server on the screen.
[0058] The user adjusts the key based on the displayed key information and prepares to sing in the optimum key.
[0059] Specific examples
[0060] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[0061] 1. Song selection and singing
[0062] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[0063] The device displays the message, "Please sing a little a cappella."
[0064] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[0065] 2. Data submission and analysis
[0066] The recording data is sent to the server.
[0067] The server receives the audio data and sends it to the music analysis API for analysis.
[0068] The music analysis API returns the result, "The optimal key is D."
[0069] 3. Displaying the results and singing
[0070] The terminal receives the analysis results from the server.
[0071] The device will display "The best key is D" on the screen.
[0072] Based on the displayed key information, the user sings "Let it be" in the key of D.
[0073] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[0074] The processing flow will be explained below.
[0075] Step 1: Song Selection
[0076] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[0077] The terminal obtains information about the selected song and displays it on the screen.
[0078] Step 2: A cappella instructions
[0079] The device displays a message to the user saying, "Please sing a little a cappella."
[0080] The user checks the message.
[0081] Step 3: Start recording
[0082] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[0083] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[0084] Step 4: End recording
[0085] When the user finishes singing, the device will automatically stop recording.
[0086] The recorded audio data is temporarily stored in memory.
[0087] Step 5: Send data
[0088] The terminal transmits the recorded audio data to the server.
[0089] The server receives the HTTP request and stores the audio file.
[0090] Step 6: Data analysis
[0091] The server generates a request to send the stored audio data to the music analysis API.
[0092] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[0093] Step 7: Obtaining analysis results
[0094] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[0095] The server stores the analysis results in a database or temporary memory.
[0096] Step 8: Send results
[0097] The server organizes the analysis results and transmits them to the user's terminal.
[0098] The optimal key information (e.g., "D key") is returned to the terminal as an HTTP response.
[0099] Step 9: View the results
[0100] The device will display the analysis results on the screen, with a message such as "The best key is D."
[0101] The user confirms this message.
[0102] Step 10: Adjust the key
[0103] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[0104] The device will retain the new settings after adjusting the keys.
[0105] Step 11: Karaoke playback
[0106] After the user completes the key setting, karaoke playback begins.
[0107] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[0108] Step 12: Enjoy singing
[0109] The user can sing songs in the optimum key and enjoy karaoke.
[0110] Example 1
[0111] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0112] Current karaoke systems require users to manually try out different keys to find the one that suits them best. This process is time-consuming and difficult, especially for users who are not familiar with music. A solution to this problem is needed to enable users to easily sing songs in the optimal key.
[0113] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0114] In this invention, the server includes a selection means for the user to select a song, a display means for instructing the user to sing a little a cappella, a recording means for recording the user's singing, a transmission means for transmitting the recorded sound data to the server, an analysis means for the server to analyze the sound data using a music analysis API, and a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means. This frees the user from the hassle of manually adjusting the key and allows them to easily enjoy karaoke in the optimal key for themselves.
[0115] The "selection means" is a function that allows the user to select a song to sing using the interface of the karaoke system.
[0116] The "means for displaying instructions" is a function that displays a message encouraging the user to sing a little a cappella.
[0117] The "recording means" is a function that records the user's singing voice through a microphone and serves to save the recorded audio data.
[0118] The "transmission means" is a function for converting the recorded sound source data into packets and transmitting them to the server.
[0119] The "analysis means" is a function that allows the server to analyze the received audio data using a music analysis API and identify the optimal key for the user.
[0120] The "display means" is a function for visually displaying to the user the optimum key information obtained by the analysis means.
[0121] To implement this invention, the following functions must be integrated into the karaoke system:
[0122] Selection method
[0123] On the device, the user operates the interface on the screen to select the song they want to sing. This selection means includes the function of displaying a song list and accepting user input. When the user selects a song, the device stores the selection information in its internal memory. For example, the user uses this selection means to select "Let it be (The Beatles)."
[0124] A means of displaying instructions
[0125] When the user selects a song, the device provides an interface that displays a message to the user saying, "Please sing a little a cappella." Specifically, a text message is displayed on the device screen.
[0126] Recording medium
[0127] When the user starts singing, the device will record the sound using the built-in microphone. This recording method will collect a certain amount of sound data (for example, 10 seconds) and save it as a digital file. The sound data will be saved in WAV or MP3 format.
[0128] Transmission method
[0129] Once the recording is complete, the device transmits the recorded audio data to the server, which encodes the audio data into an appropriate format and transmits it to the server via a network using a secure protocol (e.g., HTTPS).
[0130] Analysis means
[0131] The server analyzes the audio data received from the device. For the analysis, a music analysis API (such as Google Cloud Speech-to-Text API or IBM Watson) is used. The server sends an analysis request to the API and receives the analysis results of the audio data. For example, key information such as "The optimal key is D" is returned.
[0132] Display means
[0133] The server sends the analysis results to the device, which then visually presents the results to the user. This display includes a function that displays text such as "The best key is D" on the screen. Based on this information, the user can prepare to sing the song in the best key.
[0134] Specific examples
[0135] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[0136] 1. Song selection and singing
[0137] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[0138] The device displays the message, "Please sing a little a cappella."
[0139] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[0140] 2. Data submission and analysis
[0141] The recording data is sent to the server.
[0142] The server receives the audio data and sends it to the music analysis API for analysis.
[0143] The music analysis API returns the result, "The optimal key is D."
[0144] 3. Displaying the results and singing
[0145] The terminal receives the analysis results from the server.
[0146] The device screen will display "The best key is D."
[0147] Based on the displayed key information, the user sings "Let it be" in the key of D.
[0148] Prompt Sentence Examples
[0149] A typical example of a prompt to be input to a generative AI model is as follows:
[0150] plain
[0151] Please provide a detailed explanation of the steps a user takes to select "Let it be (The Beatles)" and optimize it for D. Please provide a detailed explanation of each step of the process, from the user selecting the song, recording the a cappella vocal, sending the audio data to the server, and receiving and displaying the analysis results.
[0152] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[0153] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0154] Step 1:
[0155] The user selects a song.
[0156] Specific operation: The user selects the song they want to sing from the song list provided on the karaoke terminal screen.
[0157] Input: The user touches or clicks on a specific song from the song list displayed on the screen.
[0158] Output: The information of the selected song is saved on the device. For example, if "Let it be (The Beatles)" is selected, the information is saved in the internal memory.
[0159] Step 2:
[0160] The device prompts the user to sing a little a cappella.
[0161] Specific operation: The device displays a message on the screen and instructs the user to "sing a little a cappella."
[0162] Input: After the user selects a song, instructions are displayed on the device screen.
[0163] Output: A visual instruction is given to the user, and the user is prepared to follow the instruction.
[0164] Step 3:
[0165] The device records the user singing a cappella.
[0166] Specific operation: When the user starts singing, the device will record the voice using the built-in microphone.
[0167] Input: The user's voice is input through the microphone.
[0168] Output: The recorded audio data is saved in a digital format (e.g., WAV or MP3). For example, the phrase "When I find myself in times of trouble" is recorded.
[0169] Step 4:
[0170] The device sends the recorded data to the server.
[0171] Specific operation: After the recording is completed, the device encodes the audio data into an appropriate format and sends it to the server over the network.
[0172] Input: Recorded audio data.
[0173] Output: The encoded audio data is split into packets and sent to the server using a secure protocol.
[0174] Step 5:
[0175] The server analyzes the audio data.
[0176] Specific operation: The server sends the received audio data to the music analysis API and analyzes the optimal key.
[0177] Input: The audio data received by the server.
[0178] Output: Analysis results from the music analysis API, e.g. "The optimal key is D."
[0179] Step 6:
[0180] The analysis results are sent from the server to the device.
[0181] Specific operation: The server converts the analysis results into a structured data format (e.g., JSON) and sends them back to the terminal using a secure protocol.
[0182] Input: Analysis result data.
[0183] Output: The analysis results are sent to the terminal in structured data format.
[0184] Step 7:
[0185] The device displays the analysis results to the user.
[0186] Specific operation: The device decodes the received analysis results and displays "The best key is D" on the screen.
[0187] Input: Analysis results received from the server.
[0188] Output: A visual display to the user, where the user obtains the best key information. For example, "The best key is D" is displayed on the terminal screen.
[0189] (Application example 1)
[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0191] Conventional karaoke systems require users to manually adjust the key through trial and error to find the optimal key for them, which is time-consuming and laborious. Furthermore, they lack a function to identify the optimal key based on the user's singing style, making it difficult for users without specialized musical knowledge. Furthermore, they lack a system that utilizes cloud technology to quickly provide analysis results.
[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0193] In this invention, the server includes a song selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded audio data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, and a communication means for the analysis means to send the analysis results to the cloud server based on a music analysis API and receive the analysis results. This allows a user to easily select a song through a smartphone application, send the singing voice to the cloud server for analysis, and quickly find the optimal key for themselves.
[0194] The "music selection means" is a function that allows the user to operate the interface to select music.
[0195] The "recording means" is a function that records the audio source sung by the user and saves the audio data as a digital file.
[0196] The "analysis means" is a function for analyzing recorded sound source data, analyzing pitch and interval, and identifying the optimum key.
[0197] The "display means" is a function that visually presents the analysis results obtained by the analysis means to the user.
[0198] The "communication means" is a function that allows the analysis means to send the analysis results to the cloud server and receive the analysis results.
[0199] A "music analysis API" is an application programming interface for analyzing sound source data and analyzing data such as pitch and interval.
[0200] A "cloud server" refers to a computer server that stores and processes data remotely over a network.
[0201] A "smartphone" is a portable information terminal that can execute multifunctional applications in addition to making calls.
[0202] In this embodiment, a system will be described in which a user can use a smartphone application to enjoy karaoke in the optimal key.
[0203] Hardware and Software Used
[0204] Smartphone: Used as a device on which users can select songs and sing a cappella.
[0205] Cloud Server: The server used to analyze the a cappella recordings.
[0206] Music Analysis API: An API that runs on a cloud server and analyzes audio data to identify the optimal key.
[0207] User interface: The interface on the smartphone application that allows users to select songs and display the results.
[0208] System Operation Details
[0209] When a user launches the smartphone application, they are first presented with an interface for selecting a song. At this stage, the user selects the song they want to sing from a list on the application. The application then prompts the user to sing a little a cappella, and records the sound using the smartphone's microphone.
[0210] Once recording is complete, the application sends the recorded audio data to a cloud server. The cloud server receives the recorded data and sends a request to the music analysis API for analysis. The music analysis API analyzes the audio data and generates key information that allows the user to sing in the optimal key.
[0211] After receiving the analysis results, the cloud server sends them to the smartphone. Based on the analysis results, the smartphone application displays a notification to the user, such as "The optimal key is G." This allows the user to enjoy singing in the key that is best for them.
[0212] Specific examples
[0213] For example, suppose a user selects "Let it be" and sings a cappella "When I find myself in times of trouble...". This audio is recorded on a smartphone and sent to a cloud server. The server uses a music analysis API to analyze the audio data and obtains the result that "the optimal key is D". The result is sent to the smartphone, which displays the message "The optimal key is D". Based on the displayed key information, the user can sing "Let it be" in the key of D.
[0214] Prompt Sentence Examples
[0215] For example, a possible prompt might be, "Please implement an API that analyzes recorded audio data and finds the optimal key."
[0216] In this way, the user can easily enjoy karaoke in the key that is most suitable for them, without requiring any specialized knowledge.
[0217] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0218] Step 1:
[0219] Song selection method
[0220] The user launches the smartphone application and selects a song.
[0221] Input: Song name specified by the user in the app
[0222] Output: Selected song data
[0223] Processing operation: A song list is displayed through the smartphone's user interface, and the user selects the song they want to sing. At this time, the application saves the data of the selected song in the internal memory.
[0224] Step 2:
[0225] Recording medium
[0226] After the user selects a song, they are prompted to sing a cappella, and when they start singing, their voice is recorded by the smartphone microphone.
[0227] Input: User's a cappella singing voice
[0228] Output: Recorded a cappella audio data (digital file)
[0229] Processing operation: When recording starts, the user's singing voice is collected through the microphone and a certain period of audio data is saved as a digital file.
[0230] Step 3:
[0231] communication means
[0232] Once recording is complete, the audio data is sent to a cloud server.
[0233] Input: Recorded a cappella audio data
[0234] Output: Audio data sent to the cloud server
[0235] Processing operation: The smartphone sends the recorded audio data to the cloud server via an HTTP request. Once the transmission is complete, the server confirms receipt of the data.
[0236] Step 4:
[0237] Analysis means
[0238] The cloud server sends the received audio data to the music analysis API and requests analysis.
[0239] Input: Received a cappella audio data
[0240] Output: Analysis results from the music analysis API (optimal key information)
[0241] Processing: The server sends the audio data to the music analysis API and makes an analysis request. The API analyzes the audio data and returns the optimal key information to the cloud server.
[0242] Step 5:
[0243] communication means
[0244] The cloud server sends the analysis results received from the music analysis API to the smartphone.
[0245] Input: Analysis results (optimal key information)
[0246] Output: Analysis results sent to your smartphone
[0247] Processing operation: After obtaining the analysis results, the cloud server returns them to the smartphone application as an HTTP response.
[0248] Step 6:
[0249] Display means
[0250] The smartphone then displays the analysis results to the user, such as "The optimal key is G."
[0251] Input: Analysis results received from the cloud server
[0252] Output: Key information that is best visible to the user
[0253] Processing operation: The smartphone displays the appropriate key information on the interface based on the analysis results received.
[0254] In this way, through a series of processing steps, the user can enjoy karaoke in the key that is most suitable for him or her.
[0255] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0256] To implement this invention, it is necessary to incorporate an emotion engine into the karaoke system and integrate the following series of functions.
[0257] 1. Song selection method
[0258] The user operates the interface on the screen of the terminal to select the song they want to sing. This song selection means includes a function for displaying a song list and accepting user input.
[0259] 2. Recording Method
[0260] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[0261] 3. Analysis method
[0262] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[0263] 4. Emotion Engine
[0264] The recorded voice data is passed through an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice data to determine the user's emotional state (e.g., joy, sadness, anger, etc.) and adjusts the analysis results accordingly.
[0265] 5. Display means
[0266] Based on the results of the analysis means and the recognition results of the emotion engine, the device displays the optimal key and the appropriate performance and advice for the emotion. This information is displayed on the screen of the device, providing visual feedback to the user.
[0267] A natural language description of the program's operation
[0268] System Overview
[0269] This system analyzes the user's singing voice and emotions, and suggests the optimal key and emotion, making karaoke more enjoyable and comfortable.
[0270] 1. Terminal Processing
[0271] The user selects the song they want to sing on the device.
[0272] The device displays a message inviting the user to sing a little a cappella.
[0273] When the user starts singing, the device records the voice through the microphone.
[0274] Once the recording is complete, the audio data is sent to the server.
[0275] 2. Server Processing
[0276] The server receives the voice data transmitted from the terminal.
[0277] To analyze the audio data, a request is sent to the music analysis API.
[0278] An emotion engine is used to recognize user emotions from voice data.
[0279] The analysis results from the music analysis API are combined with the results of the emotion engine to identify the optimal key and direction.
[0280] 3. Terminal Processing (cont.)
[0281] The analysis results received from the server and the emotion engine results are displayed on the screen.
[0282] The user checks the displayed key information and advice according to their emotions.
[0283] Specific examples
[0284] Example: A user selects "Let it be (The Beatles)" and optimizes it for D key, and the emotion engine recognizes "Joy."
[0285] 1. Song selection and singing
[0286] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[0287] The device displays the message, "Please sing a little a cappella."
[0288] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[0289] 2. Data submission and analysis
[0290] The recording data is sent to the server.
[0291] The server receives the audio data and sends it to the music analysis API for analysis.
[0292] At the same time, the server uses an emotion engine to analyze emotions and recognize "joy."
[0293] 3. Displaying the results and singing
[0294] The analysis results from the server ("The optimal key is D") and the emotion engine results ("Joy") are sent to the terminal.
[0295] The device displays on the screen, "The optimal key is D. Also, your emotion is joy, so sing with energy."
[0296] Based on the displayed key information and advice, the user can enjoy singing "Let it be" in the key of D.
[0297] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[0298] The processing flow will be explained below.
[0299] Step 1: Song Selection
[0300] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[0301] The terminal obtains information about the selected song and displays it on the screen.
[0302] Step 2: A cappella instructions
[0303] The device displays a message to the user saying, "Please sing a little a cappella."
[0304] The user checks the message.
[0305] Step 3: Start recording
[0306] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[0307] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[0308] Step 4: End recording
[0309] When the user finishes singing, the device will automatically stop recording.
[0310] The recorded audio data is temporarily stored in memory.
[0311] Step 5: Send data
[0312] The terminal transmits the recorded audio data to the server.
[0313] The server receives the HTTP request and stores the audio file.
[0314] Step 6: Data analysis
[0315] The server generates a request to send the stored audio data to the music analysis API.
[0316] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[0317] Step 7: Sentiment Analysis
[0318] The server uses an emotion engine to recognize the user's emotion from the stored audio data.
[0319] The emotion engine analyzes the voice data to determine the user's emotional state (e.g., happy, sad, angry, etc.).
[0320] Step 8: Obtaining analysis results
[0321] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[0322] The server combines the results of the music analysis and the emotion engine to generate advice based on the optimal key and emotion.
[0323] Step 9: Send results
[0324] The server organizes the analysis results and transmits them to the user's terminal.
[0325] As an HTTP response, the most appropriate key information and advice based on the emotion are returned to the device.
[0326] Step 10: View the results
[0327] The device then displays the analysis results on its screen, displaying the message, "The optimal key is D. Your emotion is joy, so sing with enthusiasm."
[0328] The user confirms this message.
[0329] Step 11: Adjust the key
[0330] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[0331] The device will retain the new settings after adjusting the keys.
[0332] Step 12: Karaoke playback
[0333] After the user completes the key setting, karaoke playback begins.
[0334] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[0335] Step 13: Enjoy singing
[0336] Users can sing songs in the optimal key and enjoy karaoke based on advice based on their emotions.
[0337] Example 2
[0338] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0339] Conventional karaoke systems can suggest the optimal key for a user's singing, but they cannot provide performances or advice that take the user's emotions into consideration. As a result, users cannot receive accurate advice or performances that reflect their emotions, limiting the quality of their karaoke experience. It is also difficult for users to enjoy a performance that matches their own emotions while singing. There is a need to solve this problem and provide a karaoke experience that is in tune with the user's emotions.
[0340] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0341] In this invention, the server includes an analysis means for analyzing the recorded sound source data, an emotion analysis means for recognizing the user's emotion from the recorded sound source data, and a display means for displaying the optimum key for the user's singing and performance and advice suited to the emotion based on the results of the analysis by the analysis means and the emotion analysis means. This makes it possible to provide not only the optimum key for the user's singing, but also appropriate advice and performance according to the emotion.
[0342] The "selection means" is a function that allows the user to select a song to sing through the interface of the karaoke system.
[0343] "Recording means" is a function for recording the user's singing through a microphone and saving it as digital audio data.
[0344] The "analysis means" is a function that uses recorded audio data to analyze pitch and interval and identify the optimal key for the user's singing.
[0345] The "emotion analysis means" is a function for recognizing and analyzing the user's emotions from the recorded audio data.
[0346] The "display means" is a function for visually presenting the analysis results obtained by the analysis means and emotion analysis means to the user, and for providing advice and presentations that correspond to the optimal key and emotion.
[0347] MODE FOR CARRYING OUT THE INVENTION
[0348] This invention improves the karaoke experience by analyzing the user's singing data and emotional data in a karaoke system and suggesting the optimal key and song based on the user's emotion. A specific method for implementing this invention will be described below.
[0349] Required Hardware and Software
[0350] To implement this system, a terminal and a server are required. The terminal has an interface for users to select songs and sing, and the server has the function of analyzing voice data and emotional data.
[0351] Terminal
[0352] Touchscreen display: The user operates this to select songs.
[0353] Microphone: Used to record the user's singing.
[0354] Internet connection: Required to send data to the server.
[0355] server
[0356] Music analysis API: Used to analyze audio data and identify the optimal key (e.g., Spotify API).
[0357] Sentiment analysis engine: Used to identify user emotions from voice data (e.g., IBM Watson).
[0358] Database: Used to store analysis results and user information.
[0359] System configuration
[0360] 1. Song selection method
[0361] The user selects the song they want to sing from the song list displayed on the terminal.
[0362] 2. Recording Method
[0363] After the user selects a song, the device displays a message saying, "Please sing a little a cappella," encouraging the user to sing a cappella.
[0364] When the user starts singing, the device records the voice through the microphone and stores it as digital data.
[0365] 3. Analysis method
[0366] The device sends the recorded audio data to a server using an internet connection.
[0367] The server uses a music analysis API to analyze the audio data and identify pitch and interval.
[0368] At the same time, the server uses an emotion analysis engine to recognize the user's emotions from the voice data.
[0369] 4. Display means
[0370] The server transmits the results obtained from the analysis means and the emotion analysis means to the terminal.
[0371] The terminal displays this information on the screen and provides the user with the most suitable key and appropriate performance and advice for their emotion.
[0372] Specific examples
[0373] If a user uses a karaoke system to select "Let it be (The Beatles)" and sing a bit a cappella, the system performs the following steps:
[0374] 1. Song Selection
[0375] The user selects "Let it be (The Beatles)" on the device's touchscreen.
[0376] 2. Singing and Recording
[0377] The device displays the message "Please sing a little a cappella," and the user sings "When I find myself in times of trouble..."
[0378] The user's singing is recorded and saved as digital data.
[0379] 3. Sending and analyzing audio data
[0380] The recording data is sent from the device to the server, which uses a music analysis API to identify the optimal key (D).
[0381] At the same time, an emotion analysis engine is used to recognize emotions (happiness).
[0382] 4. Displaying results and giving advice
[0383] The analysis results from the server (optimal key: D) and emotion analysis results (joy) are displayed on the device.
[0384] The device displays the message, "The optimal key is D. Also, your emotion is joy, so sing with enthusiasm."
[0385] Prompt Sentence Examples
[0386] "In a karaoke system, a user selects The Beatles' "Let it Be" and sings a portion a cappella. This singing voice is recorded, and the system analyzes it using an emotion engine. The system recognizes the user's emotion as "joy." The system also determines that the key that best suits the singing voice is D. Based on these results, what kind of feedback should be provided to the user?"
[0387] In this way, the present invention can provide a multi-functional karaoke system that not only allows the user to enjoy karaoke in the key that is best suited to them, but also allows them to receive advice according to their emotions at the time.
[0388] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0389] Specific explanation of processing steps
[0390] Step 1: Select a song
[0391] Terminal handling
[0392] Input: The user operates the touchscreen display to select a song from the displayed song list.
[0393] Operation:
[0394] The terminal receives touchscreen input and retrieves song information based on the user's selection.
[0395] Output: The title and artist name of the selected song will be displayed on the screen.
[0396] Step 2: Sing and record your a cappella
[0397] Terminal handling
[0398] Input: After a song is selected, the user begins singing.
[0399] Operation:
[0400] The device displays the message, "Please sing a little a cappella."
[0401] When the user starts singing, the voice is recorded through a microphone.
[0402] The recorded audio data is saved as a digital file.
[0403] Output: The recording data is saved on the device.
[0404] Step 3: Sending audio data to the server
[0405] Terminal handling
[0406] Input: Recorded data is saved.
[0407] Operation:
[0408] The device sends the recorded data to a server via the Internet.
[0409] The HTTPS protocol is used for data transmission.
[0410] Output: The recording data is sent to the server.
[0411] Step 4: Analyzing the audio data
[0412] Server Processing
[0413] Input: Recording data sent from the device.
[0414] Operation:
[0415] The server receives the recording data.
[0416] Sends a request to an audio analysis API (e.g., Spotify API) to analyze the pitch and interval of the audio data.
[0417] The voice analysis API returns the analysis results to the server.
[0418] Output: Audio analysis results (optimal key, etc.).
[0419] Step 5: Sentiment Analysis
[0420] Server Processing
[0421] Input: Recording data sent from the device.
[0422] Operation:
[0423] The server uses an emotion analysis engine (e.g., IBM Watson) to analyze the user's emotions.
[0424] The emotion analysis engine identifies emotions based on the recorded data and sends the results to the server.
[0425] Output: Sentiment analysis results (e.g., joy, sadness, anger, etc.).
[0426] Step 6: Combining the analysis results
[0427] Server Processing
[0428] Input: Voice analysis results and emotion analysis results.
[0429] Operation:
[0430] The server combines the results of the voice analysis and the emotion analysis.
[0431] Identify the best key and emotionally appropriate performance and advice.
[0432] Output: Final analysis results (optimal key and emotional advice).
[0433] Step 7: Viewing the analysis results
[0434] Terminal handling
[0435] Input: Final analysis result data from the server.
[0436] Operation:
[0437] The terminal displays the analysis results received from the server on the screen.
[0438] The user sees the best key and emotion-based advice.
[0439] Output: The screen displays "The best key is D. Your emotion is joy, so sing with enthusiasm."
[0440] In this way, the system analyzes the user's singing data at multiple stages and provides optimal feedback to the user, enhancing the karaoke experience.
[0441] (Application example 2)
[0442] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0443] Conventional karaoke systems have the problem of being unable to find the optimal key for the user when singing and providing a performance that reflects the user's emotions at that time. In particular, they are unable to provide feedback based on the user's emotions, preventing users from having a more enjoyable and personalized karaoke experience. This leaves karaoke booths and live music bars with a challenge to improve customer satisfaction.
[0444] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0445] In this invention, the server includes a selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded sound data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, an emotion engine for recognizing emotions from the user's voice data, and a means for providing the user with visual feedback based on the recognition results of the emotion engine. This not only enables the user to sing in the optimal key for themselves, but also allows them to receive advice based on their emotions at the time.
[0446] A "selection means" is a means that provides an interface for a user to select a song and sing along a bit a cappella with it.
[0447] "Recording means" means any means, including devices and software, for recording the audio of a user's singing.
[0448] The "analysis means" is a means for analyzing the recorded sound source data and performing a process to identify the appropriate key and other musical elements for the song.
[0449] The "display means" is a means for visually presenting the optimum key and other information to the user based on the analysis results obtained by the analysis means.
[0450] The "emotion engine" is a means including software and algorithms for recognizing emotions from the user's voice data and feeding back the recognition results to the analysis means.
[0451] The "means for providing visual feedback" is a means for visually displaying appropriate advice or dramatic effects to the user based on the recognition results of the emotion engine.
[0452] The present invention provides a system for enabling users to select songs and improve their singing experience in a physical venue such as a karaoke booth or live music bar, etc. The system includes a selection means, a recording means, an analysis means, a display means, an emotion engine, and a means for providing visual feedback.
[0453] System Configuration
[0454] 1. The selection means is a device such as a tablet or smartphone. The user operates the interface on the screen to select the song they want to sing. This selection means includes a function to display a song list and accept user input.
[0455] 2. The recording function provides a function to record the user's a cappella singing using a microphone. The recording method collects a certain amount of audio data and saves it as a digital file. For example, a microphone built into a smartphone or a microphone installed in a karaoke booth can be used.
[0456] 3. The analysis method involves analyzing the audio data on the server using a music analysis API (such as IBM Watson or Google Cloud Speech-to-Text API). This analysis involves analyzing pitch and interval and comparing it with existing data. Based on the analysis results, the optimal key for the user's singing is identified.
[0457] 4. The emotion engine uses software and algorithms to recognize emotions from the user's voice data. For example, it uses Microsoft Azure's Emotion API or a proprietary emotion recognition algorithm. This analyzes the user's emotional state (e.g., joy, sadness, anger, etc.).
[0458] 5. As a display method, visual feedback is provided to the user based on the analysis results and the emotion engine's recognition results. For example, optimal keys and advice based on emotions are displayed on the screen of a tablet or smartphone.
[0459] Example
[0460] Specific examples
[0461] Below is a case where the user selects the song "Let it be (The Beatles)" and analysis and emotion recognition are performed based on that.
[0462] 1. Song selection and singing:
[0463] The user selects "Let it be (The Beatles)" on the karaoke machine.
[0464] The device displays the message, "Please sing a little a cappella."
[0465] The user sings a bit of "When I find myself in times of trouble..." and their voice is recorded.
[0466] 2. Data Analysis and Emotion Recognition:
[0467] The recorded data is sent to the server and analyzed using a music analysis API.
[0468] At the same time, the emotion engine recognizes the user's emotion as "joy."
[0469] 3. Displaying the results:
[0470] The analysis results and emotion recognition results are sent to the device, and the message "The optimal key is D. Your emotion is joy, so sing with enthusiasm." is displayed.
[0471] Using the displayed key information and advice, users can enjoy singing "Let it be" in the key of D.
[0472] Prompt Sentence Examples
[0473] "User selects 'Let it be'. The recording is analyzed to determine the best key. The emotion engine recognizes 'joy' and provides advice and the best key."
[0474] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[0475] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0476] Step 1:
[0477] The user selects a song using the selection means on a tablet or smartphone. The input is the ID and title of the song selected by the user, which is provided to the system. The device displays a list of songs, accepts the user's selection, and obtains song information.
[0478] Step 2:
[0479] After the user selects a song, the device prompts the user to sing a little a cappella using the microphone. The input is the user's singing voice, and the output is the recorded audio data. When the user starts singing, the device captures the audio signal and stores it digitally as audio data.
[0480] Step 3:
[0481] The device sends the recorded audio data to the server. The input is the audio data file, and the output is the data received on the server. The server sends the audio data to the specified endpoint and confirms receipt.
[0482] Step 4:
[0483] The server sends the received audio data to the music analysis API and requests analysis. The input is audio data, and the output is analyzed musical element data such as interval and pitch. The server sends a request to the API endpoint and receives the analysis results in return.
[0484] Step 5:
[0485] The server then uses an emotion engine to recognize the user's emotion from the voice data. The input is the voice data and the output is the user's emotional state (e.g., "joy"). The emotion engine analyzes the voice data and generates extracted emotion information.
[0486] Step 6:
[0487] The server integrates the analysis results from the music analysis API and the recognition results from the emotion engine to identify the optimal key and advice according to the emotion. The input is musical element data and emotion information, and the output is the optimal key and advice. The server processes these data and generates the integrated result.
[0488] Step 7:
[0489] The server sends the integration result to the terminal and provides feedback to the user using a display means. The input is the integration result data, and the output is the optimal key and advice according to the emotion displayed on the user terminal. The terminal displays the received information on the screen and provides visual feedback to the user.
[0490] Step 8:
[0491] The user sings the song using the displayed optimal key and emotional advice. The input is the displayed information, and the output is the user's improved singing performance. Based on the feedback, the user can enjoy singing in the optimal key.
[0492] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0493] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0494] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0495] [Second embodiment]
[0496] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0497] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0498] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0499] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0500] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0501] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0502] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0503] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0504] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0505] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0506] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0507] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0508] To implement this invention, the following functions must be integrated into the karaoke system:
[0509] 1. Song selection method
[0510] The user selects the song he or she wants to sing by operating the interface on the terminal screen. This song selection means includes a function for displaying a song list and accepting user input.
[0511] 2. Recording Method
[0512] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[0513] 3. Analysis method
[0514] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[0515] 4. Display means
[0516] Based on the analysis results, the server derives the optimal key and sends the result to the terminal. The terminal's display means visually presents the analysis results to the user. For example, a notification such as "The optimal key is G" is displayed on the screen.
[0517] A natural language description of the program's operation
[0518] System Overview
[0519] The system provides functionality to allow users to sing in the optimal key for karaoke.
[0520] 1. Terminal Processing
[0521] The user selects the song they want to sing on the device.
[0522] The device displays a message inviting the user to sing a little a cappella.
[0523] When the user starts singing, the device records the voice through the microphone.
[0524] Once the recording is complete, the audio data is sent to the server.
[0525] 2. Server Processing
[0526] The server receives the voice data transmitted from the terminal.
[0527] To analyze the audio data, a request is sent to the music analysis API.
[0528] Receive analysis results containing relevant key information.
[0529] The analysis results are sent to the device.
[0530] 3. Terminal Processing (cont.)
[0531] The terminal displays the analysis results received from the server on the screen.
[0532] The user adjusts the key based on the displayed key information and prepares to sing in the optimum key.
[0533] Specific examples
[0534] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[0535] 1. Song selection and singing
[0536] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[0537] The device displays the message, "Please sing a little a cappella."
[0538] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[0539] 2. Data submission and analysis
[0540] The recording data is sent to the server.
[0541] The server receives the audio data and sends it to the music analysis API for analysis.
[0542] The music analysis API returns the result, "The optimal key is D."
[0543] 3. Displaying the results and singing
[0544] The terminal receives the analysis results from the server.
[0545] The device will display "The best key is D" on the screen.
[0546] Based on the displayed key information, the user sings "Let it be" in the key of D.
[0547] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[0548] The processing flow will be explained below.
[0549] Step 1: Song Selection
[0550] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[0551] The terminal obtains information about the selected song and displays it on the screen.
[0552] Step 2: A cappella instructions
[0553] The device displays a message to the user saying, "Please sing a little a cappella."
[0554] The user checks the message.
[0555] Step 3: Start recording
[0556] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[0557] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[0558] Step 4: End recording
[0559] When the user finishes singing, the device will automatically stop recording.
[0560] The recorded audio data is temporarily stored in memory.
[0561] Step 5: Send data
[0562] The terminal transmits the recorded audio data to the server.
[0563] The server receives the HTTP request and stores the audio file.
[0564] Step 6: Data analysis
[0565] The server generates a request to send the stored audio data to the music analysis API.
[0566] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[0567] Step 7: Obtaining analysis results
[0568] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[0569] The server stores the analysis results in a database or temporary memory.
[0570] Step 8: Send results
[0571] The server organizes the analysis results and transmits them to the user's terminal.
[0572] The optimal key information (e.g., "D key") is returned to the terminal as an HTTP response.
[0573] Step 9: View the results
[0574] The device will display the analysis results on the screen, with a message such as "The best key is D."
[0575] The user confirms this message.
[0576] Step 10: Adjust the key
[0577] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[0578] The device will retain the new settings after adjusting the keys.
[0579] Step 11: Karaoke playback
[0580] After the user completes the key setting, karaoke playback begins.
[0581] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[0582] Step 12: Enjoy singing
[0583] The user can sing songs in the optimum key and enjoy karaoke.
[0584] Example 1
[0585] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0586] Current karaoke systems require users to manually try out different keys to find the one that suits them best. This process is time-consuming and difficult, especially for users who are not familiar with music. A solution to this problem is needed to enable users to easily sing songs in the optimal key.
[0587] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0588] In this invention, the server includes a selection means for the user to select a song, a display means for instructing the user to sing a little a cappella, a recording means for recording the user's singing, a transmission means for transmitting the recorded sound data to the server, an analysis means for the server to analyze the sound data using a music analysis API, and a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means. This frees the user from the hassle of manually adjusting the key and allows them to easily enjoy karaoke in the optimal key for themselves.
[0589] The "selection means" is a function that allows the user to select a song to sing using the interface of the karaoke system.
[0590] The "means for displaying instructions" is a function that displays a message encouraging the user to sing a little a cappella.
[0591] The "recording means" is a function that records the user's singing voice through a microphone and serves to save the recorded audio data.
[0592] The "transmission means" is a function for converting the recorded sound source data into packets and transmitting them to the server.
[0593] The "analysis means" is a function that allows the server to analyze the received audio data using a music analysis API and identify the optimal key for the user.
[0594] The "display means" is a function for visually displaying to the user the optimum key information obtained by the analysis means.
[0595] To implement this invention, the following functions must be integrated into the karaoke system:
[0596] Selection method
[0597] On the device, the user operates the interface on the screen to select the song they want to sing. This selection means includes the function of displaying a song list and accepting user input. When the user selects a song, the device stores the selection information in its internal memory. For example, the user uses this selection means to select "Let it be (The Beatles)."
[0598] A means of displaying instructions
[0599] When the user selects a song, the device provides an interface that displays a message to the user saying, "Please sing a little a cappella." Specifically, a text message is displayed on the device screen.
[0600] Recording medium
[0601] When the user starts singing, the device will record the sound using the built-in microphone. This recording method will collect a certain amount of sound data (for example, 10 seconds) and save it as a digital file. The sound data will be saved in WAV or MP3 format.
[0602] Transmission method
[0603] Once the recording is complete, the device transmits the recorded audio data to the server, which encodes the audio data into an appropriate format and transmits it to the server via a network using a secure protocol (e.g., HTTPS).
[0604] Analysis means
[0605] The server analyzes the audio data received from the device. For the analysis, a music analysis API (such as Google Cloud Speech-to-Text API or IBM Watson) is used. The server sends an analysis request to the API and receives the analysis results of the audio data. For example, key information such as "The optimal key is D" is returned.
[0606] Display means
[0607] The server sends the analysis results to the device, which then visually presents the results to the user. This display includes a function that displays text such as "The best key is D" on the screen. Based on this information, the user can prepare to sing the song in the best key.
[0608] Specific examples
[0609] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[0610] 1. Song selection and singing
[0611] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[0612] The device displays the message, "Please sing a little a cappella."
[0613] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[0614] 2. Data submission and analysis
[0615] The recording data is sent to the server.
[0616] The server receives the audio data and sends it to the music analysis API for analysis.
[0617] The music analysis API returns the result, "The optimal key is D."
[0618] 3. Displaying the results and singing
[0619] The terminal receives the analysis results from the server.
[0620] The device screen will display "The best key is D."
[0621] Based on the displayed key information, the user sings "Let it be" in the key of D.
[0622] Prompt Sentence Examples
[0623] A typical example of a prompt to be input to a generative AI model is as follows:
[0624] plain
[0625] Please provide a detailed explanation of the steps a user takes to select "Let it be (The Beatles)" and optimize it for D. Please provide a detailed explanation of each step of the process, from the user selecting the song, recording the a cappella vocal, sending the audio data to the server, and receiving and displaying the analysis results.
[0626] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[0627] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0628] Step 1:
[0629] The user selects a song.
[0630] Specific operation: The user selects the song they want to sing from the song list provided on the karaoke terminal screen.
[0631] Input: The user touches or clicks on a specific song from the song list displayed on the screen.
[0632] Output: The information of the selected song is saved on the device. For example, if "Let it be (The Beatles)" is selected, the information is saved in the internal memory.
[0633] Step 2:
[0634] The device prompts the user to sing a little a cappella.
[0635] Specific operation: The device displays a message on the screen and instructs the user to "sing a little a cappella."
[0636] Input: After the user selects a song, instructions are displayed on the device screen.
[0637] Output: A visual instruction is given to the user, and the user is prepared to follow the instruction.
[0638] Step 3:
[0639] The device records the user singing a cappella.
[0640] Specific operation: When the user starts singing, the device will record the voice using the built-in microphone.
[0641] Input: The user's voice is input through the microphone.
[0642] Output: The recorded audio data is saved in a digital format (e.g., WAV or MP3). For example, the phrase "When I find myself in times of trouble" is recorded.
[0643] Step 4:
[0644] The device sends the recorded data to the server.
[0645] Specific operation: After the recording is completed, the device encodes the audio data into an appropriate format and sends it to the server over the network.
[0646] Input: Recorded audio data.
[0647] Output: The encoded audio data is split into packets and sent to the server using a secure protocol.
[0648] Step 5:
[0649] The server analyzes the audio data.
[0650] Specific operation: The server sends the received audio data to the music analysis API and analyzes the optimal key.
[0651] Input: The audio data received by the server.
[0652] Output: Analysis results from the music analysis API, e.g. "The optimal key is D."
[0653] Step 6:
[0654] The analysis results are sent from the server to the device.
[0655] Specific operation: The server converts the analysis results into a structured data format (e.g., JSON) and sends them back to the terminal using a secure protocol.
[0656] Input: Analysis result data.
[0657] Output: The analysis results are sent to the terminal in structured data format.
[0658] Step 7:
[0659] The device displays the analysis results to the user.
[0660] Specific operation: The device decodes the received analysis results and displays "The best key is D" on the screen.
[0661] Input: Analysis results received from the server.
[0662] Output: A visual display to the user, where the user obtains the best key information. For example, "The best key is D" is displayed on the terminal screen.
[0663] (Application example 1)
[0664] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0665] Conventional karaoke systems require users to manually adjust the key through trial and error to find the optimal key for them, which is time-consuming and laborious. Furthermore, they lack a function to identify the optimal key based on the user's singing style, making it difficult for users without specialized musical knowledge. Furthermore, they lack a system that utilizes cloud technology to quickly provide analysis results.
[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0667] In this invention, the server includes a song selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded audio data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, and a communication means for the analysis means to send the analysis results to the cloud server based on a music analysis API and receive the analysis results. This allows a user to easily select a song through a smartphone application, send the singing voice to the cloud server for analysis, and quickly find the optimal key for themselves.
[0668] The "music selection means" is a function that allows the user to operate the interface to select music.
[0669] The "recording means" is a function that records the audio source sung by the user and saves the audio data as a digital file.
[0670] The "analysis means" is a function for analyzing recorded sound source data, analyzing pitch and interval, and identifying the optimum key.
[0671] The "display means" is a function that visually presents the analysis results obtained by the analysis means to the user.
[0672] The "communication means" is a function that allows the analysis means to send the analysis results to the cloud server and receive the analysis results.
[0673] A "music analysis API" is an application programming interface for analyzing sound source data and analyzing data such as pitch and interval.
[0674] A "cloud server" refers to a computer server that stores and processes data remotely over a network.
[0675] A "smartphone" is a portable information terminal that can execute multifunctional applications in addition to making calls.
[0676] In this embodiment, a system will be described in which a user can use a smartphone application to enjoy karaoke in the optimal key.
[0677] Hardware and Software Used
[0678] Smartphone: Used as a device on which users can select songs and sing a cappella.
[0679] Cloud Server: The server used to analyze the a cappella recordings.
[0680] Music Analysis API: An API that runs on a cloud server and analyzes audio data to identify the optimal key.
[0681] User interface: The interface on the smartphone application that allows users to select songs and display the results.
[0682] System Operation Details
[0683] When a user launches the smartphone application, they are first presented with an interface for selecting a song. At this stage, the user selects the song they want to sing from a list on the application. The application then prompts the user to sing a little a cappella, and records the sound using the smartphone's microphone.
[0684] Once recording is complete, the application sends the recorded audio data to a cloud server. The cloud server receives the recorded data and sends a request to the music analysis API for analysis. The music analysis API analyzes the audio data and generates key information that allows the user to sing in the optimal key.
[0685] After receiving the analysis results, the cloud server sends them to the smartphone. Based on the analysis results, the smartphone application displays a notification to the user, such as "The optimal key is G." This allows the user to enjoy singing in the key that is best for them.
[0686] Specific examples
[0687] For example, suppose a user selects "Let it be" and sings a cappella "When I find myself in times of trouble...". This audio is recorded on a smartphone and sent to a cloud server. The server uses a music analysis API to analyze the audio data and obtains the result that "the optimal key is D". The result is sent to the smartphone, which displays the message "The optimal key is D". Based on the displayed key information, the user can sing "Let it be" in the key of D.
[0688] Prompt Sentence Examples
[0689] For example, a possible prompt might be, "Please implement an API that analyzes recorded audio data and finds the optimal key."
[0690] In this way, the user can easily enjoy karaoke in the key that is most suitable for them, without requiring any specialized knowledge.
[0691] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0692] Step 1:
[0693] Song selection method
[0694] The user launches the smartphone application and selects a song.
[0695] Input: Song name specified by the user in the app
[0696] Output: Selected song data
[0697] Processing operation: A song list is displayed through the smartphone's user interface, and the user selects the song they want to sing. At this time, the application saves the data of the selected song in the internal memory.
[0698] Step 2:
[0699] Recording medium
[0700] After the user selects a song, they are prompted to sing a cappella, and when they start singing, their voice is recorded by the smartphone microphone.
[0701] Input: User's a cappella singing voice
[0702] Output: Recorded a cappella audio data (digital file)
[0703] Processing operation: When recording starts, the user's singing voice is collected through the microphone and a certain period of audio data is saved as a digital file.
[0704] Step 3:
[0705] communication means
[0706] Once recording is complete, the audio data is sent to a cloud server.
[0707] Input: Recorded a cappella audio data
[0708] Output: Audio data sent to the cloud server
[0709] Processing operation: The smartphone sends the recorded audio data to the cloud server via an HTTP request. Once the transmission is complete, the server confirms receipt of the data.
[0710] Step 4:
[0711] Analysis means
[0712] The cloud server sends the received audio data to the music analysis API and requests analysis.
[0713] Input: Received a cappella audio data
[0714] Output: Analysis results from the music analysis API (optimal key information)
[0715] Processing: The server sends the audio data to the music analysis API and makes an analysis request. The API analyzes the audio data and returns the optimal key information to the cloud server.
[0716] Step 5:
[0717] communication means
[0718] The cloud server sends the analysis results received from the music analysis API to the smartphone.
[0719] Input: Analysis results (optimal key information)
[0720] Output: Analysis results sent to your smartphone
[0721] Processing operation: After obtaining the analysis results, the cloud server returns them to the smartphone application as an HTTP response.
[0722] Step 6:
[0723] Display means
[0724] The smartphone then displays the analysis results to the user, such as "The optimal key is G."
[0725] Input: Analysis results received from the cloud server
[0726] Output: Key information that is best visible to the user
[0727] Processing operation: The smartphone displays the appropriate key information on the interface based on the analysis results received.
[0728] In this way, through a series of processing steps, the user can enjoy karaoke in the key that is most suitable for him or her.
[0729] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0730] To implement this invention, it is necessary to incorporate an emotion engine into the karaoke system and integrate the following series of functions.
[0731] 1. Song selection method
[0732] The user operates the interface on the screen of the terminal to select the song they want to sing. This song selection means includes a function for displaying a song list and accepting user input.
[0733] 2. Recording Method
[0734] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[0735] 3. Analysis method
[0736] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[0737] 4. Emotion Engine
[0738] The recorded voice data is passed through an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice data to determine the user's emotional state (e.g., joy, sadness, anger, etc.) and adjusts the analysis results accordingly.
[0739] 5. Display means
[0740] Based on the results of the analysis means and the recognition results of the emotion engine, the device displays the optimal key and the appropriate performance and advice for the emotion. This information is displayed on the screen of the device, providing visual feedback to the user.
[0741] A natural language description of the program's operation
[0742] System Overview
[0743] This system analyzes the user's singing voice and emotions, and suggests the optimal key and emotion, making karaoke more enjoyable and comfortable.
[0744] 1. Terminal Processing
[0745] The user selects the song they want to sing on the device.
[0746] The device displays a message inviting the user to sing a little a cappella.
[0747] When the user starts singing, the device records the voice through the microphone.
[0748] Once the recording is complete, the audio data is sent to the server.
[0749] 2. Server Processing
[0750] The server receives the voice data transmitted from the terminal.
[0751] To analyze the audio data, a request is sent to the music analysis API.
[0752] An emotion engine is used to recognize user emotions from voice data.
[0753] The analysis results from the music analysis API are combined with the results of the emotion engine to identify the optimal key and direction.
[0754] 3. Terminal Processing (cont.)
[0755] The analysis results received from the server and the emotion engine results are displayed on the screen.
[0756] The user checks the displayed key information and advice according to their emotions.
[0757] Specific examples
[0758] Example: A user selects "Let it be (The Beatles)" and optimizes it for D key, and the emotion engine recognizes "Joy."
[0759] 1. Song selection and singing
[0760] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[0761] The device displays the message, "Please sing a little a cappella."
[0762] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[0763] 2. Data submission and analysis
[0764] The recording data is sent to the server.
[0765] The server receives the audio data and sends it to the music analysis API for analysis.
[0766] At the same time, the server uses an emotion engine to analyze emotions and recognize "joy."
[0767] 3. Displaying the results and singing
[0768] The analysis results from the server ("The optimal key is D") and the emotion engine results ("Joy") are sent to the terminal.
[0769] The device displays on the screen, "The optimal key is D. Also, your emotion is joy, so sing with energy."
[0770] Based on the displayed key information and advice, the user can enjoy singing "Let it be" in the key of D.
[0771] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[0772] The processing flow will be explained below.
[0773] Step 1: Song Selection
[0774] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[0775] The terminal obtains information about the selected song and displays it on the screen.
[0776] Step 2: A cappella instructions
[0777] The device displays a message to the user saying, "Please sing a little a cappella."
[0778] The user checks the message.
[0779] Step 3: Start recording
[0780] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[0781] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[0782] Step 4: End recording
[0783] When the user finishes singing, the device will automatically stop recording.
[0784] The recorded audio data is temporarily stored in memory.
[0785] Step 5: Send data
[0786] The terminal transmits the recorded audio data to the server.
[0787] The server receives the HTTP request and stores the audio file.
[0788] Step 6: Data analysis
[0789] The server generates a request to send the stored audio data to the music analysis API.
[0790] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[0791] Step 7: Sentiment Analysis
[0792] The server uses an emotion engine to recognize the user's emotion from the stored audio data.
[0793] The emotion engine analyzes the voice data to determine the user's emotional state (e.g., happy, sad, angry, etc.).
[0794] Step 8: Obtaining analysis results
[0795] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[0796] The server combines the results of the music analysis and the emotion engine to generate advice based on the optimal key and emotion.
[0797] Step 9: Send results
[0798] The server organizes the analysis results and transmits them to the user's terminal.
[0799] As an HTTP response, the most appropriate key information and advice based on the emotion are returned to the device.
[0800] Step 10: View the results
[0801] The device then displays the analysis results on its screen, displaying the message, "The optimal key is D. Your emotion is joy, so sing with enthusiasm."
[0802] The user confirms this message.
[0803] Step 11: Adjust the key
[0804] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[0805] The device will retain the new settings after adjusting the keys.
[0806] Step 12: Karaoke playback
[0807] After the user completes the key setting, karaoke playback begins.
[0808] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[0809] Step 13: Enjoy singing
[0810] Users can sing songs in the optimal key and enjoy karaoke based on advice based on their emotions.
[0811] Example 2
[0812] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0813] Conventional karaoke systems can suggest the optimal key for a user's singing, but they cannot provide performances or advice that take the user's emotions into consideration. As a result, users cannot receive accurate advice or performances that reflect their emotions, limiting the quality of their karaoke experience. It is also difficult for users to enjoy a performance that matches their own emotions while singing. There is a need to solve this problem and provide a karaoke experience that is in tune with the user's emotions.
[0814] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0815] In this invention, the server includes an analysis means for analyzing the recorded sound source data, an emotion analysis means for recognizing the user's emotion from the recorded sound source data, and a display means for displaying the optimum key for the user's singing and performance and advice suited to the emotion based on the results of the analysis by the analysis means and the emotion analysis means. This makes it possible to provide not only the optimum key for the user's singing, but also appropriate advice and performance according to the emotion.
[0816] The "selection means" is a function that allows the user to select a song to sing through the interface of the karaoke system.
[0817] "Recording means" is a function for recording the user's singing through a microphone and saving it as digital audio data.
[0818] The "analysis means" is a function that uses recorded audio data to analyze pitch and interval and identify the optimal key for the user's singing.
[0819] The "emotion analysis means" is a function for recognizing and analyzing the user's emotions from the recorded audio data.
[0820] The "display means" is a function for visually presenting the analysis results obtained by the analysis means and emotion analysis means to the user, and for providing advice and presentations that correspond to the optimal key and emotion.
[0821] MODE FOR CARRYING OUT THE INVENTION
[0822] This invention improves the karaoke experience by analyzing the user's singing data and emotional data in a karaoke system and suggesting the optimal key and song based on the user's emotion. A specific method for implementing this invention will be described below.
[0823] Required Hardware and Software
[0824] To implement this system, a terminal and a server are required. The terminal has an interface for users to select songs and sing, and the server has the function of analyzing voice data and emotional data.
[0825] Terminal
[0826] Touchscreen display: The user operates this to select songs.
[0827] Microphone: Used to record the user's singing.
[0828] Internet connection: Required to send data to the server.
[0829] server
[0830] Music analysis API: Used to analyze audio data and identify the optimal key (e.g., Spotify API).
[0831] Sentiment analysis engine: Used to identify user emotions from voice data (e.g., IBM Watson).
[0832] Database: Used to store analysis results and user information.
[0833] System configuration
[0834] 1. Song selection method
[0835] The user selects the song they want to sing from the song list displayed on the terminal.
[0836] 2. Recording Method
[0837] After the user selects a song, the device displays a message saying, "Please sing a little a cappella," encouraging the user to sing a cappella.
[0838] When the user starts singing, the device records the voice through the microphone and stores it as digital data.
[0839] 3. Analysis method
[0840] The device sends the recorded audio data to a server using an internet connection.
[0841] The server uses a music analysis API to analyze the audio data and identify pitch and interval.
[0842] At the same time, the server uses an emotion analysis engine to recognize the user's emotions from the voice data.
[0843] 4. Display means
[0844] The server transmits the results obtained from the analysis means and the emotion analysis means to the terminal.
[0845] The terminal displays this information on the screen and provides the user with the most suitable key and appropriate performance and advice for their emotion.
[0846] Specific examples
[0847] If a user uses a karaoke system to select "Let it be (The Beatles)" and sing a bit a cappella, the system performs the following steps:
[0848] 1. Song Selection
[0849] The user selects "Let it be (The Beatles)" on the device's touchscreen.
[0850] 2. Singing and Recording
[0851] The device displays the message "Please sing a little a cappella," and the user sings "When I find myself in times of trouble..."
[0852] The user's singing is recorded and saved as digital data.
[0853] 3. Sending and analyzing audio data
[0854] The recording data is sent from the device to the server, which uses a music analysis API to identify the optimal key (D).
[0855] At the same time, an emotion analysis engine is used to recognize emotions (happiness).
[0856] 4. Displaying results and giving advice
[0857] The analysis results from the server (optimal key: D) and emotion analysis results (joy) are displayed on the device.
[0858] The device displays the message, "The optimal key is D. Also, your emotion is joy, so sing with enthusiasm."
[0859] Prompt Sentence Examples
[0860] "In a karaoke system, a user selects The Beatles' "Let it Be" and sings a portion a cappella. This singing voice is recorded, and the system analyzes it using an emotion engine. The system recognizes the user's emotion as "joy." The system also determines that the key that best suits the singing voice is D. Based on these results, what kind of feedback should be provided to the user?"
[0861] In this way, the present invention can provide a multi-functional karaoke system that not only allows the user to enjoy karaoke in the key that is best suited to them, but also allows them to receive advice according to their emotions at the time.
[0862] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0863] Specific explanation of processing steps
[0864] Step 1: Select a song
[0865] Terminal handling
[0866] Input: The user operates the touchscreen display to select a song from the displayed song list.
[0867] Operation:
[0868] The terminal receives touchscreen input and retrieves song information based on the user's selection.
[0869] Output: The title and artist name of the selected song will be displayed on the screen.
[0870] Step 2: Sing and record your a cappella
[0871] Terminal handling
[0872] Input: After a song is selected, the user begins singing.
[0873] Operation:
[0874] The device displays the message, "Please sing a little a cappella."
[0875] When the user starts singing, the voice is recorded through a microphone.
[0876] The recorded audio data is saved as a digital file.
[0877] Output: The recording data is saved on the device.
[0878] Step 3: Sending audio data to the server
[0879] Terminal handling
[0880] Input: Recorded data is saved.
[0881] Operation:
[0882] The device sends the recorded data to a server via the Internet.
[0883] The HTTPS protocol is used for data transmission.
[0884] Output: The recording data is sent to the server.
[0885] Step 4: Analyzing the audio data
[0886] Server Processing
[0887] Input: Recording data sent from the device.
[0888] Operation:
[0889] The server receives the recording data.
[0890] Sends a request to an audio analysis API (e.g., Spotify API) to analyze the pitch and interval of the audio data.
[0891] The voice analysis API returns the analysis results to the server.
[0892] Output: Audio analysis results (optimal key, etc.).
[0893] Step 5: Sentiment Analysis
[0894] Server Processing
[0895] Input: Recording data sent from the device.
[0896] Operation:
[0897] The server uses an emotion analysis engine (e.g., IBM Watson) to analyze the user's emotions.
[0898] The emotion analysis engine identifies emotions based on the recorded data and sends the results to the server.
[0899] Output: Sentiment analysis results (e.g., joy, sadness, anger, etc.).
[0900] Step 6: Combining the analysis results
[0901] Server Processing
[0902] Input: Voice analysis results and emotion analysis results.
[0903] Operation:
[0904] The server combines the results of the voice analysis and the emotion analysis.
[0905] Identify the best key and emotionally appropriate performance and advice.
[0906] Output: Final analysis results (optimal key and emotional advice).
[0907] Step 7: Viewing the analysis results
[0908] Terminal handling
[0909] Input: Final analysis result data from the server.
[0910] Operation:
[0911] The terminal displays the analysis results received from the server on the screen.
[0912] The user sees the best key and emotion-based advice.
[0913] Output: The screen displays "The best key is D. Your emotion is joy, so sing with enthusiasm."
[0914] In this way, the system analyzes the user's singing data at multiple stages and provides optimal feedback to the user, enhancing the karaoke experience.
[0915] (Application example 2)
[0916] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0917] Conventional karaoke systems have the problem of being unable to find the optimal key for the user when singing and providing a performance that reflects the user's emotions at that time. In particular, they are unable to provide feedback based on the user's emotions, preventing users from having a more enjoyable and personalized karaoke experience. This leaves karaoke booths and live music bars with a challenge to improve customer satisfaction.
[0918] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0919] In this invention, the server includes a selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded sound data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, an emotion engine for recognizing emotions from the user's voice data, and a means for providing the user with visual feedback based on the recognition results of the emotion engine. This not only enables the user to sing in the optimal key for themselves, but also allows them to receive advice based on their emotions at the time.
[0920] A "selection means" is a means that provides an interface for a user to select a song and sing along a bit a cappella with it.
[0921] "Recording means" means any means, including devices and software, for recording the audio of a user's singing.
[0922] The "analysis means" is a means for analyzing the recorded sound source data and performing a process to identify the appropriate key and other musical elements for the song.
[0923] The "display means" is a means for visually presenting the optimum key and other information to the user based on the analysis results obtained by the analysis means.
[0924] The "emotion engine" is a means including software and algorithms for recognizing emotions from the user's voice data and feeding back the recognition results to the analysis means.
[0925] The "means for providing visual feedback" is a means for visually displaying appropriate advice or dramatic effects to the user based on the recognition results of the emotion engine.
[0926] The present invention provides a system for enabling users to select songs and improve their singing experience in a physical venue such as a karaoke booth or live music bar, etc. The system includes a selection means, a recording means, an analysis means, a display means, an emotion engine, and a means for providing visual feedback.
[0927] System Configuration
[0928] 1. The selection means is a device such as a tablet or smartphone. The user operates the interface on the screen to select the song they want to sing. This selection means includes a function to display a song list and accept user input.
[0929] 2. The recording function provides a function to record the user's a cappella singing using a microphone. The recording method collects a certain amount of audio data and saves it as a digital file. For example, a microphone built into a smartphone or a microphone installed in a karaoke booth can be used.
[0930] 3. The analysis method involves analyzing the audio data on the server using a music analysis API (such as IBM Watson or Google Cloud Speech-to-Text API). This analysis involves analyzing pitch and interval and comparing it with existing data. Based on the analysis results, the optimal key for the user's singing is identified.
[0931] 4. The emotion engine uses software and algorithms to recognize emotions from the user's voice data. For example, it uses Microsoft Azure's Emotion API or a proprietary emotion recognition algorithm. This analyzes the user's emotional state (e.g., joy, sadness, anger, etc.).
[0932] 5. As a display method, visual feedback is provided to the user based on the analysis results and the emotion engine's recognition results. For example, optimal keys and advice based on emotions are displayed on the screen of a tablet or smartphone.
[0933] Example
[0934] Specific examples
[0935] Below is a case where the user selects the song "Let it be (The Beatles)" and analysis and emotion recognition are performed based on that.
[0936] 1. Song selection and singing:
[0937] The user selects "Let it be (The Beatles)" on the karaoke machine.
[0938] The device displays the message, "Please sing a little a cappella."
[0939] The user sings a bit of "When I find myself in times of trouble..." and their voice is recorded.
[0940] 2. Data Analysis and Emotion Recognition:
[0941] The recorded data is sent to the server and analyzed using a music analysis API.
[0942] At the same time, the emotion engine recognizes the user's emotion as "joy."
[0943] 3. Displaying the results:
[0944] The analysis results and emotion recognition results are sent to the device, and the message "The optimal key is D. Your emotion is joy, so sing with enthusiasm." is displayed.
[0945] Using the displayed key information and advice, users can enjoy singing "Let it be" in the key of D.
[0946] Prompt Sentence Examples
[0947] "User selects 'Let it be'. The recording is analyzed to determine the best key. The emotion engine recognizes 'joy' and provides advice and the best key."
[0948] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[0949] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0950] Step 1:
[0951] The user selects a song using the selection means on a tablet or smartphone. The input is the ID and title of the song selected by the user, which is provided to the system. The device displays a list of songs, accepts the user's selection, and obtains song information.
[0952] Step 2:
[0953] After the user selects a song, the device prompts the user to sing a little a cappella using the microphone. The input is the user's singing voice, and the output is the recorded audio data. When the user starts singing, the device captures the audio signal and stores it digitally as audio data.
[0954] Step 3:
[0955] The device sends the recorded audio data to the server. The input is the audio data file, and the output is the data received on the server. The server sends the audio data to the specified endpoint and confirms receipt.
[0956] Step 4:
[0957] The server sends the received audio data to the music analysis API and requests analysis. The input is audio data, and the output is analyzed musical element data such as interval and pitch. The server sends a request to the API endpoint and receives the analysis results in return.
[0958] Step 5:
[0959] The server then uses an emotion engine to recognize the user's emotion from the voice data. The input is the voice data and the output is the user's emotional state (e.g., "joy"). The emotion engine analyzes the voice data and generates extracted emotion information.
[0960] Step 6:
[0961] The server integrates the analysis results from the music analysis API and the recognition results from the emotion engine to identify the optimal key and advice according to the emotion. The input is musical element data and emotion information, and the output is the optimal key and advice. The server processes these data and generates the integrated result.
[0962] Step 7:
[0963] The server sends the integration result to the terminal and provides feedback to the user using a display means. The input is the integration result data, and the output is the optimal key and advice according to the emotion displayed on the user terminal. The terminal displays the received information on the screen and provides visual feedback to the user.
[0964] Step 8:
[0965] The user sings the song using the displayed optimal key and emotional advice. The input is the displayed information, and the output is the user's improved singing performance. Based on the feedback, the user can enjoy singing in the optimal key.
[0966] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0967] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0968] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0969] [Third embodiment]
[0970] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0971] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0972] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0973] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0974] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0975] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0976] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0977] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0978] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0979] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0980] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0981] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0982] To implement this invention, the following functions must be integrated into the karaoke system:
[0983] 1. Song selection method
[0984] The user selects the song he or she wants to sing by operating the interface on the terminal screen. This song selection means includes a function for displaying a song list and accepting user input.
[0985] 2. Recording Method
[0986] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[0987] 3. Analysis method
[0988] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[0989] 4. Display means
[0990] Based on the analysis results, the server derives the optimal key and sends the result to the terminal. The terminal's display means visually presents the analysis results to the user. For example, a notification such as "The optimal key is G" is displayed on the screen.
[0991] A natural language description of the program's operation
[0992] System Overview
[0993] The system provides functionality to allow users to sing in the optimal key for karaoke.
[0994] 1. Terminal Processing
[0995] The user selects the song they want to sing on the device.
[0996] The device displays a message inviting the user to sing a little a cappella.
[0997] When the user starts singing, the device records the voice through the microphone.
[0998] Once the recording is complete, the audio data is sent to the server.
[0999] 2. Server Processing
[1000] The server receives the voice data transmitted from the terminal.
[1001] To analyze the audio data, a request is sent to the music analysis API.
[1002] Receive analysis results containing relevant key information.
[1003] The analysis results are sent to the device.
[1004] 3. Terminal Processing (cont.)
[1005] The terminal displays the analysis results received from the server on the screen.
[1006] The user adjusts the key based on the displayed key information and prepares to sing in the optimum key.
[1007] Specific examples
[1008] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[1009] 1. Song selection and singing
[1010] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[1011] The device displays the message, "Please sing a little a cappella."
[1012] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[1013] 2. Data submission and analysis
[1014] The recording data is sent to the server.
[1015] The server receives the audio data and sends it to the music analysis API for analysis.
[1016] The music analysis API returns the result, "The optimal key is D."
[1017] 3. Displaying the results and singing
[1018] The terminal receives the analysis results from the server.
[1019] The device will display "The best key is D" on the screen.
[1020] Based on the displayed key information, the user sings "Let it be" in the key of D.
[1021] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[1022] The processing flow will be explained below.
[1023] Step 1: Song Selection
[1024] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[1025] The terminal obtains information about the selected song and displays it on the screen.
[1026] Step 2: A cappella instructions
[1027] The device displays a message to the user saying, "Please sing a little a cappella."
[1028] The user checks the message.
[1029] Step 3: Start recording
[1030] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[1031] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[1032] Step 4: End recording
[1033] When the user finishes singing, the device will automatically stop recording.
[1034] The recorded audio data is temporarily stored in memory.
[1035] Step 5: Send data
[1036] The terminal transmits the recorded audio data to the server.
[1037] The server receives the HTTP request and stores the audio file.
[1038] Step 6: Data analysis
[1039] The server generates a request to send the stored audio data to the music analysis API.
[1040] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[1041] Step 7: Obtaining analysis results
[1042] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[1043] The server stores the analysis results in a database or temporary memory.
[1044] Step 8: Send results
[1045] The server organizes the analysis results and transmits them to the user's terminal.
[1046] The optimal key information (e.g., "D key") is returned to the terminal as an HTTP response.
[1047] Step 9: View the results
[1048] The device will display the analysis results on the screen, with a message such as "The best key is D."
[1049] The user confirms this message.
[1050] Step 10: Adjust the key
[1051] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[1052] The device will retain the new settings after adjusting the keys.
[1053] Step 11: Karaoke playback
[1054] After the user completes the key setting, karaoke playback begins.
[1055] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[1056] Step 12: Enjoy singing
[1057] The user can sing songs in the optimum key and enjoy karaoke.
[1058] Example 1
[1059] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1060] Current karaoke systems require users to manually try out different keys to find the one that suits them best. This process is time-consuming and difficult, especially for users who are not familiar with music. A solution to this problem is needed to enable users to easily sing songs in the optimal key.
[1061] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1062] In this invention, the server includes a selection means for the user to select a song, a display means for instructing the user to sing a little a cappella, a recording means for recording the user's singing, a transmission means for transmitting the recorded sound data to the server, an analysis means for the server to analyze the sound data using a music analysis API, and a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means. This frees the user from the hassle of manually adjusting the key and allows them to easily enjoy karaoke in the optimal key for themselves.
[1063] The "selection means" is a function that allows the user to select a song to sing using the interface of the karaoke system.
[1064] The "means for displaying instructions" is a function that displays a message encouraging the user to sing a little a cappella.
[1065] The "recording means" is a function that records the user's singing voice through a microphone and serves to save the recorded audio data.
[1066] The "transmission means" is a function for converting the recorded sound source data into packets and transmitting them to the server.
[1067] The "analysis means" is a function that allows the server to analyze the received audio data using a music analysis API and identify the optimal key for the user.
[1068] The "display means" is a function for visually displaying to the user the optimum key information obtained by the analysis means.
[1069] To implement this invention, the following functions must be integrated into the karaoke system:
[1070] Selection method
[1071] On the device, the user operates the interface on the screen to select the song they want to sing. This selection means includes the function of displaying a song list and accepting user input. When the user selects a song, the device stores the selection information in its internal memory. For example, the user uses this selection means to select "Let it be (The Beatles)."
[1072] A means of displaying instructions
[1073] When the user selects a song, the device provides an interface that displays a message to the user saying, "Please sing a little a cappella." Specifically, a text message is displayed on the device screen.
[1074] Recording medium
[1075] When the user starts singing, the device will record the sound using the built-in microphone. This recording method will collect a certain amount of sound data (for example, 10 seconds) and save it as a digital file. The sound data will be saved in WAV or MP3 format.
[1076] Transmission method
[1077] Once the recording is complete, the device transmits the recorded audio data to the server, which encodes the audio data into an appropriate format and transmits it to the server via a network using a secure protocol (e.g., HTTPS).
[1078] Analysis means
[1079] The server analyzes the audio data received from the device. For the analysis, a music analysis API (such as Google Cloud Speech-to-Text API or IBM Watson) is used. The server sends an analysis request to the API and receives the analysis results of the audio data. For example, key information such as "The optimal key is D" is returned.
[1080] Display means
[1081] The server sends the analysis results to the device, which then visually presents the results to the user. This display includes a function that displays text such as "The best key is D" on the screen. Based on this information, the user can prepare to sing the song in the best key.
[1082] Specific examples
[1083] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[1084] 1. Song selection and singing
[1085] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[1086] The device displays the message, "Please sing a little a cappella."
[1087] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[1088] 2. Data submission and analysis
[1089] The recording data is sent to the server.
[1090] The server receives the audio data and sends it to the music analysis API for analysis.
[1091] The music analysis API returns the result, "The optimal key is D."
[1092] 3. Displaying the results and singing
[1093] The terminal receives the analysis results from the server.
[1094] The device screen will display "The best key is D."
[1095] Based on the displayed key information, the user sings "Let it be" in the key of D.
[1096] Prompt Sentence Examples
[1097] A typical example of a prompt to be input to a generative AI model is as follows:
[1098] plain
[1099] Please provide a detailed explanation of the steps a user takes to select "Let it be (The Beatles)" and optimize it for D. Please provide a detailed explanation of each step of the process, from the user selecting the song, recording the a cappella vocal, sending the audio data to the server, and receiving and displaying the analysis results.
[1100] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[1101] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1102] Step 1:
[1103] The user selects a song.
[1104] Specific operation: The user selects the song they want to sing from the song list provided on the karaoke terminal screen.
[1105] Input: The user touches or clicks on a specific song from the song list displayed on the screen.
[1106] Output: The information of the selected song is saved on the device. For example, if "Let it be (The Beatles)" is selected, the information is saved in the internal memory.
[1107] Step 2:
[1108] The device prompts the user to sing a little a cappella.
[1109] Specific operation: The device displays a message on the screen and instructs the user to "sing a little a cappella."
[1110] Input: After the user selects a song, instructions are displayed on the device screen.
[1111] Output: A visual instruction is given to the user, and the user is prepared to follow the instruction.
[1112] Step 3:
[1113] The device records the user singing a cappella.
[1114] Specific operation: When the user starts singing, the device will record the voice using the built-in microphone.
[1115] Input: The user's voice is input through the microphone.
[1116] Output: The recorded audio data is saved in a digital format (e.g., WAV or MP3). For example, the phrase "When I find myself in times of trouble" is recorded.
[1117] Step 4:
[1118] The device sends the recorded data to the server.
[1119] Specific operation: After the recording is completed, the device encodes the audio data into an appropriate format and sends it to the server over the network.
[1120] Input: Recorded audio data.
[1121] Output: The encoded audio data is split into packets and sent to the server using a secure protocol.
[1122] Step 5:
[1123] The server analyzes the audio data.
[1124] Specific operation: The server sends the received audio data to the music analysis API and analyzes the optimal key.
[1125] Input: The audio data received by the server.
[1126] Output: Analysis results from the music analysis API, e.g. "The optimal key is D."
[1127] Step 6:
[1128] The analysis results are sent from the server to the device.
[1129] Specific operation: The server converts the analysis results into a structured data format (e.g., JSON) and sends them back to the terminal using a secure protocol.
[1130] Input: Analysis result data.
[1131] Output: The analysis results are sent to the terminal in structured data format.
[1132] Step 7:
[1133] The device displays the analysis results to the user.
[1134] Specific operation: The device decodes the received analysis results and displays "The best key is D" on the screen.
[1135] Input: Analysis results received from the server.
[1136] Output: A visual display to the user, where the user obtains the best key information. For example, "The best key is D" is displayed on the terminal screen.
[1137] (Application example 1)
[1138] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1139] Conventional karaoke systems require users to manually adjust the key through trial and error to find the optimal key for them, which is time-consuming and laborious. Furthermore, they lack a function to identify the optimal key based on the user's singing style, making it difficult for users without specialized musical knowledge. Furthermore, they lack a system that utilizes cloud technology to quickly provide analysis results.
[1140] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1141] In this invention, the server includes a song selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded audio data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, and a communication means for the analysis means to send the analysis results to the cloud server based on a music analysis API and receive the analysis results. This allows a user to easily select a song through a smartphone application, send the singing voice to the cloud server for analysis, and quickly find the optimal key for themselves.
[1142] The "music selection means" is a function that allows the user to operate the interface to select music.
[1143] The "recording means" is a function that records the audio source sung by the user and saves the audio data as a digital file.
[1144] The "analysis means" is a function for analyzing recorded sound source data, analyzing pitch and interval, and identifying the optimum key.
[1145] The "display means" is a function that visually presents the analysis results obtained by the analysis means to the user.
[1146] The "communication means" is a function that allows the analysis means to send the analysis results to the cloud server and receive the analysis results.
[1147] A "music analysis API" is an application programming interface for analyzing sound source data and analyzing data such as pitch and interval.
[1148] A "cloud server" refers to a computer server that stores and processes data remotely over a network.
[1149] A "smartphone" is a portable information terminal that can execute multifunctional applications in addition to making calls.
[1150] In this embodiment, a system will be described in which a user can use a smartphone application to enjoy karaoke in the optimal key.
[1151] Hardware and Software Used
[1152] Smartphone: Used as a device on which users can select songs and sing a cappella.
[1153] Cloud Server: The server used to analyze the a cappella recordings.
[1154] Music Analysis API: An API that runs on a cloud server and analyzes audio data to identify the optimal key.
[1155] User interface: The interface on the smartphone application that allows users to select songs and display the results.
[1156] System Operation Details
[1157] When a user launches the smartphone application, they are first presented with an interface for selecting a song. At this stage, the user selects the song they want to sing from a list on the application. The application then prompts the user to sing a little a cappella, and records the sound using the smartphone's microphone.
[1158] Once recording is complete, the application sends the recorded audio data to a cloud server. The cloud server receives the recorded data and sends a request to the music analysis API for analysis. The music analysis API analyzes the audio data and generates key information that allows the user to sing in the optimal key.
[1159] After receiving the analysis results, the cloud server sends them to the smartphone. Based on the analysis results, the smartphone application displays a notification to the user, such as "The optimal key is G." This allows the user to enjoy singing in the key that is best for them.
[1160] Specific examples
[1161] For example, suppose a user selects "Let it be" and sings a cappella "When I find myself in times of trouble...". This audio is recorded on a smartphone and sent to a cloud server. The server uses a music analysis API to analyze the audio data and obtains the result that "the optimal key is D". The result is sent to the smartphone, which displays the message "The optimal key is D". Based on the displayed key information, the user can sing "Let it be" in the key of D.
[1162] Prompt Sentence Examples
[1163] For example, a possible prompt might be, "Please implement an API that analyzes recorded audio data and finds the optimal key."
[1164] In this way, the user can easily enjoy karaoke in the key that is most suitable for them, without requiring any specialized knowledge.
[1165] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1166] Step 1:
[1167] Song selection method
[1168] The user launches the smartphone application and selects a song.
[1169] Input: Song name specified by the user in the app
[1170] Output: Selected song data
[1171] Processing operation: A song list is displayed through the smartphone's user interface, and the user selects the song they want to sing. At this time, the application saves the data of the selected song in the internal memory.
[1172] Step 2:
[1173] Recording medium
[1174] After the user selects a song, they are prompted to sing a cappella, and when they start singing, their voice is recorded by the smartphone microphone.
[1175] Input: User's a cappella singing voice
[1176] Output: Recorded a cappella audio data (digital file)
[1177] Processing operation: When recording starts, the user's singing voice is collected through the microphone and a certain period of audio data is saved as a digital file.
[1178] Step 3:
[1179] communication means
[1180] Once recording is complete, the audio data is sent to a cloud server.
[1181] Input: Recorded a cappella audio data
[1182] Output: Audio data sent to the cloud server
[1183] Processing operation: The smartphone sends the recorded audio data to the cloud server via an HTTP request. Once the transmission is complete, the server confirms receipt of the data.
[1184] Step 4:
[1185] Analysis means
[1186] The cloud server sends the received audio data to the music analysis API and requests analysis.
[1187] Input: Received a cappella audio data
[1188] Output: Analysis results from the music analysis API (optimal key information)
[1189] Processing: The server sends the audio data to the music analysis API and makes an analysis request. The API analyzes the audio data and returns the optimal key information to the cloud server.
[1190] Step 5:
[1191] communication means
[1192] The cloud server sends the analysis results received from the music analysis API to the smartphone.
[1193] Input: Analysis results (optimal key information)
[1194] Output: Analysis results sent to your smartphone
[1195] Processing operation: After obtaining the analysis results, the cloud server returns them to the smartphone application as an HTTP response.
[1196] Step 6:
[1197] Display means
[1198] The smartphone then displays the analysis results to the user, such as "The optimal key is G."
[1199] Input: Analysis results received from the cloud server
[1200] Output: Key information that is best visible to the user
[1201] Processing operation: The smartphone displays the appropriate key information on the interface based on the analysis results received.
[1202] In this way, through a series of processing steps, the user can enjoy karaoke in the key that is most suitable for him or her.
[1203] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1204] To implement this invention, it is necessary to incorporate an emotion engine into the karaoke system and integrate the following series of functions.
[1205] 1. Song selection method
[1206] The user operates the interface on the screen of the terminal to select the song they want to sing. This song selection means includes a function for displaying a song list and accepting user input.
[1207] 2. Recording Method
[1208] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[1209] 3. Analysis method
[1210] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[1211] 4. Emotion Engine
[1212] The recorded voice data is passed through an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice data to determine the user's emotional state (e.g., joy, sadness, anger, etc.) and adjusts the analysis results accordingly.
[1213] 5. Display means
[1214] Based on the results of the analysis means and the recognition results of the emotion engine, the device displays the optimal key and the appropriate performance and advice for the emotion. This information is displayed on the screen of the device, providing visual feedback to the user.
[1215] A natural language description of the program's operation
[1216] System Overview
[1217] This system analyzes the user's singing voice and emotions, and suggests the optimal key and emotion, making karaoke more enjoyable and comfortable.
[1218] 1. Terminal Processing
[1219] The user selects the song they want to sing on the device.
[1220] The device displays a message inviting the user to sing a little a cappella.
[1221] When the user starts singing, the device records the voice through the microphone.
[1222] Once the recording is complete, the audio data is sent to the server.
[1223] 2. Server Processing
[1224] The server receives the voice data transmitted from the terminal.
[1225] To analyze the audio data, a request is sent to the music analysis API.
[1226] An emotion engine is used to recognize user emotions from voice data.
[1227] The analysis results from the music analysis API are combined with the results of the emotion engine to identify the optimal key and direction.
[1228] 3. Terminal Processing (cont.)
[1229] The analysis results received from the server and the emotion engine results are displayed on the screen.
[1230] The user checks the displayed key information and advice according to their emotions.
[1231] Specific examples
[1232] Example: A user selects "Let it be (The Beatles)" and optimizes it for D key, and the emotion engine recognizes "Joy."
[1233] 1. Song selection and singing
[1234] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[1235] The device displays the message, "Please sing a little a cappella."
[1236] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[1237] 2. Data submission and analysis
[1238] The recording data is sent to the server.
[1239] The server receives the audio data and sends it to the music analysis API for analysis.
[1240] At the same time, the server uses an emotion engine to analyze emotions and recognize "joy."
[1241] 3. Displaying the results and singing
[1242] The analysis results from the server ("The optimal key is D") and the emotion engine results ("Joy") are sent to the terminal.
[1243] The device displays on the screen, "The optimal key is D. Also, your emotion is joy, so sing with energy."
[1244] Based on the displayed key information and advice, the user can enjoy singing "Let it be" in the key of D.
[1245] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[1246] The processing flow will be explained below.
[1247] Step 1: Song Selection
[1248] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[1249] The terminal obtains information about the selected song and displays it on the screen.
[1250] Step 2: A cappella instructions
[1251] The device displays a message to the user saying, "Please sing a little a cappella."
[1252] The user checks the message.
[1253] Step 3: Start recording
[1254] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[1255] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[1256] Step 4: End recording
[1257] When the user finishes singing, the device will automatically stop recording.
[1258] The recorded audio data is temporarily stored in memory.
[1259] Step 5: Send data
[1260] The terminal transmits the recorded audio data to the server.
[1261] The server receives the HTTP request and stores the audio file.
[1262] Step 6: Data analysis
[1263] The server generates a request to send the stored audio data to the music analysis API.
[1264] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[1265] Step 7: Sentiment Analysis
[1266] The server uses an emotion engine to recognize the user's emotion from the stored audio data.
[1267] The emotion engine analyzes the voice data to determine the user's emotional state (e.g., happy, sad, angry, etc.).
[1268] Step 8: Obtaining analysis results
[1269] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[1270] The server combines the results of the music analysis and the emotion engine to generate advice based on the optimal key and emotion.
[1271] Step 9: Send results
[1272] The server organizes the analysis results and transmits them to the user's terminal.
[1273] As an HTTP response, the most appropriate key information and advice based on the emotion are returned to the device.
[1274] Step 10: View the results
[1275] The device then displays the analysis results on its screen, displaying the message, "The optimal key is D. Your emotion is joy, so sing with enthusiasm."
[1276] The user confirms this message.
[1277] Step 11: Adjust the key
[1278] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[1279] The device will retain the new settings after adjusting the keys.
[1280] Step 12: Karaoke playback
[1281] After the user completes the key setting, karaoke playback begins.
[1282] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[1283] Step 13: Enjoy singing
[1284] Users can sing songs in the optimal key and enjoy karaoke based on advice based on their emotions.
[1285] Example 2
[1286] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1287] Conventional karaoke systems can suggest the optimal key for a user's singing, but they cannot provide performances or advice that take the user's emotions into consideration. As a result, users cannot receive accurate advice or performances that reflect their emotions, limiting the quality of their karaoke experience. It is also difficult for users to enjoy a performance that matches their own emotions while singing. There is a need to solve this problem and provide a karaoke experience that is in tune with the user's emotions.
[1288] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1289] In this invention, the server includes an analysis means for analyzing the recorded sound source data, an emotion analysis means for recognizing the user's emotion from the recorded sound source data, and a display means for displaying the optimum key for the user's singing and performance and advice suited to the emotion based on the results of the analysis by the analysis means and the emotion analysis means. This makes it possible to provide not only the optimum key for the user's singing, but also appropriate advice and performance according to the emotion.
[1290] The "selection means" is a function that allows the user to select a song to sing through the interface of the karaoke system.
[1291] "Recording means" is a function for recording the user's singing through a microphone and saving it as digital audio data.
[1292] The "analysis means" is a function that uses recorded audio data to analyze pitch and interval and identify the optimal key for the user's singing.
[1293] The "emotion analysis means" is a function for recognizing and analyzing the user's emotions from the recorded audio data.
[1294] The "display means" is a function for visually presenting the analysis results obtained by the analysis means and emotion analysis means to the user, and for providing advice and presentations that correspond to the optimal key and emotion.
[1295] MODE FOR CARRYING OUT THE INVENTION
[1296] This invention improves the karaoke experience by analyzing the user's singing data and emotional data in a karaoke system and suggesting the optimal key and song based on the user's emotion. A specific method for implementing this invention will be described below.
[1297] Required Hardware and Software
[1298] To implement this system, a terminal and a server are required. The terminal has an interface for users to select songs and sing, and the server has the function of analyzing voice data and emotional data.
[1299] Terminal
[1300] Touchscreen display: The user operates this to select songs.
[1301] Microphone: Used to record the user's singing.
[1302] Internet connection: Required to send data to the server.
[1303] server
[1304] Music analysis API: Used to analyze audio data and identify the optimal key (e.g., Spotify API).
[1305] Sentiment analysis engine: Used to identify user emotions from voice data (e.g., IBM Watson).
[1306] Database: Used to store analysis results and user information.
[1307] System configuration
[1308] 1. Song selection method
[1309] The user selects the song they want to sing from the song list displayed on the terminal.
[1310] 2. Recording Method
[1311] After the user selects a song, the device displays a message saying, "Please sing a little a cappella," encouraging the user to sing a cappella.
[1312] When the user starts singing, the device records the voice through the microphone and stores it as digital data.
[1313] 3. Analysis method
[1314] The device sends the recorded audio data to a server using an internet connection.
[1315] The server uses a music analysis API to analyze the audio data and identify pitch and interval.
[1316] At the same time, the server uses an emotion analysis engine to recognize the user's emotions from the voice data.
[1317] 4. Display means
[1318] The server transmits the results obtained from the analysis means and the emotion analysis means to the terminal.
[1319] The terminal displays this information on the screen and provides the user with the most suitable key and appropriate performance and advice for their emotion.
[1320] Specific examples
[1321] If a user uses a karaoke system to select "Let it be (The Beatles)" and sing a bit a cappella, the system performs the following steps:
[1322] 1. Song Selection
[1323] The user selects "Let it be (The Beatles)" on the device's touchscreen.
[1324] 2. Singing and Recording
[1325] The device displays the message "Please sing a little a cappella," and the user sings "When I find myself in times of trouble..."
[1326] The user's singing is recorded and saved as digital data.
[1327] 3. Sending and analyzing audio data
[1328] The recording data is sent from the device to the server, which uses a music analysis API to identify the optimal key (D).
[1329] At the same time, an emotion analysis engine is used to recognize emotions (happiness).
[1330] 4. Displaying results and giving advice
[1331] The analysis results from the server (optimal key: D) and emotion analysis results (joy) are displayed on the device.
[1332] The device displays the message, "The optimal key is D. Also, your emotion is joy, so sing with enthusiasm."
[1333] Prompt Sentence Examples
[1334] "In a karaoke system, a user selects The Beatles' "Let it Be" and sings a portion a cappella. This singing voice is recorded, and the system analyzes it using an emotion engine. The system recognizes the user's emotion as "joy." The system also determines that the key that best suits the singing voice is D. Based on these results, what kind of feedback should be provided to the user?"
[1335] In this way, the present invention can provide a multi-functional karaoke system that not only allows the user to enjoy karaoke in the key that is best suited to them, but also allows them to receive advice according to their emotions at the time.
[1336] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1337] Specific explanation of processing steps
[1338] Step 1: Select a song
[1339] Terminal handling
[1340] Input: The user operates the touchscreen display to select a song from the displayed song list.
[1341] Operation:
[1342] The terminal receives touchscreen input and retrieves song information based on the user's selection.
[1343] Output: The title and artist name of the selected song will be displayed on the screen.
[1344] Step 2: Sing and record your a cappella
[1345] Terminal handling
[1346] Input: After a song is selected, the user begins singing.
[1347] Operation:
[1348] The device displays the message, "Please sing a little a cappella."
[1349] When the user starts singing, the voice is recorded through a microphone.
[1350] The recorded audio data is saved as a digital file.
[1351] Output: The recording data is saved on the device.
[1352] Step 3: Sending audio data to the server
[1353] Terminal handling
[1354] Input: Recorded data is saved.
[1355] Operation:
[1356] The device sends the recorded data to a server via the Internet.
[1357] The HTTPS protocol is used for data transmission.
[1358] Output: The recording data is sent to the server.
[1359] Step 4: Analyzing the audio data
[1360] Server Processing
[1361] Input: Recording data sent from the device.
[1362] Operation:
[1363] The server receives the recording data.
[1364] Sends a request to an audio analysis API (e.g., Spotify API) to analyze the pitch and interval of the audio data.
[1365] The voice analysis API returns the analysis results to the server.
[1366] Output: Audio analysis results (optimal key, etc.).
[1367] Step 5: Sentiment Analysis
[1368] Server Processing
[1369] Input: Recording data sent from the device.
[1370] Operation:
[1371] The server uses an emotion analysis engine (e.g., IBM Watson) to analyze the user's emotions.
[1372] The emotion analysis engine identifies emotions based on the recorded data and sends the results to the server.
[1373] Output: Sentiment analysis results (e.g., joy, sadness, anger, etc.).
[1374] Step 6: Combining the analysis results
[1375] Server Processing
[1376] Input: Voice analysis results and emotion analysis results.
[1377] Operation:
[1378] The server combines the results of the voice analysis and the emotion analysis.
[1379] Identify the best key and emotionally appropriate performance and advice.
[1380] Output: Final analysis results (optimal key and emotional advice).
[1381] Step 7: Viewing the analysis results
[1382] Terminal handling
[1383] Input: Final analysis result data from the server.
[1384] Operation:
[1385] The terminal displays the analysis results received from the server on the screen.
[1386] The user sees the best key and emotion-based advice.
[1387] Output: The screen displays "The best key is D. Your emotion is joy, so sing with enthusiasm."
[1388] In this way, the system analyzes the user's singing data at multiple stages and provides optimal feedback to the user, enhancing the karaoke experience.
[1389] (Application example 2)
[1390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1391] Conventional karaoke systems have the problem of being unable to find the optimal key for the user when singing and providing a performance that reflects the user's emotions at that time. In particular, they are unable to provide feedback based on the user's emotions, preventing users from having a more enjoyable and personalized karaoke experience. This leaves karaoke booths and live music bars with a challenge to improve customer satisfaction.
[1392] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1393] In this invention, the server includes a selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded sound data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, an emotion engine for recognizing emotions from the user's voice data, and a means for providing the user with visual feedback based on the recognition results of the emotion engine. This not only enables the user to sing in the optimal key for themselves, but also allows them to receive advice based on their emotions at the time.
[1394] A "selection means" is a means that provides an interface for a user to select a song and sing along a bit a cappella with it.
[1395] "Recording means" means any means, including devices and software, for recording the audio of a user's singing.
[1396] The "analysis means" is a means for analyzing the recorded sound source data and performing a process to identify the appropriate key and other musical elements for the song.
[1397] The "display means" is a means for visually presenting the optimum key and other information to the user based on the analysis results obtained by the analysis means.
[1398] The "emotion engine" is a means including software and algorithms for recognizing emotions from the user's voice data and feeding back the recognition results to the analysis means.
[1399] The "means for providing visual feedback" is a means for visually displaying appropriate advice or dramatic effects to the user based on the recognition results of the emotion engine.
[1400] The present invention provides a system for enabling users to select songs and improve their singing experience in a physical venue such as a karaoke booth or live music bar, etc. The system includes a selection means, a recording means, an analysis means, a display means, an emotion engine, and a means for providing visual feedback.
[1401] System Configuration
[1402] 1. The selection means is a device such as a tablet or smartphone. The user operates the interface on the screen to select the song they want to sing. This selection means includes a function to display a song list and accept user input.
[1403] 2. The recording function provides a function to record the user's a cappella singing using a microphone. The recording method collects a certain amount of audio data and saves it as a digital file. For example, a microphone built into a smartphone or a microphone installed in a karaoke booth can be used.
[1404] 3. The analysis method involves analyzing the audio data on the server using a music analysis API (such as IBM Watson or Google Cloud Speech-to-Text API). This analysis involves analyzing pitch and interval and comparing it with existing data. Based on the analysis results, the optimal key for the user's singing is identified.
[1405] 4. The emotion engine uses software and algorithms to recognize emotions from the user's voice data. For example, it uses Microsoft Azure's Emotion API or a proprietary emotion recognition algorithm. This analyzes the user's emotional state (e.g., joy, sadness, anger, etc.).
[1406] 5. As a display method, visual feedback is provided to the user based on the analysis results and the emotion engine's recognition results. For example, optimal keys and advice based on emotions are displayed on the screen of a tablet or smartphone.
[1407] Example
[1408] Specific examples
[1409] Below is a case where the user selects the song "Let it be (The Beatles)" and analysis and emotion recognition are performed based on that.
[1410] 1. Song selection and singing:
[1411] The user selects "Let it be (The Beatles)" on the karaoke machine.
[1412] The device displays the message, "Please sing a little a cappella."
[1413] The user sings a bit of "When I find myself in times of trouble..." and their voice is recorded.
[1414] 2. Data Analysis and Emotion Recognition:
[1415] The recorded data is sent to the server and analyzed using a music analysis API.
[1416] At the same time, the emotion engine recognizes the user's emotion as "joy."
[1417] 3. Displaying the results:
[1418] The analysis results and emotion recognition results are sent to the device, and the message "The optimal key is D. Your emotion is joy, so sing with enthusiasm." is displayed.
[1419] Using the displayed key information and advice, users can enjoy singing "Let it be" in the key of D.
[1420] Prompt Sentence Examples
[1421] "User selects 'Let it be'. The recording is analyzed to determine the best key. The emotion engine recognizes 'joy' and provides advice and the best key."
[1422] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[1423] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1424] Step 1:
[1425] The user selects a song using the selection means on a tablet or smartphone. The input is the ID and title of the song selected by the user, which is provided to the system. The device displays a list of songs, accepts the user's selection, and obtains song information.
[1426] Step 2:
[1427] After the user selects a song, the device prompts the user to sing a little a cappella using the microphone. The input is the user's singing voice, and the output is the recorded audio data. When the user starts singing, the device captures the audio signal and stores it digitally as audio data.
[1428] Step 3:
[1429] The device sends the recorded audio data to the server. The input is the audio data file, and the output is the data received on the server. The server sends the audio data to the specified endpoint and confirms receipt.
[1430] Step 4:
[1431] The server sends the received audio data to the music analysis API and requests analysis. The input is audio data, and the output is analyzed musical element data such as interval and pitch. The server sends a request to the API endpoint and receives the analysis results in return.
[1432] Step 5:
[1433] The server then uses an emotion engine to recognize the user's emotion from the voice data. The input is the voice data and the output is the user's emotional state (e.g., "joy"). The emotion engine analyzes the voice data and generates extracted emotion information.
[1434] Step 6:
[1435] The server integrates the analysis results from the music analysis API and the recognition results from the emotion engine to identify the optimal key and advice according to the emotion. The input is musical element data and emotion information, and the output is the optimal key and advice. The server processes these data and generates the integrated result.
[1436] Step 7:
[1437] The server sends the integration result to the terminal and provides feedback to the user using a display means. The input is the integration result data, and the output is the optimal key and advice according to the emotion displayed on the user terminal. The terminal displays the received information on the screen and provides visual feedback to the user.
[1438] Step 8:
[1439] The user sings the song using the displayed optimal key and emotional advice. The input is the displayed information, and the output is the user's improved singing performance. Based on the feedback, the user can enjoy singing in the optimal key.
[1440] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1441] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1442] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1443] [Fourth embodiment]
[1444] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1445] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1446] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1447] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1448] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1450] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1451] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1452] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1453] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1454] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1455] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1456] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1457] To implement this invention, the following functions must be integrated into the karaoke system:
[1458] 1. Song selection method
[1459] The user selects the song he or she wants to sing by operating the interface on the terminal screen. This song selection means includes a function for displaying a song list and accepting user input.
[1460] 2. Recording Method
[1461] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[1462] 3. Analysis method
[1463] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[1464] 4. Display means
[1465] Based on the analysis results, the server derives the optimal key and sends the result to the terminal. The terminal's display means visually presents the analysis results to the user. For example, a notification such as "The optimal key is G" is displayed on the screen.
[1466] A natural language description of the program's operation
[1467] System Overview
[1468] The system provides functionality to allow users to sing in the optimal key for karaoke.
[1469] 1. Terminal Processing
[1470] The user selects the song they want to sing on the device.
[1471] The device displays a message inviting the user to sing a little a cappella.
[1472] When the user starts singing, the device records the voice through the microphone.
[1473] Once the recording is complete, the audio data is sent to the server.
[1474] 2. Server Processing
[1475] The server receives the voice data transmitted from the terminal.
[1476] To analyze the audio data, a request is sent to the music analysis API.
[1477] Receive analysis results containing relevant key information.
[1478] The analysis results are sent to the device.
[1479] 3. Terminal Processing (cont.)
[1480] The terminal displays the analysis results received from the server on the screen.
[1481] The user adjusts the key based on the displayed key information and prepares to sing in the optimum key.
[1482] Specific examples
[1483] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[1484] 1. Song selection and singing
[1485] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[1486] The device displays the message, "Please sing a little a cappella."
[1487] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[1488] 2. Data submission and analysis
[1489] The recording data is sent to the server.
[1490] The server receives the audio data and sends it to the music analysis API for analysis.
[1491] The music analysis API returns the result, "The optimal key is D."
[1492] 3. Displaying the results and singing
[1493] The terminal receives the analysis results from the server.
[1494] The device will display "The best key is D" on the screen.
[1495] Based on the displayed key information, the user sings "Let it be" in the key of D.
[1496] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[1497] The processing flow will be explained below.
[1498] Step 1: Song Selection
[1499] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[1500] The terminal obtains information about the selected song and displays it on the screen.
[1501] Step 2: A cappella instructions
[1502] The device displays a message to the user saying, "Please sing a little a cappella."
[1503] The user checks the message.
[1504] Step 3: Start recording
[1505] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[1506] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[1507] Step 4: End recording
[1508] When the user finishes singing, the device will automatically stop recording.
[1509] The recorded audio data is temporarily stored in memory.
[1510] Step 5: Send data
[1511] The terminal transmits the recorded audio data to the server.
[1512] The server receives the HTTP request and stores the audio file.
[1513] Step 6: Data analysis
[1514] The server generates a request to send the stored audio data to the music analysis API.
[1515] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[1516] Step 7: Obtaining analysis results
[1517] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[1518] The server stores the analysis results in a database or temporary memory.
[1519] Step 8: Send results
[1520] The server organizes the analysis results and transmits them to the user's terminal.
[1521] The optimal key information (e.g., "D key") is returned to the terminal as an HTTP response.
[1522] Step 9: View the results
[1523] The device will display the analysis results on the screen, with a message such as "The best key is D."
[1524] The user confirms this message.
[1525] Step 10: Adjust the key
[1526] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[1527] The device will retain the new settings after adjusting the keys.
[1528] Step 11: Karaoke playback
[1529] After the user completes the key setting, karaoke playback begins.
[1530] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[1531] Step 12: Enjoy singing
[1532] The user can sing songs in the optimum key and enjoy karaoke.
[1533] Example 1
[1534] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1535] Current karaoke systems require users to manually try out different keys to find the one that suits them best. This process is time-consuming and difficult, especially for users who are not familiar with music. A solution to this problem is needed to enable users to easily sing songs in the optimal key.
[1536] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1537] In this invention, the server includes a selection means for the user to select a song, a display means for instructing the user to sing a little a cappella, a recording means for recording the user's singing, a transmission means for transmitting the recorded sound data to the server, an analysis means for the server to analyze the sound data using a music analysis API, and a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means. This frees the user from the hassle of manually adjusting the key and allows them to easily enjoy karaoke in the optimal key for themselves.
[1538] The "selection means" is a function that allows the user to select a song to sing using the interface of the karaoke system.
[1539] The "means for displaying instructions" is a function that displays a message encouraging the user to sing a little a cappella.
[1540] The "recording means" is a function that records the user's singing voice through a microphone and serves to save the recorded audio data.
[1541] The "transmission means" is a function for converting the recorded sound source data into packets and transmitting them to the server.
[1542] The "analysis means" is a function that allows the server to analyze the received audio data using a music analysis API and identify the optimal key for the user.
[1543] The "display means" is a function for visually displaying to the user the optimum key information obtained by the analysis means.
[1544] To implement this invention, the following functions must be integrated into the karaoke system:
[1545] Selection method
[1546] On the device, the user operates the interface on the screen to select the song they want to sing. This selection means includes the function of displaying a song list and accepting user input. When the user selects a song, the device stores the selection information in its internal memory. For example, the user uses this selection means to select "Let it be (The Beatles)."
[1547] A means of displaying instructions
[1548] When the user selects a song, the device provides an interface that displays a message to the user saying, "Please sing a little a cappella." Specifically, a text message is displayed on the device screen.
[1549] Recording medium
[1550] When the user starts singing, the device will record the sound using the built-in microphone. This recording method will collect a certain amount of sound data (for example, 10 seconds) and save it as a digital file. The sound data will be saved in WAV or MP3 format.
[1551] Transmission method
[1552] Once the recording is complete, the device transmits the recorded audio data to the server, which encodes the audio data into an appropriate format and transmits it to the server via a network using a secure protocol (e.g., HTTPS).
[1553] Analysis means
[1554] The server analyzes the audio data received from the device. For the analysis, a music analysis API (such as Google Cloud Speech-to-Text API or IBM Watson) is used. The server sends an analysis request to the API and receives the analysis results of the audio data. For example, key information such as "The optimal key is D" is returned.
[1555] Display means
[1556] The server sends the analysis results to the device, which then visually presents the results to the user. This display includes a function that displays text such as "The best key is D" on the screen. Based on this information, the user can prepare to sing the song in the best key.
[1557] Specific examples
[1558] Example: User selects "Let it be (The Beatles)" and optimizes for D key
[1559] 1. Song selection and singing
[1560] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[1561] The device displays the message, "Please sing a little a cappella."
[1562] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[1563] 2. Data submission and analysis
[1564] The recording data is sent to the server.
[1565] The server receives the audio data and sends it to the music analysis API for analysis.
[1566] The music analysis API returns the result, "The optimal key is D."
[1567] 3. Displaying the results and singing
[1568] The terminal receives the analysis results from the server.
[1569] The device screen will display "The best key is D."
[1570] Based on the displayed key information, the user sings "Let it be" in the key of D.
[1571] Prompt Sentence Examples
[1572] A typical example of a prompt to be input to a generative AI model is as follows:
[1573] plain
[1574] Please provide a detailed explanation of the steps a user takes to select "Let it be (The Beatles)" and optimize it for D. Please provide a detailed explanation of each step of the process, from the user selecting the song, recording the a cappella vocal, sending the audio data to the server, and receiving and displaying the analysis results.
[1575] In this way, the present invention allows users to enjoy karaoke in the key that is most suitable for them.
[1576] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1577] Step 1:
[1578] The user selects a song.
[1579] Specific operation: The user selects the song they want to sing from the song list provided on the karaoke terminal screen.
[1580] Input: The user touches or clicks on a specific song from the song list displayed on the screen.
[1581] Output: The information of the selected song is saved on the device. For example, if "Let it be (The Beatles)" is selected, the information is saved in the internal memory.
[1582] Step 2:
[1583] The device prompts the user to sing a little a cappella.
[1584] Specific operation: The device displays a message on the screen and instructs the user to "sing a little a cappella."
[1585] Input: After the user selects a song, instructions are displayed on the device screen.
[1586] Output: A visual instruction is given to the user, and the user is prepared to follow the instruction.
[1587] Step 3:
[1588] The device records the user singing a cappella.
[1589] Specific operation: When the user starts singing, the device will record the voice using the built-in microphone.
[1590] Input: The user's voice is input through the microphone.
[1591] Output: The recorded audio data is saved in a digital format (e.g., WAV or MP3). For example, the phrase "When I find myself in times of trouble" is recorded.
[1592] Step 4:
[1593] The device sends the recorded data to the server.
[1594] Specific operation: After the recording is completed, the device encodes the audio data into an appropriate format and sends it to the server over the network.
[1595] Input: Recorded audio data.
[1596] Output: The encoded audio data is split into packets and sent to the server using a secure protocol.
[1597] Step 5:
[1598] The server analyzes the audio data.
[1599] Specific operation: The server sends the received audio data to the music analysis API and analyzes the optimal key.
[1600] Input: The audio data received by the server.
[1601] Output: Analysis results from the music analysis API, e.g. "The optimal key is D."
[1602] Step 6:
[1603] The analysis results are sent from the server to the device.
[1604] Specific operation: The server converts the analysis results into a structured data format (e.g., JSON) and sends them back to the terminal using a secure protocol.
[1605] Input: Analysis result data.
[1606] Output: The analysis results are sent to the terminal in structured data format.
[1607] Step 7:
[1608] The device displays the analysis results to the user.
[1609] Specific operation: The device decodes the received analysis results and displays "The best key is D" on the screen.
[1610] Input: Analysis results received from the server.
[1611] Output: A visual display to the user, where the user obtains the best key information. For example, "The best key is D" is displayed on the terminal screen.
[1612] (Application example 1)
[1613] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1614] Conventional karaoke systems require users to manually adjust the key through trial and error to find the optimal key for them, which is time-consuming and laborious. Furthermore, they lack a function to identify the optimal key based on the user's singing style, making it difficult for users without specialized musical knowledge. Furthermore, they lack a system that utilizes cloud technology to quickly provide analysis results.
[1615] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1616] In this invention, the server includes a song selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded audio data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, and a communication means for the analysis means to send the analysis results to the cloud server based on a music analysis API and receive the analysis results. This allows a user to easily select a song through a smartphone application, send the singing voice to the cloud server for analysis, and quickly find the optimal key for themselves.
[1617] The "music selection means" is a function that allows the user to operate the interface to select music.
[1618] The "recording means" is a function that records the audio source sung by the user and saves the audio data as a digital file.
[1619] The "analysis means" is a function for analyzing recorded sound source data, analyzing pitch and interval, and identifying the optimum key.
[1620] The "display means" is a function that visually presents the analysis results obtained by the analysis means to the user.
[1621] The "communication means" is a function that allows the analysis means to send the analysis results to the cloud server and receive the analysis results.
[1622] A "music analysis API" is an application programming interface for analyzing sound source data and analyzing data such as pitch and interval.
[1623] A "cloud server" refers to a computer server that stores and processes data remotely over a network.
[1624] A "smartphone" is a portable information terminal that can execute multifunctional applications in addition to making calls.
[1625] In this embodiment, a system will be described in which a user can use a smartphone application to enjoy karaoke in the optimal key.
[1626] Hardware and Software Used
[1627] Smartphone: Used as a device on which users can select songs and sing a cappella.
[1628] Cloud Server: The server used to analyze the a cappella recordings.
[1629] Music Analysis API: An API that runs on a cloud server and analyzes audio data to identify the optimal key.
[1630] User interface: The interface on the smartphone application that allows users to select songs and display the results.
[1631] System Operation Details
[1632] When a user launches the smartphone application, they are first presented with an interface for selecting a song. At this stage, the user selects the song they want to sing from a list on the application. The application then prompts the user to sing a little a cappella, and records the sound using the smartphone's microphone.
[1633] Once recording is complete, the application sends the recorded audio data to a cloud server. The cloud server receives the recorded data and sends a request to the music analysis API for analysis. The music analysis API analyzes the audio data and generates key information that allows the user to sing in the optimal key.
[1634] After receiving the analysis results, the cloud server sends them to the smartphone. Based on the analysis results, the smartphone application displays a notification to the user, such as "The optimal key is G." This allows the user to enjoy singing in the key that is best for them.
[1635] Specific examples
[1636] For example, suppose a user selects "Let it be" and sings a cappella "When I find myself in times of trouble...". This audio is recorded on a smartphone and sent to a cloud server. The server uses a music analysis API to analyze the audio data and obtains the result that "the optimal key is D". The result is sent to the smartphone, which displays the message "The optimal key is D". Based on the displayed key information, the user can sing "Let it be" in the key of D.
[1637] Prompt Sentence Examples
[1638] For example, a possible prompt might be, "Please implement an API that analyzes recorded audio data and finds the optimal key."
[1639] In this way, the user can easily enjoy karaoke in the key that is most suitable for them, without requiring any specialized knowledge.
[1640] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1641] Step 1:
[1642] Song selection method
[1643] The user launches the smartphone application and selects a song.
[1644] Input: Song name specified by the user in the app
[1645] Output: Selected song data
[1646] Processing operation: A song list is displayed through the smartphone's user interface, and the user selects the song they want to sing. At this time, the application saves the data of the selected song in the internal memory.
[1647] Step 2:
[1648] Recording medium
[1649] After the user selects a song, they are prompted to sing a cappella, and when they start singing, their voice is recorded by the smartphone microphone.
[1650] Input: User's a cappella singing voice
[1651] Output: Recorded a cappella audio data (digital file)
[1652] Processing operation: When recording starts, the user's singing voice is collected through the microphone and a certain period of audio data is saved as a digital file.
[1653] Step 3:
[1654] communication means
[1655] Once recording is complete, the audio data is sent to a cloud server.
[1656] Input: Recorded a cappella audio data
[1657] Output: Audio data sent to the cloud server
[1658] Processing operation: The smartphone sends the recorded audio data to the cloud server via an HTTP request. Once the transmission is complete, the server confirms receipt of the data.
[1659] Step 4:
[1660] Analysis means
[1661] The cloud server sends the received audio data to the music analysis API and requests analysis.
[1662] Input: Received a cappella audio data
[1663] Output: Analysis results from the music analysis API (optimal key information)
[1664] Processing: The server sends the audio data to the music analysis API and makes an analysis request. The API analyzes the audio data and returns the optimal key information to the cloud server.
[1665] Step 5:
[1666] communication means
[1667] The cloud server sends the analysis results received from the music analysis API to the smartphone.
[1668] Input: Analysis results (optimal key information)
[1669] Output: Analysis results sent to your smartphone
[1670] Processing operation: After obtaining the analysis results, the cloud server returns them to the smartphone application as an HTTP response.
[1671] Step 6:
[1672] Display means
[1673] The smartphone then displays the analysis results to the user, such as "The optimal key is G."
[1674] Input: Analysis results received from the cloud server
[1675] Output: Key information that is best visible to the user
[1676] Processing operation: The smartphone displays the appropriate key information on the interface based on the analysis results received.
[1677] In this way, through a series of processing steps, the user can enjoy karaoke in the key that is most suitable for him or her.
[1678] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1679] To implement this invention, it is necessary to incorporate an emotion engine into the karaoke system and integrate the following series of functions.
[1680] 1. Song selection method
[1681] The user operates the interface on the screen of the terminal to select the song they want to sing. This song selection means includes a function for displaying a song list and accepting user input.
[1682] 2. Recording Method
[1683] After the user selects a song, the device provides a function to record the user's a cappella singing using a microphone. The recording means collects a certain period of audio data and saves it as a digital file.
[1684] 3. Analysis method
[1685] The server receives the recorded audio data and analyzes it using a music analysis API. The analysis means includes processing functions to analyze the audio data and identify the optimal key for the user. This analysis includes analyzing pitch and interval and comparing it with existing data.
[1686] 4. Emotion Engine
[1687] The recorded voice data is passed through an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice data to determine the user's emotional state (e.g., joy, sadness, anger, etc.) and adjusts the analysis results accordingly.
[1688] 5. Display means
[1689] Based on the results of the analysis means and the recognition results of the emotion engine, the device displays the optimal key and the appropriate performance and advice for the emotion. This information is displayed on the screen of the device, providing visual feedback to the user.
[1690] A natural language description of the program's operation
[1691] System Overview
[1692] This system analyzes the user's singing voice and emotions, and suggests the optimal key and emotion, making karaoke more enjoyable and comfortable.
[1693] 1. Terminal Processing
[1694] The user selects the song they want to sing on the device.
[1695] The device displays a message inviting the user to sing a little a cappella.
[1696] When the user starts singing, the device records the voice through the microphone.
[1697] Once the recording is complete, the audio data is sent to the server.
[1698] 2. Server Processing
[1699] The server receives the voice data transmitted from the terminal.
[1700] To analyze the audio data, a request is sent to the music analysis API.
[1701] An emotion engine is used to recognize user emotions from voice data.
[1702] The analysis results from the music analysis API are combined with the results of the emotion engine to identify the optimal key and direction.
[1703] 3. Terminal Processing (cont.)
[1704] The analysis results received from the server and the emotion engine results are displayed on the screen.
[1705] The user checks the displayed key information and advice according to their emotions.
[1706] Specific examples
[1707] Example: A user selects "Let it be (The Beatles)" and optimizes it for D key, and the emotion engine recognizes "Joy."
[1708] 1. Song selection and singing
[1709] The user selects "Let it be (The Beatles)" on the karaoke terminal.
[1710] The device displays the message, "Please sing a little a cappella."
[1711] The user sings a bit of "When I find myself in times of trouble..." and the audio is recorded.
[1712] 2. Data submission and analysis
[1713] The recording data is sent to the server.
[1714] The server receives the audio data and sends it to the music analysis API for analysis.
[1715] At the same time, the server uses an emotion engine to analyze emotions and recognize "joy."
[1716] 3. Displaying the results and singing
[1717] The analysis results from the server ("The optimal key is D") and the emotion engine results ("Joy") are sent to the terminal.
[1718] The device displays on the screen, "The optimal key is D. Also, your emotion is joy, so sing with energy."
[1719] Based on the displayed key information and advice, the user can enjoy singing "Let it be" in the key of D.
[1720] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[1721] The processing flow will be explained below.
[1722] Step 1: Song Selection
[1723] The user operates the interface of the karaoke terminal and selects the song they want to sing.
[1724] The terminal obtains information about the selected song and displays it on the screen.
[1725] Step 2: A cappella instructions
[1726] The device displays a message to the user saying, "Please sing a little a cappella."
[1727] The user checks the message.
[1728] Step 3: Start recording
[1729] When the user starts singing, the device records the user's singing voice through the built-in microphone.
[1730] The recording means stores a certain period of singing voice (e.g., about 5 seconds) as digital data.
[1731] Step 4: End recording
[1732] When the user finishes singing, the device will automatically stop recording.
[1733] The recorded audio data is temporarily stored in memory.
[1734] Step 5: Send data
[1735] The terminal transmits the recorded audio data to the server.
[1736] The server receives the HTTP request and stores the audio file.
[1737] Step 6: Data analysis
[1738] The server generates a request to send the stored audio data to the music analysis API.
[1739] The audio data is analyzed using a music analysis API, which includes obtaining pitch and estimating intervals.
[1740] Step 7: Sentiment Analysis
[1741] The server uses an emotion engine to recognize the user's emotion from the stored audio data.
[1742] The emotion engine analyzes the voice data to determine the user's emotional state (e.g., happy, sad, angry, etc.).
[1743] Step 8: Obtaining analysis results
[1744] The music analysis API returns the analysis results (e.g., optimal key information) to the server.
[1745] The server combines the results of the music analysis and the emotion engine to generate advice based on the optimal key and emotion.
[1746] Step 9: Send results
[1747] The server organizes the analysis results and transmits them to the user's terminal.
[1748] As an HTTP response, the most appropriate key information and advice based on the emotion are returned to the device.
[1749] Step 10: View the results
[1750] The device then displays the analysis results on its screen, displaying the message, "The optimal key is D. Your emotion is joy, so sing with enthusiasm."
[1751] The user confirms this message.
[1752] Step 11: Adjust the key
[1753] The user uses the key adjustment function on the terminal to change the karaoke playback key to the optimum key displayed.
[1754] The device will retain the new settings after adjusting the keys.
[1755] Step 12: Karaoke playback
[1756] After the user completes the key setting, karaoke playback begins.
[1757] The device will play the song in the new key setting, allowing the user to sing in the optimal key.
[1758] Step 13: Enjoy singing
[1759] Users can sing songs in the optimal key and enjoy karaoke based on advice based on their emotions.
[1760] Example 2
[1761] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1762] Conventional karaoke systems can suggest the optimal key for a user's singing, but they cannot provide performances or advice that take the user's emotions into consideration. As a result, users cannot receive accurate advice or performances that reflect their emotions, limiting the quality of their karaoke experience. It is also difficult for users to enjoy a performance that matches their own emotions while singing. There is a need to solve this problem and provide a karaoke experience that is in tune with the user's emotions.
[1763] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1764] In this invention, the server includes an analysis means for analyzing the recorded sound source data, an emotion analysis means for recognizing the user's emotion from the recorded sound source data, and a display means for displaying the optimum key for the user's singing and performance and advice suited to the emotion based on the results of the analysis by the analysis means and the emotion analysis means. This makes it possible to provide not only the optimum key for the user's singing, but also appropriate advice and performance according to the emotion.
[1765] The "selection means" is a function that allows the user to select a song to sing through the interface of the karaoke system.
[1766] "Recording means" is a function for recording the user's singing through a microphone and saving it as digital audio data.
[1767] The "analysis means" is a function that uses recorded audio data to analyze pitch and interval and identify the optimal key for the user's singing.
[1768] The "emotion analysis means" is a function for recognizing and analyzing the user's emotions from the recorded audio data.
[1769] The "display means" is a function for visually presenting the analysis results obtained by the analysis means and emotion analysis means to the user, and for providing advice and presentations that correspond to the optimal key and emotion.
[1770] MODE FOR CARRYING OUT THE INVENTION
[1771] This invention improves the karaoke experience by analyzing the user's singing data and emotional data in a karaoke system and suggesting the optimal key and song based on the user's emotion. A specific method for implementing this invention will be described below.
[1772] Required Hardware and Software
[1773] To implement this system, a terminal and a server are required. The terminal has an interface for users to select songs and sing, and the server has the function of analyzing voice data and emotional data.
[1774] Terminal
[1775] Touchscreen display: The user operates this to select songs.
[1776] Microphone: Used to record the user's singing.
[1777] Internet connection: Required to send data to the server.
[1778] server
[1779] Music analysis API: Used to analyze audio data and identify the optimal key (e.g., Spotify API).
[1780] Sentiment analysis engine: Used to identify user emotions from voice data (e.g., IBM Watson).
[1781] Database: Used to store analysis results and user information.
[1782] System configuration
[1783] 1. Song selection method
[1784] The user selects the song they want to sing from the song list displayed on the terminal.
[1785] 2. Recording Method
[1786] After the user selects a song, the device displays a message saying, "Please sing a little a cappella," encouraging the user to sing a cappella.
[1787] When the user starts singing, the device records the voice through the microphone and stores it as digital data.
[1788] 3. Analysis method
[1789] The device sends the recorded audio data to a server using an internet connection.
[1790] The server uses a music analysis API to analyze the audio data and identify pitch and interval.
[1791] At the same time, the server uses an emotion analysis engine to recognize the user's emotions from the voice data.
[1792] 4. Display means
[1793] The server transmits the results obtained from the analysis means and the emotion analysis means to the terminal.
[1794] The terminal displays this information on the screen and provides the user with the most suitable key and appropriate performance and advice for their emotion.
[1795] Specific examples
[1796] If a user uses a karaoke system to select "Let it be (The Beatles)" and sing a bit a cappella, the system performs the following steps:
[1797] 1. Song Selection
[1798] The user selects "Let it be (The Beatles)" on the device's touchscreen.
[1799] 2. Singing and Recording
[1800] The device displays the message "Please sing a little a cappella," and the user sings "When I find myself in times of trouble..."
[1801] The user's singing is recorded and saved as digital data.
[1802] 3. Sending and analyzing audio data
[1803] The recording data is sent from the device to the server, which uses a music analysis API to identify the optimal key (D).
[1804] At the same time, an emotion analysis engine is used to recognize emotions (happiness).
[1805] 4. Displaying results and giving advice
[1806] The analysis results from the server (optimal key: D) and emotion analysis results (joy) are displayed on the device.
[1807] The device displays the message, "The optimal key is D. Also, your emotion is joy, so sing with enthusiasm."
[1808] Prompt Sentence Examples
[1809] "In a karaoke system, a user selects The Beatles' "Let it Be" and sings a portion a cappella. This singing voice is recorded, and the system analyzes it using an emotion engine. The system recognizes the user's emotion as "joy." The system also determines that the key that best suits the singing voice is D. Based on these results, what kind of feedback should be provided to the user?"
[1810] In this way, the present invention can provide a multi-functional karaoke system that not only allows the user to enjoy karaoke in the key that is best suited to them, but also allows them to receive advice according to their emotions at the time.
[1811] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1812] Specific explanation of processing steps
[1813] Step 1: Select a song
[1814] Terminal handling
[1815] Input: The user operates the touchscreen display to select a song from the displayed song list.
[1816] Operation:
[1817] The terminal receives touchscreen input and retrieves song information based on the user's selection.
[1818] Output: The title and artist name of the selected song will be displayed on the screen.
[1819] Step 2: Sing and record your a cappella
[1820] Terminal handling
[1821] Input: After a song is selected, the user begins singing.
[1822] Operation:
[1823] The device displays the message, "Please sing a little a cappella."
[1824] When the user starts singing, the voice is recorded through a microphone.
[1825] The recorded audio data is saved as a digital file.
[1826] Output: The recording data is saved on the device.
[1827] Step 3: Sending audio data to the server
[1828] Terminal handling
[1829] Input: Recorded data is saved.
[1830] Operation:
[1831] The device sends the recorded data to a server via the Internet.
[1832] The HTTPS protocol is used for data transmission.
[1833] Output: The recording data is sent to the server.
[1834] Step 4: Analyzing the audio data
[1835] Server Processing
[1836] Input: Recording data sent from the device.
[1837] Operation:
[1838] The server receives the recording data.
[1839] Sends a request to an audio analysis API (e.g., Spotify API) to analyze the pitch and interval of the audio data.
[1840] The voice analysis API returns the analysis results to the server.
[1841] Output: Audio analysis results (optimal key, etc.).
[1842] Step 5: Sentiment Analysis
[1843] Server Processing
[1844] Input: Recording data sent from the device.
[1845] Operation:
[1846] The server uses an emotion analysis engine (e.g., IBM Watson) to analyze the user's emotions.
[1847] The emotion analysis engine identifies emotions based on the recorded data and sends the results to the server.
[1848] Output: Sentiment analysis results (e.g., joy, sadness, anger, etc.).
[1849] Step 6: Combining the analysis results
[1850] Server Processing
[1851] Input: Voice analysis results and emotion analysis results.
[1852] Operation:
[1853] The server combines the results of the voice analysis and the emotion analysis.
[1854] Identify the best key and emotionally appropriate performance and advice.
[1855] Output: Final analysis results (optimal key and emotional advice).
[1856] Step 7: Viewing the analysis results
[1857] Terminal handling
[1858] Input: Final analysis result data from the server.
[1859] Operation:
[1860] The terminal displays the analysis results received from the server on the screen.
[1861] The user sees the best key and emotion-based advice.
[1862] Output: The screen displays "The best key is D. Your emotion is joy, so sing with enthusiasm."
[1863] In this way, the system analyzes the user's singing data at multiple stages and provides optimal feedback to the user, enhancing the karaoke experience.
[1864] (Application example 2)
[1865] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1866] Conventional karaoke systems have the problem of being unable to find the optimal key for the user when singing and providing a performance that reflects the user's emotions at that time. In particular, they are unable to provide feedback based on the user's emotions, preventing users from having a more enjoyable and personalized karaoke experience. This leaves karaoke booths and live music bars with a challenge to improve customer satisfaction.
[1867] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1868] In this invention, the server includes a selection means for allowing a user to select a song and sing a little a cappella along with it, a recording means for recording the user's singing, an analysis means for analyzing the recorded sound data, a display means for displaying the optimal key for the user's singing based on the analysis results by the analysis means, an emotion engine for recognizing emotions from the user's voice data, and a means for providing the user with visual feedback based on the recognition results of the emotion engine. This not only enables the user to sing in the optimal key for themselves, but also allows them to receive advice based on their emotions at the time.
[1869] A "selection means" is a means that provides an interface for a user to select a song and sing along a bit a cappella with it.
[1870] "Recording means" means any means, including devices and software, for recording the audio of a user's singing.
[1871] The "analysis means" is a means for analyzing the recorded sound source data and performing a process to identify the appropriate key and other musical elements for the song.
[1872] The "display means" is a means for visually presenting the optimum key and other information to the user based on the analysis results obtained by the analysis means.
[1873] The "emotion engine" is a means including software and algorithms for recognizing emotions from the user's voice data and feeding back the recognition results to the analysis means.
[1874] The "means for providing visual feedback" is a means for visually displaying appropriate advice or dramatic effects to the user based on the recognition results of the emotion engine.
[1875] The present invention provides a system for enabling users to select songs and improve their singing experience in a physical venue such as a karaoke booth or live music bar, etc. The system includes a selection means, a recording means, an analysis means, a display means, an emotion engine, and a means for providing visual feedback.
[1876] System Configuration
[1877] 1. The selection means is a device such as a tablet or smartphone. The user operates the interface on the screen to select the song they want to sing. This selection means includes a function to display a song list and accept user input.
[1878] 2. The recording function provides a function to record the user's a cappella singing using a microphone. The recording method collects a certain amount of audio data and saves it as a digital file. For example, a microphone built into a smartphone or a microphone installed in a karaoke booth can be used.
[1879] 3. The analysis method involves analyzing the audio data on the server using a music analysis API (such as IBM Watson or Google Cloud Speech-to-Text API). This analysis involves analyzing pitch and interval and comparing it with existing data. Based on the analysis results, the optimal key for the user's singing is identified.
[1880] 4. The emotion engine uses software and algorithms to recognize emotions from the user's voice data. For example, it uses Microsoft Azure's Emotion API or a proprietary emotion recognition algorithm. This analyzes the user's emotional state (e.g., joy, sadness, anger, etc.).
[1881] 5. As a display method, visual feedback is provided to the user based on the analysis results and the emotion engine's recognition results. For example, optimal keys and advice based on emotions are displayed on the screen of a tablet or smartphone.
[1882] Example
[1883] Specific examples
[1884] Below is a case where the user selects the song "Let it be (The Beatles)" and analysis and emotion recognition are performed based on that.
[1885] 1. Song selection and singing:
[1886] The user selects "Let it be (The Beatles)" on the karaoke machine.
[1887] The device displays the message, "Please sing a little a cappella."
[1888] The user sings a bit of "When I find myself in times of trouble..." and their voice is recorded.
[1889] 2. Data Analysis and Emotion Recognition:
[1890] The recorded data is sent to the server and analyzed using a music analysis API.
[1891] At the same time, the emotion engine recognizes the user's emotion as "joy."
[1892] 3. Displaying the results:
[1893] The analysis results and emotion recognition results are sent to the device, and the message "The optimal key is D. Your emotion is joy, so sing with enthusiasm." is displayed.
[1894] Using the displayed key information and advice, users can enjoy singing "Let it be" in the key of D.
[1895] Prompt Sentence Examples
[1896] "User selects 'Let it be'. The recording is analyzed to determine the best key. The emotion engine recognizes 'joy' and provides advice and the best key."
[1897] In this way, the present invention provides a system that allows users to not only enjoy karaoke in the key that is best suited to them, but also receive advice that corresponds to their emotions at the time.
[1898] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1899] Step 1:
[1900] The user selects a song using the selection means on a tablet or smartphone. The input is the ID and title of the song selected by the user, which is provided to the system. The device displays a list of songs, accepts the user's selection, and obtains song information.
[1901] Step 2:
[1902] After the user selects a song, the device prompts the user to sing a little a cappella using the microphone. The input is the user's singing voice, and the output is the recorded audio data. When the user starts singing, the device captures the audio signal and stores it digitally as audio data.
[1903] Step 3:
[1904] The device sends the recorded audio data to the server. The input is the audio data file, and the output is the data received on the server. The server sends the audio data to the specified endpoint and confirms receipt.
[1905] Step 4:
[1906] The server sends the received audio data to the music analysis API and requests analysis. The input is audio data, and the output is analyzed musical element data such as interval and pitch. The server sends a request to the API endpoint and receives the analysis results in return.
[1907] Step 5:
[1908] The server then uses an emotion engine to recognize the user's emotion from the voice data. The input is the voice data and the output is the user's emotional state (e.g., "joy"). The emotion engine analyzes the voice data and generates extracted emotion information.
[1909] Step 6:
[1910] The server integrates the analysis results from the music analysis API and the recognition results from the emotion engine to identify the optimal key and advice according to the emotion. The input is musical element data and emotion information, and the output is the optimal key and advice. The server processes these data and generates the integrated result.
[1911] Step 7:
[1912] The server sends the integration result to the terminal and provides feedback to the user using a display means. The input is the integration result data, and the output is the optimal key and advice according to the emotion displayed on the user terminal. The terminal displays the received information on the screen and provides visual feedback to the user.
[1913] Step 8:
[1914] The user sings the song using the displayed optimal key and emotional advice. The input is the displayed information, and the output is the user's improved singing performance. Based on the feedback, the user can enjoy singing in the optimal key.
[1915] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1916] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1917] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1918] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1919] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1920] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1921] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1922] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1923] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1924] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1925] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1926] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1927] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1928] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1929] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1930] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1931] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1932] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1933] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1934] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1935] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1936] The following is further disclosed regarding the above embodiment.
[1937] (Claim 1)
[1938] A means for the user to select a song and sing along a bit a cappella;
[1939] a recording means for recording the sound source sung by the user;
[1940] analysis means for analyzing the recorded sound source data;
[1941] and a display means for displaying the optimum key for the user's singing based on the analysis result obtained by the analysis means.
[1942] (Claim 2)
[1943] 2. The system according to claim 1, wherein the analysis means analyzes the sound source data using a music analysis API.
[1944] (Claim 3)
[1945] 2. The system according to claim 1, wherein the display means allows the user to change the key of the karaoke playback based on the displayed optimum key.
[1946] "Example 1"
[1947] (Claim 1)
[1948] a selection means for a user to select a song;
[1949] a means for prompting the user to sing a little a cappella;
[1950] a recording means for recording the sound source sung by the user;
[1951] a transmitting means for transmitting the recorded sound source data to a server;
[1952] an analysis means for the server to analyze the sound source data using a music analysis API;
[1953] a display means for displaying the optimum key for the user's singing based on the analysis result by the analysis means;
[1954] A system including:
[1955] (Claim 2)
[1956] 2. The system of claim 1, wherein the analysis means includes a process for transmitting audio data via a network and receiving optimal key information using a music analysis API.
[1957] (Claim 3)
[1958] 2. The system according to claim 1, wherein the display means visually displays the optimal key on the screen of the user's device based on the analysis results received from the server.
[1959] "Application Example 1"
[1960] (Claim 1)
[1961] A song selection mechanism that allows users to select a song and sing along a bit a cappella to it;
[1962] a recording means for recording the sound source sung by the user;
[1963] analysis means for analyzing the recorded sound source data;
[1964] a display means for displaying the optimum key for the user's singing based on the analysis result by the analysis means;
[1965] a communication means for transmitting the analysis result to the cloud server based on the music analysis API and receiving the analysis result;
[1966] A system including:
[1967] (Claim 2)
[1968] 2. The system of claim 1, wherein the sound source data is analyzed using a music analysis API.
[1969] (Claim 3)
[1970] 2. The system according to claim 1, wherein the display means allows the user to change the key of the karaoke playback based on the displayed optimum key.
[1971] (Claim 4)
[1972] 2. The system of claim 1, wherein the music selection means includes a function for selecting music on a smartphone via a user interface.
[1973] "Example 2: Combining Emotion Engines"
[1974] (Claim 1)
[1975] A means for the user to select a song and sing along a bit a cappella;
[1976] a recording means for recording the sound source sung by the user;
[1977] analysis means for analyzing the recorded sound source data;
[1978] An emotion analysis means for recognizing a user's emotion from recorded audio data;
[1979] a display means for displaying the optimum key for the user's singing and the performance and advice suited to the user's emotions based on the results of the analysis by the analysis means and the emotion analysis means;
[1980] A system including:
[1981] (Claim 2)
[1982] 2. The system according to claim 1, wherein the analysis means analyzes the sound source data using a sound analysis API.
[1983] (Claim 3)
[1984] 2. The system according to claim 1, wherein the display means makes it possible to change the key and performance of the karaoke playback based on the displayed optimum key and emotion.
[1985] "Application example 2 when combining emotion engines"
[1986] (Claim 1)
[1987] A means for the user to select a song and sing along a bit a cappella;
[1988] a recording means for recording the sound source sung by the user;
[1989] analysis means for analyzing the recorded sound source data;
[1990] a display means for displaying the optimum key for the user's singing based on the analysis result by the analysis means;
[1991] An emotion engine for recognizing emotions from the user's voice data;
[1992] a means for providing visual feedback to the user based on the recognition results of the emotion engine;
[1993] A system including:
[1994] (Claim 2)
[1995] 2. The system according to claim 1, wherein the analysis means analyzes the sound source data using a music analysis API.
[1996] (Claim 3)
[1997] 2. The system according to claim 1, wherein the display means allows the user to change the key of the karaoke playback based on the displayed optimum key. [Explanation of symbols]
[1998] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for the user to select a song and sing along a bit a cappella; a recording means for recording the sound source sung by the user; analysis means for analyzing the recorded sound source data; and a display means for displaying the optimum key for the user's singing based on the results of the analysis by the analysis means.
2. 2. The system according to claim 1, wherein the analysis means analyzes the sound source data using a music analysis API.
3. 2. The system of claim 1, wherein the display means allows the user to change the key of the karaoke playback based on the displayed optimum key.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A