system
The system addresses the lack of effective feedback in karaoke by analyzing pitch and timing, providing real-time guidance, and suggesting improvements, enhancing singing ability through continuous practice and feedback.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Existing karaoke systems fail to provide effective real-time analysis and feedback for improving singing ability, making it difficult for users to practice vocalization, confirm pitch, and identify areas for improvement.
A system that records user singing, analyzes pitch and vocal timing on a server, provides immediate feedback, and suggests improvement methods, while allowing real-time audio guidance and saving analysis results for future training.
Enables users to practice effectively, receive immediate feedback, and improve their singing skills through continuous evaluation and personalized guidance.
Smart Images

Figure 2026062307000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] ]]There is a problem that although one wants to improve singing ability in karaoke, it is difficult to practice vocalization alone or confirm the pitch, and it is difficult to perform self-evaluation and problem identification. In particular, since real-time analysis during singing and appropriate feedback cannot be obtained, effective practice is difficult. Also, since specific improvement methods are not known, the same problems tend to be repeated. To solve such problems, there is a need for a means for a user to easily receive an evaluation of their own singing and obtain specific improvement methods.
Means for Solving the Problems
[0005] This invention is a system for supporting the improvement of singing ability in karaoke. The system has means for recording the user's singing and transmitting the audio data to a server. Furthermore, it has means for analyzing pitch and vocal timing on the server side, and for evaluation and problem identification. It also includes means for providing feedback to the user on the evaluation and improvement methods generated based on the analysis results. In addition, it includes means for providing real-time audio guidance and means for saving the analysis results to a database for use as training data in the future. In this way, the user can practice repeatedly while receiving effective feedback, gain confidence in karaoke, and improve their singing ability.
[0006] A "user" refers to an individual who uses the system to practice singing.
[0007] "Means of recording" refers to the function that saves the user's singing as audio data.
[0008] "Audio data" refers to audio information recorded in digital format from the user's singing.
[0009] A "server" refers to a computer system that receives, analyzes, and processes audio data.
[0010] "Means of transmission" refers to the function of sending voice data from the terminal to the server.
[0011] "Pitch analysis" refers to the process of detecting and analyzing the pitch of sounds in audio data.
[0012] "Vocalization timing analysis" refers to the process of analyzing the start timing and rhythm of sounds in audio data.
[0013] "Evaluation" refers to the result that expresses the quality of a user's singing, based on analysis results, using numerical values and comments.
[0014] "Problem identification" refers to the process of identifying areas in singing that need improvement from the analysis results.
[0015] "Improvement method" refers to specific solutions or advice for the identified problems.
[0016] "Feedback means" refers to the function of conveying the evaluation from the server and the improvement method to the user.
[0017] "Voice guide" refers to voice instructions and advice provided in real time during the user's singing.
[0018] "Database" refers to an information management system for storing analysis results and other related information.
[0019] "Training data" refers to a dataset for using past data such as analysis results for learning and analysis.
Brief Description of Drawings
[0020] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0021] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0022] First, the language used in the following description will be described.
[0023] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0024] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0025] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0026] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0028] [First Embodiment]
[0029] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0030] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0033] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0036] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0040] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0041] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This allows users to efficiently improve their singing skills.
[0042] Server-side processing
[0043] Receiving audio data
[0044] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[0045] Voice analysis
[0046] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine primarily performs the following two types of analysis:
[0047] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[0048] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[0049] Evaluation and Problem Identification
[0050] The server evaluates the user's singing based on the results of the audio analysis. It quantifies deviations in pitch and timing, and identifies specific problems. For example, it may evaluate things like "the pitch is too low" or "the rhythm is off."
[0051] Suggestions for improvement
[0052] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice maintaining a consistent rhythm." These evaluation results and improvement methods are stored in a database and may be used as training data in the future.
[0053] Submission of evaluation results and improvement methods
[0054] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0055] Terminal-side processing
[0056] Voice recording
[0057] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[0058] Sending audio data to the server
[0059] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[0060] Receiving and displaying evaluation results and improvement methods
[0061] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[0062] Audio guides are available.
[0063] The device provides real-time voice guidance based on improvement suggestions received from the server. For example, it provides users with instructions such as "Sing this part a little higher" or advice such as "Keep the rhythm consistent" as voice guidance.
[0064] Specific example
[0065] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to a server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the analysis results in "the pitch is a little low and the rhythm is off", the server generates advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." The device displays this evaluation result and advice to the user, providing it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[0066] In this way, the system of the present invention provides effective support for users to enjoy karaoke with confidence and improve their singing ability.
[0067] The following describes the processing flow.
[0068] Step 1:
[0069] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[0070] Step 2:
[0071] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, and the audio data is temporarily buffered.
[0072] Step 3:
[0073] The terminal sends buffered audio data to the server in fixed time slices (e.g., every 5 seconds). The transmission is asynchronous, and data is sent while recording.
[0074] Step 4:
[0075] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[0076] Step 5:
[0077] The server identifies the pitch of each sound in the audio data through pitch analysis. It also analyzes the rhythm and timing through speech timing analysis.
[0078] Step 6:
[0079] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[0080] Step 7:
[0081] The server identifies problems based on the evaluation results. For example, it might display specific issues such as "the pitch is too low" or "the rhythm is inaccurate."
[0082] Step 8:
[0083] Based on the problems identified by the server, it generates appropriate improvement methods. For example, it creates specific advice such as, "Next time, be sure to sing in a higher voice."
[0084] Step 9:
[0085] The server saves the evaluation results and improvement methods it generates to a database, making them available as training data for the future.
[0086] Step 10:
[0087] The server sends evaluation results and improvement suggestions to the terminal. This transmission is done in real time, allowing the user to receive immediate feedback during their next singing session.
[0088] Step 11:
[0089] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display "The pitch is too low. Next time, try to hit higher notes."
[0090] Step 12:
[0091] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Raise this part a little."
[0092] Step 13:
[0093] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[0094] Through this series of steps, users can efficiently receive feedback on their singing and suggestions for improvement, thereby boosting their confidence in karaoke.
[0095] (Example 1)
[0096] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] A major drawback of conventional karaoke systems is that they do not provide concrete feedback to efficiently improve users' singing skills. Furthermore, they lack mechanisms to identify inappropriate pitch or rhythm and suggest ways to improve them, making it difficult for users to improve themselves.
[0098] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0099] In this invention, the server includes means for using a speech processing engine to analyze pitch and vocalization timing, means for using a data analysis engine to perform evaluation and problem identification based on the analysis results, and means for using a natural language generation engine to generate evaluation and improvement methods. This makes it possible to evaluate the user's singing skills in real time and provide specific improvement methods.
[0100] A "user" is a person who uses this system to sing karaoke.
[0101] "Means of recording" refers to a function that captures the user's voice with a microphone and saves it as digital audio data.
[0102] "Means of sending audio data to a server" refers to the function of uploading recorded audio files to a server via a network.
[0103] A "speech processing engine that analyzes pitch and timing" refers to software or hardware that analyzes speech data and evaluates the accuracy of pitch and timing.
[0104] A "data analysis engine that performs evaluation and problem identification based on analysis results" refers to software or hardware that uses data obtained from a speech processing engine to evaluate singing and identify specific problems.
[0105] A "natural language generation engine for generating evaluation and improvement methods" refers to software or hardware that generates improvement advice for users based on identified problems.
[0106] "Means of providing feedback" refers to a function that communicates the generated evaluation results and improvement methods to the user.
[0107] "Means of providing audio guidance" refers to a function that provides users with real-time audio advice.
[0108] "Means of saving to a database" refers to the function of saving analysis results and evaluation results in digital format, making them available as data for future use.
[0109] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This system provides specific feedback to efficiently improve the user's singing skills. The following describes specific embodiments for carrying out this invention in detail.
[0110] Server Processing
[0111] The server first receives the user's voice data sent from the terminal. Methods such as HTTP POST requests are used to receive the voice data. The received voice data is analyzed using a speech processing engine such as Google Cloud Speech-to-Text. The speech processing engine mainly performs pitch analysis and speech timing analysis. Pitch analysis evaluates whether the voice data is sung at the correct pitch, and speech timing analysis evaluates the accuracy of the rhythm and timing.
[0112] Next, the server uses a data analysis engine to evaluate the user's singing based on the results of the voice analysis and identify problems. The data analysis uses libraries such as Python's Pandas and NumPy. For example, it extracts specific data such as the percentage of pitch deviation or the number of milliseconds the rhythm is delayed, identifying problems like "the pitch is too low" or "the rhythm is off."
[0113] Based on the identified problems, the server uses a natural language generation engine (e.g., OpenAI®'s GPT-3®) to generate improvement methods. The generated evaluation results and improvement methods are stored in a database. This storage is done using a database management system such as MySQL® or PostgreSQL.
[0114] Finally, the server sends the evaluation results and improvement suggestions to the terminal in JSON format, providing feedback to the user on points to keep in mind for their next singing session.
[0115] Terminal processing
[0116] The device records the user's singing using its built-in microphone. The recording uses the Android® or iOS voice recording API. The recorded audio data is sent to the server in real time. Upon receiving a response from the server, the evaluation results and suggestions for improvement are displayed in the user interface. This utilizes mobile application development frameworks such as React Native and Flutter®.
[0117] Furthermore, the device provides real-time voice guidance based on improvement suggestions received from the server. For example, the device's TTS (text-to-speech) engine is used to provide voice instructions such as "Sing this part a little higher" or advice like "Keep the rhythm consistent." This allows users to receive immediate feedback and improve their singing skills.
[0118] Specific example
[0119] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to the server via an HTTP POST request. The server analyzes the pitch and timing using the Google Cloud Speech-to-Text engine and derives an evaluation result. If the analysis results indicate that "the pitch is a little low and the rhythm is off", the server uses GPT-3 to generate advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." This evaluation result and advice are sent to the device in JSON format, which displays it to the user and simultaneously provides it as an audio guide using the TTS engine.
[0120] Example of a prompt
[0121] "Please generate improvement methods to enhance the user's singing skills. For example, please provide advice for when a user's pitch is low and their rhythm is off."
[0122] As described above, this system evaluates the user's singing in real time and provides specific methods for improvement, thereby effectively supporting users in enjoying karaoke with confidence.
[0123] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0124] Step 1: Audio Recording
[0125] The user launches the karaoke app, selects a song, and begins singing. The device records the user's singing using its built-in microphone. The recording uses the Android or iOS voice recording API. The input is the user's singing, and the output is digital audio data. This audio data is temporarily stored in the device's memory.
[0126] Step 2: Sending audio data to the server
[0127] The terminal sends recorded audio data to the server at regular time slices. This transmission is asynchronous and uses HTTP POST requests. The input is the recorded digital audio data, and the output is the audio data sent to the server. The terminal monitors the HTTP response to confirm that the transmission was successful.
[0128] Step 3: Receiving audio data
[0129] The server receives user voice data transmitted from the terminal. The input is the voice data transmitted from the terminal, and the output is the voice data stored in the server's memory. The server temporarily stores the received data and uses it for subsequent processing.
[0130] Step 4: Voice Analysis
[0131] The server uses the Google Cloud Speech-to-Text engine to analyze the received audio data. It converts the audio data to text, and then performs pitch and timing analysis. The input is the received audio data, and the output is analysis data regarding pitch and timing. Pitch analysis identifies the pitch of each part of the audio data, while timing analysis evaluates the accuracy of rhythm and timing. The analysis results are stored as numerical data on the server.
[0132] Step 5: Assessment and Problem Identification
[0133] The server evaluates the user's singing based on the results of the voice analysis and identifies problems. Using Python data analysis libraries (e.g., Pandas or NumPy), it evaluates the analysis results and quantifies specific problems such as "pitched too low" or "rhythm off." The input is the results of the voice analysis and numerical data, and the output is the evaluation score and data identifying the problems.
[0134] Step 6: Propose improvement methods
[0135] The server generates appropriate improvement methods based on the identified problems. Using a natural language generation engine (e.g., OpenAI's GPT-3), it generates specific advice such as, "Next time, try to focus on using a higher pitch." The input is the evaluation results and identified problems, and the output is the text data of the generated improvement methods. The generated advice is stored in the server's database.
[0136] Step 7: Submit evaluation results and improvement suggestions
[0137] The server sends the generated evaluation results and improvement suggestions to the terminal. It returns JSON data to the terminal as an HTTP response. The input is text data of the evaluation results and improvement suggestions, and the output is the data sent to the terminal. The terminal receives this data in real time.
[0138] Step 8: Displaying evaluation results and improvement methods.
[0139] The terminal receives evaluation results and improvement suggestions from the server and displays them on the user interface. Evaluation scores are displayed in graph format, and specific advice is provided in text format. Input is the evaluation results and improvement suggestions data received from the server, while output is the information displayed on the user interface. Users can use this information to improve their next singing performance.
[0140] Step 9: Providing audio guides
[0141] The terminal provides real-time voice guidance based on improvement suggestions received from the server. Using a TTS (text-to-speech) engine, it communicates instructions to the user via voice, such as "Sing this part a little higher" or "Maintain a consistent rhythm." The input is text data of improvement suggestions, and the output is voice data as voice guidance. Users can listen to this and work on improving their singing.
[0142] (Application Example 1)
[0143] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0144] While existing systems provide real-time evaluation and feedback to help users efficiently improve their singing skills during karaoke sessions, dynamically delivering personalized advertisements based on users' singing habits and analysis results would further enhance its value. However, traditional karaoke systems lack these features and fail to appropriately deliver advertisements that match user interests. This makes it difficult to improve user engagement and maximize advertising revenue.
[0145] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0146] In this invention, the server includes means for analyzing pitch and vocal timing, means for evaluating the user's singing based on the analysis results and identifying problems, means for generating evaluation and improvement methods, and means for dynamically generating advertisements based on the evaluation and analysis results. This allows the user to receive personalized advertisements based on the analysis results, along with real-time feedback.
[0147] A "user" refers to an individual who uses a karaoke system to sing.
[0148] "Audio data" refers to data that digitally records the voice of a user singing.
[0149] A "server" refers to a central processing system used for analyzing, evaluating, providing feedback on, and generating advertisements for audio data.
[0150] "Pitch" refers to the element that identifies the height (frequency) of each part in audio data and evaluates whether it is accurate.
[0151] "Vocalization timing" refers to the elements used to evaluate the rhythm and timing of each part of audio data.
[0152] "Analysis results" refers to evaluation data obtained from the analysis of pitch and vocalization timing.
[0153] "Evaluation" refers to data that represents a user's singing skill using numerical values or indicators based on the analysis results.
[0154] "Problem identification" refers to clearly indicating specific areas for improvement in the user's singing based on the evaluation results.
[0155] "Improvement methods" refer to specific practice methods and advice for addressing identified problems.
[0156] "Feedback" refers to providing users with information to help them improve their performance and singing in the future.
[0157] "Advertising" refers to commercial messages that are dynamically generated and provided to users based on user analysis results and evaluations.
[0158] "Dynamically generating" refers to generating ads in real time based on the user's current rating results and analytical data.
[0159] This invention is a system that analyzes a user's singing voice in real time while they are singing karaoke, provides feedback on the evaluation and methods for improvement, and simultaneously delivers personalized advertisements based on the user's analysis results. This allows users to efficiently improve their singing skills and receive advertisements that are tailored to their interests.
[0160] This system mainly consists of the following components:
[0161] 1. Device for voice recording:
[0162] A smartphone or other device is used to record the user's singing voice. The device collects the audio data using its built-in microphone.
[0163] 2. Analysis of pitch and vocal timing:
[0164] The server receives audio data and uses a speech processing engine to analyze pitch and timing. Specifically, it uses Python and a speech analysis library (e.g., LibROSA).
[0165] 3. Evaluation and feedback based on analysis results:
[0166] Based on the analysis results, the server evaluates the user's singing and displays deviations in pitch and timing as specific numerical values. For example, it generates evaluation results such as "the pitch is too low" or "the rhythm is off."
[0167] 4. Propose ways to improve:
[0168] Based on the evaluation results, the server generates specific improvement methods for addressing the user's problems. For example, it provides advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." In addition to the aforementioned voice analysis library, it uses a generative AI model to generate even more detailed advice.
[0169] 5. Generating and delivering personalized ads:
[0170] The server uses an algorithm that dynamically generates advertisements based on user analysis results and evaluation data. The generated advertisements, along with evaluation feedback, are sent to the device and presented to the user. For example, if a user frequently sings cover songs by a particular singer, advertisements for that singer's new songs or live performance information will be displayed.
[0171] The HTTP protocol and RESTful API are used for communication between the server and the terminal, enabling real-time data transmission and reception. Voice data is sent from the terminal to the server at regular intervals, and analysis results and advertisements are fed back to the user.
[0172] Specific example:
[0173] For example, suppose a user sings many songs by their favorite singer. This audio data is recorded using the smartphone's microphone and sent to a server. The server analyzes the pitch and rhythm and evaluates it as "low pitch." Next, the server generates a suggestion for improvement, such as "Next time, try to focus on singing in a higher register," and also generates an advertisement for the singer's new album. All of this information is sent to the user's smartphone and provided as feedback.
[0174] Example of a prompt:
[0175] Use the following input prompt statements for the generative AI model:
[0176] Based on the user's singing voice data and its evaluation results (pitch, timing), generate advertisements tailored to the user's interests. For example, users who frequently sing cover songs by a particular singer should be shown advertisements for that singer's latest album, while users with low pitch should be shown advertisements for vocal training. Output the evaluation results and advertisements in text format.
[0177] This allows users to receive real-time feedback and personalized ads, further enhancing their singing experience.
[0178] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0179] Step 1:
[0180] When the user starts singing along to a karaoke song they've selected, the device uses its built-in microphone to record their voice in real time. The recorded audio data is temporarily stored in a buffer.
[0181] (Input) User's singing voice
[0182] (Output) Buffered audio data
[0183] Step 2:
[0184] The device sends recorded audio data to the server at regular time slices. This allows the server to analyze the data in real time. Because the audio data is transmitted asynchronously, the user can receive continuous feedback without delay.
[0185] (Input) Buffered audio data
[0186] (Output) Audio data sent to the server
[0187] Step 3:
[0188] The server analyzes the received audio data. Using an audio processing engine, it analyzes the pitch (frequency) and utterance timing (rhythm). An audio analysis library (e.g., LibROSA) is used for this analysis. The server generates the results of the pitch and timing analysis and saves them as evaluation data.
[0189] (Input) Audio data sent to the server
[0190] (Output) Analysis results of pitch and vocalization timing
[0191] Step 4:
[0192] The server evaluates the user's singing based on the analysis results. It quantifies deviations in pitch and timing, and identifies specific problems (e.g., pitch is too low, rhythm is off). The evaluation data is stored in a database for each user.
[0193] (Input) Analysis results of pitch and vocalization timing
[0194] (Output) Evaluation results and identification of problems
[0195] Step 5:
[0196] The server generates improvement methods based on the identified problems. Specifically, it might generate advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." A generative AI model may also be used for this process.
[0197] (Input) Evaluation results and problems
[0198] (Output) Specific improvement methods
[0199] Step 6:
[0200] The server generates personalized ads based on analysis results and evaluation data. If a user frequently sings songs by a particular singer, ads suggesting new releases or related products from that singer are dynamically generated.
[0201] (Input) Evaluation results and user singing data
[0202] (Output) Generated advertisement
[0203] Step 7:
[0204] The server sends evaluation results, improvement suggestions, and advertisements to the device. This allows users to receive feedback and advertisements in real time.
[0205] (Input) Evaluation results, improvement methods, generated ads
[0206] (Output) Evaluation results, improvement methods, and advertisements sent to the terminal.
[0207] Step 8:
[0208] The device displays evaluation results, improvement suggestions, and advertisements received from the server in its user interface. This allows users to identify areas for improvement in their singing skills and apply them to their next performance. They can also view personalized advertisements.
[0209] (Input) Evaluation results, improvement methods, advertisements
[0210] (Output) Feedback and advertisements displayed in the user interface
[0211] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0212] This invention relates to a system that recognizes the user's emotions along with their singing voice when they sing karaoke, and adjusts the evaluation and improvement methods based on those emotions. This provides feedback and guidance for improvement that is tailored to the emotional state the user is experiencing, enabling effective improvement of singing ability.
[0213] Server-side processing
[0214] Receiving audio data
[0215] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[0216] Emotion recognition by an emotion engine
[0217] The server uses an emotion engine to recognize the user's emotional state based on voice data, user facial expression data, and other factors. For example, it identifies emotions such as "joy," "sadness," and "tension" from voice tone, speaking speed, and facial expressions.
[0218] Voice analysis
[0219] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine performs the following two analyses:
[0220] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[0221] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[0222] Evaluation and Problem Identification
[0223] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. For example, it evaluates if the pitch is too low, the rhythm is off, or if the user is nervous.
[0224] Suggestions for improvement
[0225] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice singing in a relaxed state." It also provides feedback based on emotional state.
[0226] Submission of evaluation results and improvement methods
[0227] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0228] Terminal-side processing
[0229] Voice recording
[0230] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[0231] Sending audio data to the server
[0232] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[0233] Sending emotional data to the server
[0234] The device also sends emotional data, including facial recognition and voice tone analysis results, to the server. This allows the server to analyze the user's emotional state in real time while they are singing.
[0235] Receiving and displaying evaluation results and improvement methods
[0236] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[0237] Audio guides are available.
[0238] The device provides real-time voice guidance based on improvement instructions received from the server. For example, it might tell the user, "Try singing more relaxed," or "Sing this part a little higher."
[0239] Specific example
[0240] For example, suppose a user sings "song B". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data and emotion data to the server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the emotion engine detects the user's tension, the analysis results may indicate "the pitch is a little low and the rhythm is off" and "the user is tense." Based on these results, advice is generated such as, "Next time, try to sing with a higher pitch, maintain a consistent rhythm, and relax." The device displays this evaluation result and advice to the user and also provides it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[0241] In this way, the system of the present invention provides effective support that allows users to enjoy karaoke with confidence and to constantly be conscious of improving their singing ability.
[0242] The following describes the processing flow.
[0243] Step 1:
[0244] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[0245] Step 2:
[0246] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, buffering both the audio data and the user's facial expressions.
[0247] Step 3:
[0248] The terminal sends buffered audio data and user facial expression data to the server at regular time slices (e.g., every 5 seconds). Transmission is performed asynchronously, with data being sent while recording and recognition are in progress.
[0249] Step 4:
[0250] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[0251] Step 5:
[0252] The server receives facial expression data and passes it to the emotion engine, which analyzes the user's facial features to identify emotions. The emotion engine also considers factors such as voice tone and speaking speed to determine emotions such as "joy," "sadness," and "tension."
[0253] Step 6:
[0254] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[0255] Step 7:
[0256] The server incorporates the user's emotional state into the evaluation based on the results of the emotion engine. For example, if the user is feeling tense, information such as "You need to relax" will be added to the evaluation result.
[0257] Step 8:
[0258] The server identifies pitch and rhythm issues based on the evaluation results and suggests improvement methods based on emotional state. For example, it generates feedback such as "the pitch is too low" or "sing more relaxed."
[0259] Step 9:
[0260] The server saves the evaluation results and improvement methods to a database, making them available as training data for the future.
[0261] Step 10:
[0262] The server sends the evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0263] Step 11:
[0264] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display messages such as, "Your pitch is too low. Next time, try to sing higher notes," or "Try to relax while singing."
[0265] Step 12:
[0266] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Sing this part a little higher," or "Try to relax while you sing."
[0267] Step 13:
[0268] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[0269] Through this series of steps, users can efficiently receive evaluations of their singing and suggestions for improvement, thereby boosting their confidence in karaoke. Furthermore, by utilizing an emotion engine, the system can provide optimal feedback tailored to the user's emotional state.
[0270] (Example 2)
[0271] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0272] Traditional karaoke systems offered limited feedback to help users improve their singing abilities, often focusing only on pitch and rhythm. Furthermore, they failed to reflect emotional states such as tension or joy during singing, making effective improvement guidance tailored to individual user characteristics and circumstances difficult. As a result, users lacked specific advice aligned with their emotional state, hindering efficient improvement of their singing skills.
[0273] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0274] In this invention, the server includes means for recognizing emotions based on voice data and facial expression data, means for providing feedback based on emotional state, and means for performing evaluation and problem identification based on the results of voice analysis. This allows users to receive comprehensive feedback that also takes their own emotional state into account in real time, thereby more effectively supporting the improvement of individual singing abilities.
[0275] "Audio data" refers to data that records the voice of a user singing in digital format.
[0276] "Facial expression data" refers to data obtained by capturing and analyzing the user's facial expressions using a camera.
[0277] "Means for recognizing emotions" refers to a device or software that analyzes and identifies a user's emotional state based on voice data and facial expression data.
[0278] "Means of providing feedback" refers to devices or software that present evaluation results and methods for improvement to users.
[0279] "Pitch" refers to the element that describes the height of the sound produced by the user during singing.
[0280] "Vocalization timing" refers to the timing and rhythm of the user's singing.
[0281] "The result of voice analysis" refers to evaluation data of pitch and pronunciation timing analyzed by a voice processing engine.
[0282] "Means for performing evaluation and problem identification" refers to a device or software for evaluating a user's singing based on the analysis result and identifying points that need improvement.
[0283] "Real-time" refers to data transmission, reception, and processing within a time period where processing is performed almost instantaneously and the user does not feel a delay.
[0284] "Database" refers to a data storage system constructed to systematically store information for later search and use.
[0285] "Training data" refers to data stored for future use in training algorithms such as machine learning.
[0286] The system of the present invention analyzes the singing voice and expression of a user when singing karaoke, recognizes emotions based on the voice data and expression data, and provides an evaluation and improvement method. In order to effectively support the improvement of the user's singing ability, the following specific hardware and software are used.
[0287] Server-side configuration
[0288] The server includes the following hardware and software configurations:
[0289] Voice processing engine: Analyzes voice data. For example, by using the Google Cloud Speech-to-Text API, pitch analysis and pronunciation timing analysis are performed.
[0290] Emotion Recognition Engine: Recognizes user emotions based on voice and facial expression data. For example, it uses the Microsoft® Azure® Emotion API to identify emotions such as "joy," "sadness," and "tension" from the user's facial expressions and voice tone.
[0291] Database: A data storage system that stores analysis results and emotional data for future use as training data.
[0292] Evaluation Logic: Based on the results of voice analysis and emotion recognition, the system evaluates the user's singing, identifies problems, and generates appropriate feedback.
[0293] Terminal-side configuration
[0294] The device includes the following hardware and software configuration:
[0295] Built-in microphone: Used to record the user's singing.
[0296] Built-in camera: Used to acquire user facial expression data.
[0297] Karaoke application: Software that allows users to select their favorite songs and begin singing.
[0298] Asynchronous communication module: A means of communication for transmitting voice and facial expression data to a server.
[0299] Speech synthesis engine: Provides received feedback as voice guidance. For example, it can use Amazon Polly to generate voice instructions for the user.
[0300] Explanation with specific examples
[0301] For example, suppose a user wants to sing a song called "Song B". First, the user selects this song in the karaoke application and starts singing. The built-in microphone of the terminal records the user's singing voice, and the facial expression data is acquired by the built-in camera. The recorded voice data and facial expression data are transmitted to the server in real time.
[0302] On the server side, the voice processing engine performs pitch analysis and pronunciation timing analysis, and the emotion recognition engine analyzes the user's emotion. For example, evaluation results such as "the user's pitch is relatively low", "the rhythm is slow", and "the user is nervous" are obtained, and based on the results, improvement methods such as "next time, try to consciously raise the pitch, keep the rhythm constant, and sing more relaxedly" are generated.
[0303] The generated evaluation results and improvement methods are transmitted to the terminal side in real time and displayed on the user interface of the karaoke application. At the same time, using the voice synthesis engine, advice such as "sing more relaxedly" and "sing this part higher" is provided to the user as a voice guide. In this way, the user can receive real-time feedback on their singing and effectively improve their singing ability.
[0304] Examples of prompt sentences
[0305] "Next time, try to be more conscious of singing in a higher voice."
[0306] "Practice singing in a relaxed state."
[0307] "Please pay attention to keeping the rhythm of this part constant."
[0308] In the above form, according to this invention, the user can receive accurate feedback in real time, and comprehensive improvement considering the user's emotional state becomes possible.
[0309] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0310] Step 1: Audio Recording
[0311] The user launches the karaoke application, selects a song, and begins singing. The device records the user's singing using its built-in microphone. This recorded audio data becomes the input. The recorded audio data is temporarily stored in a buffer. The output is the audio data stored in the buffer.
[0312] Specific actions:
[0313] The device's audio recording function is used to capture the singing voice, which is then saved to memory as audio data.
[0314] Step 2: Acquire facial expression data
[0315] As the user continues to sing, the device's built-in camera captures facial expression data from the user. This captured facial expression data becomes the input. The user's facial features are extracted from the camera data, and the output is facial expression data.
[0316] Specific actions:
[0317] The device's camera function is used to analyze the user's facial expressions and collect facial expression data in real time.
[0318] Step 3: Send audio data to the server
[0319] The terminal sends recorded audio data to the server at regular time slice intervals. This audio data serves as input, and after transmission, it is stored on the server side. The output is the audio data sent to the server.
[0320] Specific actions:
[0321] The device uses asynchronous communication to send voice data to the server as an HTTP request.
[0322] Step 4: Send facial expression data to the server
[0323] The terminal transmits the acquired facial expression data to the server in real time. This facial expression data serves as input, and after transmission, the data is stored on the server side. The output is the facial expression data transmitted to the server.
[0324] Specific actions:
[0325] The device uses asynchronous communication to send facial expression data to the server as an HTTP request.
[0326] Step 5: Voice Analysis
[0327] The server passes the received audio data to the audio processing engine, which performs pitch analysis and speech timing analysis. This audio data serves as input, and the analysis results output as pitch data and timing data.
[0328] Specific actions:
[0329] The server uses a speech processing engine such as the Google Cloud Speech-to-Text API to analyze pitch and timing, and stores the results in a database.
[0330] Step 6: Emotion Recognition
[0331] The server passes voice data and facial expression data to the emotion recognition engine, which analyzes the user's emotions. This voice data and facial expression data serve as input, and the emotion data is output as the analysis result.
[0332] Specific actions:
[0333] The server uses an emotion recognition engine, such as the Microsoft Azure Emotion API, to identify the user's emotional state and stores the results in a database.
[0334] Step 7: Assessment and Problem Identification
[0335] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. This analysis data serves as input, and problems are extracted as part of the evaluation results. The output consists of the evaluation results and the problems identified.
[0336] Specific actions:
[0337] The server uses its internal evaluation logic to identify problems such as low pitch, delayed rhythm, and user tension as text data.
[0338] Step 8: Propose improvement methods
[0339] The server generates appropriate improvement methods based on the problems. This problem data serves as input, and the generated improvement methods are output. The output consists of specific advice.
[0340] Specific actions:
[0341] The server uses an internal suggestion algorithm to generate specific advice such as, "Next time, try to sing in a higher pitch."
[0342] Step 9: Submit evaluation results and improvement suggestions
[0343] The server sends the generated evaluation results and improvement methods to the terminal. These evaluation results and improvement methods become the input, and this data is sent to the terminal. The output is the evaluation results and improvement methods sent to the terminal.
[0344] Specific actions:
[0345] The server sends the evaluation results and improvement suggestions as an HTTP response.
[0346] Step 10: Displaying evaluation results and improvement methods.
[0347] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. These evaluation results and improvement suggestions are the input, and the displayed feedback is the output.
[0348] Specific actions:
[0349] The device parses the received evaluation results and improvement suggestions and displays them in the karaoke app's UI.
[0350] Step 11: Providing audio guides
[0351] The terminal generates audio guidance in real time based on improvement methods received from the server. These improvement methods are the input, and the audio guidance is the output.
[0352] Specific actions:
[0353] The device uses a speech synthesis engine such as Amazon Polly to generate and play voice guidance such as, "Let's sing more relaxed."
[0354] (Application Example 2)
[0355] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0356] Traditional karaoke systems offered limited feedback for users to evaluate their own singing ability, often only providing assessments of pitch and rhythm accuracy. Furthermore, evaluations were often unconsidered, meaning they could be inaccurate if the user was experiencing certain emotions. This made effective feedback and suggestions for improvement difficult. Additionally, the technology for providing real-time feedback was insufficient, causing users to miss opportunities to improve their singing on the spot.
[0357] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording the user's singing, means for transmitting the recorded voice data and facial expression data to the server, means for the server to analyze pitch, vocal timing and emotional state, means for performing evaluation and problem identification based on the analysis results, means for generating evaluation and appropriate improvement methods, and means for providing feedback to the user on the generated evaluation and improvement methods. This enables a detailed evaluation of the user's singing and appropriate feedback according to their emotional state. Furthermore, by providing real-time voice guidance, immediate improvement is possible on the spot, resulting in increased user satisfaction and improved singing ability.
[0358] "Means for recording a user's singing" refers to the function of equipment or software that records a user's voice when they sing in a karaoke box or other musical environment.
[0359] "Means for transmitting recorded audio data and facial expression data to a server" refers to a communication function for transferring recorded audio data and user facial expression data to a server via the internet or an internal network.
[0360] "Means for the server to analyze pitch, vocal timing, and emotional state" refers to a mechanism or algorithm that automatically analyzes the user's pitch, vocal timing, and emotional state based on the audio data and facial expression data received by the server.
[0361] "Means for evaluation and problem identification based on analysis results" refers to the process of evaluating the user's singing based on the analyzed data and identifying problems.
[0362] "Means for generating evaluations and appropriate improvement methods" refers to a function that generates advice and specific suggestions for improvement for the user during their next singing session, based on identified problems.
[0363] "Means for providing feedback to the user on the generated evaluation and improvement methods" refers to a display device, audio output device, or other feedback mechanism for communicating the generated evaluation results and improvement methods to the user.
[0364] "Means of providing real-time voice guidance and feedback" refers to a function that provides real-time voice guidance and feedback to the user while they are singing, offering immediate advice on areas for improvement.
[0365] "A means of saving the generated analysis results to a database and using them as training data in the future" refers to the function of saving and managing the analyzed data in a database for use in subsequent analysis and training.
[0366] This invention is a system that analyzes singing voice and facial expression data when a user sings in a physical store such as a karaoke box, and provides real-time evaluation and feedback.
[0367] Hardware and software configuration
[0368] Hardware configuration
[0369] Camera: Used to capture user facial expression data.
[0370] Microphone: Used to record the user's singing voice.
[0371] Device (smartphone or tablet): This device has the karaoke application installed and is responsible for sending voice and facial expression data to the server.
[0372] Software Configuration
[0373] Karaoke application: Installed on the device, it allows for recording, data transmission, and feedback display.
[0374] Speech processing engine: A server-side engine that analyzes speech data and evaluates pitch and timing of speech.
[0375] Emotion Recognition Engine: A server-side engine that analyzes user facial expression data to recognize their emotional state.
[0376] Evaluation and Feedback Generation Engine: A server-side engine that generates evaluations and improvement methods based on the results of voice and emotion analysis.
[0377] Database: The analysis results are stored and used as training data for the future.
[0378] System operation and data flow
[0379] User-side actions
[0380] The user launches the karaoke application, selects their favorite song, and begins singing. The device's microphone records the singing voice, and the camera captures the user's facial expressions.
[0381] Data transmission
[0382] The combined data (voice data and facial expression data) is transmitted to the server in real time. This is done via wireless communication or the internet.
[0383] Analysis and evaluation
[0384] On the server side, the speech processing engine analyzes pitch and timing of speech, and the emotion recognition engine recognizes the user's emotional state from facial expression data. Based on this, the evaluation and feedback generation engine generates an overall evaluation and improvement methods, which are then returned to the terminal.
[0385] Provide feedback
[0386] The device provides real-time feedback to the user, including evaluation results and improvement suggestions from the server. Voice guidance is also provided, allowing users to immediately become aware of areas for improvement while singing.
[0387] Specific examples and usage methods
[0388] For example, suppose a user sings "song B". The user selects this song in the karaoke application on their device and begins singing. The camera and microphone capture facial and audio data, respectively, and send them to the server in real time. The server analyzes the audio data and evaluates any inaccuracies in pitch or rhythm. In addition, an emotion recognition engine determines the user's emotional state and generates feedback accordingly, such as when the user is nervous. For example, it might provide specific advice like, "Next time, try to be more conscious of your pitch, keep your rhythm consistent, and sing in a relaxed manner."
[0389] Example of a prompt
[0390] "Please tell me what points I should focus on in my next performance and what advice you would like me to give for improvement."
[0391] "What are some specific ways to sing in a relaxed state?"
[0392] This allows users to receive real-time feedback and continuously improve their singing skills. The embodiment for carrying out the invention is designed in this way to make the user's singing experience more effective and satisfying.
[0393] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0394] Step 1:
[0395] The user launches the karaoke application, selects a song, and begins singing. The device (smartphone or tablet) uses its built-in microphone to record the singing voice in real time and uses its camera to capture the user's facial expressions.
[0396] Input: User-selected song, singing voice, facial expression data
[0397] Output: Recorded audio data, captured facial expression data
[0398] Step 2:
[0399] The device sends recorded audio data and captured facial expression data to the server at regular time slices. The transmission is asynchronous, providing continuous feedback to the user while they are singing without delay.
[0400] Input: Recorded audio data, captured facial expression data
[0401] Output: Audio data and facial expression data sent to the server
[0402] Step 3:
[0403] The server analyzes the received audio and facial expression data. First, the audio processing engine analyzes the pitch and timing of speech, and then the emotion recognition engine recognizes the emotional state from the facial expression data.
[0404] Input: Voice data and facial expression data sent to the server.
[0405] Output: Analyzed pitch, vocal timing, and emotional state data
[0406] Step 4:
[0407] The server performs evaluation and problem identification based on the results of voice analysis and emotion recognition. The evaluation and feedback generation engine then performs a comprehensive evaluation based on the analysis results and identifies the user's problems.
[0408] Input: Analyzed pitch, vocal timing, and emotional state data.
[0409] Output: Evaluation results, problems
[0410] Step 5:
[0411] The evaluation and feedback generation engine generates appropriate improvement methods based on the identified problems. The generated evaluation and improvement methods include specific advice provided to the user as points to keep in mind for the next singing performance.
[0412] Input: Evaluation results, problems
[0413] Output: Improvement methods and specific advice
[0414] Step 6:
[0415] The server sends the generated evaluation results and improvement methods to the terminal. The terminal displays the received evaluation results and improvement methods on its user interface and also provides them as audio guidance.
[0416] Input: Improvement methods and specific advice
[0417] Output: Evaluation results and audio guide displayed on the device.
[0418] Step 7:
[0419] The device provides users with real-time voice guidance. For example, it gives voice instructions such as, "Let's sing more relaxed," or "Sing this part a little higher."
[0420] Input: Evaluation results and specific advice
[0421] Output: Audio guide provided to the user
[0422] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0423] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0424] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0425] [Second Embodiment]
[0426] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0427] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0428] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0429] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0430] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0431] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0432] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0433] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0434] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0435] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0436] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0437] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0438] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This allows users to efficiently improve their singing skills.
[0439] Server-side processing
[0440] Receiving audio data
[0441] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[0442] Voice analysis
[0443] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine primarily performs the following two types of analysis:
[0444] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[0445] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[0446] Evaluation and Problem Identification
[0447] The server evaluates the user's singing based on the results of the audio analysis. It quantifies deviations in pitch and timing, and identifies specific problems. For example, it may evaluate things like "the pitch is too low" or "the rhythm is off."
[0448] Suggestions for improvement
[0449] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice maintaining a consistent rhythm." These evaluation results and improvement methods are stored in a database and may be used as training data in the future.
[0450] Submission of evaluation results and improvement methods
[0451] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0452] Terminal-side processing
[0453] Voice recording
[0454] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[0455] Sending audio data to the server
[0456] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[0457] Receiving and displaying evaluation results and improvement methods
[0458] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[0459] Audio guides are available.
[0460] The device provides real-time voice guidance based on improvement suggestions received from the server. For example, it provides users with instructions such as "Sing this part a little higher" or advice such as "Keep the rhythm consistent" as voice guidance.
[0461] Specific example
[0462] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to a server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the analysis results in "the pitch is a little low and the rhythm is off", the server generates advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." The device displays this evaluation result and advice to the user, providing it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[0463] In this way, the system of the present invention provides effective support for users to enjoy karaoke with confidence and improve their singing ability.
[0464] The following describes the processing flow.
[0465] Step 1:
[0466] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[0467] Step 2:
[0468] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, and the audio data is temporarily buffered.
[0469] Step 3:
[0470] The terminal sends buffered audio data to the server in fixed time slices (e.g., every 5 seconds). The transmission is asynchronous, and data is sent while recording.
[0471] Step 4:
[0472] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[0473] Step 5:
[0474] The server identifies the pitch of each sound in the audio data through pitch analysis. It also analyzes the rhythm and timing through speech timing analysis.
[0475] Step 6:
[0476] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[0477] Step 7:
[0478] The server identifies problems based on the evaluation results. For example, it might display specific issues such as "the pitch is too low" or "the rhythm is inaccurate."
[0479] Step 8:
[0480] Based on the problems identified by the server, it generates appropriate improvement methods. For example, it creates specific advice such as, "Next time, be sure to sing in a higher voice."
[0481] Step 9:
[0482] The server saves the evaluation results and improvement methods it generates to a database, making them available as training data for the future.
[0483] Step 10:
[0484] The server sends evaluation results and improvement suggestions to the terminal. This transmission is done in real time, allowing the user to receive immediate feedback during their next singing session.
[0485] Step 11:
[0486] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display "The pitch is too low. Next time, try to hit higher notes."
[0487] Step 12:
[0488] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Raise this part a little."
[0489] Step 13:
[0490] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[0491] Through this series of steps, users can efficiently receive feedback on their singing and suggestions for improvement, thereby boosting their confidence in karaoke.
[0492] (Example 1)
[0493] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0494] A major drawback of conventional karaoke systems is that they do not provide concrete feedback to efficiently improve users' singing skills. Furthermore, they lack mechanisms to identify inappropriate pitch or rhythm and suggest ways to improve them, making it difficult for users to improve themselves.
[0495] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0496] In this invention, the server includes means for using a speech processing engine to analyze pitch and vocalization timing, means for using a data analysis engine to perform evaluation and problem identification based on the analysis results, and means for using a natural language generation engine to generate evaluation and improvement methods. This makes it possible to evaluate the user's singing skills in real time and provide specific improvement methods.
[0497] A "user" is a person who uses this system to sing karaoke.
[0498] "Means of recording" refers to a function that captures the user's voice with a microphone and saves it as digital audio data.
[0499] "Means of sending audio data to a server" refers to the function of uploading recorded audio files to a server via a network.
[0500] A "speech processing engine that analyzes pitch and timing" refers to software or hardware that analyzes speech data and evaluates the accuracy of pitch and timing.
[0501] A "data analysis engine that performs evaluation and problem identification based on analysis results" refers to software or hardware that uses data obtained from a speech processing engine to evaluate singing and identify specific problems.
[0502] A "natural language generation engine for generating evaluation and improvement methods" refers to software or hardware that generates improvement advice for users based on identified problems.
[0503] "Means of providing feedback" refers to a function that communicates the generated evaluation results and improvement methods to the user.
[0504] "Means of providing audio guidance" refers to a function that provides users with real-time audio advice.
[0505] "Means of saving to a database" refers to the function of saving analysis results and evaluation results in digital format, making them available as data for future use.
[0506] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This system provides specific feedback to efficiently improve the user's singing skills. The following describes specific embodiments for carrying out this invention in detail.
[0507] Server Processing
[0508] The server first receives the user's voice data sent from the terminal. Methods such as HTTP POST requests are used to receive the voice data. The received voice data is analyzed using a speech processing engine such as Google Cloud Speech-to-Text. The speech processing engine mainly performs pitch analysis and speech timing analysis. Pitch analysis evaluates whether the voice data is sung at the correct pitch, and speech timing analysis evaluates the accuracy of the rhythm and timing.
[0509] Next, the server uses a data analysis engine to evaluate the user's singing based on the results of the voice analysis and identify problems. The data analysis uses libraries such as Python's Pandas and NumPy. For example, it extracts specific data such as the percentage of pitch deviation or the number of milliseconds the rhythm is delayed, identifying problems like "the pitch is too low" or "the rhythm is off."
[0510] Based on the identified problems, the server uses a natural language generation engine (e.g., OpenAI's GPT-3) to generate improvement methods. The generated evaluation results and improvement methods are stored in a database. This storage is done using a database management system such as MySQL or PostgreSQL.
[0511] Finally, the server sends the evaluation results and improvement suggestions to the terminal in JSON format, providing feedback to the user on points to keep in mind for their next singing session.
[0512] Terminal processing
[0513] The device records the user's singing using its built-in microphone. Android and iOS voice recording APIs are used for recording. The recorded audio data is sent to the server in real time. Upon receiving a response from the server, the evaluation results and suggestions for improvement are displayed in the user interface. Mobile application development frameworks such as React Native and Flutter are used for this.
[0514] Furthermore, the device provides real-time voice guidance based on improvement suggestions received from the server. For example, the device's TTS (text-to-speech) engine is used to provide voice instructions such as "Sing this part a little higher" or advice like "Keep the rhythm consistent." This allows users to receive immediate feedback and improve their singing skills.
[0515] Specific example
[0516] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to the server via an HTTP POST request. The server analyzes the pitch and timing using the Google Cloud Speech-to-Text engine and derives an evaluation result. If the analysis results indicate that "the pitch is a little low and the rhythm is off", the server uses GPT-3 to generate advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." This evaluation result and advice are sent to the device in JSON format, which displays it to the user and simultaneously provides it as an audio guide using the TTS engine.
[0517] Example of a prompt
[0518] "Please generate improvement methods to enhance the user's singing skills. For example, please provide advice for when a user's pitch is low and their rhythm is off."
[0519] As described above, this system evaluates the user's singing in real time and provides specific methods for improvement, thereby effectively supporting users in enjoying karaoke with confidence.
[0520] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0521] Step 1: Audio Recording
[0522] The user launches the karaoke app, selects a song, and begins singing. The device records the user's singing using its built-in microphone. The recording uses the Android or iOS voice recording API. The input is the user's singing, and the output is digital audio data. This audio data is temporarily stored in the device's memory.
[0523] Step 2: Sending audio data to the server
[0524] The terminal sends recorded audio data to the server at regular time slices. This transmission is asynchronous and uses HTTP POST requests. The input is the recorded digital audio data, and the output is the audio data sent to the server. The terminal monitors the HTTP response to confirm that the transmission was successful.
[0525] Step 3: Receiving audio data
[0526] The server receives user voice data transmitted from the terminal. The input is the voice data transmitted from the terminal, and the output is the voice data stored in the server's memory. The server temporarily stores the received data and uses it for subsequent processing.
[0527] Step 4: Voice Analysis
[0528] The server uses the Google Cloud Speech-to-Text engine to analyze the received audio data. It converts the audio data to text, and then performs pitch and timing analysis. The input is the received audio data, and the output is analysis data regarding pitch and timing. Pitch analysis identifies the pitch of each part of the audio data, while timing analysis evaluates the accuracy of rhythm and timing. The analysis results are stored as numerical data on the server.
[0529] Step 5: Assessment and Problem Identification
[0530] The server evaluates the user's singing based on the results of the voice analysis and identifies problems. Using Python data analysis libraries (e.g., Pandas or NumPy), it evaluates the analysis results and quantifies specific problems such as "pitched too low" or "rhythm off." The input is the results of the voice analysis and numerical data, and the output is the evaluation score and data identifying the problems.
[0531] Step 6: Propose improvement methods
[0532] The server generates appropriate improvement methods based on the identified problems. Using a natural language generation engine (e.g., OpenAI's GPT-3), it generates specific advice such as, "Next time, try to focus on using a higher pitch." The input is the evaluation results and identified problems, and the output is the text data of the generated improvement methods. The generated advice is stored in the server's database.
[0533] Step 7: Submit evaluation results and improvement suggestions
[0534] The server sends the generated evaluation results and improvement suggestions to the terminal. It returns JSON data to the terminal as an HTTP response. The input is text data of the evaluation results and improvement suggestions, and the output is the data sent to the terminal. The terminal receives this data in real time.
[0535] Step 8: Displaying evaluation results and improvement methods.
[0536] The terminal receives evaluation results and improvement suggestions from the server and displays them on the user interface. Evaluation scores are displayed in graph format, and specific advice is provided in text format. Input is the evaluation results and improvement suggestions data received from the server, while output is the information displayed on the user interface. Users can use this information to improve their next singing performance.
[0537] Step 9: Providing audio guides
[0538] The terminal provides real-time voice guidance based on improvement suggestions received from the server. Using a TTS (text-to-speech) engine, it communicates instructions to the user via voice, such as "Sing this part a little higher" or "Maintain a consistent rhythm." The input is text data of improvement suggestions, and the output is voice data as voice guidance. Users can listen to this and work on improving their singing.
[0539] (Application Example 1)
[0540] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0541] While existing systems provide real-time evaluation and feedback to help users efficiently improve their singing skills during karaoke sessions, dynamically delivering personalized advertisements based on users' singing habits and analysis results would further enhance its value. However, traditional karaoke systems lack these features and fail to appropriately deliver advertisements that match user interests. This makes it difficult to improve user engagement and maximize advertising revenue.
[0542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0543] In this invention, the server includes means for analyzing pitch and vocal timing, means for evaluating the user's singing based on the analysis results and identifying problems, means for generating evaluation and improvement methods, and means for dynamically generating advertisements based on the evaluation and analysis results. This allows the user to receive personalized advertisements based on the analysis results, along with real-time feedback.
[0544] A "user" refers to an individual who uses a karaoke system to sing.
[0545] "Audio data" refers to data that digitally records the voice of a user singing.
[0546] A "server" refers to a central processing system used for analyzing, evaluating, providing feedback on, and generating advertisements for audio data.
[0547] "Pitch" refers to the element that identifies the height (frequency) of each part in audio data and evaluates whether it is accurate.
[0548] "Vocalization timing" refers to the elements used to evaluate the rhythm and timing of each part of audio data.
[0549] "Analysis results" refers to evaluation data obtained from the analysis of pitch and vocalization timing.
[0550] "Evaluation" refers to data that represents a user's singing skill using numerical values or indicators based on the analysis results.
[0551] "Problem identification" refers to clearly indicating specific areas for improvement in the user's singing based on the evaluation results.
[0552] "Improvement methods" refer to specific practice methods and advice for addressing identified problems.
[0553] "Feedback" refers to providing users with information to help them improve their performance and singing in the future.
[0554] "Advertising" refers to commercial messages that are dynamically generated and provided to users based on user analysis results and evaluations.
[0555] "Dynamically generating" refers to generating ads in real time based on the user's current rating results and analytical data.
[0556] This invention is a system that analyzes a user's singing voice in real time while they are singing karaoke, provides feedback on the evaluation and methods for improvement, and simultaneously delivers personalized advertisements based on the user's analysis results. This allows users to efficiently improve their singing skills and receive advertisements that are tailored to their interests.
[0557] This system mainly consists of the following components:
[0558] 1. Device for voice recording:
[0559] A smartphone or other device is used to record the user's singing voice. The device collects the audio data using its built-in microphone.
[0560] 2. Analysis of pitch and vocal timing:
[0561] The server receives audio data and uses a speech processing engine to analyze pitch and timing. Specifically, it uses Python and a speech analysis library (e.g., LibROSA).
[0562] 3. Evaluation and feedback based on analysis results:
[0563] Based on the analysis results, the server evaluates the user's singing and displays deviations in pitch and timing as specific numerical values. For example, it generates evaluation results such as "the pitch is too low" or "the rhythm is off."
[0564] 4. Propose ways to improve:
[0565] Based on the evaluation results, the server generates specific improvement methods for addressing the user's problems. For example, it provides advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." In addition to the aforementioned voice analysis library, it uses a generative AI model to generate even more detailed advice.
[0566] 5. Generating and delivering personalized ads:
[0567] The server uses an algorithm that dynamically generates advertisements based on user analysis results and evaluation data. The generated advertisements, along with evaluation feedback, are sent to the device and presented to the user. For example, if a user frequently sings cover songs by a particular singer, advertisements for that singer's new songs or live performance information will be displayed.
[0568] The HTTP protocol and RESTful API are used for communication between the server and the terminal, enabling real-time data transmission and reception. Voice data is sent from the terminal to the server at regular intervals, and analysis results and advertisements are fed back to the user.
[0569] Specific example:
[0570] For example, suppose a user sings many songs by their favorite singer. This audio data is recorded using the smartphone's microphone and sent to a server. The server analyzes the pitch and rhythm and evaluates it as "low pitch." Next, the server generates a suggestion for improvement, such as "Next time, try to focus on singing in a higher register," and also generates an advertisement for the singer's new album. All of this information is sent to the user's smartphone and provided as feedback.
[0571] Example of a prompt:
[0572] Use the following input prompt statements for the generative AI model:
[0573] Based on the user's singing voice data and its evaluation results (pitch, timing), generate advertisements tailored to the user's interests. For example, users who frequently sing cover songs by a particular singer should be shown advertisements for that singer's latest album, while users with low pitch should be shown advertisements for vocal training. Output the evaluation results and advertisements in text format.
[0574] This allows users to receive real-time feedback and personalized ads, further enhancing their singing experience.
[0575] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0576] Step 1:
[0577] When the user starts singing along to a karaoke song they've selected, the device uses its built-in microphone to record their voice in real time. The recorded audio data is temporarily stored in a buffer.
[0578] (Input) User's singing voice
[0579] (Output) Buffered audio data
[0580] Step 2:
[0581] The device sends recorded audio data to the server at regular time slices. This allows the server to analyze the data in real time. Because the audio data is transmitted asynchronously, the user can receive continuous feedback without delay.
[0582] (Input) Buffered audio data
[0583] (Output) Audio data sent to the server
[0584] Step 3:
[0585] The server analyzes the received audio data. Using an audio processing engine, it analyzes the pitch (frequency) and utterance timing (rhythm). An audio analysis library (e.g., LibROSA) is used for this analysis. The server generates the results of the pitch and timing analysis and saves them as evaluation data.
[0586] (Input) Audio data sent to the server
[0587] (Output) Analysis results of pitch and vocalization timing
[0588] Step 4:
[0589] The server evaluates the user's singing based on the analysis results. It quantifies deviations in pitch and timing, and identifies specific problems (e.g., pitch is too low, rhythm is off). The evaluation data is stored in a database for each user.
[0590] (Input) Analysis results of pitch and vocalization timing
[0591] (Output) Evaluation results and identification of problems
[0592] Step 5:
[0593] The server generates improvement methods based on the identified problems. Specifically, it might generate advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." A generative AI model may also be used for this process.
[0594] (Input) Evaluation results and problems
[0595] (Output) Specific improvement methods
[0596] Step 6:
[0597] The server generates personalized ads based on analysis results and evaluation data. If a user frequently sings songs by a particular singer, ads suggesting new releases or related products from that singer are dynamically generated.
[0598] (Input) Evaluation results and user singing data
[0599] (Output) Generated advertisement
[0600] Step 7:
[0601] The server sends evaluation results, improvement suggestions, and advertisements to the device. This allows users to receive feedback and advertisements in real time.
[0602] (Input) Evaluation results, improvement methods, generated ads
[0603] (Output) Evaluation results, improvement methods, and advertisements sent to the terminal.
[0604] Step 8:
[0605] The device displays evaluation results, improvement suggestions, and advertisements received from the server in its user interface. This allows users to identify areas for improvement in their singing skills and apply them to their next performance. They can also view personalized advertisements.
[0606] (Input) Evaluation results, improvement methods, advertisements
[0607] (Output) Feedback and advertisements displayed in the user interface
[0608] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0609] This invention relates to a system that recognizes the user's emotions along with their singing voice when they sing karaoke, and adjusts the evaluation and improvement methods based on those emotions. This provides feedback and guidance for improvement that is tailored to the emotional state the user is experiencing, enabling effective improvement of singing ability.
[0610] Server-side processing
[0611] Receiving audio data
[0612] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[0613] Emotion recognition by an emotion engine
[0614] The server uses an emotion engine to recognize the user's emotional state based on voice data, user facial expression data, and other factors. For example, it identifies emotions such as "joy," "sadness," and "tension" from voice tone, speaking speed, and facial expressions.
[0615] Voice analysis
[0616] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine performs the following two analyses:
[0617] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[0618] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[0619] Evaluation and Problem Identification
[0620] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. For example, it evaluates if the pitch is too low, the rhythm is off, or if the user is nervous.
[0621] Suggestions for improvement
[0622] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice singing in a relaxed state." It also provides feedback based on emotional state.
[0623] Submission of evaluation results and improvement methods
[0624] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0625] Terminal-side processing
[0626] Voice recording
[0627] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[0628] Sending audio data to the server
[0629] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[0630] Sending emotional data to the server
[0631] The device also sends emotional data, including facial recognition and voice tone analysis results, to the server. This allows the server to analyze the user's emotional state in real time while they are singing.
[0632] Receiving and displaying evaluation results and improvement methods
[0633] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[0634] Audio guides are available.
[0635] The device provides real-time voice guidance based on improvement instructions received from the server. For example, it might tell the user, "Try singing more relaxed," or "Sing this part a little higher."
[0636] Specific example
[0637] For example, suppose a user sings "song B". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data and emotion data to the server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the emotion engine detects the user's tension, the analysis results may indicate "the pitch is a little low and the rhythm is off" and "the user is tense." Based on these results, advice is generated such as, "Next time, try to sing with a higher pitch, maintain a consistent rhythm, and relax." The device displays this evaluation result and advice to the user and also provides it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[0638] In this way, the system of the present invention provides effective support that allows users to enjoy karaoke with confidence and to constantly be conscious of improving their singing ability.
[0639] The following describes the processing flow.
[0640] Step 1:
[0641] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[0642] Step 2:
[0643] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, buffering both the audio data and the user's facial expressions.
[0644] Step 3:
[0645] The terminal sends buffered audio data and user facial expression data to the server at regular time slices (e.g., every 5 seconds). Transmission is performed asynchronously, with data being sent while recording and recognition are in progress.
[0646] Step 4:
[0647] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[0648] Step 5:
[0649] The server receives facial expression data and passes it to the emotion engine, which analyzes the user's facial features to identify emotions. The emotion engine also considers factors such as voice tone and speaking speed to determine emotions such as "joy," "sadness," and "tension."
[0650] Step 6:
[0651] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[0652] Step 7:
[0653] The server incorporates the user's emotional state into the evaluation based on the results of the emotion engine. For example, if the user is feeling tense, information such as "You need to relax" will be added to the evaluation result.
[0654] Step 8:
[0655] The server identifies pitch and rhythm issues based on the evaluation results and suggests improvement methods based on emotional state. For example, it generates feedback such as "the pitch is too low" or "sing more relaxed."
[0656] Step 9:
[0657] The server saves the evaluation results and improvement methods to a database, making them available as training data for the future.
[0658] Step 10:
[0659] The server sends the evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0660] Step 11:
[0661] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display messages such as, "Your pitch is too low. Next time, try to sing higher notes," or "Try to relax while singing."
[0662] Step 12:
[0663] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Sing this part a little higher," or "Try to relax while you sing."
[0664] Step 13:
[0665] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[0666] Through this series of steps, users can efficiently receive evaluations of their singing and suggestions for improvement, thereby boosting their confidence in karaoke. Furthermore, by utilizing an emotion engine, the system can provide optimal feedback tailored to the user's emotional state.
[0667] (Example 2)
[0668] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0669] Traditional karaoke systems offered limited feedback to help users improve their singing abilities, often focusing only on pitch and rhythm. Furthermore, they failed to reflect emotional states such as tension or joy during singing, making effective improvement guidance tailored to individual user characteristics and circumstances difficult. As a result, users lacked specific advice aligned with their emotional state, hindering efficient improvement of their singing skills.
[0670] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0671] In this invention, the server includes means for recognizing emotions based on voice data and facial expression data, means for providing feedback based on emotional state, and means for performing evaluation and problem identification based on the results of voice analysis. This allows users to receive comprehensive feedback that also takes their own emotional state into account in real time, thereby more effectively supporting the improvement of individual singing abilities.
[0672] "Audio data" refers to data that records the voice of a user singing in digital format.
[0673] "Facial expression data" refers to data obtained by capturing and analyzing the user's facial expressions using a camera.
[0674] "Means for recognizing emotions" refers to a device or software that analyzes and identifies a user's emotional state based on voice data and facial expression data.
[0675] "Means of providing feedback" refers to devices or software that present evaluation results and methods for improvement to users.
[0676] "Pitch" refers to the element that describes the height of the sound produced by the user during singing.
[0677] "Vocalization timing" refers to the timing and rhythm of the user's singing.
[0678] "Speech analysis results" refer to evaluation data of pitch and vocalization timing analyzed by the speech processing engine.
[0679] "Means for evaluation and problem identification" refers to a device or software that evaluates the user's singing based on the analysis results and identifies points that need improvement.
[0680] "Real-time" refers to data transmission and processing that occurs almost instantaneously, within a timeframe where users do not perceive any delay.
[0681] A "database" is a data storage system that systematically stores information and is structured to allow for later retrieval and use.
[0682] "Training data" refers to data that is stored for future use in training algorithms such as machine learning.
[0683] The present invention provides a system that analyzes a user's singing voice and facial expressions while they are singing karaoke, recognizes emotions based on the voice and facial expression data, and offers evaluation and improvement methods. To effectively support the improvement of the user's singing ability, this system utilizes the following specific hardware and software.
[0684] Server-side configuration
[0685] The server includes the following hardware and software configuration:
[0686] Speech processing engine: Performs analysis of speech data. For example, it uses the Google Cloud Speech-to-Text API to perform pitch analysis and speech timing analysis.
[0687] Emotion Recognition Engine: Recognizes user emotions based on voice and facial expression data. For example, it uses the Microsoft Azure Emotion API to identify emotions such as "joy," "sadness," and "tension" from the user's facial expressions and voice tone.
[0688] Database: A data storage system that stores analysis results and emotional data for future use as training data.
[0689] Evaluation Logic: Based on the results of voice analysis and emotion recognition, the system evaluates the user's singing, identifies problems, and generates appropriate feedback.
[0690] Terminal-side configuration
[0691] The device includes the following hardware and software configuration:
[0692] Built-in microphone: Used to record the user's singing.
[0693] Built-in camera: Used to acquire user facial expression data.
[0694] Karaoke application: Software that allows users to select their favorite songs and begin singing.
[0695] Asynchronous communication module: A means of communication for transmitting voice and facial expression data to a server.
[0696] Speech synthesis engine: Provides received feedback as voice guidance. For example, it can use Amazon Polly to generate voice instructions for the user.
[0697] Explanation with specific examples
[0698] For example, suppose a user wants to sing "song B". First, the user selects this song in a karaoke application and begins singing. The device's built-in microphone records the user's singing voice, and facial expression data is captured by the built-in camera. The recorded audio and facial expression data are sent to the server in real time.
[0699] On the server side, the speech processing engine performs pitch and vocal timing analysis, while the emotion recognition engine analyzes the user's emotions. For example, evaluation results such as "the user's pitch is a little low," "the rhythm is off," or "the user is nervous" may be obtained, and based on these results, improvement suggestions such as "next time, be mindful of singing in a higher pitch, keep the rhythm consistent, and sing in a relaxed manner" are generated.
[0700] The generated evaluation results and improvement methods are transmitted to the terminal in real time and displayed on the karaoke application's user interface. Simultaneously, a speech synthesis engine provides the user with voice guidance, offering advice such as "Try to relax more when you sing" or "Try singing this part higher." This allows users to receive real-time feedback on their singing and effectively improve their vocal abilities.
[0701] Example of a prompt
[0702] "Next time, try to focus on singing in a higher pitch."
[0703] "Let's practice singing in a relaxed state."
[0704] "Please be careful to maintain a consistent rhythm in this section."
[0705] In this manner, this invention allows users to receive accurate feedback in real time, enabling comprehensive improvements that take into account their own emotional state.
[0706] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0707] Step 1: Audio Recording
[0708] The user launches the karaoke application, selects a song, and begins singing. The device records the user's singing using its built-in microphone. This recorded audio data becomes the input. The recorded audio data is temporarily stored in a buffer. The output is the audio data stored in the buffer.
[0709] Specific actions:
[0710] The device's audio recording function is used to capture the singing voice, which is then saved to memory as audio data.
[0711] Step 2: Acquire facial expression data
[0712] As the user continues to sing, the device's built-in camera captures facial expression data from the user. This captured facial expression data becomes the input. The user's facial features are extracted from the camera data, and the output is facial expression data.
[0713] Specific actions:
[0714] The device's camera function is used to analyze the user's facial expressions and collect facial expression data in real time.
[0715] Step 3: Send audio data to the server
[0716] The terminal sends recorded audio data to the server at regular time slice intervals. This audio data serves as input, and after transmission, it is stored on the server side. The output is the audio data sent to the server.
[0717] Specific actions:
[0718] The device uses asynchronous communication to send voice data to the server as an HTTP request.
[0719] Step 4: Send facial expression data to the server
[0720] The terminal transmits the acquired facial expression data to the server in real time. This facial expression data serves as input, and after transmission, the data is stored on the server side. The output is the facial expression data transmitted to the server.
[0721] Specific actions:
[0722] The device uses asynchronous communication to send facial expression data to the server as an HTTP request.
[0723] Step 5: Voice Analysis
[0724] The server passes the received audio data to the audio processing engine, which performs pitch analysis and speech timing analysis. This audio data serves as input, and the analysis results output as pitch data and timing data.
[0725] Specific actions:
[0726] The server uses a speech processing engine such as the Google Cloud Speech-to-Text API to analyze pitch and timing, and stores the results in a database.
[0727] Step 6: Emotion Recognition
[0728] The server passes voice data and facial expression data to the emotion recognition engine, which analyzes the user's emotions. This voice data and facial expression data serve as input, and the emotion data is output as the analysis result.
[0729] Specific actions:
[0730] The server uses an emotion recognition engine, such as the Microsoft Azure Emotion API, to identify the user's emotional state and stores the results in a database.
[0731] Step 7: Assessment and Problem Identification
[0732] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. This analysis data serves as input, and problems are extracted as part of the evaluation results. The output consists of the evaluation results and the problems identified.
[0733] Specific actions:
[0734] The server uses its internal evaluation logic to identify problems such as low pitch, delayed rhythm, and user tension as text data.
[0735] Step 8: Propose improvement methods
[0736] The server generates appropriate improvement methods based on the problems. This problem data serves as input, and the generated improvement methods are output. The output consists of specific advice.
[0737] Specific actions:
[0738] The server uses an internal suggestion algorithm to generate specific advice such as, "Next time, try to sing in a higher pitch."
[0739] Step 9: Submit evaluation results and improvement suggestions
[0740] The server sends the generated evaluation results and improvement methods to the terminal. These evaluation results and improvement methods become the input, and this data is sent to the terminal. The output is the evaluation results and improvement methods sent to the terminal.
[0741] Specific actions:
[0742] The server sends the evaluation results and improvement suggestions as an HTTP response.
[0743] Step 10: Displaying evaluation results and improvement methods.
[0744] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. These evaluation results and improvement suggestions are the input, and the displayed feedback is the output.
[0745] Specific actions:
[0746] The device parses the received evaluation results and improvement suggestions and displays them in the karaoke app's UI.
[0747] Step 11: Providing audio guides
[0748] The terminal generates audio guidance in real time based on improvement methods received from the server. These improvement methods are the input, and the audio guidance is the output.
[0749] Specific actions:
[0750] The device uses a speech synthesis engine such as Amazon Polly to generate and play voice guidance such as, "Let's sing more relaxed."
[0751] (Application Example 2)
[0752] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0753] Traditional karaoke systems offered limited feedback for users to evaluate their own singing ability, often only providing assessments of pitch and rhythm accuracy. Furthermore, evaluations were often unconsidered, meaning they could be inaccurate if the user was experiencing certain emotions. This made effective feedback and suggestions for improvement difficult. Additionally, the technology for providing real-time feedback was insufficient, causing users to miss opportunities to improve their singing on the spot.
[0754] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording the user's singing, means for transmitting the recorded voice data and facial expression data to the server, means for the server to analyze pitch, vocal timing and emotional state, means for performing evaluation and problem identification based on the analysis results, means for generating evaluation and appropriate improvement methods, and means for providing feedback to the user on the generated evaluation and improvement methods. This enables a detailed evaluation of the user's singing and appropriate feedback according to their emotional state. Furthermore, by providing real-time voice guidance, immediate improvement is possible on the spot, resulting in increased user satisfaction and improved singing ability.
[0755] "Means for recording a user's singing" refers to the function of equipment or software that records a user's voice when they sing in a karaoke box or other musical environment.
[0756] "Means for transmitting recorded audio data and facial expression data to a server" refers to a communication function for transferring recorded audio data and user facial expression data to a server via the internet or an internal network.
[0757] "Means for the server to analyze pitch, vocal timing, and emotional state" refers to a mechanism or algorithm that automatically analyzes the user's pitch, vocal timing, and emotional state based on the audio data and facial expression data received by the server.
[0758] "Means for evaluation and problem identification based on analysis results" refers to the process of evaluating the user's singing based on the analyzed data and identifying problems.
[0759] "Means for generating evaluations and appropriate improvement methods" refers to a function that generates advice and specific suggestions for improvement for the user during their next singing session, based on identified problems.
[0760] "Means for providing feedback to the user on the generated evaluation and improvement methods" refers to a display device, audio output device, or other feedback mechanism for communicating the generated evaluation results and improvement methods to the user.
[0761] "Means of providing real-time voice guidance and feedback" refers to a function that provides real-time voice guidance and feedback to the user while they are singing, offering immediate advice on areas for improvement.
[0762] "A means of saving the generated analysis results to a database and using them as training data in the future" refers to the function of saving and managing the analyzed data in a database for use in subsequent analysis and training.
[0763] This invention is a system that analyzes singing voice and facial expression data when a user sings in a physical store such as a karaoke box, and provides real-time evaluation and feedback.
[0764] Hardware and software configuration
[0765] Hardware configuration
[0766] Camera: Used to capture user facial expression data.
[0767] Microphone: Used to record the user's singing voice.
[0768] Device (smartphone or tablet): This device has the karaoke application installed and is responsible for sending voice and facial expression data to the server.
[0769] Software Configuration
[0770] Karaoke application: Installed on the device, it allows for recording, data transmission, and feedback display.
[0771] Speech processing engine: A server-side engine that analyzes speech data and evaluates pitch and timing of speech.
[0772] Emotion Recognition Engine: A server-side engine that analyzes user facial expression data to recognize their emotional state.
[0773] Evaluation and Feedback Generation Engine: A server-side engine that generates evaluations and improvement methods based on the results of voice and emotion analysis.
[0774] Database: The analysis results are stored and used as training data for the future.
[0775] System operation and data flow
[0776] User-side actions
[0777] The user launches the karaoke application, selects their favorite song, and begins singing. The device's microphone records the singing voice, and the camera captures the user's facial expressions.
[0778] Data transmission
[0779] The combined data (voice data and facial expression data) is transmitted to the server in real time. This is done via wireless communication or the internet.
[0780] Analysis and evaluation
[0781] On the server side, the speech processing engine analyzes pitch and timing of speech, and the emotion recognition engine recognizes the user's emotional state from facial expression data. Based on this, the evaluation and feedback generation engine generates an overall evaluation and improvement methods, which are then returned to the terminal.
[0782] Provide feedback
[0783] The device provides real-time feedback to the user, including evaluation results and improvement suggestions from the server. Voice guidance is also provided, allowing users to immediately become aware of areas for improvement while singing.
[0784] Specific examples and usage methods
[0785] For example, suppose a user sings "song B". The user selects this song in the karaoke application on their device and begins singing. The camera and microphone capture facial and audio data, respectively, and send them to the server in real time. The server analyzes the audio data and evaluates any inaccuracies in pitch or rhythm. In addition, an emotion recognition engine determines the user's emotional state and generates feedback accordingly, such as when the user is nervous. For example, it might provide specific advice like, "Next time, try to be more conscious of your pitch, keep your rhythm consistent, and sing in a relaxed manner."
[0786] Example of a prompt
[0787] "Please tell me what points I should focus on in my next performance and what advice you would like me to give for improvement."
[0788] "What are some specific ways to sing in a relaxed state?"
[0789] This allows users to receive real-time feedback and continuously improve their singing skills. The embodiment for carrying out the invention is designed in this way to make the user's singing experience more effective and satisfying.
[0790] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0791] Step 1:
[0792] The user launches the karaoke application, selects a song, and begins singing. The device (smartphone or tablet) uses its built-in microphone to record the singing voice in real time and uses its camera to capture the user's facial expressions.
[0793] Input: User-selected song, singing voice, facial expression data
[0794] Output: Recorded audio data, captured facial expression data
[0795] Step 2:
[0796] The device sends recorded audio data and captured facial expression data to the server at regular time slices. The transmission is asynchronous, providing continuous feedback to the user while they are singing without delay.
[0797] Input: Recorded audio data, captured facial expression data
[0798] Output: Audio data and facial expression data sent to the server
[0799] Step 3:
[0800] The server analyzes the received audio and facial expression data. First, the audio processing engine analyzes the pitch and timing of speech, and then the emotion recognition engine recognizes the emotional state from the facial expression data.
[0801] Input: Voice data and facial expression data sent to the server.
[0802] Output: Analyzed pitch, vocal timing, and emotional state data
[0803] Step 4:
[0804] The server performs evaluation and problem identification based on the results of voice analysis and emotion recognition. The evaluation and feedback generation engine then performs a comprehensive evaluation based on the analysis results and identifies the user's problems.
[0805] Input: Analyzed pitch, vocal timing, and emotional state data.
[0806] Output: Evaluation results, problems
[0807] Step 5:
[0808] The evaluation and feedback generation engine generates appropriate improvement methods based on the identified problems. The generated evaluation and improvement methods include specific advice provided to the user as points to keep in mind for the next singing performance.
[0809] Input: Evaluation results, problems
[0810] Output: Improvement methods and specific advice
[0811] Step 6:
[0812] The server sends the generated evaluation results and improvement methods to the terminal. The terminal displays the received evaluation results and improvement methods on its user interface and also provides them as audio guidance.
[0813] Input: Improvement methods and specific advice
[0814] Output: Evaluation results and audio guide displayed on the device.
[0815] Step 7:
[0816] The device provides users with real-time voice guidance. For example, it gives voice instructions such as, "Let's sing more relaxed," or "Sing this part a little higher."
[0817] Input: Evaluation results and specific advice
[0818] Output: Audio guide provided to the user
[0819] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0820] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0821] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0822] [Third Embodiment]
[0823] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0824] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0825] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0826] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0827] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0828] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0829] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0830] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0831] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0832] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0833] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0834] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0835] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This allows users to efficiently improve their singing skills.
[0836] Server-side processing
[0837] Receiving audio data
[0838] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[0839] Voice analysis
[0840] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine primarily performs the following two types of analysis:
[0841] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[0842] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[0843] Evaluation and Problem Identification
[0844] The server evaluates the user's singing based on the results of the audio analysis. It quantifies deviations in pitch and timing, and identifies specific problems. For example, it may evaluate things like "the pitch is too low" or "the rhythm is off."
[0845] Suggestions for improvement
[0846] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice maintaining a consistent rhythm." These evaluation results and improvement methods are stored in a database and may be used as training data in the future.
[0847] Submission of evaluation results and improvement methods
[0848] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[0849] Terminal-side processing
[0850] Voice recording
[0851] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[0852] Sending audio data to the server
[0853] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[0854] Receiving and displaying evaluation results and improvement methods
[0855] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[0856] Audio guides are available.
[0857] The device provides real-time voice guidance based on improvement suggestions received from the server. For example, it provides users with instructions such as "Sing this part a little higher" or advice such as "Keep the rhythm consistent" as voice guidance.
[0858] Specific example
[0859] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to a server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the analysis results in "the pitch is a little low and the rhythm is off", the server generates advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." The device displays this evaluation result and advice to the user, providing it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[0860] In this way, the system of the present invention provides effective support for users to enjoy karaoke with confidence and improve their singing ability.
[0861] The following describes the processing flow.
[0862] Step 1:
[0863] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[0864] Step 2:
[0865] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, and the audio data is temporarily buffered.
[0866] Step 3:
[0867] The terminal sends buffered audio data to the server in fixed time slices (e.g., every 5 seconds). The transmission is asynchronous, and data is sent while recording.
[0868] Step 4:
[0869] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[0870] Step 5:
[0871] The server identifies the pitch of each sound in the audio data through pitch analysis. It also analyzes the rhythm and timing through speech timing analysis.
[0872] Step 6:
[0873] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[0874] Step 7:
[0875] The server identifies problems based on the evaluation results. For example, it might display specific issues such as "the pitch is too low" or "the rhythm is inaccurate."
[0876] Step 8:
[0877] Based on the problems identified by the server, it generates appropriate improvement methods. For example, it creates specific advice such as, "Next time, be sure to sing in a higher voice."
[0878] Step 9:
[0879] The server saves the evaluation results and improvement methods it generates to a database, making them available as training data for the future.
[0880] Step 10:
[0881] The server sends evaluation results and improvement suggestions to the terminal. This transmission is done in real time, allowing the user to receive immediate feedback during their next singing session.
[0882] Step 11:
[0883] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display "The pitch is too low. Next time, try to hit higher notes."
[0884] Step 12:
[0885] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Raise this part a little."
[0886] Step 13:
[0887] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[0888] Through this series of steps, users can efficiently receive feedback on their singing and suggestions for improvement, thereby boosting their confidence in karaoke.
[0889] (Example 1)
[0890] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0891] A major drawback of conventional karaoke systems is that they do not provide concrete feedback to efficiently improve users' singing skills. Furthermore, they lack mechanisms to identify inappropriate pitch or rhythm and suggest ways to improve them, making it difficult for users to improve themselves.
[0892] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0893] In this invention, the server includes means for using a speech processing engine to analyze pitch and vocalization timing, means for using a data analysis engine to perform evaluation and problem identification based on the analysis results, and means for using a natural language generation engine to generate evaluation and improvement methods. This makes it possible to evaluate the user's singing skills in real time and provide specific improvement methods.
[0894] A "user" is a person who uses this system to sing karaoke.
[0895] "Means of recording" refers to a function that captures the user's voice with a microphone and saves it as digital audio data.
[0896] "Means of sending audio data to a server" refers to the function of uploading recorded audio files to a server via a network.
[0897] A "speech processing engine that analyzes pitch and timing" refers to software or hardware that analyzes speech data and evaluates the accuracy of pitch and timing.
[0898] A "data analysis engine that performs evaluation and problem identification based on analysis results" refers to software or hardware that uses data obtained from a speech processing engine to evaluate singing and identify specific problems.
[0899] A "natural language generation engine for generating evaluation and improvement methods" refers to software or hardware that generates improvement advice for users based on identified problems.
[0900] "Means of providing feedback" refers to a function that communicates the generated evaluation results and improvement methods to the user.
[0901] "Means of providing audio guidance" refers to a function that provides users with real-time audio advice.
[0902] "Means of saving to a database" refers to the function of saving analysis results and evaluation results in digital format, making them available as data for future use.
[0903] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This system provides specific feedback to efficiently improve the user's singing skills. The following describes specific embodiments for carrying out this invention in detail.
[0904] Server Processing
[0905] The server first receives the user's voice data sent from the terminal. Methods such as HTTP POST requests are used to receive the voice data. The received voice data is analyzed using a speech processing engine such as Google Cloud Speech-to-Text. The speech processing engine mainly performs pitch analysis and speech timing analysis. Pitch analysis evaluates whether the voice data is sung at the correct pitch, and speech timing analysis evaluates the accuracy of the rhythm and timing.
[0906] Next, the server uses a data analysis engine to evaluate the user's singing based on the results of the voice analysis and identify problems. The data analysis uses libraries such as Python's Pandas and NumPy. For example, it extracts specific data such as the percentage of pitch deviation or the number of milliseconds the rhythm is delayed, identifying problems like "the pitch is too low" or "the rhythm is off."
[0907] Based on the identified problems, the server uses a natural language generation engine (e.g., OpenAI's GPT-3) to generate improvement methods. The generated evaluation results and improvement methods are stored in a database. This storage is done using a database management system such as MySQL or PostgreSQL.
[0908] Finally, the server sends the evaluation results and improvement suggestions to the terminal in JSON format, providing feedback to the user on points to keep in mind for their next singing session.
[0909] Terminal processing
[0910] The device records the user's singing using its built-in microphone. Android and iOS voice recording APIs are used for recording. The recorded audio data is sent to the server in real time. Upon receiving a response from the server, the evaluation results and suggestions for improvement are displayed in the user interface. Mobile application development frameworks such as React Native and Flutter are used for this.
[0911] Furthermore, the device provides real-time voice guidance based on improvement suggestions received from the server. For example, the device's TTS (text-to-speech) engine is used to provide voice instructions such as "Sing this part a little higher" or advice like "Keep the rhythm consistent." This allows users to receive immediate feedback and improve their singing skills.
[0912] Specific example
[0913] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to the server via an HTTP POST request. The server analyzes the pitch and timing using the Google Cloud Speech-to-Text engine and derives an evaluation result. If the analysis results indicate that "the pitch is a little low and the rhythm is off", the server uses GPT-3 to generate advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." This evaluation result and advice are sent to the device in JSON format, which displays it to the user and simultaneously provides it as an audio guide using the TTS engine.
[0914] Example of a prompt
[0915] "Please generate improvement methods to enhance the user's singing skills. For example, please provide advice for when a user's pitch is low and their rhythm is off."
[0916] As described above, this system evaluates the user's singing in real time and provides specific methods for improvement, thereby effectively supporting users in enjoying karaoke with confidence.
[0917] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0918] Step 1: Audio Recording
[0919] The user launches the karaoke app, selects a song, and begins singing. The device records the user's singing using its built-in microphone. The recording uses the Android or iOS voice recording API. The input is the user's singing, and the output is digital audio data. This audio data is temporarily stored in the device's memory.
[0920] Step 2: Sending audio data to the server
[0921] The terminal sends recorded audio data to the server at regular time slices. This transmission is asynchronous and uses HTTP POST requests. The input is the recorded digital audio data, and the output is the audio data sent to the server. The terminal monitors the HTTP response to confirm that the transmission was successful.
[0922] Step 3: Receiving audio data
[0923] The server receives user voice data transmitted from the terminal. The input is the voice data transmitted from the terminal, and the output is the voice data stored in the server's memory. The server temporarily stores the received data and uses it for subsequent processing.
[0924] Step 4: Voice Analysis
[0925] The server uses the Google Cloud Speech-to-Text engine to analyze the received audio data. It converts the audio data to text, and then performs pitch and timing analysis. The input is the received audio data, and the output is analysis data regarding pitch and timing. Pitch analysis identifies the pitch of each part of the audio data, while timing analysis evaluates the accuracy of rhythm and timing. The analysis results are stored as numerical data on the server.
[0926] Step 5: Assessment and Problem Identification
[0927] The server evaluates the user's singing based on the results of the voice analysis and identifies problems. Using Python data analysis libraries (e.g., Pandas or NumPy), it evaluates the analysis results and quantifies specific problems such as "pitched too low" or "rhythm off." The input is the results of the voice analysis and numerical data, and the output is the evaluation score and data identifying the problems.
[0928] Step 6: Propose improvement methods
[0929] The server generates appropriate improvement methods based on the identified problems. Using a natural language generation engine (e.g., OpenAI's GPT-3), it generates specific advice such as, "Next time, try to focus on using a higher pitch." The input is the evaluation results and identified problems, and the output is the text data of the generated improvement methods. The generated advice is stored in the server's database.
[0930] Step 7: Submit evaluation results and improvement suggestions
[0931] The server sends the generated evaluation results and improvement suggestions to the terminal. It returns JSON data to the terminal as an HTTP response. The input is text data of the evaluation results and improvement suggestions, and the output is the data sent to the terminal. The terminal receives this data in real time.
[0932] Step 8: Displaying evaluation results and improvement methods.
[0933] The terminal receives evaluation results and improvement suggestions from the server and displays them on the user interface. Evaluation scores are displayed in graph format, and specific advice is provided in text format. Input is the evaluation results and improvement suggestions data received from the server, while output is the information displayed on the user interface. Users can use this information to improve their next singing performance.
[0934] Step 9: Providing audio guides
[0935] The terminal provides real-time voice guidance based on improvement suggestions received from the server. Using a TTS (text-to-speech) engine, it communicates instructions to the user via voice, such as "Sing this part a little higher" or "Maintain a consistent rhythm." The input is text data of improvement suggestions, and the output is voice data as voice guidance. Users can listen to this and work on improving their singing.
[0936] (Application Example 1)
[0937] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0938] While existing systems provide real-time evaluation and feedback to help users efficiently improve their singing skills during karaoke sessions, dynamically delivering personalized advertisements based on users' singing habits and analysis results would further enhance its value. However, traditional karaoke systems lack these features and fail to appropriately deliver advertisements that match user interests. This makes it difficult to improve user engagement and maximize advertising revenue.
[0939] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0940] In this invention, the server includes means for analyzing pitch and vocal timing, means for evaluating the user's singing based on the analysis results and identifying problems, means for generating evaluation and improvement methods, and means for dynamically generating advertisements based on the evaluation and analysis results. This allows the user to receive personalized advertisements based on the analysis results, along with real-time feedback.
[0941] A "user" refers to an individual who uses a karaoke system to sing.
[0942] "Audio data" refers to data that digitally records the voice of a user singing.
[0943] A "server" refers to a central processing system used for analyzing, evaluating, providing feedback on, and generating advertisements for audio data.
[0944] "Pitch" refers to the element that identifies the height (frequency) of each part in audio data and evaluates whether it is accurate.
[0945] "Vocalization timing" refers to the elements used to evaluate the rhythm and timing of each part of audio data.
[0946] "Analysis results" refers to evaluation data obtained from the analysis of pitch and vocalization timing.
[0947] "Evaluation" refers to data that represents a user's singing skill using numerical values or indicators based on the analysis results.
[0948] "Problem identification" refers to clearly indicating specific areas for improvement in the user's singing based on the evaluation results.
[0949] "Improvement methods" refer to specific practice methods and advice for addressing identified problems.
[0950] "Feedback" refers to providing users with information to help them improve their performance and singing in the future.
[0951] "Advertising" refers to commercial messages that are dynamically generated and provided to users based on user analysis results and evaluations.
[0952] "Dynamically generating" refers to generating ads in real time based on the user's current rating results and analytical data.
[0953] This invention is a system that analyzes a user's singing voice in real time while they are singing karaoke, provides feedback on the evaluation and methods for improvement, and simultaneously delivers personalized advertisements based on the user's analysis results. This allows users to efficiently improve their singing skills and receive advertisements that are tailored to their interests.
[0954] This system mainly consists of the following components:
[0955] 1. Device for voice recording:
[0956] A smartphone or other device is used to record the user's singing voice. The device collects the audio data using its built-in microphone.
[0957] 2. Analysis of pitch and vocal timing:
[0958] The server receives audio data and uses a speech processing engine to analyze pitch and timing. Specifically, it uses Python and a speech analysis library (e.g., LibROSA).
[0959] 3. Evaluation and feedback based on analysis results:
[0960] Based on the analysis results, the server evaluates the user's singing and displays deviations in pitch and timing as specific numerical values. For example, it generates evaluation results such as "the pitch is too low" or "the rhythm is off."
[0961] 4. Propose ways to improve:
[0962] Based on the evaluation results, the server generates specific improvement methods for addressing the user's problems. For example, it provides advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." In addition to the aforementioned voice analysis library, it uses a generative AI model to generate even more detailed advice.
[0963] 5. Generating and delivering personalized ads:
[0964] The server uses an algorithm that dynamically generates advertisements based on user analysis results and evaluation data. The generated advertisements, along with evaluation feedback, are sent to the device and presented to the user. For example, if a user frequently sings cover songs by a particular singer, advertisements for that singer's new songs or live performance information will be displayed.
[0965] The HTTP protocol and RESTful API are used for communication between the server and the terminal, enabling real-time data transmission and reception. Voice data is sent from the terminal to the server at regular intervals, and analysis results and advertisements are fed back to the user.
[0966] Specific example:
[0967] For example, suppose a user sings many songs by their favorite singer. This audio data is recorded using the smartphone's microphone and sent to a server. The server analyzes the pitch and rhythm and evaluates it as "low pitch." Next, the server generates a suggestion for improvement, such as "Next time, try to focus on singing in a higher register," and also generates an advertisement for the singer's new album. All of this information is sent to the user's smartphone and provided as feedback.
[0968] Example of a prompt:
[0969] Use the following input prompt statements for the generative AI model:
[0970] Based on the user's singing voice data and its evaluation results (pitch, timing), generate advertisements tailored to the user's interests. For example, users who frequently sing cover songs by a particular singer should be shown advertisements for that singer's latest album, while users with low pitch should be shown advertisements for vocal training. Output the evaluation results and advertisements in text format.
[0971] This allows users to receive real-time feedback and personalized ads, further enhancing their singing experience.
[0972] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0973] Step 1:
[0974] When the user starts singing along to a karaoke song they've selected, the device uses its built-in microphone to record their voice in real time. The recorded audio data is temporarily stored in a buffer.
[0975] (Input) User's singing voice
[0976] (Output) Buffered audio data
[0977] Step 2:
[0978] The device sends recorded audio data to the server at regular time slices. This allows the server to analyze the data in real time. Because the audio data is transmitted asynchronously, the user can receive continuous feedback without delay.
[0979] (Input) Buffered audio data
[0980] (Output) Audio data sent to the server
[0981] Step 3:
[0982] The server analyzes the received audio data. Using an audio processing engine, it analyzes the pitch (frequency) and utterance timing (rhythm). An audio analysis library (e.g., LibROSA) is used for this analysis. The server generates the results of the pitch and timing analysis and saves them as evaluation data.
[0983] (Input) Audio data sent to the server
[0984] (Output) Analysis results of pitch and vocalization timing
[0985] Step 4:
[0986] The server evaluates the user's singing based on the analysis results. It quantifies deviations in pitch and timing, and identifies specific problems (e.g., pitch is too low, rhythm is off). The evaluation data is stored in a database for each user.
[0987] (Input) Analysis results of pitch and vocalization timing
[0988] (Output) Evaluation results and identification of problems
[0989] Step 5:
[0990] The server generates improvement methods based on the identified problems. Specifically, it might generate advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." A generative AI model may also be used for this process.
[0991] (Input) Evaluation results and problems
[0992] (Output) Specific improvement methods
[0993] Step 6:
[0994] The server generates personalized ads based on analysis results and evaluation data. If a user frequently sings songs by a particular singer, ads suggesting new releases or related products from that singer are dynamically generated.
[0995] (Input) Evaluation results and user singing data
[0996] (Output) Generated advertisement
[0997] Step 7:
[0998] The server sends evaluation results, improvement suggestions, and advertisements to the device. This allows users to receive feedback and advertisements in real time.
[0999] (Input) Evaluation results, improvement methods, generated ads
[1000] (Output) Evaluation results, improvement methods, and advertisements sent to the terminal.
[1001] Step 8:
[1002] The device displays evaluation results, improvement suggestions, and advertisements received from the server in its user interface. This allows users to identify areas for improvement in their singing skills and apply them to their next performance. They can also view personalized advertisements.
[1003] (Input) Evaluation results, improvement methods, advertisements
[1004] (Output) Feedback and advertisements displayed in the user interface
[1005] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1006] This invention relates to a system that recognizes the user's emotions along with their singing voice when they sing karaoke, and adjusts the evaluation and improvement methods based on those emotions. This provides feedback and guidance for improvement that is tailored to the emotional state the user is experiencing, enabling effective improvement of singing ability.
[1007] Server-side processing
[1008] Receiving audio data
[1009] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[1010] Emotion recognition by an emotion engine
[1011] The server uses an emotion engine to recognize the user's emotional state based on voice data, user facial expression data, and other factors. For example, it identifies emotions such as "joy," "sadness," and "tension" from voice tone, speaking speed, and facial expressions.
[1012] Voice analysis
[1013] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine performs the following two analyses:
[1014] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[1015] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[1016] Evaluation and Problem Identification
[1017] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. For example, it evaluates if the pitch is too low, the rhythm is off, or if the user is nervous.
[1018] Suggestions for improvement
[1019] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice singing in a relaxed state." It also provides feedback based on emotional state.
[1020] Submission of evaluation results and improvement methods
[1021] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[1022] Terminal-side processing
[1023] Voice recording
[1024] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[1025] Sending audio data to the server
[1026] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[1027] Sending emotional data to the server
[1028] The device also sends emotional data, including facial recognition and voice tone analysis results, to the server. This allows the server to analyze the user's emotional state in real time while they are singing.
[1029] Receiving and displaying evaluation results and improvement methods
[1030] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[1031] Audio guides are available.
[1032] The device provides real-time voice guidance based on improvement instructions received from the server. For example, it might tell the user, "Try singing more relaxed," or "Sing this part a little higher."
[1033] Specific example
[1034] For example, suppose a user sings "song B". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data and emotion data to the server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the emotion engine detects the user's tension, the analysis results may indicate "the pitch is a little low and the rhythm is off" and "the user is tense." Based on these results, advice is generated such as, "Next time, try to sing with a higher pitch, maintain a consistent rhythm, and relax." The device displays this evaluation result and advice to the user and also provides it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[1035] In this way, the system of the present invention provides effective support that allows users to enjoy karaoke with confidence and to constantly be conscious of improving their singing ability.
[1036] The following describes the processing flow.
[1037] Step 1:
[1038] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[1039] Step 2:
[1040] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, buffering both the audio data and the user's facial expressions.
[1041] Step 3:
[1042] The terminal sends buffered audio data and user facial expression data to the server at regular time slices (e.g., every 5 seconds). Transmission is performed asynchronously, with data being sent while recording and recognition are in progress.
[1043] Step 4:
[1044] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[1045] Step 5:
[1046] The server receives facial expression data and passes it to the emotion engine, which analyzes the user's facial features to identify emotions. The emotion engine also considers factors such as voice tone and speaking speed to determine emotions such as "joy," "sadness," and "tension."
[1047] Step 6:
[1048] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[1049] Step 7:
[1050] The server incorporates the user's emotional state into the evaluation based on the results of the emotion engine. For example, if the user is feeling tense, information such as "You need to relax" will be added to the evaluation result.
[1051] Step 8:
[1052] The server identifies pitch and rhythm issues based on the evaluation results and suggests improvement methods based on emotional state. For example, it generates feedback such as "the pitch is too low" or "sing more relaxed."
[1053] Step 9:
[1054] The server saves the evaluation results and improvement methods to a database, making them available as training data for the future.
[1055] Step 10:
[1056] The server sends the evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[1057] Step 11:
[1058] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display messages such as, "Your pitch is too low. Next time, try to sing higher notes," or "Try to relax while singing."
[1059] Step 12:
[1060] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Sing this part a little higher," or "Try to relax while you sing."
[1061] Step 13:
[1062] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[1063] Through this series of steps, users can efficiently receive evaluations of their singing and suggestions for improvement, thereby boosting their confidence in karaoke. Furthermore, by utilizing an emotion engine, the system can provide optimal feedback tailored to the user's emotional state.
[1064] (Example 2)
[1065] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1066] Traditional karaoke systems offered limited feedback to help users improve their singing abilities, often focusing only on pitch and rhythm. Furthermore, they failed to reflect emotional states such as tension or joy during singing, making effective improvement guidance tailored to individual user characteristics and circumstances difficult. As a result, users lacked specific advice aligned with their emotional state, hindering efficient improvement of their singing skills.
[1067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1068] In this invention, the server includes means for recognizing emotions based on voice data and facial expression data, means for providing feedback based on emotional state, and means for performing evaluation and problem identification based on the results of voice analysis. This allows users to receive comprehensive feedback that also takes their own emotional state into account in real time, thereby more effectively supporting the improvement of individual singing abilities.
[1069] "Audio data" refers to data that records the voice of a user singing in digital format.
[1070] "Facial expression data" refers to data obtained by capturing and analyzing the user's facial expressions using a camera.
[1071] "Means for recognizing emotions" refers to a device or software that analyzes and identifies a user's emotional state based on voice data and facial expression data.
[1072] "Means of providing feedback" refers to devices or software that present evaluation results and methods for improvement to users.
[1073] "Pitch" refers to the element that describes the height of the sound produced by the user during singing.
[1074] "Vocalization timing" refers to the timing and rhythm of the user's singing.
[1075] "Speech analysis results" refer to evaluation data of pitch and vocalization timing analyzed by the speech processing engine.
[1076] "Means for evaluation and problem identification" refers to a device or software that evaluates the user's singing based on the analysis results and identifies points that need improvement.
[1077] "Real-time" refers to data transmission and processing that occurs almost instantaneously, within a timeframe where users do not perceive any delay.
[1078] A "database" is a data storage system that systematically stores information and is structured to allow for later retrieval and use.
[1079] "Training data" refers to data that is stored for future use in training algorithms such as machine learning.
[1080] The present invention provides a system that analyzes a user's singing voice and facial expressions while they are singing karaoke, recognizes emotions based on the voice and facial expression data, and offers evaluation and improvement methods. To effectively support the improvement of the user's singing ability, this system utilizes the following specific hardware and software.
[1081] Server-side configuration
[1082] The server includes the following hardware and software configuration:
[1083] Speech processing engine: Performs analysis of speech data. For example, it uses the Google Cloud Speech-to-Text API to perform pitch analysis and speech timing analysis.
[1084] Emotion Recognition Engine: Recognizes user emotions based on voice and facial expression data. For example, it uses the Microsoft Azure Emotion API to identify emotions such as "joy," "sadness," and "tension" from the user's facial expressions and voice tone.
[1085] Database: A data storage system that stores analysis results and emotional data for future use as training data.
[1086] Evaluation Logic: Based on the results of voice analysis and emotion recognition, the system evaluates the user's singing, identifies problems, and generates appropriate feedback.
[1087] Terminal-side configuration
[1088] The device includes the following hardware and software configuration:
[1089] Built-in microphone: Used to record the user's singing.
[1090] Built-in camera: Used to acquire user facial expression data.
[1091] Karaoke application: Software that allows users to select their favorite songs and begin singing.
[1092] Asynchronous communication module: A means of communication for transmitting voice and facial expression data to a server.
[1093] Speech synthesis engine: Provides received feedback as voice guidance. For example, it can use Amazon Polly to generate voice instructions for the user.
[1094] Explanation with specific examples
[1095] For example, suppose a user wants to sing "song B". First, the user selects this song in a karaoke application and begins singing. The device's built-in microphone records the user's singing voice, and facial expression data is captured by the built-in camera. The recorded audio and facial expression data are sent to the server in real time.
[1096] On the server side, the speech processing engine performs pitch and vocal timing analysis, while the emotion recognition engine analyzes the user's emotions. For example, evaluation results such as "the user's pitch is a little low," "the rhythm is off," or "the user is nervous" may be obtained, and based on these results, improvement suggestions such as "next time, be mindful of singing in a higher pitch, keep the rhythm consistent, and sing in a relaxed manner" are generated.
[1097] The generated evaluation results and improvement methods are transmitted to the terminal in real time and displayed on the karaoke application's user interface. Simultaneously, a speech synthesis engine provides the user with voice guidance, offering advice such as "Try to relax more when you sing" or "Try singing this part higher." This allows users to receive real-time feedback on their singing and effectively improve their vocal abilities.
[1098] Example of a prompt
[1099] "Next time, try to focus on singing in a higher pitch."
[1100] "Let's practice singing in a relaxed state."
[1101] "Please be careful to maintain a consistent rhythm in this section."
[1102] In this manner, this invention allows users to receive accurate feedback in real time, enabling comprehensive improvements that take into account their own emotional state.
[1103] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1104] Step 1: Audio Recording
[1105] The user launches the karaoke application, selects a song, and begins singing. The device records the user's singing using its built-in microphone. This recorded audio data becomes the input. The recorded audio data is temporarily stored in a buffer. The output is the audio data stored in the buffer.
[1106] Specific actions:
[1107] The device's audio recording function is used to capture the singing voice, which is then saved to memory as audio data.
[1108] Step 2: Acquire facial expression data
[1109] As the user continues to sing, the device's built-in camera captures facial expression data from the user. This captured facial expression data becomes the input. The user's facial features are extracted from the camera data, and the output is facial expression data.
[1110] Specific actions:
[1111] The device's camera function is used to analyze the user's facial expressions and collect facial expression data in real time.
[1112] Step 3: Send audio data to the server
[1113] The terminal sends recorded audio data to the server at regular time slice intervals. This audio data serves as input, and after transmission, it is stored on the server side. The output is the audio data sent to the server.
[1114] Specific actions:
[1115] The device uses asynchronous communication to send voice data to the server as an HTTP request.
[1116] Step 4: Send facial expression data to the server
[1117] The terminal transmits the acquired facial expression data to the server in real time. This facial expression data serves as input, and after transmission, the data is stored on the server side. The output is the facial expression data transmitted to the server.
[1118] Specific actions:
[1119] The device uses asynchronous communication to send facial expression data to the server as an HTTP request.
[1120] Step 5: Voice Analysis
[1121] The server passes the received audio data to the audio processing engine, which performs pitch analysis and speech timing analysis. This audio data serves as input, and the analysis results output as pitch data and timing data.
[1122] Specific actions:
[1123] The server uses a speech processing engine such as the Google Cloud Speech-to-Text API to analyze pitch and timing, and stores the results in a database.
[1124] Step 6: Emotion Recognition
[1125] The server passes voice data and facial expression data to the emotion recognition engine, which analyzes the user's emotions. This voice data and facial expression data serve as input, and the emotion data is output as the analysis result.
[1126] Specific actions:
[1127] The server uses an emotion recognition engine, such as the Microsoft Azure Emotion API, to identify the user's emotional state and stores the results in a database.
[1128] Step 7: Assessment and Problem Identification
[1129] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. This analysis data serves as input, and problems are extracted as part of the evaluation results. The output consists of the evaluation results and the problems identified.
[1130] Specific actions:
[1131] The server uses its internal evaluation logic to identify problems such as low pitch, delayed rhythm, and user tension as text data.
[1132] Step 8: Propose improvement methods
[1133] The server generates appropriate improvement methods based on the problems. This problem data serves as input, and the generated improvement methods are output. The output consists of specific advice.
[1134] Specific actions:
[1135] The server uses an internal suggestion algorithm to generate specific advice such as, "Next time, try to sing in a higher pitch."
[1136] Step 9: Submit evaluation results and improvement suggestions
[1137] The server sends the generated evaluation results and improvement methods to the terminal. These evaluation results and improvement methods become the input, and this data is sent to the terminal. The output is the evaluation results and improvement methods sent to the terminal.
[1138] Specific actions:
[1139] The server sends the evaluation results and improvement suggestions as an HTTP response.
[1140] Step 10: Displaying evaluation results and improvement methods.
[1141] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. These evaluation results and improvement suggestions are the input, and the displayed feedback is the output.
[1142] Specific actions:
[1143] The device parses the received evaluation results and improvement suggestions and displays them in the karaoke app's UI.
[1144] Step 11: Providing audio guides
[1145] The terminal generates audio guidance in real time based on improvement methods received from the server. These improvement methods are the input, and the audio guidance is the output.
[1146] Specific actions:
[1147] The device uses a speech synthesis engine such as Amazon Polly to generate and play voice guidance such as, "Let's sing more relaxed."
[1148] (Application Example 2)
[1149] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1150] Traditional karaoke systems offered limited feedback for users to evaluate their own singing ability, often only providing assessments of pitch and rhythm accuracy. Furthermore, evaluations were often unconsidered, meaning they could be inaccurate if the user was experiencing certain emotions. This made effective feedback and suggestions for improvement difficult. Additionally, the technology for providing real-time feedback was insufficient, causing users to miss opportunities to improve their singing on the spot.
[1151] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording the user's singing, means for transmitting the recorded voice data and facial expression data to the server, means for the server to analyze pitch, vocal timing and emotional state, means for performing evaluation and problem identification based on the analysis results, means for generating evaluation and appropriate improvement methods, and means for providing feedback to the user on the generated evaluation and improvement methods. This enables a detailed evaluation of the user's singing and appropriate feedback according to their emotional state. Furthermore, by providing real-time voice guidance, immediate improvement is possible on the spot, resulting in increased user satisfaction and improved singing ability.
[1152] "Means for recording a user's singing" refers to the function of equipment or software that records a user's voice when they sing in a karaoke box or other musical environment.
[1153] "Means for transmitting recorded audio data and facial expression data to a server" refers to a communication function for transferring recorded audio data and user facial expression data to a server via the internet or an internal network.
[1154] "Means for the server to analyze pitch, vocal timing, and emotional state" refers to a mechanism or algorithm that automatically analyzes the user's pitch, vocal timing, and emotional state based on the audio data and facial expression data received by the server.
[1155] "Means for evaluation and problem identification based on analysis results" refers to the process of evaluating the user's singing based on the analyzed data and identifying problems.
[1156] "Means for generating evaluations and appropriate improvement methods" refers to a function that generates advice and specific suggestions for improvement for the user during their next singing session, based on identified problems.
[1157] "Means for providing feedback to the user on the generated evaluation and improvement methods" refers to a display device, audio output device, or other feedback mechanism for communicating the generated evaluation results and improvement methods to the user.
[1158] "Means of providing real-time voice guidance and feedback" refers to a function that provides real-time voice guidance and feedback to the user while they are singing, offering immediate advice on areas for improvement.
[1159] "A means of saving the generated analysis results to a database and using them as training data in the future" refers to the function of saving and managing the analyzed data in a database for use in subsequent analysis and training.
[1160] This invention is a system that analyzes singing voice and facial expression data when a user sings in a physical store such as a karaoke box, and provides real-time evaluation and feedback.
[1161] Hardware and software configuration
[1162] Hardware configuration
[1163] Camera: Used to capture user facial expression data.
[1164] Microphone: Used to record the user's singing voice.
[1165] Device (smartphone or tablet): This device has the karaoke application installed and is responsible for sending voice and facial expression data to the server.
[1166] Software Configuration
[1167] Karaoke application: Installed on the device, it allows for recording, data transmission, and feedback display.
[1168] Speech processing engine: A server-side engine that analyzes speech data and evaluates pitch and timing of speech.
[1169] Emotion Recognition Engine: A server-side engine that analyzes user facial expression data to recognize their emotional state.
[1170] Evaluation and Feedback Generation Engine: A server-side engine that generates evaluations and improvement methods based on the results of voice and emotion analysis.
[1171] Database: The analysis results are stored and used as training data for the future.
[1172] System operation and data flow
[1173] User-side actions
[1174] The user launches the karaoke application, selects their favorite song, and begins singing. The device's microphone records the singing voice, and the camera captures the user's facial expressions.
[1175] Data transmission
[1176] The combined data (voice data and facial expression data) is transmitted to the server in real time. This is done via wireless communication or the internet.
[1177] Analysis and evaluation
[1178] On the server side, the speech processing engine analyzes pitch and timing of speech, and the emotion recognition engine recognizes the user's emotional state from facial expression data. Based on this, the evaluation and feedback generation engine generates an overall evaluation and improvement methods, which are then returned to the terminal.
[1179] Provide feedback
[1180] The device provides real-time feedback to the user, including evaluation results and improvement suggestions from the server. Voice guidance is also provided, allowing users to immediately become aware of areas for improvement while singing.
[1181] Specific examples and usage methods
[1182] For example, suppose a user sings "song B". The user selects this song in the karaoke application on their device and begins singing. The camera and microphone capture facial and audio data, respectively, and send them to the server in real time. The server analyzes the audio data and evaluates any inaccuracies in pitch or rhythm. In addition, an emotion recognition engine determines the user's emotional state and generates feedback accordingly, such as when the user is nervous. For example, it might provide specific advice like, "Next time, try to be more conscious of your pitch, keep your rhythm consistent, and sing in a relaxed manner."
[1183] Example of a prompt
[1184] "Please tell me what points I should focus on in my next performance and what advice you would like me to give for improvement."
[1185] "What are some specific ways to sing in a relaxed state?"
[1186] This allows users to receive real-time feedback and continuously improve their singing skills. The embodiment for carrying out the invention is designed in this way to make the user's singing experience more effective and satisfying.
[1187] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1188] Step 1:
[1189] The user launches the karaoke application, selects a song, and begins singing. The device (smartphone or tablet) uses its built-in microphone to record the singing voice in real time and uses its camera to capture the user's facial expressions.
[1190] Input: User-selected song, singing voice, facial expression data
[1191] Output: Recorded audio data, captured facial expression data
[1192] Step 2:
[1193] The device sends recorded audio data and captured facial expression data to the server at regular time slices. The transmission is asynchronous, providing continuous feedback to the user while they are singing without delay.
[1194] Input: Recorded audio data, captured facial expression data
[1195] Output: Audio data and facial expression data sent to the server
[1196] Step 3:
[1197] The server analyzes the received audio and facial expression data. First, the audio processing engine analyzes the pitch and timing of speech, and then the emotion recognition engine recognizes the emotional state from the facial expression data.
[1198] Input: Voice data and facial expression data sent to the server.
[1199] Output: Analyzed pitch, vocal timing, and emotional state data
[1200] Step 4:
[1201] The server performs evaluation and problem identification based on the results of voice analysis and emotion recognition. The evaluation and feedback generation engine then performs a comprehensive evaluation based on the analysis results and identifies the user's problems.
[1202] Input: Analyzed pitch, vocal timing, and emotional state data.
[1203] Output: Evaluation results, problems
[1204] Step 5:
[1205] The evaluation and feedback generation engine generates appropriate improvement methods based on the identified problems. The generated evaluation and improvement methods include specific advice provided to the user as points to keep in mind for the next singing performance.
[1206] Input: Evaluation results, problems
[1207] Output: Improvement methods and specific advice
[1208] Step 6:
[1209] The server sends the generated evaluation results and improvement methods to the terminal. The terminal displays the received evaluation results and improvement methods on its user interface and also provides them as audio guidance.
[1210] Input: Improvement methods and specific advice
[1211] Output: Evaluation results and audio guide displayed on the device.
[1212] Step 7:
[1213] The device provides users with real-time voice guidance. For example, it gives voice instructions such as, "Let's sing more relaxed," or "Sing this part a little higher."
[1214] Input: Evaluation results and specific advice
[1215] Output: Audio guide provided to the user
[1216] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1217] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1219] [Fourth Embodiment]
[1220] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1221] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1223] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1227] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1228] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1229] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1230] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1231] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1232] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1233] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This allows users to efficiently improve their singing skills.
[1234] Server-side processing
[1235] Receiving audio data
[1236] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[1237] Voice analysis
[1238] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine primarily performs the following two types of analysis:
[1239] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[1240] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[1241] Evaluation and Problem Identification
[1242] The server evaluates the user's singing based on the results of the audio analysis. It quantifies deviations in pitch and timing, and identifies specific problems. For example, it may evaluate things like "the pitch is too low" or "the rhythm is off."
[1243] Suggestions for improvement
[1244] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice maintaining a consistent rhythm." These evaluation results and improvement methods are stored in a database and may be used as training data in the future.
[1245] Submission of evaluation results and improvement methods
[1246] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[1247] Terminal-side processing
[1248] Voice recording
[1249] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[1250] Sending audio data to the server
[1251] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[1252] Receiving and displaying evaluation results and improvement methods
[1253] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[1254] Audio guides are available.
[1255] The device provides real-time voice guidance based on improvement suggestions received from the server. For example, it provides users with instructions such as "Sing this part a little higher" or advice such as "Keep the rhythm consistent" as voice guidance.
[1256] Specific example
[1257] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to a server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the analysis results in "the pitch is a little low and the rhythm is off", the server generates advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." The device displays this evaluation result and advice to the user, providing it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[1258] In this way, the system of the present invention provides effective support for users to enjoy karaoke with confidence and improve their singing ability.
[1259] The following describes the processing flow.
[1260] Step 1:
[1261] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[1262] Step 2:
[1263] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, and the audio data is temporarily buffered.
[1264] Step 3:
[1265] The terminal sends buffered audio data to the server in fixed time slices (e.g., every 5 seconds). The transmission is asynchronous, and data is sent while recording.
[1266] Step 4:
[1267] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[1268] Step 5:
[1269] The server identifies the pitch of each sound in the audio data through pitch analysis. It also analyzes the rhythm and timing through speech timing analysis.
[1270] Step 6:
[1271] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[1272] Step 7:
[1273] The server identifies problems based on the evaluation results. For example, it might display specific issues such as "the pitch is too low" or "the rhythm is inaccurate."
[1274] Step 8:
[1275] Based on the problems identified by the server, it generates appropriate improvement methods. For example, it creates specific advice such as, "Next time, be sure to sing in a higher voice."
[1276] Step 9:
[1277] The server saves the evaluation results and improvement methods it generates to a database, making them available as training data for the future.
[1278] Step 10:
[1279] The server sends evaluation results and improvement suggestions to the terminal. This transmission is done in real time, allowing the user to receive immediate feedback during their next singing session.
[1280] Step 11:
[1281] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display "The pitch is too low. Next time, try to hit higher notes."
[1282] Step 12:
[1283] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Raise this part a little."
[1284] Step 13:
[1285] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[1286] Through this series of steps, users can efficiently receive feedback on their singing and suggestions for improvement, thereby boosting their confidence in karaoke.
[1287] (Example 1)
[1288] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1289] A major drawback of conventional karaoke systems is that they do not provide concrete feedback to efficiently improve users' singing skills. Furthermore, they lack mechanisms to identify inappropriate pitch or rhythm and suggest ways to improve them, making it difficult for users to improve themselves.
[1290] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1291] In this invention, the server includes means for using a speech processing engine to analyze pitch and vocalization timing, means for using a data analysis engine to perform evaluation and problem identification based on the analysis results, and means for using a natural language generation engine to generate evaluation and improvement methods. This makes it possible to evaluate the user's singing skills in real time and provide specific improvement methods.
[1292] A "user" is a person who uses this system to sing karaoke.
[1293] "Means of recording" refers to a function that captures the user's voice with a microphone and saves it as digital audio data.
[1294] "Means of sending audio data to a server" refers to the function of uploading recorded audio files to a server via a network.
[1295] A "speech processing engine that analyzes pitch and timing" refers to software or hardware that analyzes speech data and evaluates the accuracy of pitch and timing.
[1296] A "data analysis engine that performs evaluation and problem identification based on analysis results" refers to software or hardware that uses data obtained from a speech processing engine to evaluate singing and identify specific problems.
[1297] A "natural language generation engine for generating evaluation and improvement methods" refers to software or hardware that generates improvement advice for users based on identified problems.
[1298] "Means of providing feedback" refers to a function that communicates the generated evaluation results and improvement methods to the user.
[1299] "Means of providing audio guidance" refers to a function that provides users with real-time audio advice.
[1300] "Means of saving to a database" refers to the function of saving analysis results and evaluation results in digital format, making them available as data for future use.
[1301] This invention is a system that analyzes a user's singing in real time while they are singing karaoke, evaluates their performance, identifies problems, and provides feedback on how to improve. This system provides specific feedback to efficiently improve the user's singing skills. The following describes specific embodiments for carrying out this invention in detail.
[1302] Server Processing
[1303] The server first receives the user's voice data sent from the terminal. Methods such as HTTP POST requests are used to receive the voice data. The received voice data is analyzed using a speech processing engine such as Google Cloud Speech-to-Text. The speech processing engine mainly performs pitch analysis and speech timing analysis. Pitch analysis evaluates whether the voice data is sung at the correct pitch, and speech timing analysis evaluates the accuracy of the rhythm and timing.
[1304] Next, the server uses a data analysis engine to evaluate the user's singing based on the results of the voice analysis and identify problems. The data analysis uses libraries such as Python's Pandas and NumPy. For example, it extracts specific data such as the percentage of pitch deviation or the number of milliseconds the rhythm is delayed, identifying problems like "the pitch is too low" or "the rhythm is off."
[1305] Based on the identified problems, the server uses a natural language generation engine (e.g., OpenAI's GPT-3) to generate improvement methods. The generated evaluation results and improvement methods are stored in a database. This storage is done using a database management system such as MySQL or PostgreSQL.
[1306] Finally, the server sends the evaluation results and improvement suggestions to the terminal in JSON format, providing feedback to the user on points to keep in mind for their next singing session.
[1307] Terminal processing
[1308] The device records the user's singing using its built-in microphone. Android and iOS voice recording APIs are used for recording. The recorded audio data is sent to the server in real time. Upon receiving a response from the server, the evaluation results and suggestions for improvement are displayed in the user interface. Mobile application development frameworks such as React Native and Flutter are used for this.
[1309] Furthermore, the device provides real-time voice guidance based on improvement suggestions received from the server. For example, the device's TTS (text-to-speech) engine is used to provide voice instructions such as "Sing this part a little higher" or advice like "Keep the rhythm consistent." This allows users to receive immediate feedback and improve their singing skills.
[1310] Specific example
[1311] For example, suppose a user sings "song A". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data to the server via an HTTP POST request. The server analyzes the pitch and timing using the Google Cloud Speech-to-Text engine and derives an evaluation result. If the analysis results indicate that "the pitch is a little low and the rhythm is off", the server uses GPT-3 to generate advice such as "Next time, try to be more conscious of the pitch and keep the rhythm consistent." This evaluation result and advice are sent to the device in JSON format, which displays it to the user and simultaneously provides it as an audio guide using the TTS engine.
[1312] Example of a prompt
[1313] "Please generate improvement methods to enhance the user's singing skills. For example, please provide advice for when a user's pitch is low and their rhythm is off."
[1314] As described above, this system evaluates the user's singing in real time and provides specific methods for improvement, thereby effectively supporting users in enjoying karaoke with confidence.
[1315] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1316] Step 1: Audio Recording
[1317] The user launches the karaoke app, selects a song, and begins singing. The device records the user's singing using its built-in microphone. The recording uses the Android or iOS voice recording API. The input is the user's singing, and the output is digital audio data. This audio data is temporarily stored in the device's memory.
[1318] Step 2: Sending audio data to the server
[1319] The terminal sends recorded audio data to the server at regular time slices. This transmission is asynchronous and uses HTTP POST requests. The input is the recorded digital audio data, and the output is the audio data sent to the server. The terminal monitors the HTTP response to confirm that the transmission was successful.
[1320] Step 3: Receiving audio data
[1321] The server receives user voice data transmitted from the terminal. The input is the voice data transmitted from the terminal, and the output is the voice data stored in the server's memory. The server temporarily stores the received data and uses it for subsequent processing.
[1322] Step 4: Voice Analysis
[1323] The server uses the Google Cloud Speech-to-Text engine to analyze the received audio data. It converts the audio data to text, and then performs pitch and timing analysis. The input is the received audio data, and the output is analysis data regarding pitch and timing. Pitch analysis identifies the pitch of each part of the audio data, while timing analysis evaluates the accuracy of rhythm and timing. The analysis results are stored as numerical data on the server.
[1324] Step 5: Assessment and Problem Identification
[1325] The server evaluates the user's singing based on the results of the voice analysis and identifies problems. Using Python data analysis libraries (e.g., Pandas or NumPy), it evaluates the analysis results and quantifies specific problems such as "pitched too low" or "rhythm off." The input is the results of the voice analysis and numerical data, and the output is the evaluation score and data identifying the problems.
[1326] Step 6: Propose improvement methods
[1327] The server generates appropriate improvement methods based on the identified problems. Using a natural language generation engine (e.g., OpenAI's GPT-3), it generates specific advice such as, "Next time, try to focus on using a higher pitch." The input is the evaluation results and identified problems, and the output is the text data of the generated improvement methods. The generated advice is stored in the server's database.
[1328] Step 7: Submit evaluation results and improvement suggestions
[1329] The server sends the generated evaluation results and improvement suggestions to the terminal. It returns JSON data to the terminal as an HTTP response. The input is text data of the evaluation results and improvement suggestions, and the output is the data sent to the terminal. The terminal receives this data in real time.
[1330] Step 8: Displaying evaluation results and improvement methods.
[1331] The terminal receives evaluation results and improvement suggestions from the server and displays them on the user interface. Evaluation scores are displayed in graph format, and specific advice is provided in text format. Input is the evaluation results and improvement suggestions data received from the server, while output is the information displayed on the user interface. Users can use this information to improve their next singing performance.
[1332] Step 9: Providing audio guides
[1333] The terminal provides real-time voice guidance based on improvement suggestions received from the server. Using a TTS (text-to-speech) engine, it communicates instructions to the user via voice, such as "Sing this part a little higher" or "Maintain a consistent rhythm." The input is text data of improvement suggestions, and the output is voice data as voice guidance. Users can listen to this and work on improving their singing.
[1334] (Application Example 1)
[1335] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1336] While existing systems provide real-time evaluation and feedback to help users efficiently improve their singing skills during karaoke sessions, dynamically delivering personalized advertisements based on users' singing habits and analysis results would further enhance its value. However, traditional karaoke systems lack these features and fail to appropriately deliver advertisements that match user interests. This makes it difficult to improve user engagement and maximize advertising revenue.
[1337] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1338] In this invention, the server includes means for analyzing pitch and vocal timing, means for evaluating the user's singing based on the analysis results and identifying problems, means for generating evaluation and improvement methods, and means for dynamically generating advertisements based on the evaluation and analysis results. This allows the user to receive personalized advertisements based on the analysis results, along with real-time feedback.
[1339] A "user" refers to an individual who uses a karaoke system to sing.
[1340] "Audio data" refers to data that digitally records the voice of a user singing.
[1341] A "server" refers to a central processing system used for analyzing, evaluating, providing feedback on, and generating advertisements for audio data.
[1342] "Pitch" refers to the element that identifies the height (frequency) of each part in audio data and evaluates whether it is accurate.
[1343] "Vocalization timing" refers to the elements used to evaluate the rhythm and timing of each part of audio data.
[1344] "Analysis results" refers to evaluation data obtained from the analysis of pitch and vocalization timing.
[1345] "Evaluation" refers to data that represents a user's singing skill using numerical values or indicators based on the analysis results.
[1346] "Problem identification" refers to clearly indicating specific areas for improvement in the user's singing based on the evaluation results.
[1347] "Improvement methods" refer to specific practice methods and advice for addressing identified problems.
[1348] "Feedback" refers to providing users with information to help them improve their performance and singing in the future.
[1349] "Advertising" refers to commercial messages that are dynamically generated and provided to users based on user analysis results and evaluations.
[1350] "Dynamically generating" refers to generating ads in real time based on the user's current rating results and analytical data.
[1351] This invention is a system that analyzes a user's singing voice in real time while they are singing karaoke, provides feedback on the evaluation and methods for improvement, and simultaneously delivers personalized advertisements based on the user's analysis results. This allows users to efficiently improve their singing skills and receive advertisements that are tailored to their interests.
[1352] This system mainly consists of the following components:
[1353] 1. Device for voice recording:
[1354] A smartphone or other device is used to record the user's singing voice. The device collects the audio data using its built-in microphone.
[1355] 2. Analysis of pitch and vocal timing:
[1356] The server receives audio data and uses a speech processing engine to analyze pitch and timing. Specifically, it uses Python and a speech analysis library (e.g., LibROSA).
[1357] 3. Evaluation and feedback based on analysis results:
[1358] Based on the analysis results, the server evaluates the user's singing and displays deviations in pitch and timing as specific numerical values. For example, it generates evaluation results such as "the pitch is too low" or "the rhythm is off."
[1359] 4. Propose ways to improve:
[1360] Based on the evaluation results, the server generates specific improvement methods for addressing the user's problems. For example, it provides advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." In addition to the aforementioned voice analysis library, it uses a generative AI model to generate even more detailed advice.
[1361] 5. Generating and delivering personalized ads:
[1362] The server uses an algorithm that dynamically generates advertisements based on user analysis results and evaluation data. The generated advertisements, along with evaluation feedback, are sent to the device and presented to the user. For example, if a user frequently sings cover songs by a particular singer, advertisements for that singer's new songs or live performance information will be displayed.
[1363] The HTTP protocol and RESTful API are used for communication between the server and the terminal, enabling real-time data transmission and reception. Voice data is sent from the terminal to the server at regular intervals, and analysis results and advertisements are fed back to the user.
[1364] Specific example:
[1365] For example, suppose a user sings many songs by their favorite singer. This audio data is recorded using the smartphone's microphone and sent to a server. The server analyzes the pitch and rhythm and evaluates it as "low pitch." Next, the server generates a suggestion for improvement, such as "Next time, try to focus on singing in a higher register," and also generates an advertisement for the singer's new album. All of this information is sent to the user's smartphone and provided as feedback.
[1366] Example of a prompt:
[1367] Use the following input prompt statements for the generative AI model:
[1368] Based on the user's singing voice data and its evaluation results (pitch, timing), generate advertisements tailored to the user's interests. For example, users who frequently sing cover songs by a particular singer should be shown advertisements for that singer's latest album, while users with low pitch should be shown advertisements for vocal training. Output the evaluation results and advertisements in text format.
[1369] This allows users to receive real-time feedback and personalized ads, further enhancing their singing experience.
[1370] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1371] Step 1:
[1372] When the user starts singing along to a karaoke song they've selected, the device uses its built-in microphone to record their voice in real time. The recorded audio data is temporarily stored in a buffer.
[1373] (Input) User's singing voice
[1374] (Output) Buffered audio data
[1375] Step 2:
[1376] The device sends recorded audio data to the server at regular time slices. This allows the server to analyze the data in real time. Because the audio data is transmitted asynchronously, the user can receive continuous feedback without delay.
[1377] (Input) Buffered audio data
[1378] (Output) Audio data sent to the server
[1379] Step 3:
[1380] The server analyzes the received audio data. Using an audio processing engine, it analyzes the pitch (frequency) and utterance timing (rhythm). An audio analysis library (e.g., LibROSA) is used for this analysis. The server generates the results of the pitch and timing analysis and saves them as evaluation data.
[1381] (Input) Audio data sent to the server
[1382] (Output) Analysis results of pitch and vocalization timing
[1383] Step 4:
[1384] The server evaluates the user's singing based on the analysis results. It quantifies deviations in pitch and timing, and identifies specific problems (e.g., pitch is too low, rhythm is off). The evaluation data is stored in a database for each user.
[1385] (Input) Analysis results of pitch and vocalization timing
[1386] (Output) Evaluation results and identification of problems
[1387] Step 5:
[1388] The server generates improvement methods based on the identified problems. Specifically, it might generate advice such as, "Next time, try to focus on singing in a higher pitch," or "Practice maintaining a consistent rhythm." A generative AI model may also be used for this process.
[1389] (Input) Evaluation results and problems
[1390] (Output) Specific improvement methods
[1391] Step 6:
[1392] The server generates personalized ads based on analysis results and evaluation data. If a user frequently sings songs by a particular singer, ads suggesting new releases or related products from that singer are dynamically generated.
[1393] (Input) Evaluation results and user singing data
[1394] (Output) Generated advertisement
[1395] Step 7:
[1396] The server sends evaluation results, improvement suggestions, and advertisements to the device. This allows users to receive feedback and advertisements in real time.
[1397] (Input) Evaluation results, improvement methods, generated ads
[1398] (Output) Evaluation results, improvement methods, and advertisements sent to the terminal.
[1399] Step 8:
[1400] The device displays evaluation results, improvement suggestions, and advertisements received from the server in its user interface. This allows users to identify areas for improvement in their singing skills and apply them to their next performance. They can also view personalized advertisements.
[1401] (Input) Evaluation results, improvement methods, advertisements
[1402] (Output) Feedback and advertisements displayed in the user interface
[1403] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1404] This invention relates to a system that recognizes the user's emotions along with their singing voice when they sing karaoke, and adjusts the evaluation and improvement methods based on those emotions. This provides feedback and guidance for improvement that is tailored to the emotional state the user is experiencing, enabling effective improvement of singing ability.
[1405] Server-side processing
[1406] Receiving audio data
[1407] The server receives the user's singing audio data transmitted from the terminal. This audio data is synchronized with the accompaniment of the song selected by the user. The server receives this data and uses it for analysis.
[1408] Emotion recognition by an emotion engine
[1409] The server uses an emotion engine to recognize the user's emotional state based on voice data, user facial expression data, and other factors. For example, it identifies emotions such as "joy," "sadness," and "tension" from voice tone, speaking speed, and facial expressions.
[1410] Voice analysis
[1411] The server uses a dedicated speech processing engine to analyze the received audio data. The speech processing engine performs the following two analyses:
[1412] Pitch analysis: The server identifies the pitch of each part of the audio data and evaluates whether the user's pitch is accurate.
[1413] Vocal Timing Analysis: The server analyzes the user's singing timing and evaluates the accuracy of the rhythm and timing.
[1414] Evaluation and Problem Identification
[1415] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. For example, it evaluates if the pitch is too low, the rhythm is off, or if the user is nervous.
[1416] Suggestions for improvement
[1417] The server generates appropriate improvement methods based on the identified problems. For example, it provides specific advice such as, "Next time, try to focus on singing in a higher register," or "Practice singing in a relaxed state." It also provides feedback based on emotional state.
[1418] Submission of evaluation results and improvement methods
[1419] The server sends the generated evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[1420] Terminal-side processing
[1421] Voice recording
[1422] The user launches the karaoke app, selects a song, and begins singing. The device records the singing using its built-in microphone and buffers the audio data in real time.
[1423] Sending audio data to the server
[1424] The device sends the recorded audio data to the server at regular time slices. This transmission is asynchronous, providing the user with continuous feedback without delay while they are singing.
[1425] Sending emotional data to the server
[1426] The device also sends emotional data, including facial recognition and voice tone analysis results, to the server. This allows the server to analyze the user's emotional state in real time while they are singing.
[1427] Receiving and displaying evaluation results and improvement methods
[1428] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. This allows the user to identify problems with their singing and try to improve in their next performance.
[1429] Audio guides are available.
[1430] The device provides real-time voice guidance based on improvement instructions received from the server. For example, it might tell the user, "Try singing more relaxed," or "Sing this part a little higher."
[1431] Specific example
[1432] For example, suppose a user sings "song B". The user selects this song in the karaoke app on their device and begins singing. The device records the user's singing and sends the audio data and emotion data to the server. The server analyzes the audio data and evaluates whether the user's pitch is accurate and whether their rhythm is appropriate. If the emotion engine detects the user's tension, the analysis results may indicate "the pitch is a little low and the rhythm is off" and "the user is tense." Based on these results, advice is generated such as, "Next time, try to sing with a higher pitch, maintain a consistent rhythm, and relax." The device displays this evaluation result and advice to the user and also provides it as an audio guide. The user uses this as a reference to sing again and gradually improves their singing skills.
[1433] In this way, the system of the present invention provides effective support that allows users to enjoy karaoke with confidence and to constantly be conscious of improving their singing ability.
[1434] The following describes the processing flow.
[1435] Step 1:
[1436] The user launches the karaoke app and selects their favorite song. When the user presses the "play" button, the device starts playing the accompaniment.
[1437] Step 2:
[1438] The device begins recording the user's singing using its built-in microphone. Recording takes place in real time, buffering both the audio data and the user's facial expressions.
[1439] Step 3:
[1440] The terminal sends buffered audio data and user facial expression data to the server at regular time slices (e.g., every 5 seconds). Transmission is performed asynchronously, with data being sent while recording and recognition are in progress.
[1441] Step 4:
[1442] The server passes the received audio data to the audio processing engine. The audio processing engine uses techniques such as FFT to perform pitch analysis and speech timing analysis.
[1443] Step 5:
[1444] The server receives facial expression data and passes it to the emotion engine, which analyzes the user's facial features to identify emotions. The emotion engine also considers factors such as voice tone and speaking speed to determine emotions such as "joy," "sadness," and "tension."
[1445] Step 6:
[1446] The server evaluates the user's singing based on the results of pitch analysis and vocal timing analysis. For example, it evaluates cases where the pitch is too low or the rhythm is off.
[1447] Step 7:
[1448] The server incorporates the user's emotional state into the evaluation based on the results of the emotion engine. For example, if the user is feeling tense, information such as "You need to relax" will be added to the evaluation result.
[1449] Step 8:
[1450] The server identifies pitch and rhythm issues based on the evaluation results and suggests improvement methods based on emotional state. For example, it generates feedback such as "the pitch is too low" or "sing more relaxed."
[1451] Step 9:
[1452] The server saves the evaluation results and improvement methods to a database, making them available as training data for the future.
[1453] Step 10:
[1454] The server sends the evaluation results and improvement suggestions to the terminal. This information indicates points that the user should focus on in their next singing performance.
[1455] Step 11:
[1456] The terminal displays the evaluation results and improvement suggestions received from the server on the user interface. For example, it might display messages such as, "Your pitch is too low. Next time, try to sing higher notes," or "Try to relax while singing."
[1457] Step 12:
[1458] The device provides real-time voice guidance. For example, it gives voice instructions to the user such as, "Sing this part a little higher," or "Try to relax while you sing."
[1459] Step 13:
[1460] Users take feedback into consideration and sing the same song or a different song again. By practicing while being mindful of areas for improvement, users' singing skills improve.
[1461] Through this series of steps, users can efficiently receive evaluations of their singing and suggestions for improvement, thereby boosting their confidence in karaoke. Furthermore, by utilizing an emotion engine, the system can provide optimal feedback tailored to the user's emotional state.
[1462] (Example 2)
[1463] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1464] Traditional karaoke systems offered limited feedback to help users improve their singing abilities, often focusing only on pitch and rhythm. Furthermore, they failed to reflect emotional states such as tension or joy during singing, making effective improvement guidance tailored to individual user characteristics and circumstances difficult. As a result, users lacked specific advice aligned with their emotional state, hindering efficient improvement of their singing skills.
[1465] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1466] In this invention, the server includes means for recognizing emotions based on voice data and facial expression data, means for providing feedback based on emotional state, and means for performing evaluation and problem identification based on the results of voice analysis. This allows users to receive comprehensive feedback that also takes their own emotional state into account in real time, thereby more effectively supporting the improvement of individual singing abilities.
[1467] "Audio data" refers to data that records the voice of a user singing in digital format.
[1468] "Facial expression data" refers to data obtained by capturing and analyzing the user's facial expressions using a camera.
[1469] "Means for recognizing emotions" refers to a device or software that analyzes and identifies a user's emotional state based on voice data and facial expression data.
[1470] "Means of providing feedback" refers to devices or software that present evaluation results and methods for improvement to users.
[1471] "Pitch" refers to the element that describes the height of the sound produced by the user during singing.
[1472] "Vocalization timing" refers to the timing and rhythm of the user's singing.
[1473] "Speech analysis results" refer to evaluation data of pitch and vocalization timing analyzed by the speech processing engine.
[1474] "Means for evaluation and problem identification" refers to a device or software that evaluates the user's singing based on the analysis results and identifies points that need improvement.
[1475] "Real-time" refers to data transmission and processing that occurs almost instantaneously, within a timeframe where users do not perceive any delay.
[1476] A "database" is a data storage system that systematically stores information and is structured to allow for later retrieval and use.
[1477] "Training data" refers to data that is stored for future use in training algorithms such as machine learning.
[1478] The present invention provides a system that analyzes a user's singing voice and facial expressions while they are singing karaoke, recognizes emotions based on the voice and facial expression data, and offers evaluation and improvement methods. To effectively support the improvement of the user's singing ability, this system utilizes the following specific hardware and software.
[1479] Server-side configuration
[1480] The server includes the following hardware and software configuration:
[1481] Speech processing engine: Performs analysis of speech data. For example, it uses the Google Cloud Speech-to-Text API to perform pitch analysis and speech timing analysis.
[1482] Emotion Recognition Engine: Recognizes user emotions based on voice and facial expression data. For example, it uses the Microsoft Azure Emotion API to identify emotions such as "joy," "sadness," and "tension" from the user's facial expressions and voice tone.
[1483] Database: A data storage system that stores analysis results and emotional data for future use as training data.
[1484] Evaluation Logic: Based on the results of voice analysis and emotion recognition, the system evaluates the user's singing, identifies problems, and generates appropriate feedback.
[1485] Terminal-side configuration
[1486] The device includes the following hardware and software configuration:
[1487] Built-in microphone: Used to record the user's singing.
[1488] Built-in camera: Used to acquire user facial expression data.
[1489] Karaoke application: Software that allows users to select their favorite songs and begin singing.
[1490] Asynchronous communication module: A means of communication for transmitting voice and facial expression data to a server.
[1491] Speech synthesis engine: Provides received feedback as voice guidance. For example, it can use Amazon Polly to generate voice instructions for the user.
[1492] Explanation with specific examples
[1493] For example, suppose a user wants to sing "song B". First, the user selects this song in a karaoke application and begins singing. The device's built-in microphone records the user's singing voice, and facial expression data is captured by the built-in camera. The recorded audio and facial expression data are sent to the server in real time.
[1494] On the server side, the speech processing engine performs pitch and vocal timing analysis, while the emotion recognition engine analyzes the user's emotions. For example, evaluation results such as "the user's pitch is a little low," "the rhythm is off," or "the user is nervous" may be obtained, and based on these results, improvement suggestions such as "next time, be mindful of singing in a higher pitch, keep the rhythm consistent, and sing in a relaxed manner" are generated.
[1495] The generated evaluation results and improvement methods are transmitted to the terminal in real time and displayed on the karaoke application's user interface. Simultaneously, a speech synthesis engine provides the user with voice guidance, offering advice such as "Try to relax more when you sing" or "Try singing this part higher." This allows users to receive real-time feedback on their singing and effectively improve their vocal abilities.
[1496] Example of a prompt
[1497] "Next time, try to focus on singing in a higher pitch."
[1498] "Let's practice singing in a relaxed state."
[1499] "Please be careful to maintain a consistent rhythm in this section."
[1500] In this manner, this invention allows users to receive accurate feedback in real time, enabling comprehensive improvements that take into account their own emotional state.
[1501] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1502] Step 1: Audio Recording
[1503] The user launches the karaoke application, selects a song, and begins singing. The device records the user's singing using its built-in microphone. This recorded audio data becomes the input. The recorded audio data is temporarily stored in a buffer. The output is the audio data stored in the buffer.
[1504] Specific actions:
[1505] The device's audio recording function is used to capture the singing voice, which is then saved to memory as audio data.
[1506] Step 2: Acquire facial expression data
[1507] As the user continues to sing, the device's built-in camera captures facial expression data from the user. This captured facial expression data becomes the input. The user's facial features are extracted from the camera data, and the output is facial expression data.
[1508] Specific actions:
[1509] The device's camera function is used to analyze the user's facial expressions and collect facial expression data in real time.
[1510] Step 3: Send audio data to the server
[1511] The terminal sends recorded audio data to the server at regular time slice intervals. This audio data serves as input, and after transmission, it is stored on the server side. The output is the audio data sent to the server.
[1512] Specific actions:
[1513] The device uses asynchronous communication to send voice data to the server as an HTTP request.
[1514] Step 4: Send facial expression data to the server
[1515] The terminal transmits the acquired facial expression data to the server in real time. This facial expression data serves as input, and after transmission, the data is stored on the server side. The output is the facial expression data transmitted to the server.
[1516] Specific actions:
[1517] The device uses asynchronous communication to send facial expression data to the server as an HTTP request.
[1518] Step 5: Voice Analysis
[1519] The server passes the received audio data to the audio processing engine, which performs pitch analysis and speech timing analysis. This audio data serves as input, and the analysis results output as pitch data and timing data.
[1520] Specific actions:
[1521] The server uses a speech processing engine such as the Google Cloud Speech-to-Text API to analyze pitch and timing, and stores the results in a database.
[1522] Step 6: Emotion Recognition
[1523] The server passes voice data and facial expression data to the emotion recognition engine, which analyzes the user's emotions. This voice data and facial expression data serve as input, and the emotion data is output as the analysis result.
[1524] Specific actions:
[1525] The server uses an emotion recognition engine, such as the Microsoft Azure Emotion API, to identify the user's emotional state and stores the results in a database.
[1526] Step 7: Assessment and Problem Identification
[1527] The server evaluates the user's singing based on the results of voice analysis and emotion recognition. This analysis data serves as input, and problems are extracted as part of the evaluation results. The output consists of the evaluation results and the problems identified.
[1528] Specific actions:
[1529] The server uses its internal evaluation logic to identify problems such as low pitch, delayed rhythm, and user tension as text data.
[1530] Step 8: Propose improvement methods
[1531] The server generates appropriate improvement methods based on the problems. This problem data serves as input, and the generated improvement methods are output. The output consists of specific advice.
[1532] Specific actions:
[1533] The server uses an internal suggestion algorithm to generate specific advice such as, "Next time, try to sing in a higher pitch."
[1534] Step 9: Submit evaluation results and improvement suggestions
[1535] The server sends the generated evaluation results and improvement methods to the terminal. These evaluation results and improvement methods become the input, and this data is sent to the terminal. The output is the evaluation results and improvement methods sent to the terminal.
[1536] Specific actions:
[1537] The server sends the evaluation results and improvement suggestions as an HTTP response.
[1538] Step 10: Displaying evaluation results and improvement methods.
[1539] The terminal receives evaluation results and improvement suggestions sent from the server and displays them on the user interface. These evaluation results and improvement suggestions are the input, and the displayed feedback is the output.
[1540] Specific actions:
[1541] The device parses the received evaluation results and improvement suggestions and displays them in the karaoke app's UI.
[1542] Step 11: Providing audio guides
[1543] The terminal generates audio guidance in real time based on improvement methods received from the server. These improvement methods are the input, and the audio guidance is the output.
[1544] Specific actions:
[1545] The device uses a speech synthesis engine such as Amazon Polly to generate and play voice guidance such as, "Let's sing more relaxed."
[1546] (Application Example 2)
[1547] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1548] Traditional karaoke systems offered limited feedback for users to evaluate their own singing ability, often only providing assessments of pitch and rhythm accuracy. Furthermore, evaluations were often unconsidered, meaning they could be inaccurate if the user was experiencing certain emotions. This made effective feedback and suggestions for improvement difficult. Additionally, the technology for providing real-time feedback was insufficient, causing users to miss opportunities to improve their singing on the spot.
[1549] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for recording the user's singing, means for transmitting the recorded voice data and facial expression data to the server, means for the server to analyze pitch, vocal timing and emotional state, means for performing evaluation and problem identification based on the analysis results, means for generating evaluation and appropriate improvement methods, and means for providing feedback to the user on the generated evaluation and improvement methods. This enables a detailed evaluation of the user's singing and appropriate feedback according to their emotional state. Furthermore, by providing real-time voice guidance, immediate improvement is possible on the spot, resulting in increased user satisfaction and improved singing ability.
[1550] "Means for recording a user's singing" refers to the function of equipment or software that records a user's voice when they sing in a karaoke box or other musical environment.
[1551] "Means for transmitting recorded audio data and facial expression data to a server" refers to a communication function for transferring recorded audio data and user facial expression data to a server via the internet or an internal network.
[1552] "Means for the server to analyze pitch, vocal timing, and emotional state" refers to a mechanism or algorithm that automatically analyzes the user's pitch, vocal timing, and emotional state based on the audio data and facial expression data received by the server.
[1553] "Means for evaluation and problem identification based on analysis results" refers to the process of evaluating the user's singing based on the analyzed data and identifying problems.
[1554] "Means for generating evaluations and appropriate improvement methods" refers to a function that generates advice and specific suggestions for improvement for the user during their next singing session, based on identified problems.
[1555] "Means for providing feedback to the user on the generated evaluation and improvement methods" refers to a display device, audio output device, or other feedback mechanism for communicating the generated evaluation results and improvement methods to the user.
[1556] "Means of providing real-time voice guidance and feedback" refers to a function that provides real-time voice guidance and feedback to the user while they are singing, offering immediate advice on areas for improvement.
[1557] "A means of saving the generated analysis results to a database and using them as training data in the future" refers to the function of saving and managing the analyzed data in a database for use in subsequent analysis and training.
[1558] This invention is a system that analyzes singing voice and facial expression data when a user sings in a physical store such as a karaoke box, and provides real-time evaluation and feedback.
[1559] Hardware and software configuration
[1560] Hardware configuration
[1561] Camera: Used to capture user facial expression data.
[1562] Microphone: Used to record the user's singing voice.
[1563] Device (smartphone or tablet): This device has the karaoke application installed and is responsible for sending voice and facial expression data to the server.
[1564] Software Configuration
[1565] Karaoke application: Installed on the device, it allows for recording, data transmission, and feedback display.
[1566] Speech processing engine: A server-side engine that analyzes speech data and evaluates pitch and timing of speech.
[1567] Emotion Recognition Engine: A server-side engine that analyzes user facial expression data to recognize their emotional state.
[1568] Evaluation and Feedback Generation Engine: A server-side engine that generates evaluations and improvement methods based on the results of voice and emotion analysis.
[1569] Database: The analysis results are stored and used as training data for the future.
[1570] System operation and data flow
[1571] User-side actions
[1572] The user launches the karaoke application, selects their favorite song, and begins singing. The device's microphone records the singing voice, and the camera captures the user's facial expressions.
[1573] Data transmission
[1574] The combined data (voice data and facial expression data) is transmitted to the server in real time. This is done via wireless communication or the internet.
[1575] Analysis and evaluation
[1576] On the server side, the speech processing engine analyzes pitch and timing of speech, and the emotion recognition engine recognizes the user's emotional state from facial expression data. Based on this, the evaluation and feedback generation engine generates an overall evaluation and improvement methods, which are then returned to the terminal.
[1577] Provide feedback
[1578] The device provides real-time feedback to the user, including evaluation results and improvement suggestions from the server. Voice guidance is also provided, allowing users to immediately become aware of areas for improvement while singing.
[1579] Specific examples and usage methods
[1580] For example, suppose a user sings "song B". The user selects this song in the karaoke application on their device and begins singing. The camera and microphone capture facial and audio data, respectively, and send them to the server in real time. The server analyzes the audio data and evaluates any inaccuracies in pitch or rhythm. In addition, an emotion recognition engine determines the user's emotional state and generates feedback accordingly, such as when the user is nervous. For example, it might provide specific advice like, "Next time, try to be more conscious of your pitch, keep your rhythm consistent, and sing in a relaxed manner."
[1581] Example of a prompt
[1582] "Please tell me what points I should focus on in my next performance and what advice you would like me to give for improvement."
[1583] "What are some specific ways to sing in a relaxed state?"
[1584] This allows users to receive real-time feedback and continuously improve their singing skills. The embodiment for carrying out the invention is designed in this way to make the user's singing experience more effective and satisfying.
[1585] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1586] Step 1:
[1587] The user launches the karaoke application, selects a song, and begins singing. The device (smartphone or tablet) uses its built-in microphone to record the singing voice in real time and uses its camera to capture the user's facial expressions.
[1588] Input: User-selected song, singing voice, facial expression data
[1589] Output: Recorded audio data, captured facial expression data
[1590] Step 2:
[1591] The device sends recorded audio data and captured facial expression data to the server at regular time slices. The transmission is asynchronous, providing continuous feedback to the user while they are singing without delay.
[1592] Input: Recorded audio data, captured facial expression data
[1593] Output: Audio data and facial expression data sent to the server
[1594] Step 3:
[1595] The server analyzes the received audio and facial expression data. First, the audio processing engine analyzes the pitch and timing of speech, and then the emotion recognition engine recognizes the emotional state from the facial expression data.
[1596] Input: Voice data and facial expression data sent to the server.
[1597] Output: Analyzed pitch, vocal timing, and emotional state data
[1598] Step 4:
[1599] The server performs evaluation and problem identification based on the results of voice analysis and emotion recognition. The evaluation and feedback generation engine then performs a comprehensive evaluation based on the analysis results and identifies the user's problems.
[1600] Input: Analyzed pitch, vocal timing, and emotional state data.
[1601] Output: Evaluation results, problems
[1602] Step 5:
[1603] The evaluation and feedback generation engine generates appropriate improvement methods based on the identified problems. The generated evaluation and improvement methods include specific advice provided to the user as points to keep in mind for the next singing performance.
[1604] Input: Evaluation results, problems
[1605] Output: Improvement methods and specific advice
[1606] Step 6:
[1607] The server sends the generated evaluation results and improvement methods to the terminal. The terminal displays the received evaluation results and improvement methods on its user interface and also provides them as audio guidance.
[1608] Input: Improvement methods and specific advice
[1609] Output: Evaluation results and audio guide displayed on the device.
[1610] Step 7:
[1611] The device provides users with real-time voice guidance. For example, it gives voice instructions such as, "Let's sing more relaxed," or "Sing this part a little higher."
[1612] Input: Evaluation results and specific advice
[1613] Output: Audio guide provided to the user
[1614] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1615] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1616] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1617] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1618] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1619] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1620] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1621] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1622] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1623] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1624] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1625] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1626] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1627] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1628] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1629] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1630] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1631] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1632] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1633] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1634] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1635] The following is further disclosed regarding the embodiments described above.
[1636] (Claim 1)
[1637] A means of recording the user's singing,
[1638] A means of sending recorded audio data to a server,
[1639] The server has means for analyzing pitch and vocalization timing,
[1640] A means of performing evaluation and problem identification based on the analysis results,
[1641] Means for generating evaluation and improvement methods,
[1642] A means of providing feedback to the user on the generated evaluations and improvement methods,
[1643] A system that includes this.
[1644] (Claim 2)
[1645] The system according to claim 1, further comprising means for providing a user with real-time audio guidance.
[1646] (Claim 3)
[1647] The system according to claim 1, further comprising means for storing the analysis results in a database and using them as training data in the future.
[1648] "Example 1"
[1649] (Claim 1)
[1650] A means for users to record their singing,
[1651] A means of sending recorded audio data to a server,
[1652] A means by which the server uses a speech processing engine to analyze pitch and timing of speech,
[1653] A means of using a data analysis engine that performs evaluation and problem identification based on the analysis results,
[1654] Means for using a natural language generation engine to generate evaluation and improvement methods,
[1655] A means of providing feedback to the user on the generated evaluations and improvement methods,
[1656] A system that includes this.
[1657] (Claim 2)
[1658] The system according to claim 1, further comprising means for providing a user with real-time audio guidance.
[1659] (Claim 3)
[1660] The system according to claim 1, further comprising means for storing the analysis results in a database and using them as training data in the future.
[1661] "Application Example 1"
[1662] (Claim 1)
[1663] A means of recording the user's singing,
[1664] A means of sending recorded audio data to a server,
[1665] The server has means for analyzing pitch and vocalization timing,
[1666] A means of performing evaluation and problem identification based on the analysis results,
[1667] Means for generating evaluation and improvement methods,
[1668] A means of providing feedback to the user on the generated evaluations and improvement methods,
[1669] A means for dynamically generating advertisements based on evaluation results and analysis results,
[1670] Means of providing generated advertisements to users,
[1671] A system that includes this.
[1672] (Claim 2)
[1673] The system according to claim 1, further comprising means for providing a user with real-time audio guidance.
[1674] (Claim 3)
[1675] The system according to claim 1, further comprising means for storing the analysis results in a database and using them as training data in the future.
[1676] "Example 2 of combining an emotion engine"
[1677] (Claim 1)
[1678] A means of recording the user's singing,
[1679] A means of sending recorded audio data to a server,
[1680] The server has means for analyzing pitch and vocalization timing,
[1681] A means of performing evaluation and problem identification based on the analysis results,
[1682] Means for generating evaluation and improvement methods,
[1683] A means of providing feedback to the user on the generated evaluations and improvement methods,
[1684] A means of recognizing emotions based on voice data and facial expression data,
[1685] A means of providing feedback based on emotional state,
[1686] A system that includes this.
[1687] (Claim 2)
[1688] The system according to claim 1, further comprising means for providing a user with real-time audio guidance.
[1689] (Claim 3)
[1690] The system according to claim 1, further comprising means for storing analysis results and emotional data in a database and using them as training data in the future.
[1691] "Application example 2 when combining with an emotional engine"
[1692] (Claim 1)
[1693] A means of recording the user's singing,
[1694] A means for transmitting recorded audio data and facial expression data to a server,
[1695] The server has means to analyze pitch, vocal timing, and emotional state,
[1696] A means of performing evaluation and problem identification based on the analysis results,
[1697] Means for generating evaluation and appropriate improvement methods,
[1698] A means of providing feedback to the user on the generated evaluations and improvement methods,
[1699] A system that includes this.
[1700] (Claim 2)
[1701] The system according to claim 1, further comprising means for providing the user with real-time voice guidance and feedback.
[1702] (Claim 3)
[1703] The system according to claim 1, further comprising means for storing the generated analysis results in a database and using them as training data in the future. [Explanation of Symbols]
[1704] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of recording the user's singing, A means of sending recorded audio data to a server, The server has means for analyzing pitch and vocalization timing, A means of performing evaluation and problem identification based on the analysis results, Means for generating evaluation and improvement methods, A means of providing feedback to the user on the generated evaluations and improvement methods, A system that includes this.
2. The system according to claim 1, further comprising means for providing a real-time audio guide to a user.
3. The system according to claim 1, further comprising means for storing the analysis results in a database and using them as training data in the future.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A