System

The system addresses the lack of detailed feedback in karaoke systems by analyzing user voice data and generating ideal singing data for comparison, providing effective practice guidance and skill improvement.

JP2026019842APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121590
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing karaoke systems provide insufficient feedback, making it difficult for users, especially elderly people, to improve their singing skills effectively, as they lack specific and detailed guidance on areas for improvement.

Method used

A system that collects user voice data, analyzes it for characteristics like pitch, rhythm, and vibrato, generates ideal singing data using a generative AI model, compares the user's singing with the ideal data, and provides specific feedback for improvement, allowing users to practice and receive real-time guidance.

Benefits of technology

Enables users to improve their karaoke skills effectively by receiving detailed feedback tailored to their individual singing style, enhancing user satisfaction and skill development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019842000001_ABST
    Figure 2026019842000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: Means for collecting voice data of a user, means for transmitting the collected voice data to a server, means for analyzing the voice data of the user in the server and extracting features such as a pitch, a rhythm, a vibrato, and a long tone, means for generating ideal singing data generated on the basis of an analysis result, means for transmitting the generated singing data to a user terminal, means for comparing singing of the user with the generated singing data, means for generating specific feedback on the basis of a comparison result, means for displaying feedback content to the user, and means for collecting new voice data for practice of the user, the system includes means for transmitting to the server again, and means for generating new scores and improvements based on the re-analysis results.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] To get a high score in karaoke, you need skills such as pitch, rhythm, and vocal technique, but self-practice has its limitations, and it is difficult to know specific areas for improvement. Furthermore, current karaoke systems provide insufficient feedback, making it difficult for users to obtain hints for improving their skills. This situation is particularly challenging for elderly people and karaoke users who want to improve their skills. [Means for solving the problem]

[0005] This invention provides a system including means for collecting a user's voice data and transmitting it to a server, means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data based on the analysis results, means for transmitting the generated singing data to a user terminal, means for comparing the user's singing with the generated singing data, means for generating specific feedback based on the comparison results, means for displaying the feedback to the user, means for collecting new voice data for the user's practice and transmitting it again to the server, and means for generating a new score and areas for improvement based on the reanalysis results. This system allows users to effectively improve their skills while receiving specific and detailed feedback, thereby improving their karaoke scores.

[0006] "Audio data" refers to data in which the user's singing, vocalization, or other sounds are recorded in digital format.

[0007] A "server" is a computer system that analyzes, stores, and generates voice data via a network.

[0008] "Analysis" is the process of analyzing the content of audio data and extracting characteristics such as pitch, rhythm, vibrato, and long tones.

[0009] "Pitch" refers to the pitch of each note when the user sings, and accurate pitch is a factor that affects the karaoke score.

[0010] "Rhythm" refers to the length of notes, tempo, and timing of a user's singing.

[0011] "Vibrato" is a technique in which pitch fluctuates at a regular interval, and is an element that adds emotional expression and depth to singing.

[0012] "Long tone" is a technique for sustaining a constant sound stably for a long period of time.

[0013] "Generation" is the process in which the AI ​​model creates ideal singing data based on the user's voice data.

[0014] "Feedback" refers to specific evaluations of the user's singing and suggestions for improvement.

[0015] The "score" is a numerical evaluation of the user's singing, and is a score provided by the karaoke function. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[0038] Audio data collection

[0039] User

[0040] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[0041] Terminal

[0042] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0043] Analysis of audio data and generation of ideal singing data

[0044] server

[0045] The server analyzes the received audio data. Specifically, it extracts characteristics such as pitch, rhythm, vibrato, and long tones. Based on the analysis results, a generative AI model that has learned the user's singing style generates ideal singing data. This ideal singing data reflects the user's characteristics.

[0046] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0047] Generating and Providing Feedback

[0048] Terminal

[0049] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0050] server

[0051] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results, including specific suggestions and ways to improve on pitch deviations, rhythmic irregularities, and lack of vibrato.

[0052] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[0053] Practice and score improvement support

[0054] User

[0055] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0056] Terminal

[0057] The newly collected audio data is compressed again and sent to the server.

[0058] server

[0059] The server analyzes the new audio data and compares it with the previous feedback. It generates a new score and points for improvement and sends the results to the device. The user can then practice based on this feedback and aim to improve their karaoke score.

[0060] Specific examples

[0061] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, sings again, sends the audio data, and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[0062] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[0063] The processing flow will be explained below.

[0064] Step 1:

[0065] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[0066] Step 2:

[0067] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[0068] Step 3:

[0069] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[0070] Step 4:

[0071] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[0072] Step 5:

[0073] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[0074] Step 6:

[0075] Server: Based on the analysis results, the user's singing data is input into a generative AI model to generate ideal singing data.

[0076] Step 7:

[0077] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[0078] Step 8:

[0079] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[0080] Step 9:

[0081] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[0082] Step 10:

[0083] Server: Based on the analysis results and the differences between the user's singing, the server generates specific feedback in text and video format, including suggestions for improvement such as pitch deviations, rhythmic inconsistencies, and lack of vibrato.

[0084] Step 11:

[0085] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[0086] Step 12:

[0087] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[0088] Step 13:

[0089] Terminal: The newly collected voice data is recompressed and sent to the server.

[0090] Step 14:

[0091] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0092] Step 15:

[0093] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[0094] Through this process, users can continue to practice effectively while receiving detailed feedback, steadily improving their karaoke scores.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] Conventional karaoke practice systems have made it difficult for users to obtain specific feedback to effectively improve their singing skills. Due to low analysis accuracy, they were unable to provide advice tailored to specific singing styles, limiting user satisfaction and actual skill improvement. Furthermore, they did not generate ideal singing data that reflected each user's individual characteristics, making it difficult to use for comparison or practice.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for analyzing the user's voice data and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data using a generative AI model based on the analysis results, and means for comparing the user's singing with the generated singing data and generating specific feedback. This enables advanced analysis of the user's singing data and provides individually optimized feedback, thereby improving skill and increasing satisfaction.

[0100] 1. "User's voice data" refers to the waveform information of voice collected by a user using a terminal, including karaoke singing voices.

[0101] 2. "Means of collection" refers to the function of detecting the user's voice through the built-in microphone or peripheral devices and recording it as digital data.

[0102] 3. "Means for transmitting to a server" refers to the communications means for transferring audio data from a terminal to a server via the Internet.

[0103] 4. "Means of analysis" refers to algorithms or software for extracting singing characteristics such as pitch, rhythm, vibrato, and long tones from the received audio data.

[0104] 5. "Generative AI model" is an artificial intelligence model that learns the user's singing style and generates ideal singing data based on the results.

[0105] 6. "Means for generating ideal singing data" refers to the process of using the analysis results and generative AI model to generate ideal singing data that reflects the user's characteristics.

[0106] 7. "Means for transmitting generated singing data to a user terminal" refers to a communication means for transferring ideal singing data from a server to a user terminal via the Internet.

[0107] 8. "Means for comparing a user's singing with the generated singing data" refers to an algorithm or software that compares a user's actual singing data with the generated ideal singing data and evaluates the differences.

[0108] 9. "Means for generating specific feedback" refers to the process of generating advice based on the comparison results, including areas for improvement such as pitch deviations, rhythmic irregularities, and lack of vibrato.

[0109] 10. "Means for displaying feedback content to the user" refers to the functionality for displaying the generated feedback on the user's device, including text and video formats.

[0110] 11. "Means for collecting new voice data and sending it again to the server" refers to a communication means for collecting the voice of the user singing again and sending it again to the server.

[0111] 12. "Means for generating a new score and improvements based on the results of reanalysis" refers to the process of analyzing the resubmitted voice data, evaluating the user's progress, and generating a new score and improvements.

[0112] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[0113] Audio data collection

[0114] User

[0115] Users launch a dedicated app installed on their smartphone, tablet, or other device. Using the app, they select the karaoke song they want to sing and check the screen that displays the lyrics. At this time, an audio accompaniment is also played.

[0116] Terminal

[0117] As soon as the user starts singing, the device collects audio data through the built-in microphone, temporarily stores the data locally, compresses it, and transmits it over the Internet to a server.

[0118] Analysis of audio data and generation of ideal singing data

[0119] server

[0120] The server analyzes the received audio data. During this process, characteristics such as pitch, rhythm, vibrato, and long tones are extracted. Based on the analysis results, a prompt is input into the generative AI model to generate ideal singing data. Specifically, the following prompt is used:

[0121] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[0122] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0123] Generating and Providing Feedback

[0124] Terminal

[0125] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0126] server

[0127] The server compares the user's singing data with the generated ideal singing data. Based on the comparison results, it generates feedback that includes specific suggestions for improvement, such as pitch deviations, rhythm disturbances, and lack of vibrato. The feedback is created in text and video format and sent to the device.

[0128] Practice and score improvement support

[0129] User

[0130] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0131] Terminal

[0132] The newly collected audio data is compressed again and sent to a server via the Internet.

[0133] server

[0134] The server analyzes the new audio data and compares it with the previous feedback to generate a new score and areas for improvement. These results are then sent back to the device, where the user can receive further feedback and practice.

[0135] Specific examples

[0136] For example, "User A" selects a karaoke song and starts singing. The device collects the audio data, compresses it, and sends it to the server. The server analyzes the audio data and inputs the following prompt sentence into the generative AI model:

[0137] "Based on user A's voice data, generate ideal singing data using the following criteria: pitch, rhythm, vibrato, long tone. Reflect user A's singing style."

[0138] The ideal singing data generated from the model is sent from the server to User A's device, where User A can listen to the singing data. User A can also practice based on specific points for improvement provided by the server. By repeating this process, User A can improve their skills.

[0139] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and generative AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0141] Step 1: Collecting audio data

[0142] User

[0143] The user launches the dedicated app installed on their device and selects the karaoke song they want to sing. The app displays the lyrics on the screen and plays the accompanying audio. At this point, the user can begin singing.

[0144] Input: User singing voice

[0145] Output: Audio data

[0146] Terminal

[0147] The device uses a built-in microphone to collect voice data as soon as the user starts singing. This voice data is temporarily stored in the device as digital data, and then the collected voice data is compressed.

[0148] Input: User singing voice

[0149] Output: Compressed audio data

[0150] Specific actions

[0151] The device automatically activates the microphone when the user starts singing, stores the collected voice data in a buffer, and then applies a compression algorithm to reduce the data size.

[0152] Step 2: Sending audio data to the server

[0153] Terminal

[0154] The compressed audio data is transmitted to a server via the Internet.

[0155] Input: Compressed audio data

[0156] Output: Status of sending to server

[0157] Specific actions

[0158] The device establishes an Internet connection and sends the compressed audio data to the server using the HTTPS protocol.

[0159] Step 3: Analyzing the audio data

[0160] server

[0161] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones.

[0162] Input: Compressed audio data

[0163] Output: Feature extracted data

[0164] Specific actions

[0165] The analysis module on the server decodes the audio data and uses signal processing technology to extract various characteristics as numerical values, allowing for a detailed analysis of the user's singing data.

[0166] Step 4: Generate ideal singing data

[0167] server

[0168] The server generates ideal singing data using a generative AI model based on the extracted feature data, using the following prompt:

[0169] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[0170] Input: feature extraction data, prompt statement

[0171] Output: Ideal singing data

[0172] Specific actions

[0173] The generative AI model receives the prompt sentence and uses the feature extraction data to generate ideal singing data, which reflects the user's singing style and produces highly accurate singing examples.

[0174] Step 5: Send ideal singing data to your device

[0175] server

[0176] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0177] Input: Ideal singing data

[0178] Output: Sending status to user terminal

[0179] Specific actions

[0180] The server establishes an Internet connection and transmits the generated ideal singing data to the user's device using the HTTPS protocol.

[0181] Step 6: Generate feedback

[0182] server

[0183] The server compares the user's singing data with the generated ideal singing data and generates specific feedback, including areas for improvement such as pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[0184] Input: User's singing data, ideal singing data

[0185] Output: Feedback data

[0186] Specific actions

[0187] A comparison algorithm in the server analyzes the user's singing data and the ideal singing data, assesses the differences, and generates feedback based on the results.

[0188] Step 7: Displaying feedback on your device

[0189] Terminal

[0190] The device displays the received feedback to the user. Feedback is provided in text and video formats.

[0191] Input: Feedback data

[0192] Output: Feedback information to the display screen

[0193] Specific actions

[0194] The device analyzes the feedback data and displays it in text and video format on the user interface, allowing the user to continue practicing while checking the feedback.

[0195] (Application example 1)

[0196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0197] Conventional karaoke systems have difficulty providing specific feedback to improve users' singing skills, and lack real-time support to help users improve their singing. This makes it difficult for users to efficiently improve their singing skills.

[0198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0199] In this invention, the server includes means for collecting user voice data, means for transmitting the collected voice data to the server, means for analyzing the user voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data based on the analysis results, means for transmitting the generated singing data to a user terminal, means for comparing the user's singing with the generated singing data, means for generating specific feedback based on the comparison results, means for displaying the feedback to the user, means for collecting new voice data for the user's practice and transmitting it again to the server, means for generating a new score and points for improvement based on the reanalysis results, means for providing real-time support for the user's practice using a device in a physical store, and means for displaying the analysis data and feedback in real time, thereby enabling the user to specifically and effectively improve their singing technique in real time.

[0200] "User's voice data" is voice information collected when a user sings a karaoke song.

[0201] The "means of collection" refers to a device that has the function of recording the voice while singing using a device such as a smartphone or tablet.

[0202] The "transmitting means" is a communication means that has the function of sending the collected voice data to a server via the Internet.

[0203] The "analyzing means" is software that executes an algorithm to analyze the audio data received by the server and extract characteristics such as pitch, rhythm, vibrato, and long tones.

[0204] The "means of generation" is software with an algorithm that creates ideal singing data based on the analysis results.

[0205] The "means for transmitting to the user terminal" is a communication means for returning the generated singing data to the user's device.

[0206] The "comparison means" is software that has the function of comparing the user's singing data with the generated ideal singing data and checking the differences.

[0207] The "means for generating feedback" is software that generates improvements and advice for the user based on the comparison results.

[0208] The "means for displaying the feedback content" is a device having a screen display function for displaying the feedback content as text or video on the user's device.

[0209] The "means for collecting new voice data" is a recording device that collects new voice data when the user sings again.

[0210] The "means for transmitting to the server again" is a communication means for sending newly collected voice data to the server again.

[0211] The "means for generating a new score and improvements based on the reanalysis results" is software that analyzes the newly transmitted audio data, compares it with the previous feedback, and generates a new score and improvements.

[0212] "Physical store devices" are devices such as tablets and smartphones used in physical stores such as karaoke booths.

[0213] The "means for providing real-time support" is a system that uses devices in physical stores to analyze the user's singing in real time and provide immediate feedback.

[0214] The "means for displaying analysis data and feedback in real time" is a device having a screen display function for displaying analysis results and feedback content in real time while the user is singing.

[0215] The system for realizing this invention consists of a series of processes that collects user voice data, transmits the data to a server for analysis, and provides feedback based on the interpreted information. This system uses a smartphone or tablet as the main hardware and a generative AI model as the software.

[0216] 1. Collection of audio data

[0217] A user selects a karaoke song using a device in a physical store (e.g., a tablet installed in a karaoke booth) and starts singing. The device collects the user's singing voice through a built-in microphone and temporarily stores the voice data.

[0218] 2. Data compression and transmission

[0219] The collected audio data is compressed on the device, which then transmits the compressed data to a server over the internet, either via Wi-Fi or 4G / 5G networks.

[0220] 3. Analysis of audio data

[0221] The server analyzes the received audio data and extracts characteristics such as pitch, rhythm, vibrato, and long tones. This analysis uses machine learning algorithms and digital signal processing technology. The generative AI model used learns the user's characteristics and generates ideal singing data based on them.

[0222] 4. Generation and transmission of ideal singing data

[0223] The server generates ideal singing data using the generative AI model, which is then sent to the user's device via the internet.

[0224] 5. Providing Feedback

[0225] The user's device plays the ideal singing data that was sent, allowing the user to compare their singing with the ideal singing. The device also displays specific feedback to help the user understand areas for improvement. The feedback is displayed in text and video format.

[0226] 6. Practice and score improvement support

[0227] The user attempts to sing again based on the provided feedback. The device collects new audio data and again sends it to the server. The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points. This information is again sent to the user's device.

[0228] Specific examples

[0229] For example, User A sings a song of his / her choice in a karaoke booth, and the audio data is sent from the tablet to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again. Through this process, User A can reliably improve his / her technique. An example of a prompt for a generative AI model is: "User A sings a karaoke song of his / her choice, and the audio data is sent from the tablet in the karaoke booth to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again to receive feedback for the next time. Through this process, User A can reliably improve his / her technique and improve his / her karaoke score."

[0230] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0231] Step 1:

[0232] A user selects a karaoke song using a device in a brick-and-mortar store (e.g., a tablet in a karaoke booth), then launches an application on the tablet, which displays the lyrics and plays the accompanying music. As the user begins to sing, the tablet's built-in microphone collects audio data, which is temporarily stored on the device.

[0233] Input: User's voice

[0234] Output: Audio data file

[0235] Specific operation: The built-in microphone collects the user's singing voice as digital audio data and saves it as a file.

[0236] Step 2:

[0237] The device compresses the collected voice data to reduce communication load, and then transmits the compressed voice data to a server via the Internet. During this time, the device uses Wi-Fi or 4G / 5G networks.

[0238] Input: Audio data file

[0239] Output: Compressed audio data

[0240] What it does: The audio data is compressed using a standard audio compression algorithm (e.g. MP3, AAC) and uploaded to the server via the HTTP protocol.

[0241] Step 3:

[0242] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones, and then uses a generative AI model to generate ideal singing data based on the analysis results.

[0243] Input: Compressed audio data

[0244] Output: Ideal singing data

[0245] Specific operation: Extract features from audio data using analysis algorithms and generative AI models to generate ideal singing data. Feature extraction uses audio signal processing techniques (e.g., Fast Fourier Transform).

[0246] Step 4:

[0247] The server transmits the generated ideal singing data to the user's terminal again via the Internet.

[0248] Input: Ideal singing data

[0249] Output: sent to the user's terminal

[0250] Specific operation: The generated singing data is sent to the user's terminal via the HTTP protocol, and the terminal receives the data.

[0251] Step 5:

[0252] The device plays back the received ideal singing data, compares it with the user's singing data, and generates specific feedback based on the analysis results received from the server, which is displayed to the user in real time.

[0253] Input: Ideal singing data, analysis results

[0254] Output: Comparison results, feedback

[0255] Specific operation: Using the audio playback function, the ideal singing data is played back, and a comparison algorithm is used to display the difference between the user's singing data and the ideal singing data. Feedback is also displayed on the screen in text and video format.

[0256] Step 6:

[0257] The user attempts to sing again based on the provided feedback, and the device collects new audio data and sends it back to the server.

[0258] Input: User's new voice

[0259] Output: New audio data file

[0260] Specific operation: New singing voices are collected using the built-in microphone, the data is compressed and sent to the server.

[0261] Step 7:

[0262] The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points, which are then sent back to the user's device.

[0263] Input: New audio data

[0264] Output: New score, improvements

[0265] Specific operation: The analysis algorithm is applied again, compared with the previous data to generate a new score and areas for improvement, and the results are sent to the user's device.

[0266] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0267] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[0268] Audio data collection

[0269] User

[0270] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[0271] Terminal

[0272] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0273] Analysis of audio data and generation of ideal singing data

[0274] server

[0275] The server analyzes the received voice data, specifically extracting characteristics such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotions from the voice data.

[0276] Emotion Engine

[0277] The emotion engine analyzes the voice data, the user's intonation, tempo, etc. to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model and reflected in the generation of ideal singing data.

[0278] server

[0279] The ideal singing data reflects the characteristics and emotional state of the user. The generated ideal singing data is transmitted to the user's terminal via the Internet.

[0280] Generating and Providing Feedback

[0281] Terminal

[0282] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0283] server

[0284] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The server also takes into account the user's emotional information recognized by the emotion engine. For example, the server may provide specific suggestions and suggestions for improvement regarding pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[0285] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[0286] Practice and score improvement support

[0287] User

[0288] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0289] Terminal

[0290] The newly collected audio data is compressed again and sent to the server.

[0291] server

[0292] The server analyzes the new voice data, compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0293] Specific examples

[0294] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes pitch, rhythm, and emotion to generate ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, then sends new audio data again and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[0295] Feedback using emotional information

[0296] For example, if user A is nervous during practice, the emotion engine can detect this tension and suggest breathing techniques or relaxation techniques to help them relax. In this way, feedback that takes into account the user's emotional information allows the user to receive more appropriate advice and effectively improve their skills.

[0297] This system allows users to understand their singing technique in real time and practice based on specific advice. Utilizing advanced analysis technology and an emotion engine, it provides effective feedback tailored to individual singing styles and emotional states, resulting in user satisfaction and improvement of technique.

[0298] The processing flow will be explained below.

[0299] Step 1:

[0300] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[0301] Step 2:

[0302] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[0303] Step 3:

[0304] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[0305] Step 4:

[0306] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[0307] Step 5:

[0308] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[0309] Step 6:

[0310] Server: Uses an emotion engine to recognize the user's emotions from the audio data, for example, identifying emotional states such as tension, joy, or sadness from changes in intonation and tempo.

[0311] Step 7:

[0312] Server: Based on the user's singing data and the recognized emotional information, the server uses a generative AI model to generate ideal singing data. The generated singing data reflects not only the characteristics of pitch and rhythm, but also the user's emotions.

[0313] Step 8:

[0314] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[0315] Step 9:

[0316] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[0317] Step 10:

[0318] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[0319] Step 11:

[0320] Server: Based on the analysis results, the differences between the user's singing voice, and the recognized emotional information, the server generates specific feedback. For example, if a user feels nervous, the server will suggest relaxing breathing techniques, which includes advice tailored to the user's emotions.

[0321] Step 12:

[0322] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[0323] Step 13:

[0324] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[0325] Step 14:

[0326] Terminal: The newly collected voice data is recompressed and sent to the server.

[0327] Step 15:

[0328] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0329] Step 16:

[0330] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[0331] Through this process, users can receive detailed feedback and continue practicing effectively, steadily improving their karaoke scores. Taking the user's emotions into consideration also enables more personalized and effective instruction.

[0332] Example 2

[0333] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0334] While conventional singing instruction systems have the ability to provide feedback on pitch and rhythm, they have difficulty providing personalized advice that takes into account the user's emotional state. As a result, instruction is insufficient for users who are nervous or who need to express their emotions, making it difficult to achieve overall improvement in singing technique. Another problem is that users are unable to receive feedback based on their own singing characteristics or emotional state, which hinders effective practice.

[0335] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0336] In this invention, the server includes a means for collecting user voice data, a means including an emotion engine for recognizing emotion information, and a means for generating specific feedback, thereby enabling personalized feedback and advice that takes into account the user's voice characteristics and emotional state.

[0337] A "user" is an individual who uses the system to receive analysis of singing data and feedback.

[0338] "Audio data" refers to data collected in digital form from the voice of a user singing.

[0339] A "terminal" is a digital device that collects voice data and transmits it to a server.

[0340] A "server" is a remote computing device that analyzes audio data, generates feedback, and recognizes emotional information.

[0341] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotional state from voice data.

[0342] "Ideal singing data" is data that indicates an ideal singing performance that is analyzed based on the user's voice data and generated by a generative AI model.

[0343] "Feedback" is specific advice and information on areas for improvement aimed at improving the user's singing technique, based on the results of analyzing the audio data.

[0344] "Comparison" refers to evaluating the user's singing data and the ideal singing data and analyzing the differences between them.

[0345] A "score" is a numerical value calculated based on specific evaluation criteria for singing data, and quantitatively represents the user's singing performance.

[0346] A "generative AI model" is an artificial intelligence algorithm that generates an ideal singing performance based on the user's singing data.

[0347] A "prompt sentence" is text data containing instructions or questions to be input into a generative AI model.

[0348] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[0349] Hardware and software used

[0350] Device: A digital device such as a smartphone or tablet that has a built-in microphone and can be operated by the user.

[0351] Server: A remote computing device on the cloud

[0352] Emotion Engine: Software for analyzing and recognizing emotions

[0353] Generative AI model: An artificial intelligence algorithm that generates ideal singing data

[0354] Audio data collection

[0355] Users launch a dedicated app installed on their smartphone or tablet and select the karaoke song they want to sing. The app displays lyrics and plays back an audio accompaniment. As soon as the user starts singing, the device uses its built-in microphone to collect audio data. This audio data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0356] Analysis of audio data and generation of ideal singing data

[0357] The server analyzes the received voice data and extracts features such as pitch, rhythm, vibrato, and long tones. It also uses an emotion engine to recognize the user's emotions from the voice data. The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model.

[0358] Example prompt sentence:

[0359] "We have received the audio data of user A's karaoke song. Please extract the characteristics of pitch, rhythm, vibrato, and long tones, analyze the emotions with the emotion engine, and generate ideal singing data."

[0360] The server generates ideal singing data and transmits it again to the user's terminal via the Internet.

[0361] Generating and Providing Feedback

[0362] The device provides the user with the ability to play back the received ideal singing data. The user can compare their own singing with the ideal singing. The device also displays the analysis results and specific feedback from the server, allowing the user to identify areas for improvement. The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. This feedback also includes the user's emotional information recognized by the emotion engine.

[0363] Practice and score improvement support

[0364] The user practices based on the provided feedback and sings the same song again, causing the device to collect audio data. The newly collected audio data is compressed again and sent to the server. The server analyzes the new audio data, compares it with the previous feedback result, calculates a new score, and generates new feedback on areas for improvement.

[0365] Specific examples of operation

[0366] For example, User A chooses a favorite karaoke song, sings it, and sends the audio data from his / her device to the server. The server analyzes pitch, rhythm, vibrato, and emotion to generate ideal singing data. This data is sent to User A's device, where User A listens to the ideal singing and compares it with his / her own singing. User A practices based on feedback from the server, sends new audio data, and receives further feedback. By repeating this process, User A can steadily improve his / her skills and increase his / her karaoke score.

[0367] Additionally, if User A becomes nervous during practice, the emotion engine will detect this and provide advice such as, "Try some breathing techniques to relax." Feedback that utilizes emotion information in this way allows User A to effectively improve their skills.

[0368] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0369] Step 1: Collecting audio data

[0370] Users launch a dedicated karaoke app installed on their smartphone or tablet and select the karaoke song they want to sing. The input is the karaoke song selected by the user. The app displays the lyrics and plays the accompaniment audio.

[0371] The device collects voice data using a built-in microphone as soon as the user starts singing. The input is the user's singing voice, and the output is the voice data temporarily stored in the device.

[0372] The device compresses the collected audio data into a specified format (e.g., MP3 or FLAC). The compressed audio data is sent to a server via the Internet. The input is the compressed audio data, and the output is the audio data sent to the server.

[0373] Step 2: Analyzing the audio data

[0374] The server receives the voice data sent from the terminal. The input is the collected voice data. The data analysis engine in the server extracts features such as pitch, rhythm, vibrato, and long tones. The output is feature information of the voice data.

[0375] The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. The input is the voice data's intonation and tempo information, and the output is the user's emotional information. This emotional information, along with the analysis results, is input into the generative AI model.

[0376] Step 3: Generate ideal singing data

[0377] The server inputs a prompt sentence into the generative AI model based on the feature information and emotional information of the audio data. An example of a prompt sentence is, "We have received audio data of User A's karaoke song. Please extract the features of pitch, rhythm, vibrato, and long tones, and analyze the emotions using the emotion engine to generate ideal singing data." The input is the prompt sentence and the analysis result information, and the generative AI model generates ideal singing data. The output is ideal singing data.

[0378] The server then transmits the generated ideal singing data to the user's terminal via the Internet. The input is the ideal singing data, and the output is the data transmitted to the user's terminal.

[0379] Step 4: Play and compare ideal singing data

[0380] The device provides the function to play back the received ideal singing data. The input is the singing data sent from the server. The user can compare their own singing with the ideal singing. The output is the user's listening state.

[0381] The terminal displays the analysis results and feedback sent from the server. The input is the analysis results and feedback information, allowing the user to identify areas that need improvement. The output is feedback displayed to the user.

[0382] Step 5: Generate feedback

[0383] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The input is the user's singing data and the ideal singing data, and the output is specific feedback. This feedback also includes the user's emotional information recognized by the emotion engine.

[0384] Step 6: Practice and additional feedback

[0385] The user practices based on the provided feedback. They then sing the same song again, causing the device to collect audio data. The input is the feedback information. By collecting new audio data, the output is the collected new audio data.

[0386] The terminal compresses the newly collected voice data and sends it to the server again. The input is the newly collected voice data, and the output is the data sent to the server.

[0387] The server analyzes the new voice data, compares it with the previous feedback result to calculate a new score, and generates feedback on improvements. The input is the new voice data and the previous feedback result, and the output is the new score and feedback information.

[0388] By dividing the specific processing of this system into these steps, it is possible to effectively support the improvement of the user's singing technique.

[0389] (Application example 2)

[0390] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0391] In modern factories, humans and robots are increasingly working together, but there are many challenges in providing instructions on-site and improving work efficiency. In particular, the emotional state of workers can affect work efficiency and robot performance, so there is a need for a means to detect this in real time and take countermeasures. In addition, there is a lack of systems that provide real-time feedback and optimize work procedures.

[0392] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0393] In this invention, the server includes a means for analyzing the user's voice data and detecting the emotional state, a means for optimizing the work procedure based on the user's voice instruction data, and a means for generating an ideal work procedure and providing feedback, thereby making it possible to detect the emotional state of the worker in real time and improve work efficiency.

[0394] "User voice data" refers to voice data when a factory worker gives instructions to a robot.

[0395] "Collection means" refers to the device or process used to acquire and transmit audio data to a server.

[0396] "Analysis of audio data" refers to the process of extracting audio features such as pitch, rhythm, vibrato, and long tones and analyzing the data.

[0397] "Emotional state" refers to the user's emotions inferred from the intonation, tempo, and sound patterns of the user's voice.

[0398] "Ideal singing data" refers to voice data of optimal singing that a user should aim for, which is generated based on analyzed voice data and emotional state.

[0399] "Feedback" refers to specific advice and guidance information provided to the user based on the analysis results.

[0400] "Work procedure optimization" refers to the process of deriving the most efficient work procedure by taking into account the user's voice instructions and emotional state.

[0401] A "generative AI model" refers to an algorithm or system that analyzes voice data and emotional data to generate ideal data and feedback.

[0402] A "prompt" refers to an instruction or question input to a generative AI model.

[0403] To put the present invention into practice, a system is constructed in which users, terminals, and servers work in cooperation with each other. Specific embodiments are as follows.

[0404] Audio data collection

[0405] The user launches a dedicated application installed on a device such as a smartphone or tablet installed in the factory. Work instructions and confirmations are entered into the application by voice, and the device's built-in microphone collects the voice data. This voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0406] Analysis of voice data and optimization of work procedures

[0407] The server analyzes the received voice data using a voice analysis library such as librosa to extract voice features such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify emotions such as tension or relaxation.

[0408] Generate ideal work procedures

[0409] Based on the analysis results and the user's emotional state, the server uses a generative AI model to generate an ideal workflow that reflects the user's characteristics and emotional state. The server generates the results in text format or prompts and sends them back to the user's device via the Internet.

[0410] Generating and Providing Feedback

[0411] The terminal provides a function to display the received ideal work procedures. The user can compare their own work procedures with the generated ideal work procedures. The terminal also displays the analysis results and feedback from the server, allowing the user to identify specific areas for improvement. For example, if a specific task is not performed efficiently, the terminal will provide feedback on the reason and how to improve it.

[0412] Collecting new data and providing feedback

[0413] The user performs an action based on the provided feedback and has the device collect new voice data. The newly collected voice data is compressed again and sent to the server. The server analyzes the new voice data and compares it with the previous feedback result. A new score is calculated and feedback is generated again with improvements.

[0414] Specific examples

[0415] For example, "Factory Worker A" uses his smartphone to issue a voice command such as "Please proceed to the next process," and this voice data is sent from the device to a server. The server analyzes pitch, rhythm, and emotion to generate an ideal work procedure. This data is sent to Factory Worker A's device, and Worker A proceeds with the work while referring to this ideal work procedure. Factory Worker A practices based on the feedback provided by the server, sends new voice data again, and receives further feedback. By repeating this process, Factory Worker A makes improvements to ensure that his work proceeds more efficiently.

[0416] Prompt Sentence Examples

[0417] "What are the next steps to ensure the robot works efficiently? And if the worker is nervous, how should we respond?"

[0418] This system allows users to understand their own emotional state in real time and work based on specific advice. By utilizing advanced analysis technology and an emotion engine, it provides effective feedback suited to individual work styles and emotional states, thereby improving user satisfaction and work efficiency.

[0419] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0420] Step 1:

[0421] A user launches a dedicated application installed on a device such as a smartphone or tablet and inputs voice commands. The device uses a built-in microphone to collect and temporarily store voice data. The input data is the user's voice commands, and the output data is the collected voice file.

[0422] Step 2:

[0423] The device compresses the collected audio data and sends it to a server via the Internet. The input is a temporarily saved audio file, and the output is compressed audio data.

[0424] Step 3:

[0425] The server analyzes the received audio data using a voice analysis library such as librosa to extract features such as pitch, rhythm, vibrato, and long tones. The input is compressed audio data, and the output is extracted audio feature data.

[0426] Step 4:

[0427] The server uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify the emotional state. The input is voice feature data, and the output is the recognized emotional state.

[0428] Step 5:

[0429] The server uses a generative AI model based on the analysis results and emotional state to generate an ideal work procedure. The generated ideal work procedure reflects the user's characteristics and emotional state. The input is voice feature data and emotional state, and the output is ideal work procedure data.

[0430] Step 6:

[0431] The server generates the ideal work procedure data in text format and prompts, and sends them to the user's terminal via the Internet. The input is the ideal work procedure data, and the output is text-format work procedure feedback.

[0432] Step 7:

[0433] The user uses a terminal to check the received ideal work procedure and compare it with their own work procedure. The input is the textual work procedure feedback from the server, and the output is the ideal work procedure that the user checks.

[0434] Step 8:

[0435] The terminal displays the analysis results and feedback, allowing the user to see specific improvements. The input is the feedback data from the server, and the output is the feedback information displayed on the terminal screen.

[0436] Step 9:

[0437] The user performs a new task based on the feedback and causes the terminal to collect voice data. The newly collected voice data is compressed again and sent to the server. The input is the new voice instruction data, and the output is the compressed voice data sent to the server.

[0438] Step 10:

[0439] The server analyzes the new voice data and compares it with the previous feedback result. It calculates a new score based on the analysis result and generates feedback indicating areas for improvement. The input is the new voice data and the previous feedback result, and the output is the new score and feedback indicating areas for improvement.

[0440] Through the above steps, users can work efficiently and effectively, and the server and terminal work together to create a system that provides optimal feedback in real time.

[0441] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0443] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0444] [Second embodiment]

[0445] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0446] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0447] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0448] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0449] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0451] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0452] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0453] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0454] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0455] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0456] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0457] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[0458] Audio data collection

[0459] User

[0460] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[0461] Terminal

[0462] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0463] Analysis of audio data and generation of ideal singing data

[0464] server

[0465] The server analyzes the received audio data. Specifically, it extracts characteristics such as pitch, rhythm, vibrato, and long tones. Based on the analysis results, a generative AI model that has learned the user's singing style generates ideal singing data. This ideal singing data reflects the user's characteristics.

[0466] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0467] Generating and Providing Feedback

[0468] Terminal

[0469] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0470] server

[0471] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results, including specific suggestions and ways to improve on pitch deviations, rhythmic irregularities, and lack of vibrato.

[0472] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[0473] Practice and score improvement support

[0474] User

[0475] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0476] Terminal

[0477] The newly collected audio data is compressed again and sent to the server.

[0478] server

[0479] The server analyzes the new audio data and compares it with the previous feedback. It generates a new score and points for improvement and sends the results to the device. The user can then practice based on this feedback and aim to improve their karaoke score.

[0480] Specific examples

[0481] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, sings again, sends the audio data, and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[0482] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[0483] The processing flow will be explained below.

[0484] Step 1:

[0485] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[0486] Step 2:

[0487] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[0488] Step 3:

[0489] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[0490] Step 4:

[0491] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[0492] Step 5:

[0493] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[0494] Step 6:

[0495] Server: Based on the analysis results, the user's singing data is input into a generative AI model to generate ideal singing data.

[0496] Step 7:

[0497] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[0498] Step 8:

[0499] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[0500] Step 9:

[0501] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[0502] Step 10:

[0503] Server: Based on the analysis results and the differences between the user's singing, the server generates specific feedback in text and video format, including suggestions for improvement such as pitch deviations, rhythmic inconsistencies, and lack of vibrato.

[0504] Step 11:

[0505] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[0506] Step 12:

[0507] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[0508] Step 13:

[0509] Terminal: The newly collected voice data is recompressed and sent to the server.

[0510] Step 14:

[0511] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0512] Step 15:

[0513] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[0514] Through this process, users can continue to practice effectively while receiving detailed feedback, steadily improving their karaoke scores.

[0515] Example 1

[0516] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0517] Conventional karaoke practice systems have made it difficult for users to obtain specific feedback to effectively improve their singing skills. Due to low analysis accuracy, they were unable to provide advice tailored to specific singing styles, limiting user satisfaction and actual skill improvement. Furthermore, they did not generate ideal singing data that reflected each user's individual characteristics, making it difficult to use for comparison or practice.

[0518] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0519] In this invention, the server includes means for analyzing the user's voice data and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data using a generative AI model based on the analysis results, and means for comparing the user's singing with the generated singing data and generating specific feedback. This enables advanced analysis of the user's singing data and provides individually optimized feedback, thereby improving skill and increasing satisfaction.

[0520] 1. "User's voice data" refers to the waveform information of voice collected by a user using a terminal, including karaoke singing voices.

[0521] 2. "Means of collection" refers to the function of detecting the user's voice through the built-in microphone or peripheral devices and recording it as digital data.

[0522] 3. "Means for transmitting to a server" refers to the communications means for transferring audio data from a terminal to a server via the Internet.

[0523] 4. "Means of analysis" refers to algorithms or software for extracting singing characteristics such as pitch, rhythm, vibrato, and long tones from the received audio data.

[0524] 5. "Generative AI model" is an artificial intelligence model that learns the user's singing style and generates ideal singing data based on the results.

[0525] 6. "Means for generating ideal singing data" refers to the process of using the analysis results and generative AI model to generate ideal singing data that reflects the user's characteristics.

[0526] 7. "Means for transmitting generated singing data to a user terminal" refers to a communication means for transferring ideal singing data from a server to a user terminal via the Internet.

[0527] 8. "Means for comparing a user's singing with the generated singing data" refers to an algorithm or software that compares a user's actual singing data with the generated ideal singing data and evaluates the differences.

[0528] 9. "Means for generating specific feedback" refers to the process of generating advice based on the comparison results, including areas for improvement such as pitch deviations, rhythmic irregularities, and lack of vibrato.

[0529] 10. "Means for displaying feedback content to the user" refers to the functionality for displaying the generated feedback on the user's device, including text and video formats.

[0530] 11. "Means for collecting new voice data and sending it again to the server" refers to a communication means for collecting the voice of the user singing again and sending it again to the server.

[0531] 12. "Means for generating a new score and improvements based on the results of reanalysis" refers to the process of analyzing the resubmitted voice data, evaluating the user's progress, and generating a new score and improvements.

[0532] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[0533] Audio data collection

[0534] User

[0535] Users launch a dedicated app installed on their smartphone, tablet, or other device. Using the app, they select the karaoke song they want to sing and check the screen that displays the lyrics. At this time, an audio accompaniment is also played.

[0536] Terminal

[0537] As soon as the user starts singing, the device collects audio data through the built-in microphone, temporarily stores the data locally, compresses it, and transmits it over the Internet to a server.

[0538] Analysis of audio data and generation of ideal singing data

[0539] server

[0540] The server analyzes the received audio data. During this process, characteristics such as pitch, rhythm, vibrato, and long tones are extracted. Based on the analysis results, a prompt is input into the generative AI model to generate ideal singing data. Specifically, the following prompt is used:

[0541] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[0542] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0543] Generating and Providing Feedback

[0544] Terminal

[0545] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0546] server

[0547] The server compares the user's singing data with the generated ideal singing data. Based on the comparison results, it generates feedback that includes specific suggestions for improvement, such as pitch deviations, rhythm disturbances, and lack of vibrato. The feedback is created in text and video format and sent to the device.

[0548] Practice and score improvement support

[0549] User

[0550] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0551] Terminal

[0552] The newly collected audio data is compressed again and sent to a server via the Internet.

[0553] server

[0554] The server analyzes the new audio data and compares it with the previous feedback to generate a new score and areas for improvement. These results are then sent back to the device, where the user can receive further feedback and practice.

[0555] Specific examples

[0556] For example, "User A" selects a karaoke song and starts singing. The device collects the audio data, compresses it, and sends it to the server. The server analyzes the audio data and inputs the following prompt sentence into the generative AI model:

[0557] "Based on user A's voice data, generate ideal singing data using the following criteria: pitch, rhythm, vibrato, long tone. Reflect user A's singing style."

[0558] The ideal singing data generated from the model is sent from the server to User A's device, where User A can listen to the singing data. User A can also practice based on specific points for improvement provided by the server. By repeating this process, User A can improve their skills.

[0559] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and generative AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[0560] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0561] Step 1: Collecting audio data

[0562] User

[0563] The user launches the dedicated app installed on their device and selects the karaoke song they want to sing. The app displays the lyrics on the screen and plays the accompanying audio. At this point, the user can begin singing.

[0564] Input: User singing voice

[0565] Output: Audio data

[0566] Terminal

[0567] The device uses a built-in microphone to collect voice data as soon as the user starts singing. This voice data is temporarily stored in the device as digital data, and then the collected voice data is compressed.

[0568] Input: User singing voice

[0569] Output: Compressed audio data

[0570] Specific actions

[0571] The device automatically activates the microphone when the user starts singing, stores the collected voice data in a buffer, and then applies a compression algorithm to reduce the data size.

[0572] Step 2: Sending audio data to the server

[0573] Terminal

[0574] The compressed audio data is transmitted to a server via the Internet.

[0575] Input: Compressed audio data

[0576] Output: Status of sending to server

[0577] Specific actions

[0578] The device establishes an Internet connection and sends the compressed audio data to the server using the HTTPS protocol.

[0579] Step 3: Analyzing the audio data

[0580] server

[0581] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones.

[0582] Input: Compressed audio data

[0583] Output: Feature extracted data

[0584] Specific actions

[0585] The analysis module on the server decodes the audio data and uses signal processing technology to extract various characteristics as numerical values, allowing for a detailed analysis of the user's singing data.

[0586] Step 4: Generate ideal singing data

[0587] server

[0588] The server generates ideal singing data using a generative AI model based on the extracted feature data, using the following prompt:

[0589] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[0590] Input: feature extraction data, prompt statement

[0591] Output: Ideal singing data

[0592] Specific actions

[0593] The generative AI model receives the prompt sentence and uses the feature extraction data to generate ideal singing data, which reflects the user's singing style and produces highly accurate singing examples.

[0594] Step 5: Send ideal singing data to your device

[0595] server

[0596] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0597] Input: Ideal singing data

[0598] Output: Sending status to user terminal

[0599] Specific actions

[0600] The server establishes an Internet connection and transmits the generated ideal singing data to the user's device using the HTTPS protocol.

[0601] Step 6: Generate feedback

[0602] server

[0603] The server compares the user's singing data with the generated ideal singing data and generates specific feedback, including areas for improvement such as pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[0604] Input: User's singing data, ideal singing data

[0605] Output: Feedback data

[0606] Specific actions

[0607] A comparison algorithm in the server analyzes the user's singing data and the ideal singing data, assesses the differences, and generates feedback based on the results.

[0608] Step 7: Displaying feedback on your device

[0609] Terminal

[0610] The device displays the received feedback to the user. Feedback is provided in text and video formats.

[0611] Input: Feedback data

[0612] Output: Feedback information to the display screen

[0613] Specific actions

[0614] The device analyzes the feedback data and displays it in text and video format on the user interface, allowing the user to continue practicing while checking the feedback.

[0615] (Application example 1)

[0616] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0617] Conventional karaoke systems have difficulty providing specific feedback to improve users' singing skills, and lack real-time support to help users improve their singing. This makes it difficult for users to efficiently improve their singing skills.

[0618] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0619] In this invention, the server includes means for collecting user voice data, means for transmitting the collected voice data to the server, means for analyzing the user voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data based on the analysis results, means for transmitting the generated singing data to a user terminal, means for comparing the user's singing with the generated singing data, means for generating specific feedback based on the comparison results, means for displaying the feedback to the user, means for collecting new voice data for the user's practice and transmitting it again to the server, means for generating a new score and points for improvement based on the reanalysis results, means for providing real-time support for the user's practice using a device in a physical store, and means for displaying the analysis data and feedback in real time, thereby enabling the user to specifically and effectively improve their singing technique in real time.

[0620] "User's voice data" is voice information collected when a user sings a karaoke song.

[0621] The "means of collection" refers to a device that has the function of recording the voice while singing using a device such as a smartphone or tablet.

[0622] The "transmitting means" is a communication means that has the function of sending the collected voice data to a server via the Internet.

[0623] The "analyzing means" is software that executes an algorithm to analyze the audio data received by the server and extract characteristics such as pitch, rhythm, vibrato, and long tones.

[0624] The "means of generation" is software with an algorithm that creates ideal singing data based on the analysis results.

[0625] The "means for transmitting to the user terminal" is a communication means for returning the generated singing data to the user's device.

[0626] The "comparison means" is software that has the function of comparing the user's singing data with the generated ideal singing data and checking the differences.

[0627] The "means for generating feedback" is software that generates improvements and advice for the user based on the comparison results.

[0628] The "means for displaying the feedback content" is a device having a screen display function for displaying the feedback content as text or video on the user's device.

[0629] The "means for collecting new voice data" is a recording device that collects new voice data when the user sings again.

[0630] The "means for transmitting to the server again" is a communication means for sending newly collected voice data to the server again.

[0631] The "means for generating a new score and improvements based on the reanalysis results" is software that analyzes the newly transmitted audio data, compares it with the previous feedback, and generates a new score and improvements.

[0632] "Physical store devices" are devices such as tablets and smartphones used in physical stores such as karaoke booths.

[0633] The "means for providing real-time support" is a system that uses devices in physical stores to analyze the user's singing in real time and provide immediate feedback.

[0634] The "means for displaying analysis data and feedback in real time" is a device having a screen display function for displaying analysis results and feedback content in real time while the user is singing.

[0635] The system for realizing this invention consists of a series of processes that collects user voice data, transmits the data to a server for analysis, and provides feedback based on the interpreted information. This system uses a smartphone or tablet as the main hardware and a generative AI model as the software.

[0636] 1. Collection of audio data

[0637] A user selects a karaoke song using a device in a physical store (e.g., a tablet installed in a karaoke booth) and starts singing. The device collects the user's singing voice through a built-in microphone and temporarily stores the voice data.

[0638] 2. Data compression and transmission

[0639] The collected audio data is compressed on the device, which then transmits the compressed data to a server over the internet, either via Wi-Fi or 4G / 5G networks.

[0640] 3. Analysis of audio data

[0641] The server analyzes the received audio data and extracts characteristics such as pitch, rhythm, vibrato, and long tones. This analysis uses machine learning algorithms and digital signal processing technology. The generative AI model used learns the user's characteristics and generates ideal singing data based on them.

[0642] 4. Generation and transmission of ideal singing data

[0643] The server generates ideal singing data using the generative AI model, which is then sent to the user's device via the internet.

[0644] 5. Providing Feedback

[0645] The user's device plays the ideal singing data that was sent, allowing the user to compare their singing with the ideal singing. The device also displays specific feedback to help the user understand areas for improvement. The feedback is displayed in text and video format.

[0646] 6. Practice and score improvement support

[0647] The user attempts to sing again based on the provided feedback. The device collects new audio data and again sends it to the server. The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points. This information is again sent to the user's device.

[0648] Specific examples

[0649] For example, User A sings a song of his / her choice in a karaoke booth, and the audio data is sent from the tablet to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again. Through this process, User A can reliably improve his / her technique. An example of a prompt for a generative AI model is: "User A sings a karaoke song of his / her choice, and the audio data is sent from the tablet in the karaoke booth to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again to receive feedback for the next time. Through this process, User A can reliably improve his / her technique and improve his / her karaoke score."

[0650] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0651] Step 1:

[0652] A user selects a karaoke song using a device in a brick-and-mortar store (e.g., a tablet in a karaoke booth), then launches an application on the tablet, which displays the lyrics and plays the accompanying music. As the user begins to sing, the tablet's built-in microphone collects audio data, which is temporarily stored on the device.

[0653] Input: User's voice

[0654] Output: Audio data file

[0655] Specific operation: The built-in microphone collects the user's singing voice as digital audio data and saves it as a file.

[0656] Step 2:

[0657] The device compresses the collected voice data to reduce communication load, and then transmits the compressed voice data to a server via the Internet. During this time, the device uses Wi-Fi or 4G / 5G networks.

[0658] Input: Audio data file

[0659] Output: Compressed audio data

[0660] What it does: The audio data is compressed using a standard audio compression algorithm (e.g. MP3, AAC) and uploaded to the server via the HTTP protocol.

[0661] Step 3:

[0662] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones, and then uses a generative AI model to generate ideal singing data based on the analysis results.

[0663] Input: Compressed audio data

[0664] Output: Ideal singing data

[0665] Specific operation: Extract features from audio data using analysis algorithms and generative AI models to generate ideal singing data. Feature extraction uses audio signal processing techniques (e.g., Fast Fourier Transform).

[0666] Step 4:

[0667] The server transmits the generated ideal singing data to the user's terminal again via the Internet.

[0668] Input: Ideal singing data

[0669] Output: sent to the user's terminal

[0670] Specific operation: The generated singing data is sent to the user's terminal via the HTTP protocol, and the terminal receives the data.

[0671] Step 5:

[0672] The device plays back the received ideal singing data, compares it with the user's singing data, and generates specific feedback based on the analysis results received from the server, which is displayed to the user in real time.

[0673] Input: Ideal singing data, analysis results

[0674] Output: Comparison results, feedback

[0675] Specific operation: Using the audio playback function, the ideal singing data is played back, and a comparison algorithm is used to display the difference between the user's singing data and the ideal singing data. Feedback is also displayed on the screen in text and video format.

[0676] Step 6:

[0677] The user attempts to sing again based on the provided feedback, and the device collects new audio data and sends it back to the server.

[0678] Input: User's new voice

[0679] Output: New audio data file

[0680] Specific operation: New singing voices are collected using the built-in microphone, the data is compressed and sent to the server.

[0681] Step 7:

[0682] The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points, which are then sent back to the user's device.

[0683] Input: New audio data

[0684] Output: New score, improvements

[0685] Specific operation: The analysis algorithm is applied again, compared with the previous data to generate a new score and areas for improvement, and the results are sent to the user's device.

[0686] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0687] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[0688] Audio data collection

[0689] User

[0690] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[0691] Terminal

[0692] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0693] Analysis of audio data and generation of ideal singing data

[0694] server

[0695] The server analyzes the received voice data, specifically extracting characteristics such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotions from the voice data.

[0696] Emotion Engine

[0697] The emotion engine analyzes the voice data, the user's intonation, tempo, etc. to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model and reflected in the generation of ideal singing data.

[0698] server

[0699] The ideal singing data reflects the characteristics and emotional state of the user. The generated ideal singing data is transmitted to the user's terminal via the Internet.

[0700] Generating and Providing Feedback

[0701] Terminal

[0702] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0703] server

[0704] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The server also takes into account the user's emotional information recognized by the emotion engine. For example, the server may provide specific suggestions and suggestions for improvement regarding pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[0705] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[0706] Practice and score improvement support

[0707] User

[0708] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0709] Terminal

[0710] The newly collected audio data is compressed again and sent to the server.

[0711] server

[0712] The server analyzes the new voice data, compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0713] Specific examples

[0714] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes pitch, rhythm, and emotion to generate ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, then sends new audio data again and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[0715] Feedback using emotional information

[0716] For example, if user A is nervous during practice, the emotion engine can detect this tension and suggest breathing techniques or relaxation techniques to help them relax. In this way, feedback that takes into account the user's emotional information allows the user to receive more appropriate advice and effectively improve their skills.

[0717] This system allows users to understand their singing technique in real time and practice based on specific advice. Utilizing advanced analysis technology and an emotion engine, it provides effective feedback tailored to individual singing styles and emotional states, resulting in user satisfaction and improvement of technique.

[0718] The processing flow will be explained below.

[0719] Step 1:

[0720] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[0721] Step 2:

[0722] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[0723] Step 3:

[0724] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[0725] Step 4:

[0726] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[0727] Step 5:

[0728] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[0729] Step 6:

[0730] Server: Uses an emotion engine to recognize the user's emotions from the audio data, for example, identifying emotional states such as tension, joy, or sadness from changes in intonation and tempo.

[0731] Step 7:

[0732] Server: Based on the user's singing data and the recognized emotional information, the server uses a generative AI model to generate ideal singing data. The generated singing data reflects not only the characteristics of pitch and rhythm, but also the user's emotions.

[0733] Step 8:

[0734] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[0735] Step 9:

[0736] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[0737] Step 10:

[0738] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[0739] Step 11:

[0740] Server: Based on the analysis results, the differences between the user's singing voice, and the recognized emotional information, the server generates specific feedback. For example, if a user feels nervous, the server will suggest relaxing breathing techniques, which includes advice tailored to the user's emotions.

[0741] Step 12:

[0742] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[0743] Step 13:

[0744] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[0745] Step 14:

[0746] Terminal: The newly collected voice data is recompressed and sent to the server.

[0747] Step 15:

[0748] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0749] Step 16:

[0750] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[0751] Through this process, users can receive detailed feedback and continue practicing effectively, steadily improving their karaoke scores. Taking the user's emotions into consideration also enables more personalized and effective instruction.

[0752] Example 2

[0753] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0754] While conventional singing instruction systems have the ability to provide feedback on pitch and rhythm, they have difficulty providing personalized advice that takes into account the user's emotional state. As a result, instruction is insufficient for users who are nervous or who need to express their emotions, making it difficult to achieve overall improvement in singing technique. Another problem is that users are unable to receive feedback based on their own singing characteristics or emotional state, which hinders effective practice.

[0755] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0756] In this invention, the server includes a means for collecting user voice data, a means including an emotion engine for recognizing emotion information, and a means for generating specific feedback, thereby enabling personalized feedback and advice that takes into account the user's voice characteristics and emotional state.

[0757] A "user" is an individual who uses the system to receive analysis of singing data and feedback.

[0758] "Audio data" refers to data collected in digital form from the voice of a user singing.

[0759] A "terminal" is a digital device that collects voice data and transmits it to a server.

[0760] A "server" is a remote computing device that analyzes audio data, generates feedback, and recognizes emotional information.

[0761] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotional state from voice data.

[0762] "Ideal singing data" is data that indicates an ideal singing performance that is analyzed based on the user's voice data and generated by a generative AI model.

[0763] "Feedback" is specific advice and information on areas for improvement aimed at improving the user's singing technique, based on the results of analyzing the audio data.

[0764] "Comparison" refers to evaluating the user's singing data and the ideal singing data and analyzing the differences between them.

[0765] A "score" is a numerical value calculated based on specific evaluation criteria for singing data, and quantitatively represents the user's singing performance.

[0766] A "generative AI model" is an artificial intelligence algorithm that generates an ideal singing performance based on the user's singing data.

[0767] A "prompt sentence" is text data containing instructions or questions to be input into a generative AI model.

[0768] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[0769] Hardware and software used

[0770] Device: A digital device such as a smartphone or tablet that has a built-in microphone and can be operated by the user.

[0771] Server: A remote computing device on the cloud

[0772] Emotion Engine: Software for analyzing and recognizing emotions

[0773] Generative AI model: An artificial intelligence algorithm that generates ideal singing data

[0774] Audio data collection

[0775] Users launch a dedicated app installed on their smartphone or tablet and select the karaoke song they want to sing. The app displays lyrics and plays back an audio accompaniment. As soon as the user starts singing, the device uses its built-in microphone to collect audio data. This audio data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0776] Analysis of audio data and generation of ideal singing data

[0777] The server analyzes the received voice data and extracts features such as pitch, rhythm, vibrato, and long tones. It also uses an emotion engine to recognize the user's emotions from the voice data. The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model.

[0778] Example prompt sentence:

[0779] "We have received the audio data of user A's karaoke song. Please extract the characteristics of pitch, rhythm, vibrato, and long tones, analyze the emotions with the emotion engine, and generate ideal singing data."

[0780] The server generates ideal singing data and transmits it again to the user's terminal via the Internet.

[0781] Generating and Providing Feedback

[0782] The device provides the user with the ability to play back the received ideal singing data. The user can compare their own singing with the ideal singing. The device also displays the analysis results and specific feedback from the server, allowing the user to identify areas for improvement. The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. This feedback also includes the user's emotional information recognized by the emotion engine.

[0783] Practice and score improvement support

[0784] The user practices based on the provided feedback and sings the same song again, causing the device to collect audio data. The newly collected audio data is compressed again and sent to the server. The server analyzes the new audio data, compares it with the previous feedback result, calculates a new score, and generates new feedback on areas for improvement.

[0785] Specific examples of operation

[0786] For example, User A chooses a favorite karaoke song, sings it, and sends the audio data from his / her device to the server. The server analyzes pitch, rhythm, vibrato, and emotion to generate ideal singing data. This data is sent to User A's device, where User A listens to the ideal singing and compares it with his / her own singing. User A practices based on feedback from the server, sends new audio data, and receives further feedback. By repeating this process, User A can steadily improve his / her skills and increase his / her karaoke score.

[0787] Additionally, if User A becomes nervous during practice, the emotion engine will detect this and provide advice such as, "Try some breathing techniques to relax." Feedback that utilizes emotion information in this way allows User A to effectively improve their skills.

[0788] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0789] Step 1: Collecting audio data

[0790] Users launch a dedicated karaoke app installed on their smartphone or tablet and select the karaoke song they want to sing. The input is the karaoke song selected by the user. The app displays the lyrics and plays the accompaniment audio.

[0791] The device collects voice data using a built-in microphone as soon as the user starts singing. The input is the user's singing voice, and the output is the voice data temporarily stored in the device.

[0792] The device compresses the collected audio data into a specified format (e.g., MP3 or FLAC). The compressed audio data is sent to a server via the Internet. The input is the compressed audio data, and the output is the audio data sent to the server.

[0793] Step 2: Analyzing the audio data

[0794] The server receives the voice data sent from the terminal. The input is the collected voice data. The data analysis engine in the server extracts features such as pitch, rhythm, vibrato, and long tones. The output is feature information of the voice data.

[0795] The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. The input is the voice data's intonation and tempo information, and the output is the user's emotional information. This emotional information, along with the analysis results, is input into the generative AI model.

[0796] Step 3: Generate ideal singing data

[0797] The server inputs a prompt sentence into the generative AI model based on the feature information and emotional information of the audio data. An example of a prompt sentence is, "We have received audio data of User A's karaoke song. Please extract the features of pitch, rhythm, vibrato, and long tones, and analyze the emotions using the emotion engine to generate ideal singing data." The input is the prompt sentence and the analysis result information, and the generative AI model generates ideal singing data. The output is ideal singing data.

[0798] The server then transmits the generated ideal singing data to the user's terminal via the Internet. The input is the ideal singing data, and the output is the data transmitted to the user's terminal.

[0799] Step 4: Play and compare ideal singing data

[0800] The device provides the function to play back the received ideal singing data. The input is the singing data sent from the server. The user can compare their own singing with the ideal singing. The output is the user's listening state.

[0801] The terminal displays the analysis results and feedback sent from the server. The input is the analysis results and feedback information, allowing the user to identify areas that need improvement. The output is feedback displayed to the user.

[0802] Step 5: Generate feedback

[0803] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The input is the user's singing data and the ideal singing data, and the output is specific feedback. This feedback also includes the user's emotional information recognized by the emotion engine.

[0804] Step 6: Practice and additional feedback

[0805] The user practices based on the provided feedback. They then sing the same song again, causing the device to collect audio data. The input is the feedback information. By collecting new audio data, the output is the collected new audio data.

[0806] The terminal compresses the newly collected voice data and sends it to the server again. The input is the newly collected voice data, and the output is the data sent to the server.

[0807] The server analyzes the new voice data, compares it with the previous feedback result to calculate a new score, and generates feedback on improvements. The input is the new voice data and the previous feedback result, and the output is the new score and feedback information.

[0808] By dividing the specific processing of this system into these steps, it is possible to effectively support the improvement of the user's singing technique.

[0809] (Application example 2)

[0810] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0811] In modern factories, humans and robots are increasingly working together, but there are many challenges in providing instructions on-site and improving work efficiency. In particular, the emotional state of workers can affect work efficiency and robot performance, so there is a need for a means to detect this in real time and take countermeasures. In addition, there is a lack of systems that provide real-time feedback and optimize work procedures.

[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0813] In this invention, the server includes a means for analyzing the user's voice data and detecting the emotional state, a means for optimizing the work procedure based on the user's voice instruction data, and a means for generating an ideal work procedure and providing feedback, thereby making it possible to detect the emotional state of the worker in real time and improve work efficiency.

[0814] "User voice data" refers to voice data when a factory worker gives instructions to a robot.

[0815] "Collection means" refers to the device or process used to acquire and transmit audio data to a server.

[0816] "Analysis of audio data" refers to the process of extracting audio features such as pitch, rhythm, vibrato, and long tones and analyzing the data.

[0817] "Emotional state" refers to the user's emotions inferred from the intonation, tempo, and sound patterns of the user's voice.

[0818] "Ideal singing data" refers to voice data of optimal singing that a user should aim for, which is generated based on analyzed voice data and emotional state.

[0819] "Feedback" refers to specific advice and guidance information provided to the user based on the analysis results.

[0820] "Work procedure optimization" refers to the process of deriving the most efficient work procedure by taking into account the user's voice instructions and emotional state.

[0821] A "generative AI model" refers to an algorithm or system that analyzes voice data and emotional data to generate ideal data and feedback.

[0822] A "prompt" refers to an instruction or question input to a generative AI model.

[0823] To put the present invention into practice, a system is constructed in which users, terminals, and servers work in cooperation with each other. Specific embodiments are as follows.

[0824] Audio data collection

[0825] The user launches a dedicated application installed on a device such as a smartphone or tablet installed in the factory. Work instructions and confirmations are entered into the application by voice, and the device's built-in microphone collects the voice data. This voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0826] Analysis of voice data and optimization of work procedures

[0827] The server analyzes the received voice data using a voice analysis library such as librosa to extract voice features such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify emotions such as tension or relaxation.

[0828] Generate ideal work procedures

[0829] Based on the analysis results and the user's emotional state, the server uses a generative AI model to generate an ideal workflow that reflects the user's characteristics and emotional state. The server generates the results in text format or prompts and sends them back to the user's device via the Internet.

[0830] Generating and Providing Feedback

[0831] The terminal provides a function to display the received ideal work procedures. The user can compare their own work procedures with the generated ideal work procedures. The terminal also displays the analysis results and feedback from the server, allowing the user to identify specific areas for improvement. For example, if a specific task is not performed efficiently, the terminal will provide feedback on the reason and how to improve it.

[0832] Collecting new data and providing feedback

[0833] The user performs an action based on the provided feedback and has the device collect new voice data. The newly collected voice data is compressed again and sent to the server. The server analyzes the new voice data and compares it with the previous feedback result. A new score is calculated and feedback is generated again with improvements.

[0834] Specific examples

[0835] For example, "Factory Worker A" uses his smartphone to issue a voice command such as "Please proceed to the next process," and this voice data is sent from the device to a server. The server analyzes pitch, rhythm, and emotion to generate an ideal work procedure. This data is sent to Factory Worker A's device, and Worker A proceeds with the work while referring to this ideal work procedure. Factory Worker A practices based on the feedback provided by the server, sends new voice data again, and receives further feedback. By repeating this process, Factory Worker A makes improvements to ensure that his work proceeds more efficiently.

[0836] Prompt Sentence Examples

[0837] "What are the next steps to ensure the robot works efficiently? And if the worker is nervous, how should we respond?"

[0838] This system allows users to understand their own emotional state in real time and work based on specific advice. By utilizing advanced analysis technology and an emotion engine, it provides effective feedback suited to individual work styles and emotional states, thereby improving user satisfaction and work efficiency.

[0839] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0840] Step 1:

[0841] A user launches a dedicated application installed on a device such as a smartphone or tablet and inputs voice commands. The device uses a built-in microphone to collect and temporarily store voice data. The input data is the user's voice commands, and the output data is the collected voice file.

[0842] Step 2:

[0843] The device compresses the collected audio data and sends it to a server via the Internet. The input is a temporarily saved audio file, and the output is compressed audio data.

[0844] Step 3:

[0845] The server analyzes the received audio data using a voice analysis library such as librosa to extract features such as pitch, rhythm, vibrato, and long tones. The input is compressed audio data, and the output is extracted audio feature data.

[0846] Step 4:

[0847] The server uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify the emotional state. The input is voice feature data, and the output is the recognized emotional state.

[0848] Step 5:

[0849] The server uses a generative AI model based on the analysis results and emotional state to generate an ideal work procedure. The generated ideal work procedure reflects the user's characteristics and emotional state. The input is voice feature data and emotional state, and the output is ideal work procedure data.

[0850] Step 6:

[0851] The server generates the ideal work procedure data in text format and prompts, and sends them to the user's terminal via the Internet. The input is the ideal work procedure data, and the output is text-format work procedure feedback.

[0852] Step 7:

[0853] The user uses a terminal to check the received ideal work procedure and compare it with their own work procedure. The input is the textual work procedure feedback from the server, and the output is the ideal work procedure that the user checks.

[0854] Step 8:

[0855] The terminal displays the analysis results and feedback, allowing the user to see specific improvements. The input is the feedback data from the server, and the output is the feedback information displayed on the terminal screen.

[0856] Step 9:

[0857] The user performs a new task based on the feedback and causes the terminal to collect voice data. The newly collected voice data is compressed again and sent to the server. The input is the new voice instruction data, and the output is the compressed voice data sent to the server.

[0858] Step 10:

[0859] The server analyzes the new voice data and compares it with the previous feedback result. It calculates a new score based on the analysis result and generates feedback indicating areas for improvement. The input is the new voice data and the previous feedback result, and the output is the new score and feedback indicating areas for improvement.

[0860] Through the above steps, users can work efficiently and effectively, and the server and terminal work together to create a system that provides optimal feedback in real time.

[0861] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0862] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0863] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0864] [Third embodiment]

[0865] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0866] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0867] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0868] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0869] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0870] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0871] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0872] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0873] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0874] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0875] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0876] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0877] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[0878] Audio data collection

[0879] User

[0880] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[0881] Terminal

[0882] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[0883] Analysis of audio data and generation of ideal singing data

[0884] server

[0885] The server analyzes the received audio data. Specifically, it extracts characteristics such as pitch, rhythm, vibrato, and long tones. Based on the analysis results, a generative AI model that has learned the user's singing style generates ideal singing data. This ideal singing data reflects the user's characteristics.

[0886] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0887] Generating and Providing Feedback

[0888] Terminal

[0889] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0890] server

[0891] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results, including specific suggestions and ways to improve on pitch deviations, rhythmic irregularities, and lack of vibrato.

[0892] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[0893] Practice and score improvement support

[0894] User

[0895] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0896] Terminal

[0897] The newly collected audio data is compressed again and sent to the server.

[0898] server

[0899] The server analyzes the new audio data and compares it with the previous feedback. It generates a new score and points for improvement and sends the results to the device. The user can then practice based on this feedback and aim to improve their karaoke score.

[0900] Specific examples

[0901] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, sings again, sends the audio data, and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[0902] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[0903] The processing flow will be explained below.

[0904] Step 1:

[0905] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[0906] Step 2:

[0907] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[0908] Step 3:

[0909] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[0910] Step 4:

[0911] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[0912] Step 5:

[0913] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[0914] Step 6:

[0915] Server: Based on the analysis results, the user's singing data is input into a generative AI model to generate ideal singing data.

[0916] Step 7:

[0917] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[0918] Step 8:

[0919] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[0920] Step 9:

[0921] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[0922] Step 10:

[0923] Server: Based on the analysis results and the differences between the user's singing, the server generates specific feedback in text and video format, including suggestions for improvement such as pitch deviations, rhythmic inconsistencies, and lack of vibrato.

[0924] Step 11:

[0925] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[0926] Step 12:

[0927] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[0928] Step 13:

[0929] Terminal: The newly collected voice data is recompressed and sent to the server.

[0930] Step 14:

[0931] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[0932] Step 15:

[0933] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[0934] Through this process, users can continue to practice effectively while receiving detailed feedback, steadily improving their karaoke scores.

[0935] Example 1

[0936] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0937] Conventional karaoke practice systems have made it difficult for users to obtain specific feedback to effectively improve their singing skills. Due to low analysis accuracy, they were unable to provide advice tailored to specific singing styles, limiting user satisfaction and actual skill improvement. Furthermore, they did not generate ideal singing data that reflected each user's individual characteristics, making it difficult to use for comparison or practice.

[0938] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0939] In this invention, the server includes means for analyzing the user's voice data and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data using a generative AI model based on the analysis results, and means for comparing the user's singing with the generated singing data and generating specific feedback. This enables advanced analysis of the user's singing data and provides individually optimized feedback, thereby improving skill and increasing satisfaction.

[0940] 1. "User's voice data" refers to the waveform information of voice collected by a user using a terminal, including karaoke singing voices.

[0941] 2. "Means of collection" refers to the function of detecting the user's voice through the built-in microphone or peripheral devices and recording it as digital data.

[0942] 3. "Means for transmitting to a server" refers to the communications means for transferring audio data from a terminal to a server via the Internet.

[0943] 4. "Means of analysis" refers to algorithms or software for extracting singing characteristics such as pitch, rhythm, vibrato, and long tones from the received audio data.

[0944] 5. "Generative AI model" is an artificial intelligence model that learns the user's singing style and generates ideal singing data based on the results.

[0945] 6. "Means for generating ideal singing data" refers to the process of using the analysis results and generative AI model to generate ideal singing data that reflects the user's characteristics.

[0946] 7. "Means for transmitting generated singing data to a user terminal" refers to a communication means for transferring ideal singing data from a server to a user terminal via the Internet.

[0947] 8. "Means for comparing a user's singing with the generated singing data" refers to an algorithm or software that compares a user's actual singing data with the generated ideal singing data and evaluates the differences.

[0948] 9. "Means for generating specific feedback" refers to the process of generating advice based on the comparison results, including areas for improvement such as pitch deviations, rhythmic irregularities, and lack of vibrato.

[0949] 10. "Means for displaying feedback content to the user" refers to the functionality for displaying the generated feedback on the user's device, including text and video formats.

[0950] 11. "Means for collecting new voice data and sending it again to the server" refers to a communication means for collecting the voice of the user singing again and sending it again to the server.

[0951] 12. "Means for generating a new score and improvements based on the results of reanalysis" refers to the process of analyzing the resubmitted voice data, evaluating the user's progress, and generating a new score and improvements.

[0952] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[0953] Audio data collection

[0954] User

[0955] Users launch a dedicated app installed on their smartphone, tablet, or other device. Using the app, they select the karaoke song they want to sing and check the screen that displays the lyrics. At this time, an audio accompaniment is also played.

[0956] Terminal

[0957] As soon as the user starts singing, the device collects audio data through the built-in microphone, temporarily stores the data locally, compresses it, and transmits it over the Internet to a server.

[0958] Analysis of audio data and generation of ideal singing data

[0959] server

[0960] The server analyzes the received audio data. During this process, characteristics such as pitch, rhythm, vibrato, and long tones are extracted. Based on the analysis results, a prompt is input into the generative AI model to generate ideal singing data. Specifically, the following prompt is used:

[0961] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[0962] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[0963] Generating and Providing Feedback

[0964] Terminal

[0965] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[0966] server

[0967] The server compares the user's singing data with the generated ideal singing data. Based on the comparison results, it generates feedback that includes specific suggestions for improvement, such as pitch deviations, rhythm disturbances, and lack of vibrato. The feedback is created in text and video format and sent to the device.

[0968] Practice and score improvement support

[0969] User

[0970] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[0971] Terminal

[0972] The newly collected audio data is compressed again and sent to a server via the Internet.

[0973] server

[0974] The server analyzes the new audio data and compares it with the previous feedback to generate a new score and areas for improvement. These results are then sent back to the device, where the user can receive further feedback and practice.

[0975] Specific examples

[0976] For example, "User A" selects a karaoke song and starts singing. The device collects the audio data, compresses it, and sends it to the server. The server analyzes the audio data and inputs the following prompt sentence into the generative AI model:

[0977] "Based on user A's voice data, generate ideal singing data using the following criteria: pitch, rhythm, vibrato, long tone. Reflect user A's singing style."

[0978] The ideal singing data generated from the model is sent from the server to User A's device, where User A can listen to the singing data. User A can also practice based on specific points for improvement provided by the server. By repeating this process, User A can improve their skills.

[0979] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and generative AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[0980] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0981] Step 1: Collecting audio data

[0982] User

[0983] The user launches the dedicated app installed on their device and selects the karaoke song they want to sing. The app displays the lyrics on the screen and plays the accompanying audio. At this point, the user can begin singing.

[0984] Input: User singing voice

[0985] Output: Audio data

[0986] Terminal

[0987] The device uses a built-in microphone to collect voice data as soon as the user starts singing. This voice data is temporarily stored in the device as digital data, and then the collected voice data is compressed.

[0988] Input: User singing voice

[0989] Output: Compressed audio data

[0990] Specific actions

[0991] The device automatically activates the microphone when the user starts singing, stores the collected voice data in a buffer, and then applies a compression algorithm to reduce the data size.

[0992] Step 2: Sending audio data to the server

[0993] Terminal

[0994] The compressed audio data is transmitted to a server via the Internet.

[0995] Input: Compressed audio data

[0996] Output: Status of sending to server

[0997] Specific actions

[0998] The device establishes an Internet connection and sends the compressed audio data to the server using the HTTPS protocol.

[0999] Step 3: Analyzing the audio data

[1000] server

[1001] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones.

[1002] Input: Compressed audio data

[1003] Output: Feature extracted data

[1004] Specific actions

[1005] The analysis module on the server decodes the audio data and uses signal processing technology to extract various characteristics as numerical values, allowing for a detailed analysis of the user's singing data.

[1006] Step 4: Generate ideal singing data

[1007] server

[1008] The server generates ideal singing data using a generative AI model based on the extracted feature data, using the following prompt:

[1009] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[1010] Input: feature extraction data, prompt statement

[1011] Output: Ideal singing data

[1012] Specific actions

[1013] The generative AI model receives the prompt sentence and uses the feature extraction data to generate ideal singing data, which reflects the user's singing style and produces highly accurate singing examples.

[1014] Step 5: Send ideal singing data to your device

[1015] server

[1016] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[1017] Input: Ideal singing data

[1018] Output: Sending status to user terminal

[1019] Specific actions

[1020] The server establishes an Internet connection and transmits the generated ideal singing data to the user's device using the HTTPS protocol.

[1021] Step 6: Generate feedback

[1022] server

[1023] The server compares the user's singing data with the generated ideal singing data and generates specific feedback, including areas for improvement such as pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[1024] Input: User's singing data, ideal singing data

[1025] Output: Feedback data

[1026] Specific actions

[1027] A comparison algorithm in the server analyzes the user's singing data and the ideal singing data, assesses the differences, and generates feedback based on the results.

[1028] Step 7: Displaying feedback on your device

[1029] Terminal

[1030] The device displays the received feedback to the user. Feedback is provided in text and video formats.

[1031] Input: Feedback data

[1032] Output: Feedback information to the display screen

[1033] Specific actions

[1034] The device analyzes the feedback data and displays it in text and video format on the user interface, allowing the user to continue practicing while checking the feedback.

[1035] (Application example 1)

[1036] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1037] Conventional karaoke systems have difficulty providing specific feedback to improve users' singing skills, and lack real-time support to help users improve their singing. This makes it difficult for users to efficiently improve their singing skills.

[1038] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1039] In this invention, the server includes means for collecting user voice data, means for transmitting the collected voice data to the server, means for analyzing the user voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data based on the analysis results, means for transmitting the generated singing data to a user terminal, means for comparing the user's singing with the generated singing data, means for generating specific feedback based on the comparison results, means for displaying the feedback to the user, means for collecting new voice data for the user's practice and transmitting it again to the server, means for generating a new score and points for improvement based on the reanalysis results, means for providing real-time support for the user's practice using a device in a physical store, and means for displaying the analysis data and feedback in real time, thereby enabling the user to specifically and effectively improve their singing technique in real time.

[1040] "User's voice data" is voice information collected when a user sings a karaoke song.

[1041] The "means of collection" refers to a device that has the function of recording the voice while singing using a device such as a smartphone or tablet.

[1042] The "transmitting means" is a communication means that has the function of sending the collected voice data to a server via the Internet.

[1043] The "analyzing means" is software that executes an algorithm to analyze the audio data received by the server and extract characteristics such as pitch, rhythm, vibrato, and long tones.

[1044] The "means of generation" is software with an algorithm that creates ideal singing data based on the analysis results.

[1045] The "means for transmitting to the user terminal" is a communication means for returning the generated singing data to the user's device.

[1046] The "comparison means" is software that has the function of comparing the user's singing data with the generated ideal singing data and checking the differences.

[1047] The "means for generating feedback" is software that generates improvements and advice for the user based on the comparison results.

[1048] The "means for displaying the feedback content" is a device having a screen display function for displaying the feedback content as text or video on the user's device.

[1049] The "means for collecting new voice data" is a recording device that collects new voice data when the user sings again.

[1050] The "means for transmitting to the server again" is a communication means for sending newly collected voice data to the server again.

[1051] The "means for generating a new score and improvements based on the reanalysis results" is software that analyzes the newly transmitted audio data, compares it with the previous feedback, and generates a new score and improvements.

[1052] "Physical store devices" are devices such as tablets and smartphones used in physical stores such as karaoke booths.

[1053] The "means for providing real-time support" is a system that uses devices in physical stores to analyze the user's singing in real time and provide immediate feedback.

[1054] The "means for displaying analysis data and feedback in real time" is a device having a screen display function for displaying analysis results and feedback content in real time while the user is singing.

[1055] The system for realizing this invention consists of a series of processes that collects user voice data, transmits the data to a server for analysis, and provides feedback based on the interpreted information. This system uses a smartphone or tablet as the main hardware and a generative AI model as the software.

[1056] 1. Collection of audio data

[1057] A user selects a karaoke song using a device in a physical store (e.g., a tablet installed in a karaoke booth) and starts singing. The device collects the user's singing voice through a built-in microphone and temporarily stores the voice data.

[1058] 2. Data compression and transmission

[1059] The collected audio data is compressed on the device, which then transmits the compressed data to a server over the internet, either via Wi-Fi or 4G / 5G networks.

[1060] 3. Analysis of audio data

[1061] The server analyzes the received audio data and extracts characteristics such as pitch, rhythm, vibrato, and long tones. This analysis uses machine learning algorithms and digital signal processing technology. The generative AI model used learns the user's characteristics and generates ideal singing data based on them.

[1062] 4. Generation and transmission of ideal singing data

[1063] The server generates ideal singing data using the generative AI model, which is then sent to the user's device via the internet.

[1064] 5. Providing Feedback

[1065] The user's device plays the ideal singing data that was sent, allowing the user to compare their singing with the ideal singing. The device also displays specific feedback to help the user understand areas for improvement. The feedback is displayed in text and video format.

[1066] 6. Practice and score improvement support

[1067] The user attempts to sing again based on the provided feedback. The device collects new audio data and again sends it to the server. The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points. This information is again sent to the user's device.

[1068] Specific examples

[1069] For example, User A sings a song of his / her choice in a karaoke booth, and the audio data is sent from the tablet to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again. Through this process, User A can reliably improve his / her technique. An example of a prompt for a generative AI model is: "User A sings a karaoke song of his / her choice, and the audio data is sent from the tablet in the karaoke booth to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again to receive feedback for the next time. Through this process, User A can reliably improve his / her technique and improve his / her karaoke score."

[1070] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1071] Step 1:

[1072] A user selects a karaoke song using a device in a brick-and-mortar store (e.g., a tablet in a karaoke booth), then launches an application on the tablet, which displays the lyrics and plays the accompanying music. As the user begins to sing, the tablet's built-in microphone collects audio data, which is temporarily stored on the device.

[1073] Input: User's voice

[1074] Output: Audio data file

[1075] Specific operation: The built-in microphone collects the user's singing voice as digital audio data and saves it as a file.

[1076] Step 2:

[1077] The device compresses the collected voice data to reduce communication load, and then transmits the compressed voice data to a server via the Internet. During this time, the device uses Wi-Fi or 4G / 5G networks.

[1078] Input: Audio data file

[1079] Output: Compressed audio data

[1080] What it does: The audio data is compressed using a standard audio compression algorithm (e.g. MP3, AAC) and uploaded to the server via the HTTP protocol.

[1081] Step 3:

[1082] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones, and then uses a generative AI model to generate ideal singing data based on the analysis results.

[1083] Input: Compressed audio data

[1084] Output: Ideal singing data

[1085] Specific operation: Extract features from audio data using analysis algorithms and generative AI models to generate ideal singing data. Feature extraction uses audio signal processing techniques (e.g., Fast Fourier Transform).

[1086] Step 4:

[1087] The server transmits the generated ideal singing data to the user's terminal again via the Internet.

[1088] Input: Ideal singing data

[1089] Output: sent to the user's terminal

[1090] Specific operation: The generated singing data is sent to the user's terminal via the HTTP protocol, and the terminal receives the data.

[1091] Step 5:

[1092] The device plays back the received ideal singing data, compares it with the user's singing data, and generates specific feedback based on the analysis results received from the server, which is displayed to the user in real time.

[1093] Input: Ideal singing data, analysis results

[1094] Output: Comparison results, feedback

[1095] Specific operation: Using the audio playback function, the ideal singing data is played back, and a comparison algorithm is used to display the difference between the user's singing data and the ideal singing data. Feedback is also displayed on the screen in text and video format.

[1096] Step 6:

[1097] The user attempts to sing again based on the provided feedback, and the device collects new audio data and sends it back to the server.

[1098] Input: User's new voice

[1099] Output: New audio data file

[1100] Specific operation: New singing voices are collected using the built-in microphone, the data is compressed and sent to the server.

[1101] Step 7:

[1102] The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points, which are then sent back to the user's device.

[1103] Input: New audio data

[1104] Output: New score, improvements

[1105] Specific operation: The analysis algorithm is applied again, compared with the previous data to generate a new score and areas for improvement, and the results are sent to the user's device.

[1106] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1107] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[1108] Audio data collection

[1109] User

[1110] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[1111] Terminal

[1112] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1113] Analysis of audio data and generation of ideal singing data

[1114] server

[1115] The server analyzes the received voice data, specifically extracting characteristics such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotions from the voice data.

[1116] Emotion Engine

[1117] The emotion engine analyzes the voice data, the user's intonation, tempo, etc. to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model and reflected in the generation of ideal singing data.

[1118] server

[1119] The ideal singing data reflects the characteristics and emotional state of the user. The generated ideal singing data is transmitted to the user's terminal via the Internet.

[1120] Generating and Providing Feedback

[1121] Terminal

[1122] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[1123] server

[1124] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The server also takes into account the user's emotional information recognized by the emotion engine. For example, the server may provide specific suggestions and suggestions for improvement regarding pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[1125] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[1126] Practice and score improvement support

[1127] User

[1128] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[1129] Terminal

[1130] The newly collected audio data is compressed again and sent to the server.

[1131] server

[1132] The server analyzes the new voice data, compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[1133] Specific examples

[1134] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes pitch, rhythm, and emotion to generate ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, then sends new audio data again and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[1135] Feedback using emotional information

[1136] For example, if user A is nervous during practice, the emotion engine can detect this tension and suggest breathing techniques or relaxation techniques to help them relax. In this way, feedback that takes into account the user's emotional information allows the user to receive more appropriate advice and effectively improve their skills.

[1137] This system allows users to understand their singing technique in real time and practice based on specific advice. Utilizing advanced analysis technology and an emotion engine, it provides effective feedback tailored to individual singing styles and emotional states, resulting in user satisfaction and improvement of technique.

[1138] The processing flow will be explained below.

[1139] Step 1:

[1140] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[1141] Step 2:

[1142] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[1143] Step 3:

[1144] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[1145] Step 4:

[1146] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[1147] Step 5:

[1148] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[1149] Step 6:

[1150] Server: Uses an emotion engine to recognize the user's emotions from the audio data, for example, identifying emotional states such as tension, joy, or sadness from changes in intonation and tempo.

[1151] Step 7:

[1152] Server: Based on the user's singing data and the recognized emotional information, the server uses a generative AI model to generate ideal singing data. The generated singing data reflects not only the characteristics of pitch and rhythm, but also the user's emotions.

[1153] Step 8:

[1154] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[1155] Step 9:

[1156] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[1157] Step 10:

[1158] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[1159] Step 11:

[1160] Server: Based on the analysis results, the differences between the user's singing voice, and the recognized emotional information, the server generates specific feedback. For example, if a user feels nervous, the server will suggest relaxing breathing techniques, which includes advice tailored to the user's emotions.

[1161] Step 12:

[1162] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[1163] Step 13:

[1164] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[1165] Step 14:

[1166] Terminal: The newly collected voice data is recompressed and sent to the server.

[1167] Step 15:

[1168] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[1169] Step 16:

[1170] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[1171] Through this process, users can receive detailed feedback and continue practicing effectively, steadily improving their karaoke scores. Taking the user's emotions into consideration also enables more personalized and effective instruction.

[1172] Example 2

[1173] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1174] While conventional singing instruction systems have the ability to provide feedback on pitch and rhythm, they have difficulty providing personalized advice that takes into account the user's emotional state. As a result, instruction is insufficient for users who are nervous or who need to express their emotions, making it difficult to achieve overall improvement in singing technique. Another problem is that users are unable to receive feedback based on their own singing characteristics or emotional state, which hinders effective practice.

[1175] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1176] In this invention, the server includes a means for collecting user voice data, a means including an emotion engine for recognizing emotion information, and a means for generating specific feedback, thereby enabling personalized feedback and advice that takes into account the user's voice characteristics and emotional state.

[1177] A "user" is an individual who uses the system to receive analysis of singing data and feedback.

[1178] "Audio data" refers to data collected in digital form from the voice of a user singing.

[1179] A "terminal" is a digital device that collects voice data and transmits it to a server.

[1180] A "server" is a remote computing device that analyzes audio data, generates feedback, and recognizes emotional information.

[1181] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotional state from voice data.

[1182] "Ideal singing data" is data that indicates an ideal singing performance that is analyzed based on the user's voice data and generated by a generative AI model.

[1183] "Feedback" is specific advice and information on areas for improvement aimed at improving the user's singing technique, based on the results of analyzing the audio data.

[1184] "Comparison" refers to evaluating the user's singing data and the ideal singing data and analyzing the differences between them.

[1185] A "score" is a numerical value calculated based on specific evaluation criteria for singing data, and quantitatively represents the user's singing performance.

[1186] A "generative AI model" is an artificial intelligence algorithm that generates an ideal singing performance based on the user's singing data.

[1187] A "prompt sentence" is text data containing instructions or questions to be input into a generative AI model.

[1188] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[1189] Hardware and software used

[1190] Device: A digital device such as a smartphone or tablet that has a built-in microphone and can be operated by the user.

[1191] Server: A remote computing device on the cloud

[1192] Emotion Engine: Software for analyzing and recognizing emotions

[1193] Generative AI model: An artificial intelligence algorithm that generates ideal singing data

[1194] Audio data collection

[1195] Users launch a dedicated app installed on their smartphone or tablet and select the karaoke song they want to sing. The app displays lyrics and plays back an audio accompaniment. As soon as the user starts singing, the device uses its built-in microphone to collect audio data. This audio data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1196] Analysis of audio data and generation of ideal singing data

[1197] The server analyzes the received voice data and extracts features such as pitch, rhythm, vibrato, and long tones. It also uses an emotion engine to recognize the user's emotions from the voice data. The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model.

[1198] Example prompt sentence:

[1199] "We have received the audio data of user A's karaoke song. Please extract the characteristics of pitch, rhythm, vibrato, and long tones, analyze the emotions with the emotion engine, and generate ideal singing data."

[1200] The server generates ideal singing data and transmits it again to the user's terminal via the Internet.

[1201] Generating and Providing Feedback

[1202] The device provides the user with the ability to play back the received ideal singing data. The user can compare their own singing with the ideal singing. The device also displays the analysis results and specific feedback from the server, allowing the user to identify areas for improvement. The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. This feedback also includes the user's emotional information recognized by the emotion engine.

[1203] Practice and score improvement support

[1204] The user practices based on the provided feedback and sings the same song again, causing the device to collect audio data. The newly collected audio data is compressed again and sent to the server. The server analyzes the new audio data, compares it with the previous feedback result, calculates a new score, and generates new feedback on areas for improvement.

[1205] Specific examples of operation

[1206] For example, User A chooses a favorite karaoke song, sings it, and sends the audio data from his / her device to the server. The server analyzes pitch, rhythm, vibrato, and emotion to generate ideal singing data. This data is sent to User A's device, where User A listens to the ideal singing and compares it with his / her own singing. User A practices based on feedback from the server, sends new audio data, and receives further feedback. By repeating this process, User A can steadily improve his / her skills and increase his / her karaoke score.

[1207] Additionally, if User A becomes nervous during practice, the emotion engine will detect this and provide advice such as, "Try some breathing techniques to relax." Feedback that utilizes emotion information in this way allows User A to effectively improve their skills.

[1208] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1209] Step 1: Collecting audio data

[1210] Users launch a dedicated karaoke app installed on their smartphone or tablet and select the karaoke song they want to sing. The input is the karaoke song selected by the user. The app displays the lyrics and plays the accompaniment audio.

[1211] The device collects voice data using a built-in microphone as soon as the user starts singing. The input is the user's singing voice, and the output is the voice data temporarily stored in the device.

[1212] The device compresses the collected audio data into a specified format (e.g., MP3 or FLAC). The compressed audio data is sent to a server via the Internet. The input is the compressed audio data, and the output is the audio data sent to the server.

[1213] Step 2: Analyzing the audio data

[1214] The server receives the voice data sent from the terminal. The input is the collected voice data. The data analysis engine in the server extracts features such as pitch, rhythm, vibrato, and long tones. The output is feature information of the voice data.

[1215] The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. The input is the voice data's intonation and tempo information, and the output is the user's emotional information. This emotional information, along with the analysis results, is input into the generative AI model.

[1216] Step 3: Generate ideal singing data

[1217] The server inputs a prompt sentence into the generative AI model based on the feature information and emotional information of the audio data. An example of a prompt sentence is, "We have received audio data of User A's karaoke song. Please extract the features of pitch, rhythm, vibrato, and long tones, and analyze the emotions using the emotion engine to generate ideal singing data." The input is the prompt sentence and the analysis result information, and the generative AI model generates ideal singing data. The output is ideal singing data.

[1218] The server then transmits the generated ideal singing data to the user's terminal via the Internet. The input is the ideal singing data, and the output is the data transmitted to the user's terminal.

[1219] Step 4: Play and compare ideal singing data

[1220] The device provides the function to play back the received ideal singing data. The input is the singing data sent from the server. The user can compare their own singing with the ideal singing. The output is the user's listening state.

[1221] The terminal displays the analysis results and feedback sent from the server. The input is the analysis results and feedback information, allowing the user to identify areas that need improvement. The output is feedback displayed to the user.

[1222] Step 5: Generate feedback

[1223] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The input is the user's singing data and the ideal singing data, and the output is specific feedback. This feedback also includes the user's emotional information recognized by the emotion engine.

[1224] Step 6: Practice and additional feedback

[1225] The user practices based on the provided feedback. They then sing the same song again, causing the device to collect audio data. The input is the feedback information. By collecting new audio data, the output is the collected new audio data.

[1226] The terminal compresses the newly collected voice data and sends it to the server again. The input is the newly collected voice data, and the output is the data sent to the server.

[1227] The server analyzes the new voice data, compares it with the previous feedback result to calculate a new score, and generates feedback on improvements. The input is the new voice data and the previous feedback result, and the output is the new score and feedback information.

[1228] By dividing the specific processing of this system into these steps, it is possible to effectively support the improvement of the user's singing technique.

[1229] (Application example 2)

[1230] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1231] In modern factories, humans and robots are increasingly working together, but there are many challenges in providing instructions on-site and improving work efficiency. In particular, the emotional state of workers can affect work efficiency and robot performance, so there is a need for a means to detect this in real time and take countermeasures. In addition, there is a lack of systems that provide real-time feedback and optimize work procedures.

[1232] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1233] In this invention, the server includes a means for analyzing the user's voice data and detecting the emotional state, a means for optimizing the work procedure based on the user's voice instruction data, and a means for generating an ideal work procedure and providing feedback, thereby making it possible to detect the emotional state of the worker in real time and improve work efficiency.

[1234] "User voice data" refers to voice data when a factory worker gives instructions to a robot.

[1235] "Collection means" refers to the device or process used to acquire and transmit audio data to a server.

[1236] "Analysis of audio data" refers to the process of extracting audio features such as pitch, rhythm, vibrato, and long tones and analyzing the data.

[1237] "Emotional state" refers to the user's emotions inferred from the intonation, tempo, and sound patterns of the user's voice.

[1238] "Ideal singing data" refers to voice data of optimal singing that a user should aim for, which is generated based on analyzed voice data and emotional state.

[1239] "Feedback" refers to specific advice and guidance information provided to the user based on the analysis results.

[1240] "Work procedure optimization" refers to the process of deriving the most efficient work procedure by taking into account the user's voice instructions and emotional state.

[1241] A "generative AI model" refers to an algorithm or system that analyzes voice data and emotional data to generate ideal data and feedback.

[1242] A "prompt" refers to an instruction or question input to a generative AI model.

[1243] To put the present invention into practice, a system is constructed in which users, terminals, and servers work in cooperation with each other. Specific embodiments are as follows.

[1244] Audio data collection

[1245] The user launches a dedicated application installed on a device such as a smartphone or tablet installed in the factory. Work instructions and confirmations are entered into the application by voice, and the device's built-in microphone collects the voice data. This voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1246] Analysis of voice data and optimization of work procedures

[1247] The server analyzes the received voice data using a voice analysis library such as librosa to extract voice features such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify emotions such as tension or relaxation.

[1248] Generate ideal work procedures

[1249] Based on the analysis results and the user's emotional state, the server uses a generative AI model to generate an ideal workflow that reflects the user's characteristics and emotional state. The server generates the results in text format or prompts and sends them back to the user's device via the Internet.

[1250] Generating and Providing Feedback

[1251] The terminal provides a function to display the received ideal work procedures. The user can compare their own work procedures with the generated ideal work procedures. The terminal also displays the analysis results and feedback from the server, allowing the user to identify specific areas for improvement. For example, if a specific task is not performed efficiently, the terminal will provide feedback on the reason and how to improve it.

[1252] Collecting new data and providing feedback

[1253] The user performs an action based on the provided feedback and has the device collect new voice data. The newly collected voice data is compressed again and sent to the server. The server analyzes the new voice data and compares it with the previous feedback result. A new score is calculated and feedback is generated again with improvements.

[1254] Specific examples

[1255] For example, "Factory Worker A" uses his smartphone to issue a voice command such as "Please proceed to the next process," and this voice data is sent from the device to a server. The server analyzes pitch, rhythm, and emotion to generate an ideal work procedure. This data is sent to Factory Worker A's device, and Worker A proceeds with the work while referring to this ideal work procedure. Factory Worker A practices based on the feedback provided by the server, sends new voice data again, and receives further feedback. By repeating this process, Factory Worker A makes improvements to ensure that his work proceeds more efficiently.

[1256] Prompt Sentence Examples

[1257] "What are the next steps to ensure the robot works efficiently? And if the worker is nervous, how should we respond?"

[1258] This system allows users to understand their own emotional state in real time and work based on specific advice. By utilizing advanced analysis technology and an emotion engine, it provides effective feedback suited to individual work styles and emotional states, thereby improving user satisfaction and work efficiency.

[1259] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1260] Step 1:

[1261] A user launches a dedicated application installed on a device such as a smartphone or tablet and inputs voice commands. The device uses a built-in microphone to collect and temporarily store voice data. The input data is the user's voice commands, and the output data is the collected voice file.

[1262] Step 2:

[1263] The device compresses the collected audio data and sends it to a server via the Internet. The input is a temporarily saved audio file, and the output is compressed audio data.

[1264] Step 3:

[1265] The server analyzes the received audio data using a voice analysis library such as librosa to extract features such as pitch, rhythm, vibrato, and long tones. The input is compressed audio data, and the output is extracted audio feature data.

[1266] Step 4:

[1267] The server uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify the emotional state. The input is voice feature data, and the output is the recognized emotional state.

[1268] Step 5:

[1269] The server uses a generative AI model based on the analysis results and emotional state to generate an ideal work procedure. The generated ideal work procedure reflects the user's characteristics and emotional state. The input is voice feature data and emotional state, and the output is ideal work procedure data.

[1270] Step 6:

[1271] The server generates the ideal work procedure data in text format and prompts, and sends them to the user's terminal via the Internet. The input is the ideal work procedure data, and the output is text-format work procedure feedback.

[1272] Step 7:

[1273] The user uses a terminal to check the received ideal work procedure and compare it with their own work procedure. The input is the textual work procedure feedback from the server, and the output is the ideal work procedure that the user checks.

[1274] Step 8:

[1275] The terminal displays the analysis results and feedback, allowing the user to see specific improvements. The input is the feedback data from the server, and the output is the feedback information displayed on the terminal screen.

[1276] Step 9:

[1277] The user performs a new task based on the feedback and causes the terminal to collect voice data. The newly collected voice data is compressed again and sent to the server. The input is the new voice instruction data, and the output is the compressed voice data sent to the server.

[1278] Step 10:

[1279] The server analyzes the new voice data and compares it with the previous feedback result. It calculates a new score based on the analysis result and generates feedback indicating areas for improvement. The input is the new voice data and the previous feedback result, and the output is the new score and feedback indicating areas for improvement.

[1280] Through the above steps, users can work efficiently and effectively, and the server and terminal work together to create a system that provides optimal feedback in real time.

[1281] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1282] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1283] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1284] [Fourth embodiment]

[1285] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1286] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1287] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1288] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1289] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1290] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1291] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1292] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1293] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1294] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1295] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1296] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1297] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1298] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[1299] Audio data collection

[1300] User

[1301] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[1302] Terminal

[1303] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1304] Analysis of audio data and generation of ideal singing data

[1305] server

[1306] The server analyzes the received audio data. Specifically, it extracts characteristics such as pitch, rhythm, vibrato, and long tones. Based on the analysis results, a generative AI model that has learned the user's singing style generates ideal singing data. This ideal singing data reflects the user's characteristics.

[1307] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[1308] Generating and Providing Feedback

[1309] Terminal

[1310] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[1311] server

[1312] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results, including specific suggestions and ways to improve on pitch deviations, rhythmic irregularities, and lack of vibrato.

[1313] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[1314] Practice and score improvement support

[1315] User

[1316] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[1317] Terminal

[1318] The newly collected audio data is compressed again and sent to the server.

[1319] server

[1320] The server analyzes the new audio data and compares it with the previous feedback. It generates a new score and points for improvement and sends the results to the device. The user can then practice based on this feedback and aim to improve their karaoke score.

[1321] Specific examples

[1322] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, sings again, sends the audio data, and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[1323] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[1324] The processing flow will be explained below.

[1325] Step 1:

[1326] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[1327] Step 2:

[1328] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[1329] Step 3:

[1330] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[1331] Step 4:

[1332] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[1333] Step 5:

[1334] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[1335] Step 6:

[1336] Server: Based on the analysis results, the user's singing data is input into a generative AI model to generate ideal singing data.

[1337] Step 7:

[1338] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[1339] Step 8:

[1340] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[1341] Step 9:

[1342] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[1343] Step 10:

[1344] Server: Based on the analysis results and the differences between the user's singing, the server generates specific feedback in text and video format, including suggestions for improvement such as pitch deviations, rhythmic inconsistencies, and lack of vibrato.

[1345] Step 11:

[1346] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[1347] Step 12:

[1348] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[1349] Step 13:

[1350] Terminal: The newly collected voice data is recompressed and sent to the server.

[1351] Step 14:

[1352] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[1353] Step 15:

[1354] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[1355] Through this process, users can continue to practice effectively while receiving detailed feedback, steadily improving their karaoke scores.

[1356] Example 1

[1357] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1358] Conventional karaoke practice systems have made it difficult for users to obtain specific feedback to effectively improve their singing skills. Due to low analysis accuracy, they were unable to provide advice tailored to specific singing styles, limiting user satisfaction and actual skill improvement. Furthermore, they did not generate ideal singing data that reflected each user's individual characteristics, making it difficult to use for comparison or practice.

[1359] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1360] In this invention, the server includes means for analyzing the user's voice data and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data using a generative AI model based on the analysis results, and means for comparing the user's singing with the generated singing data and generating specific feedback. This enables advanced analysis of the user's singing data and provides individually optimized feedback, thereby improving skill and increasing satisfaction.

[1361] 1. "User's voice data" refers to the waveform information of voice collected by a user using a terminal, including karaoke singing voices.

[1362] 2. "Means of collection" refers to the function of detecting the user's voice through the built-in microphone or peripheral devices and recording it as digital data.

[1363] 3. "Means for transmitting to a server" refers to the communications means for transferring audio data from a terminal to a server via the Internet.

[1364] 4. "Means of analysis" refers to algorithms or software for extracting singing characteristics such as pitch, rhythm, vibrato, and long tones from the received audio data.

[1365] 5. "Generative AI model" is an artificial intelligence model that learns the user's singing style and generates ideal singing data based on the results.

[1366] 6. "Means for generating ideal singing data" refers to the process of using the analysis results and generative AI model to generate ideal singing data that reflects the user's characteristics.

[1367] 7. "Means for transmitting generated singing data to a user terminal" refers to a communication means for transferring ideal singing data from a server to a user terminal via the Internet.

[1368] 8. "Means for comparing a user's singing with the generated singing data" refers to an algorithm or software that compares a user's actual singing data with the generated ideal singing data and evaluates the differences.

[1369] 9. "Means for generating specific feedback" refers to the process of generating advice based on the comparison results, including areas for improvement such as pitch deviations, rhythmic irregularities, and lack of vibrato.

[1370] 10. "Means for displaying feedback content to the user" refers to the functionality for displaying the generated feedback on the user's device, including text and video formats.

[1371] 11. "Means for collecting new voice data and sending it again to the server" refers to a communication means for collecting the voice of the user singing again and sending it again to the server.

[1372] 12. "Means for generating a new score and improvements based on the results of reanalysis" refers to the process of analyzing the resubmitted voice data, evaluating the user's progress, and generating a new score and improvements.

[1373] The present invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Specific embodiments of the system will be described below.

[1374] Audio data collection

[1375] User

[1376] Users launch a dedicated app installed on their smartphone, tablet, or other device. Using the app, they select the karaoke song they want to sing and check the screen that displays the lyrics. At this time, an audio accompaniment is also played.

[1377] Terminal

[1378] As soon as the user starts singing, the device collects audio data through the built-in microphone, temporarily stores the data locally, compresses it, and transmits it over the Internet to a server.

[1379] Analysis of audio data and generation of ideal singing data

[1380] server

[1381] The server analyzes the received audio data. During this process, characteristics such as pitch, rhythm, vibrato, and long tones are extracted. Based on the analysis results, a prompt is input into the generative AI model to generate ideal singing data. Specifically, the following prompt is used:

[1382] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[1383] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[1384] Generating and Providing Feedback

[1385] Terminal

[1386] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[1387] server

[1388] The server compares the user's singing data with the generated ideal singing data. Based on the comparison results, it generates feedback that includes specific suggestions for improvement, such as pitch deviations, rhythm disturbances, and lack of vibrato. The feedback is created in text and video format and sent to the device.

[1389] Practice and score improvement support

[1390] User

[1391] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[1392] Terminal

[1393] The newly collected audio data is compressed again and sent to a server via the Internet.

[1394] server

[1395] The server analyzes the new audio data and compares it with the previous feedback to generate a new score and areas for improvement. These results are then sent back to the device, where the user can receive further feedback and practice.

[1396] Specific examples

[1397] For example, "User A" selects a karaoke song and starts singing. The device collects the audio data, compresses it, and sends it to the server. The server analyzes the audio data and inputs the following prompt sentence into the generative AI model:

[1398] "Based on user A's voice data, generate ideal singing data using the following criteria: pitch, rhythm, vibrato, long tone. Reflect user A's singing style."

[1399] The ideal singing data generated from the model is sent from the server to User A's device, where User A can listen to the singing data. User A can also practice based on specific points for improvement provided by the server. By repeating this process, User A can improve their skills.

[1400] This system allows users to understand their singing technique in real time and practice based on specific advice. By utilizing advanced analysis technology and generative AI models, it provides effective feedback tailored to each individual singing style, resulting in user satisfaction and improvement of their technique.

[1401] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1402] Step 1: Collecting audio data

[1403] User

[1404] The user launches the dedicated app installed on their device and selects the karaoke song they want to sing. The app displays the lyrics on the screen and plays the accompanying audio. At this point, the user can begin singing.

[1405] Input: User singing voice

[1406] Output: Audio data

[1407] Terminal

[1408] The device uses a built-in microphone to collect voice data as soon as the user starts singing. This voice data is temporarily stored in the device as digital data, and then the collected voice data is compressed.

[1409] Input: User singing voice

[1410] Output: Compressed audio data

[1411] Specific actions

[1412] The device automatically activates the microphone when the user starts singing, stores the collected voice data in a buffer, and then applies a compression algorithm to reduce the data size.

[1413] Step 2: Sending audio data to the server

[1414] Terminal

[1415] The compressed audio data is transmitted to a server via the Internet.

[1416] Input: Compressed audio data

[1417] Output: Status of sending to server

[1418] Specific actions

[1419] The device establishes an Internet connection and sends the compressed audio data to the server using the HTTPS protocol.

[1420] Step 3: Analyzing the audio data

[1421] server

[1422] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones.

[1423] Input: Compressed audio data

[1424] Output: Feature extracted data

[1425] Specific actions

[1426] The analysis module on the server decodes the audio data and uses signal processing technology to extract various characteristics as numerical values, allowing for a detailed analysis of the user's singing data.

[1427] Step 4: Generate ideal singing data

[1428] server

[1429] The server generates ideal singing data using a generative AI model based on the extracted feature data, using the following prompt:

[1430] "Based on the user's voice data, generate ideal singing data based on the following criteria: pitch, rhythm, vibrato, long tone. Reflect the user's singing style."

[1431] Input: feature extraction data, prompt statement

[1432] Output: Ideal singing data

[1433] Specific actions

[1434] The generative AI model receives the prompt sentence and uses the feature extraction data to generate ideal singing data, which reflects the user's singing style and produces highly accurate singing examples.

[1435] Step 5: Send ideal singing data to your device

[1436] server

[1437] The generated ideal singing data is again transmitted to the user's terminal via the Internet.

[1438] Input: Ideal singing data

[1439] Output: Sending status to user terminal

[1440] Specific actions

[1441] The server establishes an Internet connection and transmits the generated ideal singing data to the user's device using the HTTPS protocol.

[1442] Step 6: Generate feedback

[1443] server

[1444] The server compares the user's singing data with the generated ideal singing data and generates specific feedback, including areas for improvement such as pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[1445] Input: User's singing data, ideal singing data

[1446] Output: Feedback data

[1447] Specific actions

[1448] A comparison algorithm in the server analyzes the user's singing data and the ideal singing data, assesses the differences, and generates feedback based on the results.

[1449] Step 7: Displaying feedback on your device

[1450] Terminal

[1451] The device displays the received feedback to the user. Feedback is provided in text and video formats.

[1452] Input: Feedback data

[1453] Output: Feedback information to the display screen

[1454] Specific actions

[1455] The device analyzes the feedback data and displays it in text and video format on the user interface, allowing the user to continue practicing while checking the feedback.

[1456] (Application example 1)

[1457] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1458] Conventional karaoke systems have difficulty providing specific feedback to improve users' singing skills, and lack real-time support to help users improve their singing. This makes it difficult for users to efficiently improve their singing skills.

[1459] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1460] In this invention, the server includes means for collecting user voice data, means for transmitting the collected voice data to the server, means for analyzing the user voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones, means for generating ideal singing data based on the analysis results, means for transmitting the generated singing data to a user terminal, means for comparing the user's singing with the generated singing data, means for generating specific feedback based on the comparison results, means for displaying the feedback to the user, means for collecting new voice data for the user's practice and transmitting it again to the server, means for generating a new score and points for improvement based on the reanalysis results, means for providing real-time support for the user's practice using a device in a physical store, and means for displaying the analysis data and feedback in real time, thereby enabling the user to specifically and effectively improve their singing technique in real time.

[1461] "User's voice data" is voice information collected when a user sings a karaoke song.

[1462] The "means of collection" refers to a device that has the function of recording the voice while singing using a device such as a smartphone or tablet.

[1463] The "transmitting means" is a communication means that has the function of sending the collected voice data to a server via the Internet.

[1464] The "analyzing means" is software that executes an algorithm to analyze the audio data received by the server and extract characteristics such as pitch, rhythm, vibrato, and long tones.

[1465] The "means of generation" is software with an algorithm that creates ideal singing data based on the analysis results.

[1466] The "means for transmitting to the user terminal" is a communication means for returning the generated singing data to the user's device.

[1467] The "comparison means" is software that has the function of comparing the user's singing data with the generated ideal singing data and checking the differences.

[1468] The "means for generating feedback" is software that generates improvements and advice for the user based on the comparison results.

[1469] The "means for displaying the feedback content" is a device having a screen display function for displaying the feedback content as text or video on the user's device.

[1470] The "means for collecting new voice data" is a recording device that collects new voice data when the user sings again.

[1471] The "means for transmitting to the server again" is a communication means for sending newly collected voice data to the server again.

[1472] The "means for generating a new score and improvements based on the reanalysis results" is software that analyzes the newly transmitted audio data, compares it with the previous feedback, and generates a new score and improvements.

[1473] "Physical store devices" are devices such as tablets and smartphones used in physical stores such as karaoke booths.

[1474] The "means for providing real-time support" is a system that uses devices in physical stores to analyze the user's singing in real time and provide immediate feedback.

[1475] The "means for displaying analysis data and feedback in real time" is a device having a screen display function for displaying analysis results and feedback content in real time while the user is singing.

[1476] The system for realizing this invention consists of a series of processes that collects user voice data, transmits the data to a server for analysis, and provides feedback based on the interpreted information. This system uses a smartphone or tablet as the main hardware and a generative AI model as the software.

[1477] 1. Collection of audio data

[1478] A user selects a karaoke song using a device in a physical store (e.g., a tablet installed in a karaoke booth) and starts singing. The device collects the user's singing voice through a built-in microphone and temporarily stores the voice data.

[1479] 2. Data compression and transmission

[1480] The collected audio data is compressed on the device, which then transmits the compressed data to a server over the internet, either via Wi-Fi or 4G / 5G networks.

[1481] 3. Analysis of audio data

[1482] The server analyzes the received audio data and extracts characteristics such as pitch, rhythm, vibrato, and long tones. This analysis uses machine learning algorithms and digital signal processing technology. The generative AI model used learns the user's characteristics and generates ideal singing data based on them.

[1483] 4. Generation and transmission of ideal singing data

[1484] The server generates ideal singing data using the generative AI model, which is then sent to the user's device via the internet.

[1485] 5. Providing Feedback

[1486] The user's device plays the ideal singing data that was sent, allowing the user to compare their singing with the ideal singing. The device also displays specific feedback to help the user understand areas for improvement. The feedback is displayed in text and video format.

[1487] 6. Practice and score improvement support

[1488] The user attempts to sing again based on the provided feedback. The device collects new audio data and again sends it to the server. The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points. This information is again sent to the user's device.

[1489] Specific examples

[1490] For example, User A sings a song of his / her choice in a karaoke booth, and the audio data is sent from the tablet to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again. Through this process, User A can reliably improve his / her technique. An example of a prompt for a generative AI model is: "User A sings a karaoke song of his / her choice, and the audio data is sent from the tablet in the karaoke booth to the server. The server analyzes the pitch and rhythm and generates ideal singing data. This data is sent back to the tablet, and User A can listen to the example. User A practices based on the feedback provided by the server and sings again to receive feedback for the next time. Through this process, User A can reliably improve his / her technique and improve his / her karaoke score."

[1491] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1492] Step 1:

[1493] A user selects a karaoke song using a device in a brick-and-mortar store (e.g., a tablet in a karaoke booth), then launches an application on the tablet, which displays the lyrics and plays the accompanying music. As the user begins to sing, the tablet's built-in microphone collects audio data, which is temporarily stored on the device.

[1494] Input: User's voice

[1495] Output: Audio data file

[1496] Specific operation: The built-in microphone collects the user's singing voice as digital audio data and saves it as a file.

[1497] Step 2:

[1498] The device compresses the collected voice data to reduce communication load, and then transmits the compressed voice data to a server via the Internet. During this time, the device uses Wi-Fi or 4G / 5G networks.

[1499] Input: Audio data file

[1500] Output: Compressed audio data

[1501] What it does: The audio data is compressed using a standard audio compression algorithm (e.g. MP3, AAC) and uploaded to the server via the HTTP protocol.

[1502] Step 3:

[1503] The server analyzes the received audio data, extracting characteristics such as pitch, rhythm, vibrato, and long tones, and then uses a generative AI model to generate ideal singing data based on the analysis results.

[1504] Input: Compressed audio data

[1505] Output: Ideal singing data

[1506] Specific operation: Extract features from audio data using analysis algorithms and generative AI models to generate ideal singing data. Feature extraction uses audio signal processing techniques (e.g., Fast Fourier Transform).

[1507] Step 4:

[1508] The server transmits the generated ideal singing data to the user's terminal again via the Internet.

[1509] Input: Ideal singing data

[1510] Output: sent to the user's terminal

[1511] Specific operation: The generated singing data is sent to the user's terminal via the HTTP protocol, and the terminal receives the data.

[1512] Step 5:

[1513] The device plays back the received ideal singing data, compares it with the user's singing data, and generates specific feedback based on the analysis results received from the server, which is displayed to the user in real time.

[1514] Input: Ideal singing data, analysis results

[1515] Output: Comparison results, feedback

[1516] Specific operation: Using the audio playback function, the ideal singing data is played back, and a comparison algorithm is used to display the difference between the user's singing data and the ideal singing data. Feedback is also displayed on the screen in text and video format.

[1517] Step 6:

[1518] The user attempts to sing again based on the provided feedback, and the device collects new audio data and sends it back to the server.

[1519] Input: User's new voice

[1520] Output: New audio data file

[1521] Specific operation: New singing voices are collected using the built-in microphone, the data is compressed and sent to the server.

[1522] Step 7:

[1523] The server analyzes the newly received audio data and compares it with the previous feedback to generate a new score and improvement points, which are then sent back to the user's device.

[1524] Input: New audio data

[1525] Output: New score, improvements

[1526] Specific operation: The analysis algorithm is applied again, compared with the previous data to generate a new score and areas for improvement, and the results are sent to the user's device.

[1527] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1528] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[1529] Audio data collection

[1530] User

[1531] Users launch a dedicated app installed on their smartphone, tablet, or other device, select the karaoke song they want to sing, and a screen will appear displaying the lyrics and playing the accompanying audio.

[1532] Terminal

[1533] The device uses a built-in microphone to collect voice data as soon as the user starts singing. The collected voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1534] Analysis of audio data and generation of ideal singing data

[1535] server

[1536] The server analyzes the received voice data, specifically extracting characteristics such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotions from the voice data.

[1537] Emotion Engine

[1538] The emotion engine analyzes the voice data, the user's intonation, tempo, etc. to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model and reflected in the generation of ideal singing data.

[1539] server

[1540] The ideal singing data reflects the characteristics and emotional state of the user. The generated ideal singing data is transmitted to the user's terminal via the Internet.

[1541] Generating and Providing Feedback

[1542] Terminal

[1543] The device provides a function to play back the received ideal singing data, allowing users to compare their own singing with the ideal singing. The device also displays the analysis results and feedback from the server, allowing users to identify specific areas for improvement.

[1544] server

[1545] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The server also takes into account the user's emotional information recognized by the emotion engine. For example, the server may provide specific suggestions and suggestions for improvement regarding pitch inaccuracies, rhythmic irregularities, and lack of vibrato.

[1546] The generated feedback content is sent to the terminal in text and video format, allowing the user to receive specific guidance.

[1547] Practice and score improvement support

[1548] User

[1549] The user practices based on the provided feedback, sings the same song again, and has the device collect audio data.

[1550] Terminal

[1551] The newly collected audio data is compressed again and sent to the server.

[1552] server

[1553] The server analyzes the new voice data, compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[1554] Specific examples

[1555] For example, "User A" sings a karaoke song of their choice, and the audio data is sent from their device to the server. The server analyzes pitch, rhythm, and emotion to generate ideal singing data. This data is sent to User A's device, where User A can listen to the example. User A practices based on the feedback provided by the server, then sends new audio data again and receives further feedback. By repeating this process, User A can steadily improve their skills and increase their karaoke score.

[1556] Feedback using emotional information

[1557] For example, if user A is nervous during practice, the emotion engine can detect this tension and suggest breathing techniques or relaxation techniques to help them relax. In this way, feedback that takes into account the user's emotional information allows the user to receive more appropriate advice and effectively improve their skills.

[1558] This system allows users to understand their singing technique in real time and practice based on specific advice. Utilizing advanced analysis technology and an emotion engine, it provides effective feedback tailored to individual singing styles and emotional states, resulting in user satisfaction and improvement of technique.

[1559] The processing flow will be explained below.

[1560] Step 1:

[1561] User: Launch the dedicated app installed on a device such as a smartphone or tablet. Select the karaoke song you want to sing from the app.

[1562] Step 2:

[1563] On your device: Displays the lyrics of the selected song, plays the accompaniment audio, and enables the microphone to record the user's singing.

[1564] Step 3:

[1565] User: Start singing along with the accompaniment of the song. The device will record the user's singing in real time.

[1566] Step 4:

[1567] Terminal: Temporarily stores recorded audio data and compresses it. The compressed audio data is sent to a server via the Internet.

[1568] Step 5:

[1569] Server: Decodes the received audio data and begins analysis. Extracts features such as pitch, rhythm, vibrato, and long tones.

[1570] Step 6:

[1571] Server: Uses an emotion engine to recognize the user's emotions from the audio data, for example, identifying emotional states such as tension, joy, or sadness from changes in intonation and tempo.

[1572] Step 7:

[1573] Server: Based on the user's singing data and the recognized emotional information, the server uses a generative AI model to generate ideal singing data. The generated singing data reflects not only the characteristics of pitch and rhythm, but also the user's emotions.

[1574] Step 8:

[1575] Server: The server transmits the generated ideal singing data to the user's device via the Internet.

[1576] Step 9:

[1577] Device: The device saves the received ideal singing data and makes it playable. The user plays this voice data as a model.

[1578] Step 10:

[1579] User: Compares their singing with the ideal singing data. The device provides a comparison function and visually displays any deviations in pitch or rhythm.

[1580] Step 11:

[1581] Server: Based on the analysis results, the differences between the user's singing voice, and the recognized emotional information, the server generates specific feedback. For example, if a user feels nervous, the server will suggest relaxing breathing techniques, which includes advice tailored to the user's emotions.

[1582] Step 12:

[1583] Terminal: Displays the feedback sent from the server to the user. Video feedback is also provided, showing specific practice methods.

[1584] Step 13:

[1585] User: Practice at home based on the provided feedback. Sing the same song again and collect new audio data.

[1586] Step 14:

[1587] Terminal: The newly collected voice data is recompressed and sent to the server.

[1588] Step 15:

[1589] Server: Analyzes the new voice data and compares it with the previous feedback result, calculates a new score, and generates feedback on improvements.

[1590] Step 16:

[1591] Device: The newly calculated score and feedback are displayed to the user, who then practices again based on this feedback.

[1592] Through this process, users can receive detailed feedback and continue practicing effectively, steadily improving their karaoke scores. Taking the user's emotions into consideration also enables more personalized and effective instruction.

[1593] Example 2

[1594] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1595] While conventional singing instruction systems have the ability to provide feedback on pitch and rhythm, they have difficulty providing personalized advice that takes into account the user's emotional state. As a result, instruction is insufficient for users who are nervous or who need to express their emotions, making it difficult to achieve overall improvement in singing technique. Another problem is that users are unable to receive feedback based on their own singing characteristics or emotional state, which hinders effective practice.

[1596] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1597] In this invention, the server includes a means for collecting user voice data, a means including an emotion engine for recognizing emotion information, and a means for generating specific feedback, thereby enabling personalized feedback and advice that takes into account the user's voice characteristics and emotional state.

[1598] A "user" is an individual who uses the system to receive analysis of singing data and feedback.

[1599] "Audio data" refers to data collected in digital form from the voice of a user singing.

[1600] A "terminal" is a digital device that collects voice data and transmits it to a server.

[1601] A "server" is a remote computing device that analyzes audio data, generates feedback, and recognizes emotional information.

[1602] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotional state from voice data.

[1603] "Ideal singing data" is data that indicates an ideal singing performance that is analyzed based on the user's voice data and generated by a generative AI model.

[1604] "Feedback" is specific advice and information on areas for improvement aimed at improving the user's singing technique, based on the results of analyzing the audio data.

[1605] "Comparison" refers to evaluating the user's singing data and the ideal singing data and analyzing the differences between them.

[1606] A "score" is a numerical value calculated based on specific evaluation criteria for singing data, and quantitatively represents the user's singing performance.

[1607] A "generative AI model" is an artificial intelligence algorithm that generates an ideal singing performance based on the user's singing data.

[1608] A "prompt sentence" is text data containing instructions or questions to be input into a generative AI model.

[1609] This invention relates to a system that analyzes a user's voice data, generates ideal singing data, and provides feedback to the user based on the results. Furthermore, by incorporating an emotion engine that recognizes the user's emotions and taking the user's emotions into consideration when generating the feedback and singing data, it is possible to provide more effective and personalized advice. A specific embodiment of the system is described below.

[1610] Hardware and software used

[1611] Device: A digital device such as a smartphone or tablet that has a built-in microphone and can be operated by the user.

[1612] Server: A remote computing device on the cloud

[1613] Emotion Engine: Software for analyzing and recognizing emotions

[1614] Generative AI model: An artificial intelligence algorithm that generates ideal singing data

[1615] Audio data collection

[1616] Users launch a dedicated app installed on their smartphone or tablet and select the karaoke song they want to sing. The app displays lyrics and plays back an audio accompaniment. As soon as the user starts singing, the device uses its built-in microphone to collect audio data. This audio data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1617] Analysis of audio data and generation of ideal singing data

[1618] The server analyzes the received voice data and extracts features such as pitch, rhythm, vibrato, and long tones. It also uses an emotion engine to recognize the user's emotions from the voice data. The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. This emotional information, along with the analysis results, is input into the generative AI model.

[1619] Example prompt sentence:

[1620] "We have received the audio data of user A's karaoke song. Please extract the characteristics of pitch, rhythm, vibrato, and long tones, analyze the emotions with the emotion engine, and generate ideal singing data."

[1621] The server generates ideal singing data and transmits it again to the user's terminal via the Internet.

[1622] Generating and Providing Feedback

[1623] The device provides the user with the ability to play back the received ideal singing data. The user can compare their own singing with the ideal singing. The device also displays the analysis results and specific feedback from the server, allowing the user to identify areas for improvement. The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. This feedback also includes the user's emotional information recognized by the emotion engine.

[1624] Practice and score improvement support

[1625] The user practices based on the provided feedback and sings the same song again, causing the device to collect audio data. The newly collected audio data is compressed again and sent to the server. The server analyzes the new audio data, compares it with the previous feedback result, calculates a new score, and generates new feedback on areas for improvement.

[1626] Specific examples of operation

[1627] For example, User A chooses a favorite karaoke song, sings it, and sends the audio data from his / her device to the server. The server analyzes pitch, rhythm, vibrato, and emotion to generate ideal singing data. This data is sent to User A's device, where User A listens to the ideal singing and compares it with his / her own singing. User A practices based on feedback from the server, sends new audio data, and receives further feedback. By repeating this process, User A can steadily improve his / her skills and increase his / her karaoke score.

[1628] Additionally, if User A becomes nervous during practice, the emotion engine will detect this and provide advice such as, "Try some breathing techniques to relax." Feedback that utilizes emotion information in this way allows User A to effectively improve their skills.

[1629] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1630] Step 1: Collecting audio data

[1631] Users launch a dedicated karaoke app installed on their smartphone or tablet and select the karaoke song they want to sing. The input is the karaoke song selected by the user. The app displays the lyrics and plays the accompaniment audio.

[1632] The device collects voice data using a built-in microphone as soon as the user starts singing. The input is the user's singing voice, and the output is the voice data temporarily stored in the device.

[1633] The device compresses the collected audio data into a specified format (e.g., MP3 or FLAC). The compressed audio data is sent to a server via the Internet. The input is the compressed audio data, and the output is the audio data sent to the server.

[1634] Step 2: Analyzing the audio data

[1635] The server receives the voice data sent from the terminal. The input is the collected voice data. The data analysis engine in the server extracts features such as pitch, rhythm, vibrato, and long tones. The output is feature information of the voice data.

[1636] The emotion engine analyzes the voice data and the user's intonation and tempo to identify the user's emotional state. The input is the voice data's intonation and tempo information, and the output is the user's emotional information. This emotional information, along with the analysis results, is input into the generative AI model.

[1637] Step 3: Generate ideal singing data

[1638] The server inputs a prompt sentence into the generative AI model based on the feature information and emotional information of the audio data. An example of a prompt sentence is, "We have received audio data of User A's karaoke song. Please extract the features of pitch, rhythm, vibrato, and long tones, and analyze the emotions using the emotion engine to generate ideal singing data." The input is the prompt sentence and the analysis result information, and the generative AI model generates ideal singing data. The output is ideal singing data.

[1639] The server then transmits the generated ideal singing data to the user's terminal via the Internet. The input is the ideal singing data, and the output is the data transmitted to the user's terminal.

[1640] Step 4: Play and compare ideal singing data

[1641] The device provides the function to play back the received ideal singing data. The input is the singing data sent from the server. The user can compare their own singing with the ideal singing. The output is the user's listening state.

[1642] The terminal displays the analysis results and feedback sent from the server. The input is the analysis results and feedback information, allowing the user to identify areas that need improvement. The output is feedback displayed to the user.

[1643] Step 5: Generate feedback

[1644] The server compares the user's singing data with the ideal singing data and generates feedback based on the analysis results. The input is the user's singing data and the ideal singing data, and the output is specific feedback. This feedback also includes the user's emotional information recognized by the emotion engine.

[1645] Step 6: Practice and additional feedback

[1646] The user practices based on the provided feedback. They then sing the same song again, causing the device to collect audio data. The input is the feedback information. By collecting new audio data, the output is the collected new audio data.

[1647] The terminal compresses the newly collected voice data and sends it to the server again. The input is the newly collected voice data, and the output is the data sent to the server.

[1648] The server analyzes the new voice data, compares it with the previous feedback result to calculate a new score, and generates feedback on improvements. The input is the new voice data and the previous feedback result, and the output is the new score and feedback information.

[1649] By dividing the specific processing of this system into these steps, it is possible to effectively support the improvement of the user's singing technique.

[1650] (Application example 2)

[1651] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1652] In modern factories, humans and robots are increasingly working together, but there are many challenges in providing instructions on-site and improving work efficiency. In particular, the emotional state of workers can affect work efficiency and robot performance, so there is a need for a means to detect this in real time and take countermeasures. In addition, there is a lack of systems that provide real-time feedback and optimize work procedures.

[1653] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1654] In this invention, the server includes a means for analyzing the user's voice data and detecting the emotional state, a means for optimizing the work procedure based on the user's voice instruction data, and a means for generating an ideal work procedure and providing feedback, thereby making it possible to detect the emotional state of the worker in real time and improve work efficiency.

[1655] "User voice data" refers to voice data when a factory worker gives instructions to a robot.

[1656] "Collection means" refers to the device or process used to acquire and transmit audio data to a server.

[1657] "Analysis of audio data" refers to the process of extracting audio features such as pitch, rhythm, vibrato, and long tones and analyzing the data.

[1658] "Emotional state" refers to the user's emotions inferred from the intonation, tempo, and sound patterns of the user's voice.

[1659] "Ideal singing data" refers to voice data of optimal singing that a user should aim for, which is generated based on analyzed voice data and emotional state.

[1660] "Feedback" refers to specific advice and guidance information provided to the user based on the analysis results.

[1661] "Work procedure optimization" refers to the process of deriving the most efficient work procedure by taking into account the user's voice instructions and emotional state.

[1662] A "generative AI model" refers to an algorithm or system that analyzes voice data and emotional data to generate ideal data and feedback.

[1663] A "prompt" refers to an instruction or question input to a generative AI model.

[1664] To put the present invention into practice, a system is constructed in which users, terminals, and servers work in cooperation with each other. Specific embodiments are as follows.

[1665] Audio data collection

[1666] The user launches a dedicated application installed on a device such as a smartphone or tablet installed in the factory. Work instructions and confirmations are entered into the application by voice, and the device's built-in microphone collects the voice data. This voice data is temporarily stored on the device, compressed, and then sent to a server via the Internet.

[1667] Analysis of voice data and optimization of work procedures

[1668] The server analyzes the received voice data using a voice analysis library such as librosa to extract voice features such as pitch, rhythm, vibrato, and long tones. It then uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify emotions such as tension or relaxation.

[1669] Generate ideal work procedures

[1670] Based on the analysis results and the user's emotional state, the server uses a generative AI model to generate an ideal workflow that reflects the user's characteristics and emotional state. The server generates the results in text format or prompts and sends them back to the user's device via the Internet.

[1671] Generating and Providing Feedback

[1672] The terminal provides a function to display the received ideal work procedures. The user can compare their own work procedures with the generated ideal work procedures. The terminal also displays the analysis results and feedback from the server, allowing the user to identify specific areas for improvement. For example, if a specific task is not performed efficiently, the terminal will provide feedback on the reason and how to improve it.

[1673] Collecting new data and providing feedback

[1674] The user performs an action based on the provided feedback and has the device collect new voice data. The newly collected voice data is compressed again and sent to the server. The server analyzes the new voice data and compares it with the previous feedback result. A new score is calculated and feedback is generated again with improvements.

[1675] Specific examples

[1676] For example, "Factory Worker A" uses his smartphone to issue a voice command such as "Please proceed to the next process," and this voice data is sent from the device to a server. The server analyzes pitch, rhythm, and emotion to generate an ideal work procedure. This data is sent to Factory Worker A's device, and Worker A proceeds with the work while referring to this ideal work procedure. Factory Worker A practices based on the feedback provided by the server, sends new voice data again, and receives further feedback. By repeating this process, Factory Worker A makes improvements to ensure that his work proceeds more efficiently.

[1677] Prompt Sentence Examples

[1678] "What are the next steps to ensure the robot works efficiently? And if the worker is nervous, how should we respond?"

[1679] This system allows users to understand their own emotional state in real time and work based on specific advice. By utilizing advanced analysis technology and an emotion engine, it provides effective feedback suited to individual work styles and emotional states, thereby improving user satisfaction and work efficiency.

[1680] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1681] Step 1:

[1682] A user launches a dedicated application installed on a device such as a smartphone or tablet and inputs voice commands. The device uses a built-in microphone to collect and temporarily store voice data. The input data is the user's voice commands, and the output data is the collected voice file.

[1683] Step 2:

[1684] The device compresses the collected audio data and sends it to a server via the Internet. The input is a temporarily saved audio file, and the output is compressed audio data.

[1685] Step 3:

[1686] The server analyzes the received audio data using a voice analysis library such as librosa to extract features such as pitch, rhythm, vibrato, and long tones. The input is compressed audio data, and the output is extracted audio feature data.

[1687] Step 4:

[1688] The server uses an emotion engine to recognize the user's emotional state from the voice data. The emotion engine analyzes the intonation and tempo of the user's voice to identify the emotional state. The input is voice feature data, and the output is the recognized emotional state.

[1689] Step 5:

[1690] The server uses a generative AI model based on the analysis results and emotional state to generate an ideal work procedure. The generated ideal work procedure reflects the user's characteristics and emotional state. The input is voice feature data and emotional state, and the output is ideal work procedure data.

[1691] Step 6:

[1692] The server generates the ideal work procedure data in text format and prompts, and sends them to the user's terminal via the Internet. The input is the ideal work procedure data, and the output is text-format work procedure feedback.

[1693] Step 7:

[1694] The user uses a terminal to check the received ideal work procedure and compare it with their own work procedure. The input is the textual work procedure feedback from the server, and the output is the ideal work procedure that the user checks.

[1695] Step 8:

[1696] The terminal displays the analysis results and feedback, allowing the user to see specific improvements. The input is the feedback data from the server, and the output is the feedback information displayed on the terminal screen.

[1697] Step 9:

[1698] The user performs a new task based on the feedback and causes the terminal to collect voice data. The newly collected voice data is compressed again and sent to the server. The input is the new voice instruction data, and the output is the compressed voice data sent to the server.

[1699] Step 10:

[1700] The server analyzes the new voice data and compares it with the previous feedback result. It calculates a new score based on the analysis result and generates feedback indicating areas for improvement. The input is the new voice data and the previous feedback result, and the output is the new score and feedback indicating areas for improvement.

[1701] Through the above steps, users can work efficiently and effectively, and the server and terminal work together to create a system that provides optimal feedback in real time.

[1702] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1703] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1704] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1705] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1706] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1707] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1708] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1709] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1710] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1711] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1712] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1713] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1714] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1715] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1716] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1717] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1718] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1719] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1720] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1721] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1722] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1723] The following is further disclosed regarding the above embodiment.

[1724] (Claim 1)

[1725] means for collecting user voice data;

[1726] means for transmitting the collected voice data to a server;

[1727] A means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones;

[1728] A means for generating ideal singing data based on the analysis results;

[1729] means for transmitting the generated singing data to a user terminal;

[1730] means for comparing the user's singing with the generated singing data;

[1731] a means for generating specific feedback based on the comparison results;

[1732] a means for displaying the feedback content to the user;

[1733] means for collecting new audio data for the user's practice and transmitting it again to the server;

[1734] The system includes a means for generating new scores and improvements based on the reanalysis results.

[1735] (Claim 2)

[1736] 2. The system according to claim 1, further comprising means for learning vocal characteristics based on the generated singing data.

[1737] (Claim 3)

[1738] 10. The system of claim 1, further comprising means for providing the generated feedback content in a video format.

[1739] "Example 1"

[1740] (Claim 1)

[1741] means for collecting user voice data;

[1742] means for transmitting the collected voice data to a server;

[1743] A means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones;

[1744] A means to generate ideal singing data using a generative AI model based on the analysis results, and

[1745] means for transmitting the generated singing data to a user terminal;

[1746] means for comparing the user's singing with the generated singing data;

[1747] a means for generating specific feedback based on the comparison results;

[1748] a means for displaying the feedback content to the user;

[1749] means for collecting new audio data for the user's practice and transmitting it again to the server;

[1750] The system includes a means for generating new scores and improvements based on the reanalysis results.

[1751] (Claim 2)

[1752] 2. The system according to claim 1, further comprising means for learning the user's singing style based on the generated singing data.

[1753] (Claim 3)

[1754] 10. The system of claim 1, further comprising means for providing the generated feedback content in text and video format.

[1755] "Application Example 1"

[1756] (Claim 1)

[1757] means for collecting user voice data;

[1758] means for transmitting the collected voice data to a server;

[1759] A means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones;

[1760] A means for generating ideal singing data based on the analysis results;

[1761] means for transmitting the generated singing data to a user terminal;

[1762] means for comparing the user's singing with the generated singing data;

[1763] a means for generating specific feedback based on the comparison results;

[1764] a means for displaying the feedback content to the user;

[1765] means for collecting new audio data for the user's practice and transmitting it again to the server;

[1766] a means for generating new scores and improvements based on the reanalysis results;

[1767] a means for providing real-time support of a user's practice using a device in a physical store;

[1768] A system that includes a means for displaying analytical data and feedback in real time.

[1769] (Claim 2)

[1770] 2. The system according to claim 1, further comprising means for learning vocal characteristics based on the generated singing data.

[1771] (Claim 3)

[1772] 10. The system of claim 1, further comprising means for providing the generated feedback content in a video format.

[1773] "Example 2: Combining Emotion Engines"

[1774] (Claim 1)

[1775] means for collecting user voice data;

[1776] means for transmitting the collected voice data to a server;

[1777] A means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones;

[1778] A means for generating ideal singing data based on the analysis results;

[1779] means for transmitting the generated singing data to a user terminal;

[1780] means including an emotion engine for recognizing emotion information from the user's voice data;

[1781] means for comparing the user's singing with the generated singing data;

[1782] a means for generating specific feedback based on the comparison results;

[1783] a means for displaying the feedback content to the user;

[1784] means for detecting an emotional state such as tension of a user and providing advice for relaxation;

[1785] means for collecting new audio data for the user's practice and transmitting it again to the server;

[1786] The system includes a means for generating new scores and improvements based on the reanalysis results.

[1787] (Claim 2)

[1788] 2. The system according to claim 1, further comprising means for learning voice characteristics based on the generated singing data.

[1789] (Claim 3)

[1790] 10. The system of claim 1, further comprising means for providing the generated feedback content in a video format.

[1791] "Application example 2 when combining emotion engines"

[1792] (Claim 1)

[1793] means for collecting user voice data;

[1794] means for transmitting the collected voice data to a server;

[1795] A means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones;

[1796] A means for generating ideal singing data based on the analysis results;

[1797] means for transmitting the generated singing data to a user terminal;

[1798] means for comparing the user's singing with the generated singing data;

[1799] a means for generating specific feedback based on the comparison results;

[1800] a means for displaying the feedback content to the user;

[1801] means for collecting new audio data for the user's practice and transmitting it again to the server;

[1802] a means for generating new scores and improvements based on the reanalysis results;

[1803] means for detecting the emotional state of a user and providing feedback according to the emotion;

[1804] A means for optimizing a work procedure based on the emotional state of a user;

[1805] A system including:

[1806] (Claim 2)

[1807] 2. The system according to claim 1, further comprising means for learning vocal characteristics based on the generated singing data.

[1808] (Claim 3)

[1809] 10. The system of claim 1, further comprising means for providing the generated feedback content in a video format. [Explanation of symbols]

[1810] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for collecting user voice data; means for transmitting the collected voice data to a server; A means for analyzing the user's voice data in the server and extracting characteristics such as pitch, rhythm, vibrato, and long tones; A means for generating ideal singing data based on the analysis results; means for transmitting the generated singing data to a user terminal; means for comparing the user's singing with the generated singing data; a means for generating specific feedback based on the comparison results; a means for displaying the feedback content to the user; means for collecting new audio data for the user's practice and transmitting it again to the server; The system includes a means for generating new scores and improvements based on the reanalysis results.

2. The system according to claim 1, further comprising means for learning vocal characteristics based on the generated singing data.

3. The system of claim 1 , further comprising means for providing the generated feedback content in a video format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A