system

The system addresses the limitations of conventional karaoke by using AI to provide real-time feedback and virtual reality experiences, enhancing skill improvement and community interaction.

JP2026071677APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional karaoke systems lack immersive experiences and specific feedback for skill improvement, limiting opportunities for musical expression and community formation.

Method used

A system that uses AI to analyze user voice data in real-time, providing immersive visual and auditory feedback through virtual reality, and a community platform for skill sharing and evaluation.

Benefits of technology

Enables users to improve their singing skills through real-time feedback and immersive experiences, fostering a sense of accomplishment and community engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071677000001_ABST
    Figure 2026071677000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] The system acquires user voice data in real time. A means for analyzing the aforementioned audio data and providing feedback to the user, A means of providing users with real-time visual and auditory feedback in a virtual reality environment, A means of providing a community platform on which the aforementioned feedback can be shared with other users, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system. [[ID=�]]

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional karaoke systems have been limited to enjoying individual music performances and have been difficult to provide an immersive experience like a live concert. Furthermore, there has been a lack of sufficient means to obtain specific feedback for users who want to improve their singing ability, and skill sharing with others has also been limited, resulting in few opportunities for new forms of expression and community formation.

Means for Solving the Problems

[0005] This invention provides a system that acquires user voice data in real time, analyzes it in detail using AI technology, and provides feedback. It also provides users with immersive visual and auditory feedback, similar to that of a live concert, using virtual reality technology. Furthermore, by incorporating a community platform for sharing the feedback with other users, it offers opportunities for mutual evaluation and skill development among users.

[0006] "Audio data" refers to digital audio information obtained from a user's speech.

[0007] "Analysis" is the process of thoroughly analyzing acquired audio data and evaluating specific elements.

[0008] "Feedback" refers to information about areas for improvement and evaluations provided to users based on the analysis results.

[0009] A "virtual reality environment" is a virtual space in a computer-generated three-dimensional space that allows users to immerse themselves in it intuitively.

[0010] "Visual and auditory feedback" refers to the provision of visual and auditory information that users can perceive in a virtual reality environment.

[0011] "Real-time" refers to a state where processing or responses occur almost instantly, with virtually no time delay.

[0012] A "community platform" is an online space where users can share their achievements and information and interact with each other.

[0013] A "user" is a person who uses this system to sing or play an instrument. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides a system for users to receive real-time feedback during a karaoke experience and improve their singing skills. The system mainly consists of a terminal, a server, and a virtual reality device.

[0036] The user first launches a karaoke application on their device and begins voice input. The device acquires the user's singing voice as digital audio data and sends it to the server in real time. The server then analyzes this audio data. Specifically, it uses AI technology to evaluate elements such as pitch, rhythm, and sound quality, and performs a detailed analysis of the user's performance.

[0037] After analysis, the server generates specific feedback for the user and sends it to their terminal. The feedback points out things like pitch discrepancies and rhythm inconsistencies and provides advice on how to improve.

[0038] Furthermore, when a user uses a virtual reality device, the device constructs a virtual live stage and generates audience reactions that follow the user's performance. This allows the user to have an immersive experience as if they were actually at a concert. The server provides and transmits real-time visual and auditory feedback, such as audience applause and cheers, to the device.

[0039] The system also includes a community platform where users can share their performance with other users. Users upload their performance results to the platform and receive comments and ratings from other users. Through this feedback, users can gain information to further improve their skills.

[0040] As a concrete example, consider a scenario where a user sings "Happy Birthday." In this case, the device sends audio data to a server, which analyzes the parts where the pitch is too high and provides feedback such as, "Next time, try singing this part a little lower." Simultaneously, in a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when they reach the chorus, promoting a sense of accomplishment.

[0041] Thus, the present invention provides a practical and enjoyable means for users to not only enjoy karaoke but also to improve their musical skills.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The device acquires audio data from the microphone when the user launches a karaoke application and begins voice input.

[0045] Step 2:

[0046] The terminal transmits the acquired audio data to the server in real time. This transmission is performed using a protocol that minimizes data latency.

[0047] Step 3:

[0048] The server analyzes the received audio data using AI technology. Specifically, it evaluates pitch, rhythm, and timbre to determine the quality of the performance.

[0049] Step 4:

[0050] The server generates feedback based on the analysis results. For example, it identifies off-key notes and rhythmic discrepancies and compiles advice on areas for improvement.

[0051] Step 5:

[0052] The server sends the generated feedback to the terminal.

[0053] Step 6:

[0054] The device visualizes the received feedback for the user. This is displayed in text and simple graph formats, allowing the user to see specific areas for improvement.

[0055] Step 7:

[0056] When a user uses a VR device, the device constructs a virtual live stage. While tracking the user's movements, it prepares virtual audience reactions that synchronize with the performance.

[0057] Step 8:

[0058] The server generates visual and auditory feedback, such as applause and cheers, in accordance with the user's performance and sends it to the device in real time.

[0059] Step 9:

[0060] The device presents the user with these visual and auditory feedbacks, providing an immersive experience that makes them feel as if they are actually at a concert.

[0061] Step 10:

[0062] After the performance is complete, the device prepares to upload the recorded results to the community platform. This will only happen if the user allows it.

[0063] Step 11:

[0064] Users can watch performances uploaded by other users on the community platform and leave comments and ratings.

[0065] Step 12:

[0066] The server collects feedback from other users and notifies the original user. This allows the user to gain the insights necessary for self-improvement.

[0067] (Example 1)

[0068] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0069] Traditionally, entertainment activities like karaoke have lacked concrete feedback that allows users to objectively evaluate and improve their singing skills. Furthermore, opportunities for immersive experiences using virtual reality technology and for skill improvement through communication with other users are limited. This makes it difficult to improve users' singing skills and create a richer entertainment experience.

[0070] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0071] In this invention, the server includes means for acquiring and analyzing the user's acoustic data in real time and providing the user with suggestions for improvement; means for providing the user with visual and auditory feedback in real time in a virtual reality environment and generating virtual responses based on the user's actions; and means for providing an information exchange platform on which the suggestions for improvement can be shared with other participants. This allows the user to improve their singing skills through real-time feedback and enjoy responses that are just like those on a real live stage through an immersive experience in a virtual reality environment.

[0072] "Audio data" refers to information that digitally represents a user's vocalizations or musical expressions.

[0073] "Analysis" is the process of processing acoustic data and evaluating elements such as pitch, rhythm, and voice quality.

[0074] An "improvement suggestion" is a specific proposal provided to the user based on the analysis results, aimed at improving their singing skills.

[0075] A "virtual reality environment" is a computer-generated virtual space created using digital technology, into which a user can immerse themselves.

[0076] "Visual and auditory feedback" refers to responses provided to a user based on visual or auditory information.

[0077] "Virtual response" refers to the simulated audience or environment's reaction to a user's actions in a virtual reality environment.

[0078] An "information exchange platform" is a digital space where users can share analysis results and improvement suggestions, and interact with each other.

[0079] This invention is a system aimed at enabling users to improve their singing skills through karaoke. The system mainly consists of a terminal, a server, and a virtual reality device.

[0080] The user first launches a karaoke application on their device. This application has a voice input function and captures the user's singing voice through the microphone. The captured voice is converted into digital audio data. This digital audio data is then transmitted from the device to the server in real time.

[0081] The server uses AI technology to analyze the received digital audio data. Specifically, it uses AI libraries such as "TENSORFLOW®" and "PyTorch" to evaluate pitch, rhythm, and voice quality. This analysis generates specific improvement suggestions for the user. These improvement suggestions are then provided to the user via their device.

[0082] When a user is using a virtual reality device, the terminal constructs a virtual live stage and generates virtual reactions based on the user's performance. This includes visual and auditory feedback such as applause and cheers from a virtual audience. The server generates and sends this feedback to the terminal in real time, providing the user with an immersive experience.

[0083] Furthermore, the system includes an information exchange platform where users can share their performance results with other users. Through this platform, users can receive comments and evaluations from other users. This allows users to receive more detailed feedback and obtain information to improve their singing skills.

[0084] As a concrete example, consider a case where a user sings "Happy Birthday" and the system analyzes the performance. In this case, the terminal sends the audio data to a server, which analyzes parts that are too high in pitch and provides specific suggestions for improvement, such as "Next time, try singing this part a little lower." In a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when the chorus is reached, providing an experience that promotes a sense of accomplishment.

[0085] An example of a prompt for a generative AI model would be, "Please tell me how to provide real-time voice feedback in a karaoke system."

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The user launches the karaoke application on their device. The application displays a song list, allowing the user to select a song. The input is the user's song selection, and the output is information about the selected song.

[0089] Step 2:

[0090] The user selects a song and begins singing. The device acquires the user's singing voice in real time via the microphone and converts it into digital audio data. In this step, the input is the user's analog voice, and the output is digital audio data.

[0091] Step 3:

[0092] The terminal transmits the acquired digital audio data to the server via the network. Data compression is performed during transmission to reduce communication load. The input is digital audio data, and the output is a notification to the server indicating successful transmission.

[0093] Step 4:

[0094] The server inputs the received acoustic data into the AI ​​analysis system. Here, a generated AI model is used to analyze the pitch, rhythm, and sound quality. The input for this step is the acoustic data, and the output is the analysis result.

[0095] Step 5:

[0096] The server generates improvement suggestions for the user based on the analysis results. Specifically, it detects pitch discrepancies and rhythmic inconsistencies and outputs the areas for improvement as text. The input is the analysis results, and the output is the text of the improvement suggestions.

[0097] Step 6:

[0098] The server sends the suggested improvements to the terminal. Data compression is also performed here to ensure users receive feedback quickly. The input is the text of the suggested improvements, and the output is a notification that the transmission to the terminal is complete.

[0099] Step 7:

[0100] When the device is connected to a virtual reality system, it constructs a virtual live stage. It generates visual and auditory feedback, or virtual reactions, that respond to the user's performance. The input is user performance data, and the output is visual and auditory feedback in the virtual space.

[0101] Step 8:

[0102] Users can upload their singing performances and feedback to the community platform. They can also receive comments and ratings from other users. The input for this step is improvement suggestions and performance data, and the output is feedback from other users.

[0103] (Application Example 1)

[0104] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0105] Traditional karaoke systems lack the real-time, specific feedback necessary for users to effectively improve their singing skills. Furthermore, the virtual reality experience is limited, making it difficult to achieve the same level of immersion as attending an actual concert.

[0106] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0107] In this invention, the server includes means for acquiring the user's voice signal in real time, means for analyzing the voice signal and providing information to the user, means for providing the user with visual and auditory information in a virtual environment in real time, and means for providing real-time visualized feedback to the user while they are singing using a smart device. This allows the user to evaluate their singing skills in real time and make concrete improvements while enjoying an immersive experience in a virtual environment.

[0108] "User voice signal" refers to data obtained by converting the voice spoken by the user into an electrical signal.

[0109] "Real-time acquisition" refers to the instantaneous collection of the user's voice signal as they speak.

[0110] "Analysis" is the process of analyzing collected audio signals based on certain criteria and detecting their characteristics.

[0111] An "information-providing device" is a device that presents analysis results to the user and facilitates their understanding.

[0112] A "virtual environment" is a computer-generated space, distinct from the real world, constructed using digital technology.

[0113] "Visual and auditory information" refers to data and feedback, such as responses, that are provided to the user visually or aurally.

[0114] A "smart device" is a portable electronic device equipped with a processor and communication devices that enables interaction with the user.

[0115] "Visualized feedback" refers to information that displays the results obtained through analysis in a way that allows users to intuitively understand them.

[0116] "Audience reaction" refers to the actions and emotional expressions of fictional people in response to a particular event.

[0117] The system for realizing this invention includes a series of processes for acquiring and analyzing the user's voice signal and providing real-time feedback. First, the user launches a karaoke application using a smart device and begins singing. The smart device acquires the user's voice signal in real time using its built-in microphone. The acquired voice signal is transmitted to a server in the cloud.

[0118] The server analyzes the audio data using a speech recognition API (for example, a common API for speech recognition) and evaluates the pitch, rhythm, and sound quality. Based on the analysis results, it generates specific suggestions for improvement regarding the user's musical performance and sends them to the smart device. This feedback is visualized on the smart device's display for intuitive understanding by the user and is also provided as audio feedback.

[0119] Furthermore, in the virtual environment, a computer-generated audience reacts in sync with the user's singing. The virtual environment uses a 3D engine such as Unity to provide real-time visual and auditory feedback to the user. The server pre-generates these audience reactions and plays them back in real-time in sync with the user's singing, enhancing the performance effect.

[0120] For example, when a user sings "Happy Birthday," a pitch bar is displayed on the smart device's screen, and any off-key sections are highlighted in red. Furthermore, during the chorus, the sound of a virtual audience beginning to applaud is played to enhance the user's sense of accomplishment.

[0121] An example of a prompt might be, "While singing the chorus of 'Happy Birthday,' provide real-time feedback on pitch deviations and visualize the audience's reaction."

[0122] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0123] Step 1:

[0124] The user launches a karaoke application on their smart device and begins singing. The input is the user's voice signal, and the output is recorded as audio data on the smart device. The smart device acquires this audio signal in real time through its microphone.

[0125] Step 2:

[0126] The smart device transmits the acquired voice data to a server in the cloud. The input is voice data, and the output is data transfer to the server. The smart device uses a communication module to upload the data to the server in real time.

[0127] Step 3:

[0128] The server analyzes audio data using a speech recognition API. The input is audio data sent from a smart device, and the output is an evaluation result regarding pitch, rhythm, and sound quality. The server processes the audio based on specific musical criteria and generates the analysis results.

[0129] Step 4:

[0130] The server generates improvement suggestions based on the analysis results and sends them to the smart device. The input is the result of the voice analysis, and the output is a specific feedback message. The server uses an AI algorithm to point out deviations in pitch and rhythm and create advice for improvement.

[0131] Step 5:

[0132] Smart devices receive feedback from a server and visualize it on their display to present to the user. The input is improvement suggestion messages, and the output is visualized feedback. Smart devices display graphical elements on their screens to help users intuitively understand the areas for improvement.

[0133] Step 6:

[0134] The server generates a virtual environment, creates audience reactions synchronized with the user's singing, and plays them back in real time. The input is real-time user singing data, and the output is visual and auditory audience reactions. The server constructs the virtual environment using Unity or similar software and stages scenarios where the audience applauds and cheers.

[0135] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0136] This invention provides a karaoke system that utilizes an emotion engine to analyze user voice data and recognize emotions. The system consists of a terminal, a server, and a virtual reality device.

[0137] The user launches a karaoke application on their device and begins voice input. The device acquires the user's voice data in real time and sends it to the server. The server analyzes the voice data and uses AI technology to evaluate musical elements such as pitch, rhythm, and tone quality. In addition, an emotion engine identifies the user's emotions based on the voice data.

[0138] The server generates feedback based on the analysis results and recognized emotional information, and sends it to the terminal. This feedback can include suggestions for musical technical improvements, as well as responses to and encouragement regarding the user's emotions.

[0139] Furthermore, when using a virtual reality device, the terminal constructs a virtual live stage in real time. The reactions of the virtual audience, synchronized with the user's singing, are adjusted according to the user's emotions. For example, if the user is expressing joyful emotions, the audience's reactions will be set to be more lively and the cheers louder.

[0140] The system also integrates a community platform for users to share the results of their voice data analysis, including emotional information, with other users. Users can upload their performance results to the platform and receive emotionally-based feedback and comments from other users, which can help them feel empathy and boost their motivation.

[0141] As a concrete example, suppose the emotion engine detects a sense of exhilaration when a user sings a "celebratory song." In this case, the server includes emotionally appealing feedback such as "convey that joy even more." Furthermore, the audience in the VR environment will be shown an animation that further heightens the user's emotions.

[0142] This means that the present invention allows users to receive various emotionally appealing feedback while singing karaoke, enabling them to improve their musical skills and gain a platform for emotional expression.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] When the user launches a karaoke application and begins voice input, the device acquires audio data from the microphone in real time.

[0146] Step 2:

[0147] The terminal sends the acquired audio data to the server. This transmission occurs in real time, and a protocol that minimizes data latency is used.

[0148] Step 3:

[0149] The server analyzes the received audio data using AI technology. The analysis includes evaluation of pitch, rhythm, and sound quality to determine musical skill.

[0150] Step 4:

[0151] The server uses an emotion engine to recognize the user's emotions from the voice data. The recognized emotions include, for example, joy, sadness, and excitement.

[0152] Step 5:

[0153] The server generates specific and emotionally responsive feedback based on analysis results and emotion recognition results. This feedback includes not only suggestions for musical improvements but also messages that resonate with the user's emotions.

[0154] Step 6:

[0155] The server sends the generated feedback to the terminal.

[0156] Step 7:

[0157] The device displays the received feedback to the user. The feedback is visualized in text and graph formats to help the user understand it.

[0158] Step 8:

[0159] When a user uses a virtual reality device, the device constructs a virtual live stage and adjusts the audience's reactions in sync with the user's performance.

[0160] Step 9:

[0161] The server generates virtual audience reactions based on the recognized user's emotions and sends them to the terminal.

[0162] Step 10:

[0163] The device plays these reactions in a virtual environment, providing the user with visual and auditory feedback.

[0164] Step 11:

[0165] The device prepares to upload performance recording data to the community platform.

[0166] Step 12:

[0167] On the community platform, other users can view uploaded performances and leave comments and ratings based on their emotions.

[0168] Step 13:

[0169] The server collects feedback from other users and notifies the original user. This allows users to gain emotional resonance and useful feedback.

[0170] (Example 2)

[0171] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0172] While conventional karaoke systems allowed users to receive feedback on their singing technique, they lacked mechanisms to improve emotional expression and the quality of the virtual reality experience. Furthermore, there was a challenge in that users had difficulty receiving emotion-based feedback through interaction with other users.

[0173] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0174] In this invention, the server includes means for acquiring user voice information and analyzing said voice information to identify musical elements and emotions; means for generating a virtual reality environment that is dynamically adjusted based on the emotions extracted from said voice information and providing visual and auditory feedback in said environment; and means for providing an electronic platform on which the analysis results and feedback can be shared with other users. This makes it possible for users to receive feedback on emotional expression in real time and to share emotions through interaction with others.

[0175] "User voice information" refers to the collection of acoustic data containing the user's voice, acquired through a voice input device.

[0176] "Analysis" refers to the computational process of identifying musical elements and emotions based on acquired audio information.

[0177] "Musical elements" refer to characteristics related to music that are extracted from audio information, such as pitch, rhythm, and tone quality.

[0178] "Emotions" refer to information that indicates the user's mental state or feelings, as identified from audio information.

[0179] A "virtual reality environment" refers to a three-dimensional space or audiovisual experience generated by a computer that a user can immerse themselves in.

[0180] "Visual and auditory feedback" refers to information returned to the user through visual images and auditory sounds, intended to enhance the quality of the experience.

[0181] An "electronic platform" is a technological foundation that enables the sharing and exchange of information through user communities and digital networks.

[0182] This invention relates to a system that analyzes voice information input by a user and provides musical guidance and emotional feedback. The system consists of multiple technical elements, including a terminal, a server, and a virtual reality device.

[0183] Users input voice information via the microphone using a karaoke application installed on their device. This voice information is digitized on the device and transmitted to the server in real time as audio data. WebSocket is used to maintain real-time transmission of the audio data.

[0184] The server uses a machine learning framework to analyze the received audio information. Specifically, it evaluates musical elements such as pitch, rhythm, and tone quality using pitch analysis algorithms and beat tracking. The emotion engine also uses libraries such as OpenSmile to classify emotions from the intonation and tempo of the voice. Through this analysis, not only individual musical characteristics but also the user's emotional state is evaluated.

[0185] The analysis results are compiled into feedback provided to the user using natural language generation technology. This feedback is delivered to the user visually and audibly through the device. For example, messages such as "Sing with a clearer pitch" or "Try to convey the enjoyment more effectively" may be displayed. Real-time feedback is achieved through dynamic page rendering using HTML templates and JavaScript®.

[0186] When using a virtual reality device, the terminal also has the capability to build a realistic virtual live stage based on the emotions of the voice, using Unity or Unreal Engine. It generates virtual audience reactions that adjust according to the user's expression, providing immersion in both sight and sound.

[0187] Furthermore, users can upload the generated feedback to an electronic platform and share it with other users. Through interaction with other users, they can receive emotion-based feedback, which can boost their motivation.

[0188] For example, when a user sings a "celebratory song," the emotion engine recognizes the feeling of exhilaration, and the server provides feedback to the user saying, "Please convey that joy more." An example of a prompt to input into the generative AI model would be, "Please tell me how to analyze the audio data sung by the user, recognize the user's emotions, and generate feedback."

[0189] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0190] Step 1:

[0191] The terminal receives voice input from the karaoke application launched by the user and collects it through the microphone. The voice data as input is converted into a digital signal using an audio library and sent to the server in real time. This provides the voice data in a digital format that can be used for subsequent analysis processing.

[0192] Step 2:

[0193] The server receives audio data transmitted from the terminal. The received audio data is first subjected to pitch analysis algorithms and beat tracking to evaluate pitch, rhythm, and sound quality. These analyses produce evaluation data indicating the user's musical skills, and the process then proceeds to the next step of sentiment analysis.

[0194] Step 3:

[0195] The server performs sentiment analysis on the audio data in parallel with evaluating the musical elements. The sentiment engine uses the OpenSmile library to analyze the intonation and tempo of the voice and classify the user's emotions (e.g., exhilaration or sadness). This classified sentiment data is used to generate feedback.

[0196] Step 4:

[0197] The server integrates analysis data of musical elements and emotions to generate feedback. This process utilizes natural language generation technology to create specific improvement suggestions for the user (e.g., "Let's stabilize the pitch") and emotional messages (e.g., "Let's convey the enjoyment more"). The generated feedback is then formatted for visualization.

[0198] Step 5:

[0199] The device receives feedback from the server and presents it to the user visually and audibly. The input feedback data is displayed on the screen using HTML and JavaScript, allowing the user to see the results in real time. This phase aims to improve the user's singing experience through feedback.

[0200] Step 6:

[0201] When a virtual reality device is connected, the terminal uses a game engine to construct a virtual live stage. Based on the emotional data of the input audio, the reactions of the virtual audience are dynamically generated. For example, if the emotion is heightened, an animation of the virtual audience becoming excited will be displayed. This output provides the user with an immersive experience.

[0202] Step 7:

[0203] Users upload the analyzed data and generated feedback to an electronic platform. This step facilitates information sharing and communication with other users via their devices. The outputted feedback is used to solicit sentiment-based comments from other community members.

[0204] (Application Example 2)

[0205] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0206] Conventional voice analysis systems have been unable to provide real-time feedback that takes user emotions into account, posing a challenge in fostering emotional engagement with the audience, especially during live performances. Furthermore, there was a lack of specific feedback necessary for users to evaluate and improve their own performance.

[0207] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0208] In this invention, the server includes means for acquiring and analyzing user voice information in real time, means for providing the user with visual and auditory responses in a virtual reality space in real time, and means for analyzing the user's emotional expression during live performance streaming and acquiring and presenting emotional responses from viewers in real time. This allows the user to deepen their emotional interaction with viewers while receiving rich feedback in real time to evaluate and improve their performance.

[0209] "User voice information" refers to all voice data emitted by the user, and is usually collected in digital format via a microphone.

[0210] "Real-time acquisition" refers to a technology that minimizes the delay between transmission and acquisition, allowing for the immediate recording of voice information on the spot.

[0211] "Analysis" refers to the process of performing data processing on collected audio information to recognize specific patterns or emotions.

[0212] "Response" refers to feedback or reactions generated based on the analysis of audio information, and is the information or reaction provided to the user or audience.

[0213] A "virtual reality space" is a three-dimensional digital environment generated by a computer that can be experienced visually and aurally.

[0214] "Emotional expression" refers to the emotions and feelings that users express through words, music, and other means.

[0215] "Emotional responses from viewers" refer to the emotional reactions and feedback provided by people watching a broadcast or performance.

[0216] The system for carrying out this invention is configured as follows: The user first transmits their voice information on a live performance distribution platform via a mobile device or head-mounted display. Based on this, the terminal acquires the voice information in real time using a microphone and performs initial digital processing.

[0217] The server receives audio information transmitted from the aforementioned dev device and analyzes the audio data through an audio analysis engine. The analysis includes musical elements such as pitch, rhythm, and tone quality, as well as understanding emotional expression using emotion recognition AI. The specific software used here consists of existing cloud services and emotion analysis platforms for automatic speech recognition.

[0218] The analysis results are used as data to generate visual and auditory responses in the virtual reality environment. The server uses this data to generate virtual audience reactions. The audience reactions are synchronized in real time with the user's voice and displayed on the viewer's screen.

[0219] For example, when a user sings an emotionally moving song during a live performance, if emotion analysis detects that the user was "moved," gestures indicating emotion will be reproduced in the virtual reality space as a reaction from the virtual audience. Simultaneously, viewer reactions are collected, and the user is provided with information such as, "It appears the audience was moved."

[0220] A concrete example of a prompt using a generative AI model would be the instruction, "Generate emotional feedback from viewers based on sentiment analysis." This allows users to intuitively receive feedback on their performance and gain further motivation.

[0221] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0222] Step 1:

[0223] The terminal acquires the user's voice information in real time through the microphone and formats it as digital voice data. This formatted data is input data prepared for transmission to the server.

[0224] Step 2:

[0225] The terminal sends formatted audio data to the server. The transmitted audio data becomes input data for audio analysis on the server.

[0226] Step 3:

[0227] The server analyzes the received audio data. It uses an audio analysis engine to evaluate musical elements (pitch, rhythm, tone quality) and an emotion recognition AI to perform emotion analysis. The results of this analysis are output as foundational data for generating visual and auditory responses.

[0228] Step 4:

[0229] The server generates visual and auditory responses in the virtual reality space based on the analysis results. The generated virtual audience reaction data is output as feedback to the user.

[0230] Step 5:

[0231] Through generated feedback, users receive real-time emotional responses from the audience and visually experience the reactions of a virtual audience on screen. This provides an emotional interaction with the user's performance.

[0232] Step 6:

[0233] The server processes user rhyme performance data in a format that can be uploaded to the community platform, facilitating sharing among users. This process allows for the incorporation of sentiment-based feedback from other users.

[0234] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0235] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0237] [Second Embodiment]

[0238] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0239] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0241] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0245] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0246] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0247] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0248] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0249] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0250] This invention provides a system for users to receive real-time feedback during a karaoke experience and improve their singing skills. The system mainly consists of a terminal, a server, and a virtual reality device.

[0251] The user first launches a karaoke application on their device and begins voice input. The device acquires the user's singing voice as digital audio data and sends it to the server in real time. The server then analyzes this audio data. Specifically, it uses AI technology to evaluate elements such as pitch, rhythm, and sound quality, and performs a detailed analysis of the user's performance.

[0252] After analysis, the server generates specific feedback for the user and sends it to their terminal. The feedback points out things like pitch discrepancies and rhythm inconsistencies and provides advice on how to improve.

[0253] Furthermore, when a user uses a virtual reality device, the device constructs a virtual live stage and generates audience reactions that follow the user's performance. This allows the user to have an immersive experience as if they were actually at a concert. The server provides and transmits real-time visual and auditory feedback, such as audience applause and cheers, to the device.

[0254] The system also includes a community platform where users can share their performance with other users. Users upload their performance results to the platform and receive comments and ratings from other users. Through this feedback, users can gain information to further improve their skills.

[0255] As a concrete example, consider a scenario where a user sings "Happy Birthday." In this case, the device sends audio data to a server, which analyzes the parts where the pitch is too high and provides feedback such as, "Next time, try singing this part a little lower." Simultaneously, in a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when they reach the chorus, promoting a sense of accomplishment.

[0256] Thus, the present invention provides a practical and enjoyable means for users to not only enjoy karaoke but also to improve their musical skills.

[0257] The following describes the processing flow.

[0258] Step 1:

[0259] The device acquires audio data from the microphone when the user launches a karaoke application and begins voice input.

[0260] Step 2:

[0261] The terminal transmits the acquired audio data to the server in real time. This transmission is performed using a protocol that minimizes data latency.

[0262] Step 3:

[0263] The server analyzes the received audio data using AI technology. Specifically, it evaluates pitch, rhythm, and timbre to determine the quality of the performance.

[0264] Step 4:

[0265] The server generates feedback based on the analysis results. For example, it identifies off-key notes and rhythmic discrepancies and compiles advice on areas for improvement.

[0266] Step 5:

[0267] The server sends the generated feedback to the terminal.

[0268] Step 6:

[0269] The device visualizes the received feedback for the user. This is displayed in text and simple graph formats, allowing the user to see specific areas for improvement.

[0270] Step 7:

[0271] When a user uses a VR device, the device constructs a virtual live stage. While tracking the user's movements, it prepares virtual audience reactions that synchronize with the performance.

[0272] Step 8:

[0273] The server generates visual and auditory feedback, such as applause and cheers, in accordance with the user's performance and sends it to the device in real time.

[0274] Step 9:

[0275] The device presents the user with these visual and auditory feedbacks, providing an immersive experience that makes them feel as if they are actually at a concert.

[0276] Step 10:

[0277] After the performance is complete, the device prepares to upload the recorded results to the community platform. This will only happen if the user allows it.

[0278] Step 11:

[0279] Users can watch performances uploaded by other users on the community platform and leave comments and ratings.

[0280] Step 12:

[0281] The server collects feedback from other users and notifies the original user. This allows the user to gain the insights necessary for self-improvement.

[0282] (Example 1)

[0283] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0284] Conventionally, in entertainment activities such as karaoke, there has been a lack of specific feedback for users to objectively evaluate and improve their singing skills. In addition, the provision of immersive experiences using virtual reality technology and the opportunity to improve skills through communication with other users have also been limited. As a result, it has been difficult to link the improvement of users' singing skills to a richer entertainment experience.

[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0286] In this invention, the server includes means for acquiring and analyzing the acoustic data of the user in real time and providing improvement plans to the user, means for providing visual and auditory feedback to the user in real time in a virtual reality environment and generating a virtual reaction based on the user's actions, and means for providing an information exchange platform through which the improvement plans can be shared with other participants. As a result, the user can improve their singing skills through real-time feedback and enjoy a reaction similar to an actual live stage through an immersive experience in a virtual reality environment.

[0287] <000​​​​​​​​​​​​​​​​​

[0292] "Virtual response" refers to the simulated audience or environment's reaction to a user's actions in a virtual reality environment.

[0293] An "information exchange platform" is a digital space where users can share analysis results and improvement suggestions, and interact with each other.

[0294] This invention is a system aimed at enabling users to improve their singing skills through karaoke. The system mainly consists of a terminal, a server, and a virtual reality device.

[0295] The user first launches a karaoke application on their device. This application has a voice input function and captures the user's singing voice through the microphone. The captured voice is converted into digital audio data. This digital audio data is then transmitted from the device to the server in real time.

[0296] The server uses AI technology to analyze the received digital audio data. Specifically, it uses AI libraries such as "TensorFlow" and "PyTorch" to evaluate pitch, rhythm, and audio quality. This analysis generates specific improvement suggestions for the user. These improvement suggestions are then provided to the user via their device.

[0297] When a user is using a virtual reality device, the terminal constructs a virtual live stage and generates virtual reactions based on the user's performance. This includes visual and auditory feedback such as applause and cheers from a virtual audience. The server generates and sends this feedback to the terminal in real time, providing the user with an immersive experience.

[0298] Furthermore, the system includes an information exchange platform where users can share their performance results with other users. Through this platform, users can receive comments and evaluations from other users. This allows users to receive more detailed feedback and obtain information to improve their singing skills.

[0299] As a concrete example, consider a case where a user sings "Happy Birthday" and the system analyzes the performance. In this case, the terminal sends the audio data to a server, which analyzes parts that are too high in pitch and provides specific suggestions for improvement, such as "Next time, try singing this part a little lower." In a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when the chorus is reached, providing an experience that promotes a sense of accomplishment.

[0300] An example of a prompt for a generative AI model would be, "Please tell me how to provide real-time voice feedback in a karaoke system."

[0301] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0302] Step 1:

[0303] The user launches the karaoke application on their device. The application displays a song list, allowing the user to select a song. The input is the user's song selection, and the output is information about the selected song.

[0304] Step 2:

[0305] The user selects a song and begins singing. The device acquires the user's singing voice in real time via the microphone and converts it into digital audio data. In this step, the input is the user's analog voice, and the output is digital audio data.

[0306] Step 3:

[0307] The terminal transmits the acquired digital audio data to the server via the network. When transmitting, data compression is performed to reduce the communication load. The input is digital audio data, and the output is a notification of successful transmission to the server.

[0308] Step 4:

[0309] The server inputs the received audio data into the AI analysis system. Here, the generated AI model is used to analyze pitch, rhythm, and sound quality. The input for this step is audio data, and the output is the analysis result.

[0310] Step 5:

[0311] Based on the analysis result, the server generates improvement suggestions for the user. Specifically, it detects pitch deviations and rhythm inconsistencies and outputs the improvement points as text. The input is the analysis result, and the output is the text of the improvement suggestions.

[0312] Step 6:

[0313] The server transmits the improvement suggestions to the terminal. Here, data compression is also performed to enable the user to receive feedback quickly. The input is the text of the improvement suggestions, and the output is a notification of successful transmission to the terminal.

[0314] Step 7:

[0315] If the terminal is connected to a virtual reality device, it constructs a virtual live stage. It generates visual and auditory feedback according to the user's performance, that is, virtual reactions. The input is the user's performance data, and the output is the visual and auditory feedback in the virtual space.

[0316] Step 8:

[0317] Users can upload their singing performances and feedback to the community platform. They can also receive comments and ratings from other users. The input for this step is improvement suggestions and performance data, and the output is feedback from other users.

[0318] (Application Example 1)

[0319] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0320] Traditional karaoke systems lack the real-time, specific feedback necessary for users to effectively improve their singing skills. Furthermore, the virtual reality experience is limited, making it difficult to achieve the same level of immersion as attending an actual concert.

[0321] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0322] In this invention, the server includes means for acquiring the user's voice signal in real time, means for analyzing the voice signal and providing information to the user, means for providing the user with visual and auditory information in a virtual environment in real time, and means for providing real-time visualized feedback to the user while they are singing using a smart device. This allows the user to evaluate their singing skills in real time and make concrete improvements while enjoying an immersive experience in a virtual environment.

[0323] "User voice signal" refers to data obtained by converting the voice spoken by the user into an electrical signal.

[0324] "Real-time acquisition" refers to the instantaneous collection of the user's voice signal as they speak.

[0325] "Analysis" is the process of analyzing collected audio signals based on certain criteria and detecting their characteristics.

[0326] An "information-providing device" is a device that presents analysis results to the user and facilitates their understanding.

[0327] A "virtual environment" is a computer-generated space, distinct from the real world, constructed using digital technology.

[0328] "Visual and auditory information" refers to data and feedback, such as responses, that are provided to the user visually or aurally.

[0329] A "smart device" is a portable electronic device equipped with a processor and communication devices that enables interaction with the user.

[0330] "Visualized feedback" refers to information that displays the results obtained through analysis in a way that allows users to intuitively understand them.

[0331] "Audience reaction" refers to the actions and emotional expressions of fictional people in response to a particular event.

[0332] The system for realizing this invention includes a series of processes for acquiring and analyzing the user's voice signal and providing real-time feedback. First, the user launches a karaoke application using a smart device and begins singing. The smart device acquires the user's voice signal in real time using its built-in microphone. The acquired voice signal is transmitted to a server in the cloud.

[0333] The server analyzes the audio data using a speech recognition API (for example, a common API for speech recognition) and evaluates the pitch, rhythm, and sound quality. Based on the analysis results, it generates specific suggestions for improvement regarding the user's musical performance and sends them to the smart device. This feedback is visualized on the smart device's display for intuitive understanding by the user and is also provided as audio feedback.

[0334] Furthermore, in the virtual environment, a computer-generated audience reacts in sync with the user's singing. The virtual environment uses a 3D engine such as Unity to provide real-time visual and auditory feedback to the user. The server pre-generates these audience reactions and plays them back in real-time in sync with the user's singing, enhancing the performance effect.

[0335] For example, when a user sings "Happy Birthday," a pitch bar is displayed on the smart device's screen, and any off-key sections are highlighted in red. Furthermore, during the chorus, the sound of a virtual audience beginning to applaud is played to enhance the user's sense of accomplishment.

[0336] An example of a prompt might be, "While singing the chorus of 'Happy Birthday,' provide real-time feedback on pitch deviations and visualize the audience's reaction."

[0337] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0338] Step 1:

[0339] The user launches a karaoke application on their smart device and begins singing. The input is the user's voice signal, and the output is recorded as audio data on the smart device. The smart device acquires this audio signal in real time through its microphone.

[0340] Step 2:

[0341] The smart device transmits the acquired voice data to a server in the cloud. The input is voice data, and the output is data transfer to the server. The smart device uses a communication module to upload the data to the server in real time.

[0342] Step 3:

[0343] The server analyzes audio data using a speech recognition API. The input is audio data sent from a smart device, and the output is an evaluation result regarding pitch, rhythm, and sound quality. The server processes the audio based on specific musical criteria and generates the analysis results.

[0344] Step 4:

[0345] The server generates improvement suggestions based on the analysis results and sends them to the smart device. The input is the result of the voice analysis, and the output is a specific feedback message. The server uses an AI algorithm to point out deviations in pitch and rhythm and create advice for improvement.

[0346] Step 5:

[0347] Smart devices receive feedback from a server and visualize it on their display to present to the user. The input is improvement suggestion messages, and the output is visualized feedback. Smart devices display graphical elements on their screens to help users intuitively understand the areas for improvement.

[0348] Step 6:

[0349] The server generates a virtual environment, creates audience reactions synchronized with the user's singing, and plays them back in real time. The input is real-time user singing data, and the output is visual and auditory audience reactions. The server constructs the virtual environment using Unity or similar software and stages scenarios where the audience applauds and cheers.

[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0351] This invention provides a karaoke system that utilizes an emotion engine to analyze user voice data and recognize emotions. The system consists of a terminal, a server, and a virtual reality device.

[0352] The user launches a karaoke application on their device and begins voice input. The device acquires the user's voice data in real time and sends it to the server. The server analyzes the voice data and uses AI technology to evaluate musical elements such as pitch, rhythm, and tone quality. In addition, an emotion engine identifies the user's emotions based on the voice data.

[0353] The server generates feedback based on the analysis results and recognized emotional information, and sends it to the terminal. This feedback can include suggestions for musical technical improvements, as well as responses to and encouragement regarding the user's emotions.

[0354] Furthermore, when using a virtual reality device, the terminal constructs a virtual live stage in real time. The reactions of the virtual audience, synchronized with the user's singing, are adjusted according to the user's emotions. For example, if the user is expressing joyful emotions, the audience's reactions will be set to be more lively and the cheers louder.

[0355] The system also integrates a community platform for users to share the results of their voice data analysis, including emotional information, with other users. Users can upload their performance results to the platform and receive emotionally-based feedback and comments from other users, which can help them feel empathy and boost their motivation.

[0356] As a concrete example, suppose the emotion engine detects a sense of exhilaration when a user sings a "celebratory song." In this case, the server includes emotionally appealing feedback such as "convey that joy even more." Furthermore, the audience in the VR environment will be shown an animation that further heightens the user's emotions.

[0357] This means that the present invention allows users to receive various emotionally appealing feedback while singing karaoke, enabling them to improve their musical skills and gain a platform for emotional expression.

[0358] The following describes the processing flow.

[0359] Step 1:

[0360] When the user launches a karaoke application and begins voice input, the device acquires audio data from the microphone in real time.

[0361] Step 2:

[0362] The terminal sends the acquired audio data to the server. This transmission occurs in real time, and a protocol that minimizes data latency is used.

[0363] Step 3:

[0364] The server analyzes the received audio data using AI technology. The analysis includes evaluation of pitch, rhythm, and sound quality to determine musical skill.

[0365] Step 4:

[0366] The server uses an emotion engine to recognize the user's emotions from the voice data. The recognized emotions include, for example, joy, sadness, and excitement.

[0367] Step 5:

[0368] The server generates specific and emotionally responsive feedback based on analysis results and emotion recognition results. This feedback includes not only suggestions for musical improvements but also messages that resonate with the user's emotions.

[0369] Step 6:

[0370] The server sends the generated feedback to the terminal.

[0371] Step 7:

[0372] The device displays the received feedback to the user. The feedback is visualized in text and graph formats to help the user understand it.

[0373] Step 8:

[0374] When a user uses a virtual reality device, the device constructs a virtual live stage and adjusts the audience's reactions in sync with the user's performance.

[0375] Step 9:

[0376] The server generates virtual audience reactions based on the recognized user's emotions and sends them to the terminal.

[0377] Step 10:

[0378] The device plays these reactions in a virtual environment, providing the user with visual and auditory feedback.

[0379] Step 11:

[0380] The device prepares to upload performance recording data to the community platform.

[0381] Step 12:

[0382] On the community platform, other users can view uploaded performances and leave comments and ratings based on their emotions.

[0383] Step 13:

[0384] The server collects feedback from other users and notifies the original user. This allows users to gain emotional resonance and useful feedback.

[0385] (Example 2)

[0386] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0387] While conventional karaoke systems allowed users to receive feedback on their singing technique, they lacked mechanisms to improve emotional expression and the quality of the virtual reality experience. Furthermore, there was a challenge in that users had difficulty receiving emotion-based feedback through interaction with other users.

[0388] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0389] In this invention, the server includes means for acquiring user voice information and analyzing said voice information to identify musical elements and emotions; means for generating a virtual reality environment that is dynamically adjusted based on the emotions extracted from said voice information and providing visual and auditory feedback in said environment; and means for providing an electronic platform on which the analysis results and feedback can be shared with other users. This makes it possible for users to receive feedback on emotional expression in real time and to share emotions through interaction with others.

[0390] "User voice information" refers to the collection of acoustic data containing the user's voice, acquired through a voice input device.

[0391] "Analysis" refers to the computational process of identifying musical elements and emotions based on acquired audio information.

[0392] "Musical elements" refer to characteristics related to music that are extracted from audio information, such as pitch, rhythm, and tone quality.

[0393] "Emotions" refer to information that indicates the user's mental state or feelings, as identified from audio information.

[0394] A "virtual reality environment" refers to a three-dimensional space or audiovisual experience generated by a computer that a user can immerse themselves in.

[0395] "Visual and auditory feedback" refers to information returned to the user through visual images and auditory sounds, intended to enhance the quality of the experience.

[0396] An "electronic platform" is a technological foundation that enables the sharing and exchange of information through user communities and digital networks.

[0397] This invention relates to a system that analyzes voice information input by a user and provides musical guidance and emotional feedback. The system consists of multiple technical elements, including a terminal, a server, and a virtual reality device.

[0398] Users input voice information via the microphone using a karaoke application installed on their device. This voice information is digitized on the device and transmitted to the server in real time as audio data. WebSocket is used to maintain real-time transmission of the audio data.

[0399] The server uses a machine learning framework to analyze the received audio information. Specifically, it evaluates musical elements such as pitch, rhythm, and tone quality using pitch analysis algorithms and beat tracking. The emotion engine also uses libraries such as OpenSmile to classify emotions from the intonation and tempo of the voice. Through this analysis, not only individual musical characteristics but also the user's emotional state is evaluated.

[0400] The analysis results are compiled into feedback provided to the user using natural language generation technology. This feedback is delivered to the user visually and audibly through the device. For example, messages such as "Sing with a clearer pitch" or "Try to convey the enjoyment more effectively" may be displayed. Real-time feedback is achieved through dynamic page rendering using HTML templates and JavaScript.

[0401] When using a virtual reality device, the terminal also has the capability to build a realistic virtual live stage based on the emotions of the voice, using Unity or Unreal Engine. It generates virtual audience reactions that adjust according to the user's expression, providing immersion in both sight and sound.

[0402] Furthermore, users can upload the generated feedback to an electronic platform and share it with other users. Through interaction with other users, they can receive emotion-based feedback, which can boost their motivation.

[0403] For example, when a user sings a "celebratory song," the emotion engine recognizes the feeling of exhilaration, and the server provides feedback to the user saying, "Please convey that joy more." An example of a prompt to input into the generative AI model would be, "Please tell me how to analyze the audio data sung by the user, recognize the user's emotions, and generate feedback."

[0404] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0405] Step 1:

[0406] The terminal receives voice input from the karaoke application launched by the user and collects it through the microphone. The voice data as input is converted into a digital signal using an audio library and sent to the server in real time. This provides the voice data in a digital format that can be used for subsequent analysis processing.

[0407] Step 2:

[0408] The server receives audio data transmitted from the terminal. The received audio data is first subjected to pitch analysis algorithms and beat tracking to evaluate pitch, rhythm, and sound quality. These analyses produce evaluation data indicating the user's musical skills, and the process then proceeds to the next step of sentiment analysis.

[0409] Step 3:

[0410] The server performs sentiment analysis on the audio data in parallel with evaluating the musical elements. The sentiment engine uses the OpenSmile library to analyze the intonation and tempo of the voice and classify the user's emotions (e.g., exhilaration or sadness). This classified sentiment data is used to generate feedback.

[0411] Step 4:

[0412] The server integrates analysis data of musical elements and emotions to generate feedback. This process utilizes natural language generation technology to create specific improvement suggestions for the user (e.g., "Let's stabilize the pitch") and emotional messages (e.g., "Let's convey the enjoyment more"). The generated feedback is then formatted for visualization.

[0413] Step 5:

[0414] The device receives feedback from the server and presents it to the user visually and audibly. The input feedback data is displayed on the screen using HTML and JavaScript, allowing the user to see the results in real time. This phase aims to improve the user's singing experience through feedback.

[0415] Step 6:

[0416] When a virtual reality device is connected, the terminal uses a game engine to construct a virtual live stage. Based on the emotional data of the input audio, the reactions of the virtual audience are dynamically generated. For example, if the emotion is heightened, an animation of the virtual audience becoming excited will be displayed. This output provides the user with an immersive experience.

[0417] Step 7:

[0418] Users upload the analyzed data and generated feedback to an electronic platform. This step facilitates information sharing and communication with other users via their devices. The outputted feedback is used to solicit sentiment-based comments from other community members.

[0419] (Application Example 2)

[0420] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0421] Conventional voice analysis systems have been unable to provide real-time feedback that takes user emotions into account, posing a challenge in fostering emotional engagement with the audience, especially during live performances. Furthermore, there was a lack of specific feedback necessary for users to evaluate and improve their own performance.

[0422] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0423] In this invention, the server includes means for acquiring and analyzing user voice information in real time, means for providing the user with visual and auditory responses in a virtual reality space in real time, and means for analyzing the user's emotional expression during live performance streaming and acquiring and presenting emotional responses from viewers in real time. This allows the user to deepen their emotional interaction with viewers while receiving rich feedback in real time to evaluate and improve their performance.

[0424] "User voice information" refers to all voice data emitted by the user, and is usually collected in digital format via a microphone.

[0425] "Real-time acquisition" refers to a technology that minimizes the delay between transmission and acquisition, allowing for the immediate recording of voice information on the spot.

[0426] "Analysis" refers to the process of performing data processing on collected audio information to recognize specific patterns or emotions.

[0427] "Response" refers to feedback or reactions generated based on the analysis of audio information, and is the information or reaction provided to the user or audience.

[0428] A "virtual reality space" is a three-dimensional digital environment generated by a computer that can be experienced visually and aurally.

[0429] "Emotional expression" refers to the emotions and feelings that users express through words, music, and other means.

[0430] "Emotional responses from viewers" refer to the emotional reactions and feedback provided by people watching a broadcast or performance.

[0431] The system for carrying out this invention is configured as follows: The user first transmits their voice information on a live performance distribution platform via a mobile device or head-mounted display. Based on this, the terminal acquires the voice information in real time using a microphone and performs initial digital processing.

[0432] The server receives audio information transmitted from the aforementioned dev device and analyzes the audio data through an audio analysis engine. The analysis includes musical elements such as pitch, rhythm, and tone quality, as well as understanding emotional expression using emotion recognition AI. The specific software used here consists of existing cloud services and emotion analysis platforms for automatic speech recognition.

[0433] The analysis results are used as data to generate visual and auditory responses in the virtual reality environment. The server uses this data to generate virtual audience reactions. The audience reactions are synchronized in real time with the user's voice and displayed on the viewer's screen.

[0434] For example, when a user sings an emotionally moving song during a live performance, if emotion analysis detects that the user was "moved," gestures indicating emotion will be reproduced in the virtual reality space as a reaction from the virtual audience. Simultaneously, viewer reactions are collected, and the user is provided with information such as, "It appears the audience was moved."

[0435] A concrete example of a prompt using a generative AI model would be the instruction, "Generate emotional feedback from viewers based on sentiment analysis." This allows users to intuitively receive feedback on their performance and gain further motivation.

[0436] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0437] Step 1:

[0438] The terminal acquires the user's voice information in real time through the microphone and formats it as digital voice data. This formatted data is input data prepared for transmission to the server.

[0439] Step 2:

[0440] The terminal sends formatted audio data to the server. The transmitted audio data becomes input data for audio analysis on the server.

[0441] Step 3:

[0442] The server analyzes the received audio data. It uses an audio analysis engine to evaluate musical elements (pitch, rhythm, tone quality) and an emotion recognition AI to perform emotion analysis. The results of this analysis are output as foundational data for generating visual and auditory responses.

[0443] Step 4:

[0444] The server generates visual and auditory responses in the virtual reality space based on the analysis results. The generated virtual audience reaction data is output as feedback to the user.

[0445] Step 5:

[0446] Through generated feedback, users receive real-time emotional responses from the audience and visually experience the reactions of a virtual audience on screen. This provides an emotional interaction with the user's performance.

[0447] Step 6:

[0448] The server processes user rhyme performance data in a format that can be uploaded to the community platform, facilitating sharing among users. This process allows for the incorporation of sentiment-based feedback from other users.

[0449] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0450] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0451] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0452] [Third Embodiment]

[0453] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0454] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0455] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0456] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0457] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0458] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0459] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0460] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0461] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0462] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0463] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0464] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0465] This invention provides a system for users to receive real-time feedback during a karaoke experience and improve their singing skills. The system mainly consists of a terminal, a server, and a virtual reality device.

[0466] The user first launches a karaoke application on their device and begins voice input. The device acquires the user's singing voice as digital audio data and sends it to the server in real time. The server then analyzes this audio data. Specifically, it uses AI technology to evaluate elements such as pitch, rhythm, and sound quality, and performs a detailed analysis of the user's performance.

[0467] After analysis, the server generates specific feedback for the user and sends it to their terminal. The feedback points out things like pitch discrepancies and rhythm inconsistencies and provides advice on how to improve.

[0468] Furthermore, when a user uses a virtual reality device, the device constructs a virtual live stage and generates audience reactions that follow the user's performance. This allows the user to have an immersive experience as if they were actually at a concert. The server provides and transmits real-time visual and auditory feedback, such as audience applause and cheers, to the device.

[0469] The system also includes a community platform where users can share their performance with other users. Users upload their performance results to the platform and receive comments and ratings from other users. Through this feedback, users can gain information to further improve their skills.

[0470] As a concrete example, consider a scenario where a user sings "Happy Birthday." In this case, the device sends audio data to a server, which analyzes the parts where the pitch is too high and provides feedback such as, "Next time, try singing this part a little lower." Simultaneously, in a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when they reach the chorus, promoting a sense of accomplishment.

[0471] Thus, the present invention provides a practical and enjoyable means for users to not only enjoy karaoke but also to improve their musical skills.

[0472] The following describes the processing flow.

[0473] Step 1:

[0474] The device acquires audio data from the microphone when the user launches a karaoke application and begins voice input.

[0475] Step 2:

[0476] The terminal transmits the acquired audio data to the server in real time. This transmission is performed using a protocol that minimizes data latency.

[0477] Step 3:

[0478] The server analyzes the received audio data using AI technology. Specifically, it evaluates pitch, rhythm, and timbre to determine the quality of the performance.

[0479] Step 4:

[0480] The server generates feedback based on the analysis results. For example, it identifies off-key notes and rhythmic discrepancies and compiles advice on areas for improvement.

[0481] Step 5:

[0482] The server sends the generated feedback to the terminal.

[0483] Step 6:

[0484] The device visualizes the received feedback for the user. This is displayed in text and simple graph formats, allowing the user to see specific areas for improvement.

[0485] Step 7:

[0486] When a user uses a VR device, the device constructs a virtual live stage. While tracking the user's movements, it prepares virtual audience reactions that synchronize with the performance.

[0487] Step 8:

[0488] The server generates visual and auditory feedback, such as applause and cheers, in accordance with the user's performance and sends it to the device in real time.

[0489] Step 9:

[0490] The device presents the user with these visual and auditory feedbacks, providing an immersive experience that makes them feel as if they are actually at a concert.

[0491] Step 10:

[0492] After the performance is complete, the device prepares to upload the recorded results to the community platform. This will only happen if the user allows it.

[0493] Step 11:

[0494] Users can watch performances uploaded by other users on the community platform and leave comments and ratings.

[0495] Step 12:

[0496] The server collects feedback from other users and notifies the original user. This allows the user to gain the insights necessary for self-improvement.

[0497] (Example 1)

[0498] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0499] Traditionally, entertainment activities like karaoke have lacked concrete feedback that allows users to objectively evaluate and improve their singing skills. Furthermore, opportunities for immersive experiences using virtual reality technology and for skill improvement through communication with other users are limited. This makes it difficult to improve users' singing skills and create a richer entertainment experience.

[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0501] In this invention, the server includes means for acquiring and analyzing the user's acoustic data in real time and providing the user with suggestions for improvement; means for providing the user with visual and auditory feedback in real time in a virtual reality environment and generating virtual responses based on the user's actions; and means for providing an information exchange platform on which the suggestions for improvement can be shared with other participants. This allows the user to improve their singing skills through real-time feedback and enjoy responses that are just like those on a real live stage through an immersive experience in a virtual reality environment.

[0502] "Audio data" refers to information that digitally represents a user's vocalizations or musical expressions.

[0503] "Analysis" is the process of processing acoustic data and evaluating elements such as pitch, rhythm, and voice quality.

[0504] An "improvement suggestion" is a specific proposal provided to the user based on the analysis results, aimed at improving their singing skills.

[0505] A "virtual reality environment" is a computer-generated virtual space created using digital technology, into which a user can immerse themselves.

[0506] "Visual and auditory feedback" refers to responses provided to a user based on visual or auditory information.

[0507] "Virtual response" refers to the simulated audience or environment's reaction to a user's actions in a virtual reality environment.

[0508] An "information exchange platform" is a digital space where users can share analysis results and improvement suggestions, and interact with each other.

[0509] This invention is a system aimed at enabling users to improve their singing skills through karaoke. The system mainly consists of a terminal, a server, and a virtual reality device.

[0510] The user first launches a karaoke application on their device. This application has a voice input function and captures the user's singing voice through the microphone. The captured voice is converted into digital audio data. This digital audio data is then transmitted from the device to the server in real time.

[0511] The server uses AI technology to analyze the received digital audio data. Specifically, it uses AI libraries such as "TensorFlow" and "PyTorch" to evaluate pitch, rhythm, and audio quality. This analysis generates specific improvement suggestions for the user. These improvement suggestions are then provided to the user via their device.

[0512] When a user is using a virtual reality device, the terminal constructs a virtual live stage and generates virtual reactions based on the user's performance. This includes visual and auditory feedback such as applause and cheers from a virtual audience. The server generates and sends this feedback to the terminal in real time, providing the user with an immersive experience.

[0513] Furthermore, the system includes an information exchange platform where users can share their performance results with other users. Through this platform, users can receive comments and evaluations from other users. This allows users to receive more detailed feedback and obtain information to improve their singing skills.

[0514] As a concrete example, consider a case where a user sings "Happy Birthday" and the system analyzes the performance. In this case, the terminal sends the audio data to a server, which analyzes parts that are too high in pitch and provides specific suggestions for improvement, such as "Next time, try singing this part a little lower." In a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when the chorus is reached, providing an experience that promotes a sense of accomplishment.

[0515] An example of a prompt for a generative AI model would be, "Please tell me how to provide real-time voice feedback in a karaoke system."

[0516] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0517] Step 1:

[0518] The user launches the karaoke application on their device. The application displays a song list, allowing the user to select a song. The input is the user's song selection, and the output is information about the selected song.

[0519] Step 2:

[0520] The user selects a song and begins singing. The device acquires the user's singing voice in real time via the microphone and converts it into digital audio data. In this step, the input is the user's analog voice, and the output is digital audio data.

[0521] Step 3:

[0522] The terminal transmits the acquired digital audio data to the server via the network. Data compression is performed during transmission to reduce communication load. The input is digital audio data, and the output is a notification to the server indicating successful transmission.

[0523] Step 4:

[0524] The server inputs the received acoustic data into the AI ​​analysis system. Here, a generated AI model is used to analyze the pitch, rhythm, and sound quality. The input for this step is the acoustic data, and the output is the analysis result.

[0525] Step 5:

[0526] The server generates improvement suggestions for the user based on the analysis results. Specifically, it detects pitch discrepancies and rhythmic inconsistencies and outputs the areas for improvement as text. The input is the analysis results, and the output is the text of the improvement suggestions.

[0527] Step 6:

[0528] The server sends the suggested improvements to the terminal. Data compression is also performed here to ensure users receive feedback quickly. The input is the text of the suggested improvements, and the output is a notification that the transmission to the terminal is complete.

[0529] Step 7:

[0530] When the device is connected to a virtual reality system, it constructs a virtual live stage. It generates visual and auditory feedback, or virtual reactions, that respond to the user's performance. The input is user performance data, and the output is visual and auditory feedback in the virtual space.

[0531] Step 8:

[0532] Users can upload their singing performances and feedback to the community platform. They can also receive comments and ratings from other users. The input for this step is improvement suggestions and performance data, and the output is feedback from other users.

[0533] (Application Example 1)

[0534] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0535] Traditional karaoke systems lack the real-time, specific feedback necessary for users to effectively improve their singing skills. Furthermore, the virtual reality experience is limited, making it difficult to achieve the same level of immersion as attending an actual concert.

[0536] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0537] In this invention, the server includes means for acquiring the user's voice signal in real time, means for analyzing the voice signal and providing information to the user, means for providing the user with visual and auditory information in a virtual environment in real time, and means for providing real-time visualized feedback to the user while they are singing using a smart device. This allows the user to evaluate their singing skills in real time and make concrete improvements while enjoying an immersive experience in a virtual environment.

[0538] "User voice signal" refers to data obtained by converting the voice spoken by the user into an electrical signal.

[0539] "Real-time acquisition" refers to the instantaneous collection of the user's voice signal as they speak.

[0540] "Analysis" is the process of analyzing collected audio signals based on certain criteria and detecting their characteristics.

[0541] An "information-providing device" is a device that presents analysis results to the user and facilitates their understanding.

[0542] A "virtual environment" is a computer-generated space, distinct from the real world, constructed using digital technology.

[0543] "Visual and auditory information" refers to data and feedback, such as responses, that are provided to the user visually or aurally.

[0544] A "smart device" is a portable electronic device equipped with a processor and communication devices that enables interaction with the user.

[0545] "Visualized feedback" refers to information that displays the results obtained through analysis in a way that allows users to intuitively understand them.

[0546] "Audience reaction" refers to the actions and emotional expressions of fictional people in response to a particular event.

[0547] The system for realizing this invention includes a series of processes for acquiring and analyzing the user's voice signal and providing real-time feedback. First, the user launches a karaoke application using a smart device and begins singing. The smart device acquires the user's voice signal in real time using its built-in microphone. The acquired voice signal is transmitted to a server in the cloud.

[0548] The server analyzes the audio data using a speech recognition API (for example, a common API for speech recognition) and evaluates the pitch, rhythm, and sound quality. Based on the analysis results, it generates specific suggestions for improvement regarding the user's musical performance and sends them to the smart device. This feedback is visualized on the smart device's display for intuitive understanding by the user and is also provided as audio feedback.

[0549] Furthermore, in the virtual environment, a computer-generated audience reacts in sync with the user's singing. The virtual environment uses a 3D engine such as Unity to provide real-time visual and auditory feedback to the user. The server pre-generates these audience reactions and plays them back in real-time in sync with the user's singing, enhancing the performance effect.

[0550] For example, when a user sings "Happy Birthday," a pitch bar is displayed on the smart device's screen, and any off-key sections are highlighted in red. Furthermore, during the chorus, the sound of a virtual audience beginning to applaud is played to enhance the user's sense of accomplishment.

[0551] An example of a prompt might be, "While singing the chorus of 'Happy Birthday,' provide real-time feedback on pitch deviations and visualize the audience's reaction."

[0552] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0553] Step 1:

[0554] The user launches a karaoke application on their smart device and begins singing. The input is the user's voice signal, and the output is recorded as audio data on the smart device. The smart device acquires this audio signal in real time through its microphone.

[0555] Step 2:

[0556] The smart device transmits the acquired voice data to a server in the cloud. The input is voice data, and the output is data transfer to the server. The smart device uses a communication module to upload the data to the server in real time.

[0557] Step 3:

[0558] The server analyzes audio data using a speech recognition API. The input is audio data sent from a smart device, and the output is an evaluation result regarding pitch, rhythm, and sound quality. The server processes the audio based on specific musical criteria and generates the analysis results.

[0559] Step 4:

[0560] The server generates improvement suggestions based on the analysis results and sends them to the smart device. The input is the result of the voice analysis, and the output is a specific feedback message. The server uses an AI algorithm to point out deviations in pitch and rhythm and create advice for improvement.

[0561] Step 5:

[0562] Smart devices receive feedback from a server and visualize it on their display to present to the user. The input is improvement suggestion messages, and the output is visualized feedback. Smart devices display graphical elements on their screens to help users intuitively understand the areas for improvement.

[0563] Step 6:

[0564] The server generates a virtual environment, creates audience reactions synchronized with the user's singing, and plays them back in real time. The input is real-time user singing data, and the output is visual and auditory audience reactions. The server constructs the virtual environment using Unity or similar software and stages scenarios where the audience applauds and cheers.

[0565] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0566] This invention provides a karaoke system that utilizes an emotion engine to analyze user voice data and recognize emotions. The system consists of a terminal, a server, and a virtual reality device.

[0567] The user launches a karaoke application on their device and begins voice input. The device acquires the user's voice data in real time and sends it to the server. The server analyzes the voice data and uses AI technology to evaluate musical elements such as pitch, rhythm, and tone quality. In addition, an emotion engine identifies the user's emotions based on the voice data.

[0568] The server generates feedback based on the analysis results and recognized emotional information, and sends it to the terminal. This feedback can include suggestions for musical technical improvements, as well as responses to and encouragement regarding the user's emotions.

[0569] Furthermore, when using a virtual reality device, the terminal constructs a virtual live stage in real time. The reactions of the virtual audience, synchronized with the user's singing, are adjusted according to the user's emotions. For example, if the user is expressing joyful emotions, the audience's reactions will be set to be more lively and the cheers louder.

[0570] The system also integrates a community platform for users to share the results of their voice data analysis, including emotional information, with other users. Users can upload their performance results to the platform and receive emotionally-based feedback and comments from other users, which can help them feel empathy and boost their motivation.

[0571] As a concrete example, suppose the emotion engine detects a sense of exhilaration when a user sings a "celebratory song." In this case, the server includes emotionally appealing feedback such as "convey that joy even more." Furthermore, the audience in the VR environment will be shown an animation that further heightens the user's emotions.

[0572] This means that the present invention allows users to receive various emotionally appealing feedback while singing karaoke, enabling them to improve their musical skills and gain a platform for emotional expression.

[0573] The following describes the processing flow.

[0574] Step 1:

[0575] When the user launches a karaoke application and begins voice input, the device acquires audio data from the microphone in real time.

[0576] Step 2:

[0577] The terminal sends the acquired audio data to the server. This transmission occurs in real time, and a protocol that minimizes data latency is used.

[0578] Step 3:

[0579] The server analyzes the received audio data using AI technology. The analysis includes evaluation of pitch, rhythm, and sound quality to determine musical skill.

[0580] Step 4:

[0581] The server uses an emotion engine to recognize the user's emotions from the voice data. The recognized emotions include, for example, joy, sadness, and excitement.

[0582] Step 5:

[0583] The server generates specific and emotionally responsive feedback based on analysis results and emotion recognition results. This feedback includes not only suggestions for musical improvements but also messages that resonate with the user's emotions.

[0584] Step 6:

[0585] The server sends the generated feedback to the terminal.

[0586] Step 7:

[0587] The device displays the received feedback to the user. The feedback is visualized in text and graph formats to help the user understand it.

[0588] Step 8:

[0589] When a user uses a virtual reality device, the device constructs a virtual live stage and adjusts the audience's reactions in sync with the user's performance.

[0590] Step 9:

[0591] The server generates virtual audience reactions based on the recognized user's emotions and sends them to the terminal.

[0592] Step 10:

[0593] The device plays these reactions in a virtual environment, providing the user with visual and auditory feedback.

[0594] Step 11:

[0595] The device prepares to upload performance recording data to the community platform.

[0596] Step 12:

[0597] On the community platform, other users can view uploaded performances and leave comments and ratings based on their emotions.

[0598] Step 13:

[0599] The server collects feedback from other users and notifies the original user. This allows users to gain emotional resonance and useful feedback.

[0600] (Example 2)

[0601] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0602] While conventional karaoke systems allowed users to receive feedback on their singing technique, they lacked mechanisms to improve emotional expression and the quality of the virtual reality experience. Furthermore, there was a challenge in that users had difficulty receiving emotion-based feedback through interaction with other users.

[0603] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0604] In this invention, the server includes means for acquiring user voice information and analyzing said voice information to identify musical elements and emotions; means for generating a virtual reality environment that is dynamically adjusted based on the emotions extracted from said voice information and providing visual and auditory feedback in said environment; and means for providing an electronic platform on which the analysis results and feedback can be shared with other users. This makes it possible for users to receive feedback on emotional expression in real time and to share emotions through interaction with others.

[0605] "User voice information" refers to the collection of acoustic data containing the user's voice, acquired through a voice input device.

[0606] "Analysis" refers to the computational process of identifying musical elements and emotions based on acquired audio information.

[0607] "Musical elements" refer to characteristics related to music that are extracted from audio information, such as pitch, rhythm, and tone quality.

[0608] "Emotions" refer to information that indicates the user's mental state or feelings, as identified from audio information.

[0609] A "virtual reality environment" refers to a three-dimensional space or audiovisual experience generated by a computer that a user can immerse themselves in.

[0610] "Visual and auditory feedback" refers to information returned to the user through visual images and auditory sounds, intended to enhance the quality of the experience.

[0611] An "electronic platform" is a technological foundation that enables the sharing and exchange of information through user communities and digital networks.

[0612] This invention relates to a system that analyzes voice information input by a user and provides musical guidance and emotional feedback. The system consists of multiple technical elements, including a terminal, a server, and a virtual reality device.

[0613] Users input voice information via the microphone using a karaoke application installed on their device. This voice information is digitized on the device and transmitted to the server in real time as audio data. WebSocket is used to maintain real-time transmission of the audio data.

[0614] The server uses a machine learning framework to analyze the received audio information. Specifically, it evaluates musical elements such as pitch, rhythm, and tone quality using pitch analysis algorithms and beat tracking. The emotion engine also uses libraries such as OpenSmile to classify emotions from the intonation and tempo of the voice. Through this analysis, not only individual musical characteristics but also the user's emotional state is evaluated.

[0615] The analysis results are compiled into feedback provided to the user using natural language generation technology. This feedback is delivered to the user visually and audibly through the device. For example, messages such as "Sing with a clearer pitch" or "Try to convey the enjoyment more effectively" may be displayed. Real-time feedback is achieved through dynamic page rendering using HTML templates and JavaScript.

[0616] When using a virtual reality device, the terminal also has the capability to build a realistic virtual live stage based on the emotions of the voice, using Unity or Unreal Engine. It generates virtual audience reactions that adjust according to the user's expression, providing immersion in both sight and sound.

[0617] Furthermore, users can upload the generated feedback to an electronic platform and share it with other users. Through interaction with other users, they can receive emotion-based feedback, which can boost their motivation.

[0618] For example, when a user sings a "celebratory song," the emotion engine recognizes the feeling of exhilaration, and the server provides feedback to the user saying, "Please convey that joy more." An example of a prompt to input into the generative AI model would be, "Please tell me how to analyze the audio data sung by the user, recognize the user's emotions, and generate feedback."

[0619] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0620] Step 1:

[0621] The terminal receives voice input from the karaoke application launched by the user and collects it through the microphone. The voice data as input is converted into a digital signal using an audio library and sent to the server in real time. This provides the voice data in a digital format that can be used for subsequent analysis processing.

[0622] Step 2:

[0623] The server receives audio data transmitted from the terminal. The received audio data is first subjected to pitch analysis algorithms and beat tracking to evaluate pitch, rhythm, and sound quality. These analyses produce evaluation data indicating the user's musical skills, and the process then proceeds to the next step of sentiment analysis.

[0624] Step 3:

[0625] The server performs sentiment analysis on the audio data in parallel with evaluating the musical elements. The sentiment engine uses the OpenSmile library to analyze the intonation and tempo of the voice and classify the user's emotions (e.g., exhilaration or sadness). This classified sentiment data is used to generate feedback.

[0626] Step 4:

[0627] The server integrates analysis data of musical elements and emotions to generate feedback. This process utilizes natural language generation technology to create specific improvement suggestions for the user (e.g., "Let's stabilize the pitch") and emotional messages (e.g., "Let's convey the enjoyment more"). The generated feedback is then formatted for visualization.

[0628] Step 5:

[0629] The device receives feedback from the server and presents it to the user visually and audibly. The input feedback data is displayed on the screen using HTML and JavaScript, allowing the user to see the results in real time. This phase aims to improve the user's singing experience through feedback.

[0630] Step 6:

[0631] When a virtual reality device is connected, the terminal uses a game engine to construct a virtual live stage. Based on the emotional data of the input audio, the reactions of the virtual audience are dynamically generated. For example, if the emotion is heightened, an animation of the virtual audience becoming excited will be displayed. This output provides the user with an immersive experience.

[0632] Step 7:

[0633] Users upload the analyzed data and generated feedback to an electronic platform. This step facilitates information sharing and communication with other users via their devices. The outputted feedback is used to solicit sentiment-based comments from other community members.

[0634] (Application Example 2)

[0635] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0636] Conventional voice analysis systems have been unable to provide real-time feedback that takes user emotions into account, posing a challenge in fostering emotional engagement with the audience, especially during live performances. Furthermore, there was a lack of specific feedback necessary for users to evaluate and improve their own performance.

[0637] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0638] In this invention, the server includes means for acquiring and analyzing user voice information in real time, means for providing the user with visual and auditory responses in a virtual reality space in real time, and means for analyzing the user's emotional expression during live performance streaming and acquiring and presenting emotional responses from viewers in real time. This allows the user to deepen their emotional interaction with viewers while receiving rich feedback in real time to evaluate and improve their performance.

[0639] "User voice information" refers to all voice data emitted by the user, and is usually collected in digital format via a microphone.

[0640] "Real-time acquisition" refers to a technology that minimizes the delay between transmission and acquisition, allowing for the immediate recording of voice information on the spot.

[0641] "Analysis" refers to the process of performing data processing on collected audio information to recognize specific patterns or emotions.

[0642] "Response" refers to feedback or reactions generated based on the analysis of audio information, and is the information or reaction provided to the user or audience.

[0643] A "virtual reality space" is a three-dimensional digital environment generated by a computer that can be experienced visually and aurally.

[0644] "Emotional expression" refers to the emotions and feelings that users express through words, music, and other means.

[0645] "Emotional responses from viewers" refer to the emotional reactions and feedback provided by people watching a broadcast or performance.

[0646] The system for carrying out this invention is configured as follows: The user first transmits their voice information on a live performance distribution platform via a mobile device or head-mounted display. Based on this, the terminal acquires the voice information in real time using a microphone and performs initial digital processing.

[0647] The server receives audio information transmitted from the aforementioned dev device and analyzes the audio data through an audio analysis engine. The analysis includes musical elements such as pitch, rhythm, and tone quality, as well as understanding emotional expression using emotion recognition AI. The specific software used here consists of existing cloud services and emotion analysis platforms for automatic speech recognition.

[0648] The analysis results are used as data to generate visual and auditory responses in the virtual reality environment. The server uses this data to generate virtual audience reactions. The audience reactions are synchronized in real time with the user's voice and displayed on the viewer's screen.

[0649] For example, when a user sings an emotionally moving song during a live performance, if emotion analysis detects that the user was "moved," gestures indicating emotion will be reproduced in the virtual reality space as a reaction from the virtual audience. Simultaneously, viewer reactions are collected, and the user is provided with information such as, "It appears the audience was moved."

[0650] A concrete example of a prompt using a generative AI model would be the instruction, "Generate emotional feedback from viewers based on sentiment analysis." This allows users to intuitively receive feedback on their performance and gain further motivation.

[0651] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0652] Step 1:

[0653] The terminal acquires the user's voice information in real time through the microphone and formats it as digital voice data. This formatted data is input data prepared for transmission to the server.

[0654] Step 2:

[0655] The terminal sends formatted audio data to the server. The transmitted audio data becomes input data for audio analysis on the server.

[0656] Step 3:

[0657] The server analyzes the received audio data. It uses an audio analysis engine to evaluate musical elements (pitch, rhythm, tone quality) and an emotion recognition AI to perform emotion analysis. The results of this analysis are output as foundational data for generating visual and auditory responses.

[0658] Step 4:

[0659] The server generates visual and auditory responses in the virtual reality space based on the analysis results. The generated virtual audience reaction data is output as feedback to the user.

[0660] Step 5:

[0661] Through generated feedback, users receive real-time emotional responses from the audience and visually experience the reactions of a virtual audience on screen. This provides an emotional interaction with the user's performance.

[0662] Step 6:

[0663] The server processes user rhyme performance data in a format that can be uploaded to the community platform, facilitating sharing among users. This process allows for the incorporation of sentiment-based feedback from other users.

[0664] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0665] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0666] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0667] [Fourth Embodiment]

[0668] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0669] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0670] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0671] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0672] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0673] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0674] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0675] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0676] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0677] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0678] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0679] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0680] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0681] This invention provides a system for users to receive real-time feedback during a karaoke experience and improve their singing skills. The system mainly consists of a terminal, a server, and a virtual reality device.

[0682] The user first launches a karaoke application on their device and begins voice input. The device acquires the user's singing voice as digital audio data and sends it to the server in real time. The server then analyzes this audio data. Specifically, it uses AI technology to evaluate elements such as pitch, rhythm, and sound quality, and performs a detailed analysis of the user's performance.

[0683] After analysis, the server generates specific feedback for the user and sends it to their terminal. The feedback points out things like pitch discrepancies and rhythm inconsistencies and provides advice on how to improve.

[0684] Furthermore, when a user uses a virtual reality device, the device constructs a virtual live stage and generates audience reactions that follow the user's performance. This allows the user to have an immersive experience as if they were actually at a concert. The server provides and transmits real-time visual and auditory feedback, such as audience applause and cheers, to the device.

[0685] The system also includes a community platform where users can share their performance with other users. Users upload their performance results to the platform and receive comments and ratings from other users. Through this feedback, users can gain information to further improve their skills.

[0686] As a concrete example, consider a scenario where a user sings "Happy Birthday." In this case, the device sends audio data to a server, which analyzes the parts where the pitch is too high and provides feedback such as, "Next time, try singing this part a little lower." Simultaneously, in a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when they reach the chorus, promoting a sense of accomplishment.

[0687] Thus, the present invention provides a practical and enjoyable means for users to not only enjoy karaoke but also to improve their musical skills.

[0688] The following describes the processing flow.

[0689] Step 1:

[0690] The device acquires audio data from the microphone when the user launches a karaoke application and begins voice input.

[0691] Step 2:

[0692] The terminal transmits the acquired audio data to the server in real time. This transmission is performed using a protocol that minimizes data latency.

[0693] Step 3:

[0694] The server analyzes the received audio data using AI technology. Specifically, it evaluates pitch, rhythm, and timbre to determine the quality of the performance.

[0695] Step 4:

[0696] The server generates feedback based on the analysis results. For example, it identifies off-key notes and rhythmic discrepancies and compiles advice on areas for improvement.

[0697] Step 5:

[0698] The server sends the generated feedback to the terminal.

[0699] Step 6:

[0700] The device visualizes the received feedback for the user. This is displayed in text and simple graph formats, allowing the user to see specific areas for improvement.

[0701] Step 7:

[0702] When a user uses a VR device, the device constructs a virtual live stage. While tracking the user's movements, it prepares virtual audience reactions that synchronize with the performance.

[0703] Step 8:

[0704] The server generates visual and auditory feedback, such as applause and cheers, in accordance with the user's performance and sends it to the device in real time.

[0705] Step 9:

[0706] The device presents the user with these visual and auditory feedbacks, providing an immersive experience that makes them feel as if they are actually at a concert.

[0707] Step 10:

[0708] After the performance is complete, the device prepares to upload the recorded results to the community platform. This will only happen if the user allows it.

[0709] Step 11:

[0710] Users can watch performances uploaded by other users on the community platform and leave comments and ratings.

[0711] Step 12:

[0712] The server collects feedback from other users and notifies the original user. This allows the user to gain the insights necessary for self-improvement.

[0713] (Example 1)

[0714] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0715] Traditionally, entertainment activities like karaoke have lacked concrete feedback that allows users to objectively evaluate and improve their singing skills. Furthermore, opportunities for immersive experiences using virtual reality technology and for skill improvement through communication with other users are limited. This makes it difficult to improve users' singing skills and create a richer entertainment experience.

[0716] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0717] In this invention, the server includes means for acquiring and analyzing the user's acoustic data in real time and providing the user with suggestions for improvement; means for providing the user with visual and auditory feedback in real time in a virtual reality environment and generating virtual responses based on the user's actions; and means for providing an information exchange platform on which the suggestions for improvement can be shared with other participants. This allows the user to improve their singing skills through real-time feedback and enjoy responses that are just like those on a real live stage through an immersive experience in a virtual reality environment.

[0718] "Audio data" refers to information that digitally represents a user's vocalizations or musical expressions.

[0719] "Analysis" is the process of processing acoustic data and evaluating elements such as pitch, rhythm, and voice quality.

[0720] An "improvement suggestion" is a specific proposal provided to the user based on the analysis results, aimed at improving their singing skills.

[0721] A "virtual reality environment" is a computer-generated virtual space created using digital technology, into which a user can immerse themselves.

[0722] "Visual and auditory feedback" refers to responses provided to a user based on visual or auditory information.

[0723] "Virtual response" refers to the simulated audience or environment's reaction to a user's actions in a virtual reality environment.

[0724] An "information exchange platform" is a digital space where users can share analysis results and improvement suggestions, and interact with each other.

[0725] This invention is a system aimed at enabling users to improve their singing skills through karaoke. The system mainly consists of a terminal, a server, and a virtual reality device.

[0726] The user first launches a karaoke application on their device. This application has a voice input function and captures the user's singing voice through the microphone. The captured voice is converted into digital audio data. This digital audio data is then transmitted from the device to the server in real time.

[0727] The server uses AI technology to analyze the received digital audio data. Specifically, it uses AI libraries such as "TensorFlow" and "PyTorch" to evaluate pitch, rhythm, and audio quality. This analysis generates specific improvement suggestions for the user. These improvement suggestions are then provided to the user via their device.

[0728] When a user is using a virtual reality device, the terminal constructs a virtual live stage and generates virtual reactions based on the user's performance. This includes visual and auditory feedback such as applause and cheers from a virtual audience. The server generates and sends this feedback to the terminal in real time, providing the user with an immersive experience.

[0729] Furthermore, the system includes an information exchange platform where users can share their performance results with other users. Through this platform, users can receive comments and evaluations from other users. This allows users to receive more detailed feedback and obtain information to improve their singing skills.

[0730] As a concrete example, consider a case where a user sings "Happy Birthday" and the system analyzes the performance. In this case, the terminal sends the audio data to a server, which analyzes parts that are too high in pitch and provides specific suggestions for improvement, such as "Next time, try singing this part a little lower." In a virtual reality live performance, the user is shown a scene where a virtual audience stands up and applauds when the chorus is reached, providing an experience that promotes a sense of accomplishment.

[0731] An example of a prompt for a generative AI model would be, "Please tell me how to provide real-time voice feedback in a karaoke system."

[0732] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0733] Step 1:

[0734] The user launches the karaoke application on their device. The application displays a song list, allowing the user to select a song. The input is the user's song selection, and the output is information about the selected song.

[0735] Step 2:

[0736] The user selects a song and begins singing. The device acquires the user's singing voice in real time via the microphone and converts it into digital audio data. In this step, the input is the user's analog voice, and the output is digital audio data.

[0737] Step 3:

[0738] The terminal transmits the acquired digital audio data to the server via the network. Data compression is performed during transmission to reduce communication load. The input is digital audio data, and the output is a notification to the server indicating successful transmission.

[0739] Step 4:

[0740] The server inputs the received acoustic data into the AI ​​analysis system. Here, a generated AI model is used to analyze the pitch, rhythm, and sound quality. The input for this step is the acoustic data, and the output is the analysis result.

[0741] Step 5:

[0742] The server generates improvement suggestions for the user based on the analysis results. Specifically, it detects pitch discrepancies and rhythmic inconsistencies and outputs the areas for improvement as text. The input is the analysis results, and the output is the text of the improvement suggestions.

[0743] Step 6:

[0744] The server sends the suggested improvements to the terminal. Data compression is also performed here to ensure users receive feedback quickly. The input is the text of the suggested improvements, and the output is a notification that the transmission to the terminal is complete.

[0745] Step 7:

[0746] When the device is connected to a virtual reality system, it constructs a virtual live stage. It generates visual and auditory feedback, or virtual reactions, that respond to the user's performance. The input is user performance data, and the output is visual and auditory feedback in the virtual space.

[0747] Step 8:

[0748] Users can upload their singing performances and feedback to the community platform. They can also receive comments and ratings from other users. The input for this step is improvement suggestions and performance data, and the output is feedback from other users.

[0749] (Application Example 1)

[0750] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0751] Traditional karaoke systems lack the real-time, specific feedback necessary for users to effectively improve their singing skills. Furthermore, the virtual reality experience is limited, making it difficult to achieve the same level of immersion as attending an actual concert.

[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0753] In this invention, the server includes means for acquiring the user's voice signal in real time, means for analyzing the voice signal and providing information to the user, means for providing the user with visual and auditory information in a virtual environment in real time, and means for providing real-time visualized feedback to the user while they are singing using a smart device. This allows the user to evaluate their singing skills in real time and make concrete improvements while enjoying an immersive experience in a virtual environment.

[0754] "User voice signal" refers to data obtained by converting the voice spoken by the user into an electrical signal.

[0755] "Real-time acquisition" refers to the instantaneous collection of the user's voice signal as they speak.

[0756] "Analysis" is the process of analyzing collected audio signals based on certain criteria and detecting their characteristics.

[0757] An "information-providing device" is a device that presents analysis results to the user and facilitates their understanding.

[0758] A "virtual environment" is a computer-generated space, distinct from the real world, constructed using digital technology.

[0759] "Visual and auditory information" refers to data and feedback, such as responses, that are provided to the user visually or aurally.

[0760] A "smart device" is a portable electronic device equipped with a processor and communication devices that enables interaction with the user.

[0761] "Visualized feedback" refers to information that displays the results obtained through analysis in a way that allows users to intuitively understand them.

[0762] "Audience reaction" refers to the actions and emotional expressions of fictional people in response to a particular event.

[0763] The system for realizing this invention includes a series of processes for acquiring and analyzing the user's voice signal and providing real-time feedback. First, the user launches a karaoke application using a smart device and begins singing. The smart device acquires the user's voice signal in real time using its built-in microphone. The acquired voice signal is transmitted to a server in the cloud.

[0764] The server analyzes the audio data using a speech recognition API (for example, a common API for speech recognition) and evaluates the pitch, rhythm, and sound quality. Based on the analysis results, it generates specific suggestions for improvement regarding the user's musical performance and sends them to the smart device. This feedback is visualized on the smart device's display for intuitive understanding by the user and is also provided as audio feedback.

[0765] Furthermore, in the virtual environment, a computer-generated audience reacts in sync with the user's singing. The virtual environment uses a 3D engine such as Unity to provide real-time visual and auditory feedback to the user. The server pre-generates these audience reactions and plays them back in real-time in sync with the user's singing, enhancing the performance effect.

[0766] For example, when a user sings "Happy Birthday," a pitch bar is displayed on the smart device's screen, and any off-key sections are highlighted in red. Furthermore, during the chorus, the sound of a virtual audience beginning to applaud is played to enhance the user's sense of accomplishment.

[0767] An example of a prompt might be, "While singing the chorus of 'Happy Birthday,' provide real-time feedback on pitch deviations and visualize the audience's reaction."

[0768] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0769] Step 1:

[0770] The user launches a karaoke application on their smart device and begins singing. The input is the user's voice signal, and the output is recorded as audio data on the smart device. The smart device acquires this audio signal in real time through its microphone.

[0771] Step 2:

[0772] The smart device transmits the acquired voice data to a server in the cloud. The input is voice data, and the output is data transfer to the server. The smart device uses a communication module to upload the data to the server in real time.

[0773] Step 3:

[0774] The server analyzes audio data using a speech recognition API. The input is audio data sent from a smart device, and the output is an evaluation result regarding pitch, rhythm, and sound quality. The server processes the audio based on specific musical criteria and generates the analysis results.

[0775] Step 4:

[0776] The server generates improvement suggestions based on the analysis results and sends them to the smart device. The input is the result of the voice analysis, and the output is a specific feedback message. The server uses an AI algorithm to point out deviations in pitch and rhythm and create advice for improvement.

[0777] Step 5:

[0778] Smart devices receive feedback from a server and visualize it on their display to present to the user. The input is improvement suggestion messages, and the output is visualized feedback. Smart devices display graphical elements on their screens to help users intuitively understand the areas for improvement.

[0779] Step 6:

[0780] The server generates a virtual environment, creates audience reactions synchronized with the user's singing, and plays them back in real time. The input is real-time user singing data, and the output is visual and auditory audience reactions. The server constructs the virtual environment using Unity or similar software and stages scenarios where the audience applauds and cheers.

[0781] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0782] This invention provides a karaoke system that utilizes an emotion engine to analyze user voice data and recognize emotions. The system consists of a terminal, a server, and a virtual reality device.

[0783] The user launches a karaoke application on their device and begins voice input. The device acquires the user's voice data in real time and sends it to the server. The server analyzes the voice data and uses AI technology to evaluate musical elements such as pitch, rhythm, and tone quality. In addition, an emotion engine identifies the user's emotions based on the voice data.

[0784] The server generates feedback based on the analysis results and recognized emotional information, and sends it to the terminal. This feedback can include suggestions for musical technical improvements, as well as responses to and encouragement regarding the user's emotions.

[0785] Furthermore, when using a virtual reality device, the terminal constructs a virtual live stage in real time. The reactions of the virtual audience, synchronized with the user's singing, are adjusted according to the user's emotions. For example, if the user is expressing joyful emotions, the audience's reactions will be set to be more lively and the cheers louder.

[0786] The system also integrates a community platform for users to share the results of their voice data analysis, including emotional information, with other users. Users can upload their performance results to the platform and receive emotionally-based feedback and comments from other users, which can help them feel empathy and boost their motivation.

[0787] As a concrete example, suppose the emotion engine detects a sense of exhilaration when a user sings a "celebratory song." In this case, the server includes emotionally appealing feedback such as "convey that joy even more." Furthermore, the audience in the VR environment will be shown an animation that further heightens the user's emotions.

[0788] This means that the present invention allows users to receive various emotionally appealing feedback while singing karaoke, enabling them to improve their musical skills and gain a platform for emotional expression.

[0789] The following describes the processing flow.

[0790] Step 1:

[0791] When the user launches a karaoke application and begins voice input, the device acquires audio data from the microphone in real time.

[0792] Step 2:

[0793] The terminal sends the acquired audio data to the server. This transmission occurs in real time, and a protocol that minimizes data latency is used.

[0794] Step 3:

[0795] The server analyzes the received audio data using AI technology. The analysis includes evaluation of pitch, rhythm, and sound quality to determine musical skill.

[0796] Step 4:

[0797] The server uses an emotion engine to recognize the user's emotions from the voice data. The recognized emotions include, for example, joy, sadness, and excitement.

[0798] Step 5:

[0799] The server generates specific and emotionally responsive feedback based on analysis results and emotion recognition results. This feedback includes not only suggestions for musical improvements but also messages that resonate with the user's emotions.

[0800] Step 6:

[0801] The server sends the generated feedback to the terminal.

[0802] Step 7:

[0803] The device displays the received feedback to the user. The feedback is visualized in text and graph formats to help the user understand it.

[0804] Step 8:

[0805] When a user uses a virtual reality device, the device constructs a virtual live stage and adjusts the audience's reactions in sync with the user's performance.

[0806] Step 9:

[0807] The server generates virtual audience reactions based on the recognized user's emotions and sends them to the terminal.

[0808] Step 10:

[0809] The device plays these reactions in a virtual environment, providing the user with visual and auditory feedback.

[0810] Step 11:

[0811] The device prepares to upload performance recording data to the community platform.

[0812] Step 12:

[0813] On the community platform, other users can view uploaded performances and leave comments and ratings based on their emotions.

[0814] Step 13:

[0815] The server collects feedback from other users and notifies the original user. This allows users to gain emotional resonance and useful feedback.

[0816] (Example 2)

[0817] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0818] While conventional karaoke systems allowed users to receive feedback on their singing technique, they lacked mechanisms to improve emotional expression and the quality of the virtual reality experience. Furthermore, there was a challenge in that users had difficulty receiving emotion-based feedback through interaction with other users.

[0819] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0820] In this invention, the server includes means for acquiring user voice information and analyzing said voice information to identify musical elements and emotions; means for generating a virtual reality environment that is dynamically adjusted based on the emotions extracted from said voice information and providing visual and auditory feedback in said environment; and means for providing an electronic platform on which the analysis results and feedback can be shared with other users. This makes it possible for users to receive feedback on emotional expression in real time and to share emotions through interaction with others.

[0821] "User voice information" refers to the collection of acoustic data containing the user's voice, acquired through a voice input device.

[0822] "Analysis" refers to the computational process of identifying musical elements and emotions based on acquired audio information.

[0823] "Musical elements" refer to characteristics related to music that are extracted from audio information, such as pitch, rhythm, and tone quality.

[0824] "Emotions" refer to information that indicates the user's mental state or feelings, as identified from audio information.

[0825] A "virtual reality environment" refers to a three-dimensional space or audiovisual experience generated by a computer that a user can immerse themselves in.

[0826] "Visual and auditory feedback" refers to information returned to the user through visual images and auditory sounds, intended to enhance the quality of the experience.

[0827] An "electronic platform" is a technological foundation that enables the sharing and exchange of information through user communities and digital networks.

[0828] This invention relates to a system that analyzes voice information input by a user and provides musical guidance and emotional feedback. The system consists of multiple technical elements, including a terminal, a server, and a virtual reality device.

[0829] Users input voice information via the microphone using a karaoke application installed on their device. This voice information is digitized on the device and transmitted to the server in real time as audio data. WebSocket is used to maintain real-time transmission of the audio data.

[0830] The server uses a machine learning framework to analyze the received audio information. Specifically, it evaluates musical elements such as pitch, rhythm, and tone quality using pitch analysis algorithms and beat tracking. The emotion engine also uses libraries such as OpenSmile to classify emotions from the intonation and tempo of the voice. Through this analysis, not only individual musical characteristics but also the user's emotional state is evaluated.

[0831] The analysis results are compiled into feedback provided to the user using natural language generation technology. This feedback is delivered to the user visually and audibly through the device. For example, messages such as "Sing with a clearer pitch" or "Try to convey the enjoyment more effectively" may be displayed. Real-time feedback is achieved through dynamic page rendering using HTML templates and JavaScript.

[0832] When using a virtual reality device, the terminal also has the capability to build a realistic virtual live stage based on the emotions of the voice, using Unity or Unreal Engine. It generates virtual audience reactions that adjust according to the user's expression, providing immersion in both sight and sound.

[0833] Furthermore, users can upload the generated feedback to an electronic platform and share it with other users. Through interaction with other users, they can receive emotion-based feedback, which can boost their motivation.

[0834] For example, when a user sings a "celebratory song," the emotion engine recognizes the feeling of exhilaration, and the server provides feedback to the user saying, "Please convey that joy more." An example of a prompt to input into the generative AI model would be, "Please tell me how to analyze the audio data sung by the user, recognize the user's emotions, and generate feedback."

[0835] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0836] Step 1:

[0837] The terminal receives voice input from the karaoke application launched by the user and collects it through the microphone. The voice data as input is converted into a digital signal using an audio library and sent to the server in real time. This provides the voice data in a digital format that can be used for subsequent analysis processing.

[0838] Step 2:

[0839] The server receives audio data transmitted from the terminal. The received audio data is first subjected to pitch analysis algorithms and beat tracking to evaluate pitch, rhythm, and sound quality. These analyses produce evaluation data indicating the user's musical skills, and the process then proceeds to the next step of sentiment analysis.

[0840] Step 3:

[0841] The server performs sentiment analysis on the audio data in parallel with evaluating the musical elements. The sentiment engine uses the OpenSmile library to analyze the intonation and tempo of the voice and classify the user's emotions (e.g., exhilaration or sadness). This classified sentiment data is used to generate feedback.

[0842] Step 4:

[0843] The server integrates analysis data of musical elements and emotions to generate feedback. This process utilizes natural language generation technology to create specific improvement suggestions for the user (e.g., "Let's stabilize the pitch") and emotional messages (e.g., "Let's convey the enjoyment more"). The generated feedback is then formatted for visualization.

[0844] Step 5:

[0845] The device receives feedback from the server and presents it to the user visually and audibly. The input feedback data is displayed on the screen using HTML and JavaScript, allowing the user to see the results in real time. This phase aims to improve the user's singing experience through feedback.

[0846] Step 6:

[0847] When a virtual reality device is connected, the terminal uses a game engine to construct a virtual live stage. Based on the emotional data of the input audio, the reactions of the virtual audience are dynamically generated. For example, if the emotion is heightened, an animation of the virtual audience becoming excited will be displayed. This output provides the user with an immersive experience.

[0848] Step 7:

[0849] Users upload the analyzed data and generated feedback to an electronic platform. This step facilitates information sharing and communication with other users via their devices. The outputted feedback is used to solicit sentiment-based comments from other community members.

[0850] (Application Example 2)

[0851] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0852] Conventional voice analysis systems have been unable to provide real-time feedback that takes user emotions into account, posing a challenge in fostering emotional engagement with the audience, especially during live performances. Furthermore, there was a lack of specific feedback necessary for users to evaluate and improve their own performance.

[0853] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0854] In this invention, the server includes means for acquiring and analyzing user voice information in real time, means for providing the user with visual and auditory responses in a virtual reality space in real time, and means for analyzing the user's emotional expression during live performance streaming and acquiring and presenting emotional responses from viewers in real time. This allows the user to deepen their emotional interaction with viewers while receiving rich feedback in real time to evaluate and improve their performance.

[0855] "User voice information" refers to all voice data emitted by the user, and is usually collected in digital format via a microphone.

[0856] "Real-time acquisition" refers to a technology that minimizes the delay between transmission and acquisition, allowing for the immediate recording of voice information on the spot.

[0857] "Analysis" refers to the process of performing data processing on collected audio information to recognize specific patterns or emotions.

[0858] "Response" refers to feedback or reactions generated based on the analysis of audio information, and is the information or reaction provided to the user or audience.

[0859] A "virtual reality space" is a three-dimensional digital environment generated by a computer that can be experienced visually and aurally.

[0860] "Emotional expression" refers to the emotions and feelings that users express through words, music, and other means.

[0861] "Emotional responses from viewers" refer to the emotional reactions and feedback provided by people watching a broadcast or performance.

[0862] The system for carrying out this invention is configured as follows: The user first transmits their voice information on a live performance distribution platform via a mobile device or head-mounted display. Based on this, the terminal acquires the voice information in real time using a microphone and performs initial digital processing.

[0863] The server receives audio information transmitted from the aforementioned dev device and analyzes the audio data through an audio analysis engine. The analysis includes musical elements such as pitch, rhythm, and tone quality, as well as understanding emotional expression using emotion recognition AI. The specific software used here consists of existing cloud services and emotion analysis platforms for automatic speech recognition.

[0864] The analysis results are used as data to generate visual and auditory responses in the virtual reality environment. The server uses this data to generate virtual audience reactions. The audience reactions are synchronized in real time with the user's voice and displayed on the viewer's screen.

[0865] For example, when a user sings an emotionally moving song during a live performance, if emotion analysis detects that the user was "moved," gestures indicating emotion will be reproduced in the virtual reality space as a reaction from the virtual audience. Simultaneously, viewer reactions are collected, and the user is provided with information such as, "It appears the audience was moved."

[0866] A concrete example of a prompt using a generative AI model would be the instruction, "Generate emotional feedback from viewers based on sentiment analysis." This allows users to intuitively receive feedback on their performance and gain further motivation.

[0867] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0868] Step 1:

[0869] The terminal acquires the user's voice information in real time through the microphone and formats it as digital voice data. This formatted data is input data prepared for transmission to the server.

[0870] Step 2:

[0871] The terminal sends formatted audio data to the server. The transmitted audio data becomes input data for audio analysis on the server.

[0872] Step 3:

[0873] The server analyzes the received audio data. It uses an audio analysis engine to evaluate musical elements (pitch, rhythm, tone quality) and an emotion recognition AI to perform emotion analysis. The results of this analysis are output as foundational data for generating visual and auditory responses.

[0874] Step 4:

[0875] The server generates visual and auditory responses in the virtual reality space based on the analysis results. The generated virtual audience reaction data is output as feedback to the user.

[0876] Step 5:

[0877] Through generated feedback, users receive real-time emotional responses from the audience and visually experience the reactions of a virtual audience on screen. This provides an emotional interaction with the user's performance.

[0878] Step 6:

[0879] The server processes user rhyme performance data in a format that can be uploaded to the community platform, facilitating sharing among users. This process allows for the incorporation of sentiment-based feedback from other users.

[0880] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0881] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0882] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0883] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0884] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0885] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0886] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0887] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0888] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0889] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0890] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0891] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0892] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0893] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0894] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0895] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0896] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0897] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0898] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0899] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0900] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0901] The following is further disclosed regarding the embodiments described above.

[0902] (Claim 1)

[0903] The system acquires user voice data in real time.

[0904] A means for analyzing the aforementioned audio data and providing feedback to the user,

[0905] A means of providing users with real-time visual and auditory feedback in a virtual reality environment,

[0906] A means of providing a community platform on which the aforementioned feedback can be shared with other users,

[0907] A system that includes this.

[0908] (Claim 2)

[0909] The system according to claim 1, comprising means for analyzing user voice data, evaluating pitch, rhythm, and sound quality, and generating specific improvement suggestions.

[0910] (Claim 3)

[0911] The system according to claim 1, comprising means for generating and playing back in real time audience reactions synchronized with a user's singing in a virtual reality environment.

[0912] "Example 1"

[0913] (Claim 1)

[0914] Acquire user audio data in real time,

[0915] A means for analyzing the aforementioned acoustic data and providing improvement suggestions to the user,

[0916] A means for providing users with real-time visual and auditory feedback in a virtual reality environment and generating virtual responses based on user actions,

[0917] A means of providing an information exchange platform where the aforementioned improvement proposals can be shared with other participants,

[0918] A system that includes this.

[0919] (Claim 2)

[0920] The system according to claim 1, comprising means for analyzing user acoustic data, evaluating pitch, rhythm, and voice quality, and generating specific improvement proposals.

[0921] (Claim 3)

[0922] The system according to claim 1, comprising means for generating and playing back in real time audience reactions synchronized with the user's musical expression in a virtual reality environment.

[0923] "Application Example 1"

[0924] (Claim 1)

[0925] The system acquires the user's voice signal in real time.

[0926] A device that analyzes the aforementioned audio signal and provides information to the user,

[0927] A device that provides users with visual and auditory information in real time within a virtual environment,

[0928] A device that provides an exchange platform that allows the aforementioned information to be shared with other users,

[0929] A device that provides real-time visualized feedback to a user singing using a smart device,

[0930] A system that includes this.

[0931] (Claim 2)

[0932] The system according to claim 1, comprising a device that analyzes the user's voice signal, evaluates the pitch, rhythm, and sound quality, and generates specific improvement suggestions.

[0933] (Claim 3)

[0934] The system according to claim 1, comprising a device that generates and plays back in real time visualized audience reactions synchronized with a user's singing in a virtual environment.

[0935] "Example 2 of combining an emotion engine"

[0936] (Claim 1)

[0937] A means for acquiring user voice information and analyzing said voice information to identify musical elements and emotions,

[0938] A means for generating a virtual reality environment that is dynamically adjusted based on emotions extracted from the aforementioned audio information, and for providing visual and auditory feedback in the environment,

[0939] A means of providing an electronic platform that allows users to share analysis results and feedback with other users,

[0940] A system that includes this.

[0941] (Claim 2)

[0942] The system according to claim 1, comprising means for analyzing audio information, evaluating pitch, rhythm, and sound quality as musical elements, and generating specific improvement proposals based on these evaluations.

[0943] (Claim 3)

[0944] The system according to claim 1, comprising means for generating a response synchronized with a user's expression in a virtual reality environment and for reproducing the response in real time.

[0945] "Application example 2 when combining with an emotional engine"

[0946] (Claim 1)

[0947] The system acquires user voice information in real time.

[0948] A device that analyzes the aforementioned voice information and provides a response to the user,

[0949] A device that provides users with real-time visual and auditory responses in a virtual reality space,

[0950] A device that provides a community platform on which the aforementioned response can be shared with other users,

[0951] A device that analyzes users' emotional expressions during live performance streaming and acquires and presents emotional responses from viewers in real time,

[0952] A system that includes this.

[0953] (Claim 2)

[0954] The system according to claim 1, comprising a device that analyzes the user's voice information, evaluates the pitch, rhythm, and sound quality of the sound, and generates specific improvement suggestions.

[0955] (Claim 3)

[0956] The system according to claim 1, comprising a device that generates and reproduces in real time the reactions of a virtual audience synchronized with the user's expression in a virtual reality space. [Explanation of symbols]

[0957] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. The system acquires user voice data in real time. A means for analyzing the aforementioned audio data and providing feedback to the user, A means of providing users with real-time visual and auditory feedback in a virtual reality environment, A means of providing a community platform on which the aforementioned feedback can be shared with other users, A system that includes this.

2. The system according to claim 1, comprising means for analyzing user voice data, evaluating pitch, rhythm, and sound quality, and generating specific improvement suggestions.

3. The system according to claim 1, comprising means for generating and playing back in real time audience reactions synchronized with the user's singing in a virtual reality environment.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A