system

The system addresses the lack of effective feedback for tone-deaf singers by analyzing their voice in real-time, offering audio and visual feedback, and generating corrected versions, enabling them to improve their singing skills.

JP2026072754APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

There is a lack of effective feedback and practice methods for tone-deaf people to sing with confidence.

Method used

A system comprising an analysis unit, feedback unit, visual feedback unit, and evaluation unit that analyzes singing voice in real-time, provides audio and visual feedback, and automatically generates a version with correct pitch, along with periodic evaluations and reports.

Benefits of technology

Enables tone-deaf individuals to practice singing with confidence by identifying specific areas for improvement and providing immediate feedback, allowing them to overcome tone-deafness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072754000001_ABST
    Figure 2026072754000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to enable tone-deaf people to sing with confidence. [Solution] The system according to the embodiment comprises an analysis unit, a feedback unit, a visual feedback unit, a generation unit, and an evaluation unit. The analysis unit analyzes the singing voice in real time. The feedback unit provides audio feedback based on the results analyzed by the analysis unit. The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. The evaluation unit provides periodic evaluations and reports based on the version with correct pitch generated by the generation unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that there is a lack of effective feedback and practice methods for tone-deaf people to be able to sing with confidence.

[0005] The system according to the embodiment aims to enable tone-deaf people to sing with confidence.

Means for Solving the Problems

[0006] The system according to this embodiment comprises an analysis unit, a feedback unit, a visual feedback unit, a generation unit, and an evaluation unit. The analysis unit analyzes the singing voice in real time. The feedback unit provides audio feedback based on the results analyzed by the analysis unit. The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. The evaluation unit provides periodic evaluations and reports based on the version with correct pitch generated by the generation unit. [Effects of the Invention]

[0007] The system according to this embodiment can enable tone-deaf people to sing with confidence. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The AI ​​voice training system according to an embodiment of the present invention is a system for people who are self-conscious about being tone-deaf and lack the courage to sing in front of others. This system provides an environment in which users can practice secretly at home. The AI ​​analyzes the singing voice in real time and points out specific areas for improvement regarding pitch, rhythm, and vocalization. This allows users to practice at their own pace without worrying about what others think. Next, in addition to audio feedback, visual feedback is also provided. For example, pitch deviations are shown in graphs, and correct vocalization methods are explained in videos to provide easy-to-understand instruction. Furthermore, a version with correct pitch is automatically generated, allowing users to understand specific areas for improvement. In addition, by visualizing progress through regular evaluations and reports, users can feel their progress and gain confidence. This will enable everyone to overcome tone-deafness and realize a future where everyone can sing with confidence. Thus, the AI ​​voice training system analyzes the user's singing voice in real time, provides audio and visual feedback, automatically generates a version with correct pitch, and provides regular evaluations and reports, enabling users to overcome tone-deafness and sing with confidence.

[0029] The AI ​​voice training system according to this embodiment comprises an analysis unit, a feedback unit, a visual feedback unit, a generation unit, and an evaluation unit. The analysis unit analyzes the user's singing voice in real time. The analysis unit points out specific areas for improvement, such as pitch, rhythm, and vocalization. The feedback unit provides audio feedback based on the results analyzed by the analysis unit. The feedback unit, for example, shows pitch deviations in a graph. The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. The visual feedback unit, for example, explains the correct vocalization method in a video. The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. The generation unit generates a version with correct pitch using, for example, a pitch correction algorithm. The evaluation unit provides periodic evaluations and reports based on the version with correct pitch generated by the generation unit. The evaluation unit performs, for example, weekly and monthly evaluations to visualize the user's progress. As a result, the AI ​​voice training system according to this embodiment analyzes the user's singing voice in real time, provides audio and visual feedback, automatically generates a version with correct pitch, and provides regular evaluations and reports, thereby enabling the user to overcome tone-deafness and sing with confidence.

[0030] The analysis unit analyzes the user's singing voice in real time. Specifically, it instantly processes the audio data input through the microphone when the user sings, and analyzes elements such as pitch, rhythm, and vocal technique in detail. In pitch analysis, the audio signal is decomposed into frequency components to identify the pitch of each note. In rhythm analysis, the timing and tempo of the voice are measured, and the accuracy to the beat of the song is evaluated. In vocal technique analysis, the waveform and spectrum of the voice are analyzed, and the volume, quality, and resonance of the voice are evaluated. These analysis results are used to understand the current state of the user's singing technique and to point out specific areas for improvement. For example, it identifies parts where the pitch is off, parts where the rhythm is off, and parts where there are problems with the vocal technique, and provides the user with specific advice. Furthermore, the analysis unit evaluates the user's progress by comparing it with past data, supporting long-term growth. This allows the user to clearly understand the weaknesses in their singing technique and practice effectively.

[0031] The feedback unit provides voice feedback based on the results analyzed by the analysis unit. Specifically, it analyzes the voice data sung by the user and points out pitch inaccuracies, rhythmic errors, and problems with vocal technique. For example, it displays pitch inaccuracies in a graph, visually indicating which notes are too high or too low. For rhythmic errors, it compares the correct timing on the musical score with the actual timing, clearly showing where the rhythm is off. For problems with vocal technique, it analyzes the waveform and spectrum of the voice, evaluates the volume, tone quality, and resonance, and suggests specific methods for improvement. The feedback unit provides this information to the user as voice feedback, encouraging real-time correction. For example, if the user sings off-key, it provides voice messages such as "The pitch is too high" or "The rhythm is too fast" on the spot. This allows the user to immediately correct their singing technique and practice effectively.

[0032] The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. Specifically, it analyzes the audio data sung by the user and visually displays pitch deviations, rhythm errors, and problems with vocal technique. For example, it shows pitch deviations as a graph, visually indicating which notes are too high or too low. For rhythm errors, it compares the correct timing on the musical score with the actual timing, clearly indicating where the rhythm is off. For problems with vocal technique, it analyzes the waveform and spectrum of the voice, evaluates the volume, tone quality, and resonance, and suggests specific methods for improvement. The visual feedback unit visually displays this information, showing the user specific areas for improvement. For example, it explains the correct vocal technique with a video, specifically showing how the user should sing. It also shows pitch deviations as a graph, visually indicating which notes are too high or too low. This allows the user to clearly understand the weaknesses in their singing technique and practice effectively.

[0033] The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. Specifically, it analyzes the audio data sung by the user and applies algorithms to correct pitch deviations and rhythmic errors. For example, it uses a pitch correction algorithm to generate a version with correct pitch. This algorithm decomposes the audio signal into frequency components, identifies the pitch of each note, and corrects it to the correct pitch. It also corrects rhythmic errors by adjusting the timing to achieve the correct rhythm. The generation unit performs these corrections in real time and provides immediate feedback to the user. For example, if the user sings off-key, the unit corrects the pitch on the spot and generates a version with the correct pitch. This allows the user to immediately correct weaknesses in their singing technique and practice effectively.

[0034] The evaluation unit provides regular evaluations and reports based on the pitch-correct version generated by the generation unit. Specifically, it analyzes the audio data sung by the user and evaluates pitch deviations, rhythm errors, and problems with vocal technique. Based on this information, the evaluation unit grasps the current state of the user's singing technique and points out specific areas for improvement. For example, it conducts weekly and monthly evaluations to visualize the user's progress. The evaluation unit provides this information as a report, showing the user specific areas for improvement. For example, it shows pitch deviations in a graph, visually indicating which notes are too high or too low. Regarding rhythm errors, it compares the correct timing on the musical score with the actual timing to clearly show where the rhythm is off. Regarding problems with vocal technique, it analyzes the waveform and spectrum of the voice to evaluate the volume, tone quality, and resonance, and proposes specific methods for improvement. This allows the user to clearly understand the weaknesses in their singing technique and practice effectively.

[0035] The analysis unit can point out specific areas for improvement regarding pitch, rhythm, and vocal technique. For example, the analysis unit can point out pitch inaccuracies, rhythmic irregularities, and errors in vocal technique. This allows users to practice more effectively by providing specific areas for improvement. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's singing voice data into an AI, which can then point out areas for improvement in pitch, rhythm, and vocal technique.

[0036] The feedback unit can display pitch deviations graphically. For example, the feedback unit can display pitch deviations in cents. The feedback unit can also display pitch deviations with color changes. Furthermore, the feedback unit can display pitch deviations with animations. This makes it easier for users to understand areas for improvement by visually representing pitch deviations. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can generate graphs based on pitch deviation data generated by AI.

[0037] The visual feedback unit can explain correct vocalization techniques through videos. For example, the visual feedback unit can explain diaphragmatic breathing techniques through videos. It can also explain how to use the vocal cords through videos. Furthermore, the visual feedback unit can explain proper vocalization posture through videos. This allows users to visually learn correct vocalization techniques through video explanations. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can provide videos of vocalization techniques generated by AI.

[0038] The generation unit can automatically generate a version with correct pitch. For example, the generation unit can generate a version with correct pitch using a pitch correction algorithm. Alternatively, the generation unit can generate a version with correct pitch using a reference pitch. Furthermore, the generation unit can generate a version with correct pitch using pitch correction software. This makes it easier for users to understand specific areas for improvement by automatically generating a version with correct pitch. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can generate a version with correct pitch based on pitch correction data generated by AI.

[0039] The evaluation department can provide periodic evaluations and reports. For example, the evaluation department can conduct weekly evaluations to visualize user growth. It can also conduct monthly evaluations to report on user progress. Furthermore, the evaluation department can conduct evaluations based on evaluation items and provide reports. This makes it easier for users to feel their growth by providing periodic evaluations and reports. Some or all of the above processes in the evaluation department may be performed using AI, for example, or not. For example, the evaluation department can create reports based on evaluation data generated by AI.

[0040] The analysis unit can identify individual areas for improvement by referring to the user's past singing data. For example, the analysis unit can identify specific pitch deviations based on the user's past singing data and point out areas for improvement. The analysis unit can also analyze the user's past rhythm patterns and identify areas for improvement in rhythm. Furthermore, the analysis unit can refer to the user's past vocalization techniques and identify areas for improvement in vocalization techniques. This makes it easier to identify individual areas for improvement by referring to past data. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's past singing data into AI, which can then identify individual areas for improvement.

[0041] The analysis unit can propose the optimal vocalization method by considering the user's voice quality and vocal cord condition. For example, the analysis unit can analyze the user's voice quality and propose the optimal vocalization method. It can also consider the user's vocal cord condition and propose a vocalization method that does not strain the vocal cords. Furthermore, the analysis unit can comprehensively analyze the user's voice quality and vocal cord condition to propose the optimal vocalization method. This allows for effective practice by proposing a vocalization method tailored to the user's voice quality and vocal cord condition. Some or all of the above processing in the analysis unit may be performed using AI, or without AI. For example, the analysis unit can input the user's voice quality data into AI, which can then propose the optimal vocalization method.

[0042] The analysis unit can correct the analysis results by taking into account the user's singing environment. For example, the analysis unit can analyze the acoustic characteristics of the room and correct the analysis results based on those characteristics. The analysis unit can also take into account the user's singing environment and provide analysis results that are appropriate for that environment. Furthermore, the analysis unit can comprehensively analyze the acoustic characteristics of the room and the user's singing environment to provide optimal analysis results. This allows for more accurate analysis results by taking the singing environment into consideration. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's singing environment data into the AI, and the AI ​​can correct the analysis results.

[0043] The analysis unit can apply different analysis algorithms depending on the user's singing style. For example, if the user's singing style is pop, the analysis unit will apply an analysis algorithm suitable for pop music. Similarly, if the user's singing style is classical, the analysis unit can apply an analysis algorithm suitable for classical music. Furthermore, the analysis unit can select and apply the optimal analysis algorithm based on the user's singing style. This allows for more appropriate feedback by applying an analysis algorithm tailored to the singing style. Some or all of the above processing in the analysis unit may be performed using AI, or without AI. For example, the analysis unit can input the user's singing style data into an AI, which can then apply the optimal analysis algorithm.

[0044] The feedback unit can provide personalized advice by referring to the user's past feedback history. For example, the feedback unit can point out specific areas for improvement based on the user's past feedback history. The feedback unit can also refer to the user's past feedback history and provide ongoing areas for improvement. Furthermore, the feedback unit can analyze the user's past feedback history and provide optimal advice. This makes it easier to provide personalized advice by referring to past feedback history. Some or all of the above processing in the feedback unit may be performed using AI, for example, or not using AI. For example, the feedback unit can input the user's past feedback history data into AI, and the AI ​​can provide personalized advice.

[0045] The feedback unit can suggest specific areas for improvement based on the user's singing goals. For example, if the user's singing goal is to sing a particular song perfectly, the feedback unit will suggest areas for improvement specific to that song. The feedback unit can also provide specific areas for improvement based on the user's singing goals. Furthermore, the feedback unit can consider the user's singing goals and suggest the most suitable areas for improvement. This allows the user to practice effectively towards their goals by suggesting areas for improvement based on their singing goals. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's singing goal data into an AI, which can then suggest specific areas for improvement.

[0046] The feedback unit can adjust the feedback content by taking into account the user's singing environment. For example, the feedback unit can adjust the feedback content by considering the type of microphone the user is using. The feedback unit can also provide optimal feedback by considering the user's singing environment. Furthermore, the feedback unit can comprehensively analyze the microphone type and singing environment and adjust the feedback content. This allows for more accurate feedback by considering the singing environment. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's singing environment data into the AI, which can then adjust the feedback content.

[0047] The feedback unit can provide different feedback formats depending on the user's singing style. For example, if the user's singing style is pop, the feedback unit provides audio feedback. Alternatively, if the user's singing style is classical, the feedback unit can provide text feedback. Furthermore, the feedback unit can select and provide the most appropriate feedback format based on the user's singing style. This allows for more appropriate instruction by providing feedback formats tailored to the singing style. Some or all of the above processing in the feedback unit may be performed using AI, or without AI. For example, the feedback unit can input the user's singing style data into an AI, which can then provide the most appropriate feedback format.

[0048] The visual feedback unit can provide personalized advice by referring to the user's past visual feedback history. For example, the visual feedback unit can point out specific areas for improvement based on the user's past visual feedback history. The visual feedback unit can also refer to the user's past visual feedback history and provide ongoing areas for improvement. Furthermore, the visual feedback unit can analyze the user's past visual feedback history and provide optimal advice. This makes it easier to provide personalized advice by referring to past visual feedback history. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's past visual feedback history data into AI, and the AI ​​can provide personalized advice.

[0049] The visual feedback unit can indicate specific areas for improvement based on the user's singing goals. For example, if the user's singing goal is to sing a particular song perfectly, the visual feedback unit will indicate areas for improvement specific to that song. The visual feedback unit can also provide specific areas for improvement based on the user's singing goals. Furthermore, the visual feedback unit can consider the user's singing goals and indicate the most suitable areas for improvement. This allows the user to practice effectively towards their goals by indicating areas for improvement based on their singing goals. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's singing goal data into AI, which can then indicate specific areas for improvement.

[0050] The visual feedback unit can correct the visual feedback content by taking into account the user's singing environment. For example, the visual feedback unit can provide optimal visual feedback by considering the user's singing environment. The visual feedback unit can also analyze the acoustic characteristics of the room and correct the visual feedback content based on those characteristics. Furthermore, the visual feedback unit can comprehensively analyze the user's singing environment and the acoustic characteristics of the room to provide optimal visual feedback. This allows for more accurate visual feedback by taking the singing environment into consideration. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's singing environment data into the AI, which can then correct the visual feedback content.

[0051] The visual feedback unit can provide different visual feedback formats depending on the user's singing style. For example, if the user's singing style is pop, the visual feedback unit can provide graph-based visual feedback. Alternatively, if the user's singing style is classical, the visual feedback unit can provide video-based visual feedback. Furthermore, the visual feedback unit can select and provide the most appropriate visual feedback format based on the user's singing style. This allows for more appropriate instruction by providing visual feedback formats tailored to the singing style. Some or all of the above processing in the visual feedback unit may be performed using AI, or without AI. For example, the visual feedback unit can input the user's singing style data into an AI, which can then provide the most appropriate visual feedback format.

[0052] The generation unit can refer to the user's past singing data and reflect individual areas for improvement. For example, the generation unit can reflect specific pitch deviations based on the user's past singing data. The generation unit can also analyze the user's past rhythm patterns and reflect areas for improvement in rhythm. Furthermore, the generation unit can refer to the user's past vocalization techniques and reflect areas for improvement in vocalization techniques. This makes it easier to reflect individual areas for improvement by referring to past data. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's past singing data into AI, and the AI ​​can reflect individual areas for improvement.

[0053] The generation unit can generate the optimal pitch by considering the user's voice quality and vocal cord condition. For example, the generation unit analyzes the user's voice quality and generates the optimal pitch. The generation unit can also consider the user's vocal cord condition and generate a pitch that does not strain the vocal cords. Furthermore, the generation unit can comprehensively analyze the user's voice quality and vocal cord condition to generate the optimal pitch. This enables effective practice by generating a pitch that matches the user's voice quality and vocal cord condition. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's voice quality data into AI, and the AI ​​can generate the optimal pitch.

[0054] The generation unit can correct the generation results by taking into account the user's singing environment. For example, the generation unit can consider the user's singing environment and provide the optimal generation result. The generation unit can also analyze the acoustic characteristics of the room and correct the generation result based on those characteristics. Furthermore, the generation unit can comprehensively analyze the user's singing environment and the acoustic characteristics of the room to provide the optimal generation result. This allows for the provision of more accurate generation results by considering the singing environment. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's singing environment data into the AI, and the AI ​​can correct the generation result.

[0055] The generation unit can apply different generation algorithms depending on the user's singing style. For example, if the user's singing style is pop, the generation unit will apply a generation algorithm suitable for pop music. Similarly, if the user's singing style is classical, the generation unit can apply a generation algorithm suitable for classical music. Furthermore, the generation unit can select and apply the optimal generation algorithm according to the user's singing style. This allows for more appropriate feedback by applying a generation algorithm tailored to the singing style. Some or all of the above-described processes in the generation unit may be performed using AI, or without AI. For example, the generation unit can input the user's singing style data into an AI, which can then apply the optimal generation algorithm.

[0056] The evaluation unit can provide individual evaluations by referring to the user's past evaluation data. For example, the evaluation unit can provide specific evaluation points based on the user's past evaluation data. The evaluation unit can also refer to the user's past evaluation data and provide continuous evaluation points. Furthermore, the evaluation unit can analyze the user's past evaluation data and provide optimal evaluation points. This makes it easier to provide individual evaluations by referring to past evaluation data. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's past evaluation data into AI, and the AI ​​can provide individual evaluations.

[0057] The evaluation unit can provide specific evaluation points based on the user's singing goals. For example, if the user's singing goal is to sing a particular song perfectly, the evaluation unit will provide evaluation points specific to that song. The evaluation unit can also provide specific evaluation points based on the user's singing goals. Furthermore, the evaluation unit can consider the user's singing goals and provide optimal evaluation points. This allows the user to practice effectively towards their goals by providing evaluation points based on their singing goals. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's singing goal data into AI, which can then provide specific evaluation points.

[0058] The evaluation unit can adjust the evaluation results by taking into account the user's singing environment. For example, the evaluation unit can provide an optimal evaluation by considering the user's singing environment. The evaluation unit can also analyze the acoustic characteristics of the room and adjust the evaluation results based on those characteristics. Furthermore, the evaluation unit can comprehensively analyze the user's singing environment and the acoustic characteristics of the room to provide an optimal evaluation. This allows for a more accurate evaluation by considering the singing environment. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's singing environment data into the AI, which can then adjust the evaluation results.

[0059] The evaluation unit can provide different evaluation formats depending on the user's singing style. For example, if the user's singing style is pop, the evaluation unit can provide a score-based evaluation. Alternatively, if the user's singing style is classical, the evaluation unit can provide a comment-based evaluation. Furthermore, the evaluation unit can select and provide the most appropriate evaluation format based on the user's singing style. This allows for more appropriate guidance by providing evaluation formats tailored to the singing style. Some or all of the above processing in the evaluation unit may be performed using AI, or without AI. For example, the evaluation unit can input the user's singing style data into an AI, which can then provide the most appropriate evaluation format.

[0060] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0061] An AI voice training system can have the functionality to save users' singing data to the cloud and share it with other users. For example, a user can upload their singing data to the cloud, and other users can download that data for reference. Users can also provide comments and feedback on other users' singing data. Furthermore, a ranking function can be provided on the cloud where users can compete with each other. This allows users to interact with other users and learn from each other.

[0062] The AI ​​voice training system can automatically generate personalized training plans based on the user's singing data. For example, it can identify the user's weaknesses in pitch and rhythm and provide a training menu accordingly. It can also adjust the training plan according to the user's progress. Furthermore, it can provide a training plan tailored to the user's goals. This allows users to receive the most suitable training for themselves.

[0063] An AI voice training system can feature a function that allows users to perform with a virtual band based on their singing data. For example, a user can play their own singing data along with a virtual band and perform together. The virtual band can also adjust its performance to match the user's singing. Furthermore, the virtual band can provide feedback on the user's singing. This allows users to practice while experiencing the feeling of performing with a band.

[0064] An AI voice training system can feature a virtual audience that reacts in real time based on the user's singing data. For example, the virtual audience can applaud or cheer in response to the user's singing. The virtual audience can also provide feedback on the user's singing. Furthermore, the virtual audience can change its reactions according to the user's progress. This allows the user to practice while feeling the audience's reactions.

[0065] An AI voice training system can have the functionality to simulate performance on a virtual stage based on the user's singing data. For example, the user can sing on a virtual stage and have their performance evaluated. It can also simulate movements and facial expressions on the virtual stage. Furthermore, it can provide feedback on the performance on the virtual stage. This allows users to practice while simulating their performance on a real stage.

[0066] The following briefly describes the processing flow for example form 1.

[0067] Step 1: The analysis unit analyzes the user's singing voice in real time. The analysis unit points out specific areas for improvement, such as pitch, rhythm, and vocal technique. Step 2: The feedback unit provides audio feedback based on the results analyzed by the analysis unit. For example, the feedback unit displays pitch deviations graphically. Step 3: The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. For example, the visual feedback unit might explain the correct pronunciation technique with a video. Step 4: The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. The generation unit generates a version with correct pitch using, for example, a pitch correction algorithm. Step 5: The evaluation unit provides periodic evaluations and reports based on the correct version of the pitch generated by the generation unit. The evaluation unit performs weekly and monthly evaluations, for example, to visualize user growth.

[0068] (Example of form 2) The AI ​​voice training system according to an embodiment of the present invention is a system for people who are self-conscious about being tone-deaf and lack the courage to sing in front of others. This system provides an environment in which users can practice secretly at home. The AI ​​analyzes the singing voice in real time and points out specific areas for improvement regarding pitch, rhythm, and vocalization. This allows users to practice at their own pace without worrying about what others think. Next, in addition to audio feedback, visual feedback is also provided. For example, pitch deviations are shown in graphs, and correct vocalization methods are explained in videos to provide easy-to-understand instruction. Furthermore, a version with correct pitch is automatically generated, allowing users to understand specific areas for improvement. In addition, by visualizing progress through regular evaluations and reports, users can feel their progress and gain confidence. This will enable everyone to overcome tone-deafness and realize a future where everyone can sing with confidence. Thus, the AI ​​voice training system analyzes the user's singing voice in real time, provides audio and visual feedback, automatically generates a version with correct pitch, and provides regular evaluations and reports, enabling users to overcome tone-deafness and sing with confidence.

[0069] The AI ​​voice training system according to this embodiment comprises an analysis unit, a feedback unit, a visual feedback unit, a generation unit, and an evaluation unit. The analysis unit analyzes the user's singing voice in real time. The analysis unit points out specific areas for improvement, such as pitch, rhythm, and vocalization. The feedback unit provides audio feedback based on the results analyzed by the analysis unit. The feedback unit, for example, shows pitch deviations in a graph. The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. The visual feedback unit, for example, explains the correct vocalization method in a video. The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. The generation unit generates a version with correct pitch using, for example, a pitch correction algorithm. The evaluation unit provides periodic evaluations and reports based on the version with correct pitch generated by the generation unit. The evaluation unit performs, for example, weekly and monthly evaluations to visualize the user's progress. As a result, the AI ​​voice training system according to this embodiment analyzes the user's singing voice in real time, provides audio and visual feedback, automatically generates a version with correct pitch, and provides regular evaluations and reports, thereby enabling the user to overcome tone-deafness and sing with confidence.

[0070] The analysis unit analyzes the user's singing voice in real time. Specifically, it instantly processes the audio data input through the microphone when the user sings, and analyzes elements such as pitch, rhythm, and vocal technique in detail. In pitch analysis, the audio signal is decomposed into frequency components to identify the pitch of each note. In rhythm analysis, the timing and tempo of the voice are measured, and the accuracy to the beat of the song is evaluated. In vocal technique analysis, the waveform and spectrum of the voice are analyzed, and the volume, quality, and resonance of the voice are evaluated. These analysis results are used to understand the current state of the user's singing technique and to point out specific areas for improvement. For example, it identifies parts where the pitch is off, parts where the rhythm is off, and parts where there are problems with the vocal technique, and provides the user with specific advice. Furthermore, the analysis unit evaluates the user's progress by comparing it with past data, supporting long-term growth. This allows the user to clearly understand the weaknesses in their singing technique and practice effectively.

[0071] The feedback unit provides voice feedback based on the results analyzed by the analysis unit. Specifically, it analyzes the voice data sung by the user and points out pitch inaccuracies, rhythmic errors, and problems with vocal technique. For example, it displays pitch inaccuracies in a graph, visually indicating which notes are too high or too low. For rhythmic errors, it compares the correct timing on the musical score with the actual timing, clearly showing where the rhythm is off. For problems with vocal technique, it analyzes the waveform and spectrum of the voice, evaluates the volume, tone quality, and resonance, and suggests specific methods for improvement. The feedback unit provides this information to the user as voice feedback, encouraging real-time correction. For example, if the user sings off-key, it provides voice messages such as "The pitch is too high" or "The rhythm is too fast" on the spot. This allows the user to immediately correct their singing technique and practice effectively.

[0072] The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. Specifically, it analyzes the audio data sung by the user and visually displays pitch deviations, rhythm errors, and problems with vocal technique. For example, it shows pitch deviations as a graph, visually indicating which notes are too high or too low. For rhythm errors, it compares the correct timing on the musical score with the actual timing, clearly indicating where the rhythm is off. For problems with vocal technique, it analyzes the waveform and spectrum of the voice, evaluates the volume, tone quality, and resonance, and suggests specific methods for improvement. The visual feedback unit visually displays this information, showing the user specific areas for improvement. For example, it explains the correct vocal technique with a video, specifically showing how the user should sing. It also shows pitch deviations as a graph, visually indicating which notes are too high or too low. This allows the user to clearly understand the weaknesses in their singing technique and practice effectively.

[0073] The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. Specifically, it analyzes the audio data sung by the user and applies algorithms to correct pitch deviations and rhythmic errors. For example, it uses a pitch correction algorithm to generate a version with correct pitch. This algorithm decomposes the audio signal into frequency components, identifies the pitch of each note, and corrects it to the correct pitch. It also corrects rhythmic errors by adjusting the timing to achieve the correct rhythm. The generation unit performs these corrections in real time and provides immediate feedback to the user. For example, if the user sings off-key, the unit corrects the pitch on the spot and generates a version with the correct pitch. This allows the user to immediately correct weaknesses in their singing technique and practice effectively.

[0074] The evaluation unit provides regular evaluations and reports based on the pitch-correct version generated by the generation unit. Specifically, it analyzes the audio data sung by the user and evaluates pitch deviations, rhythm errors, and problems with vocal technique. Based on this information, the evaluation unit grasps the current state of the user's singing technique and points out specific areas for improvement. For example, it conducts weekly and monthly evaluations to visualize the user's progress. The evaluation unit provides this information as a report, showing the user specific areas for improvement. For example, it shows pitch deviations in a graph, visually indicating which notes are too high or too low. Regarding rhythm errors, it compares the correct timing on the musical score with the actual timing to clearly show where the rhythm is off. Regarding problems with vocal technique, it analyzes the waveform and spectrum of the voice to evaluate the volume, tone quality, and resonance, and proposes specific methods for improvement. This allows the user to clearly understand the weaknesses in their singing technique and practice effectively.

[0075] The analysis unit can point out specific areas for improvement regarding pitch, rhythm, and vocal technique. For example, the analysis unit can point out pitch inaccuracies, rhythmic irregularities, and errors in vocal technique. This allows users to practice more effectively by providing specific areas for improvement. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's singing voice data into an AI, which can then point out areas for improvement in pitch, rhythm, and vocal technique.

[0076] The feedback unit can display pitch deviations graphically. For example, the feedback unit can display pitch deviations in cents. The feedback unit can also display pitch deviations with color changes. Furthermore, the feedback unit can display pitch deviations with animations. This makes it easier for users to understand areas for improvement by visually representing pitch deviations. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can generate graphs based on pitch deviation data generated by AI.

[0077] The visual feedback unit can explain correct vocalization techniques through videos. For example, the visual feedback unit can explain diaphragmatic breathing techniques through videos. It can also explain how to use the vocal cords through videos. Furthermore, the visual feedback unit can explain proper vocalization posture through videos. This allows users to visually learn correct vocalization techniques through video explanations. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can provide videos of vocalization techniques generated by AI.

[0078] The generation unit can automatically generate a version with correct pitch. For example, the generation unit can generate a version with correct pitch using a pitch correction algorithm. Alternatively, the generation unit can generate a version with correct pitch using a reference pitch. Furthermore, the generation unit can generate a version with correct pitch using pitch correction software. This makes it easier for users to understand specific areas for improvement by automatically generating a version with correct pitch. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can generate a version with correct pitch based on pitch correction data generated by AI.

[0079] The evaluation department can provide periodic evaluations and reports. For example, the evaluation department can conduct weekly evaluations to visualize user growth. It can also conduct monthly evaluations to report on user progress. Furthermore, the evaluation department can conduct evaluations based on evaluation items and provide reports. This makes it easier for users to feel their growth by providing periodic evaluations and reports. Some or all of the above processes in the evaluation department may be performed using AI, for example, or not. For example, the evaluation department can create reports based on evaluation data generated by AI.

[0080] The analysis unit can estimate the user's emotions and adjust the accuracy of the analysis based on the estimated emotions. For example, if the user is tense, the AI ​​in the analysis unit can loosen the accuracy of the analysis to help the user relax. Conversely, if the user is relaxed, the AI ​​in the analysis unit can increase the accuracy of the analysis and provide more detailed feedback. Furthermore, if the user is excited, the AI ​​in the analysis unit can adjust the accuracy of the analysis and provide appropriate feedback. This allows for more appropriate feedback to be provided by adjusting the accuracy of the analysis according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI or not using AI. For example, the analysis unit can input the user's voice data into a generative AI, which can estimate the user's emotions and adjust the accuracy of the analysis.

[0081] The analysis unit can identify individual areas for improvement by referring to the user's past singing data. For example, the analysis unit can identify specific pitch deviations based on the user's past singing data and point out areas for improvement. The analysis unit can also analyze the user's past rhythm patterns and identify areas for improvement in rhythm. Furthermore, the analysis unit can refer to the user's past vocalization techniques and identify areas for improvement in vocalization techniques. This makes it easier to identify individual areas for improvement by referring to past data. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's past singing data into AI, which can then identify individual areas for improvement.

[0082] The analysis unit can propose the optimal vocalization method by considering the user's voice quality and vocal cord condition. For example, the analysis unit can analyze the user's voice quality and propose the optimal vocalization method. It can also consider the user's vocal cord condition and propose a vocalization method that does not strain the vocal cords. Furthermore, the analysis unit can comprehensively analyze the user's voice quality and vocal cord condition to propose the optimal vocalization method. This allows for effective practice by proposing a vocalization method tailored to the user's voice quality and vocal cord condition. Some or all of the above processing in the analysis unit may be performed using AI, or without AI. For example, the analysis unit can input the user's voice quality data into AI, which can then propose the optimal vocalization method.

[0083] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. For example, if the user is nervous, the analysis unit can provide a simple and highly visible display method. If the user is relaxed, the analysis unit can also provide a display method that includes detailed information. If the user is in a hurry, the analysis unit can also provide a display method that gets straight to the point. By providing a display method that matches the user's emotions, it becomes easier for the user to understand. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or not using AI. For example, the analysis unit can input the user's voice data into the generative AI, which can estimate the user's emotions and adjust the display method.

[0084] The analysis unit can correct the analysis results by taking into account the user's singing environment. For example, the analysis unit can analyze the acoustic characteristics of the room and correct the analysis results based on those characteristics. The analysis unit can also take into account the user's singing environment and provide analysis results that are appropriate for that environment. Furthermore, the analysis unit can comprehensively analyze the acoustic characteristics of the room and the user's singing environment to provide optimal analysis results. This allows for more accurate analysis results by taking the singing environment into consideration. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's singing environment data into the AI, and the AI ​​can correct the analysis results.

[0085] The analysis unit can apply different analysis algorithms depending on the user's singing style. For example, if the user's singing style is pop, the analysis unit will apply an analysis algorithm suitable for pop music. Similarly, if the user's singing style is classical, the analysis unit can apply an analysis algorithm suitable for classical music. Furthermore, the analysis unit can select and apply the optimal analysis algorithm based on the user's singing style. This allows for more appropriate feedback by applying an analysis algorithm tailored to the singing style. Some or all of the above processing in the analysis unit may be performed using AI, or without AI. For example, the analysis unit can input the user's singing style data into an AI, which can then apply the optimal analysis algorithm.

[0086] The feedback unit can estimate the user's emotions and adjust the content of the feedback based on the estimated emotions. For example, if the user is nervous, the feedback unit will provide feedback in a gentle tone. If the user is relaxed, the feedback unit can also provide detailed feedback. If the user is in a hurry, the feedback unit can provide concise and to-the-point feedback. This allows for more effective instruction by providing feedback that is tailored to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the feedback unit may be performed using AI, or not using AI. For example, the feedback unit can input the user's voice data into a generative AI, which can estimate the user's emotions and adjust the content of the feedback.

[0087] The feedback unit can provide personalized advice by referring to the user's past feedback history. For example, the feedback unit can point out specific areas for improvement based on the user's past feedback history. The feedback unit can also refer to the user's past feedback history and provide ongoing areas for improvement. Furthermore, the feedback unit can analyze the user's past feedback history and provide optimal advice. This makes it easier to provide personalized advice by referring to past feedback history. Some or all of the above processing in the feedback unit may be performed using AI, for example, or not using AI. For example, the feedback unit can input the user's past feedback history data into AI, and the AI ​​can provide personalized advice.

[0088] The feedback unit can suggest specific areas for improvement based on the user's singing goals. For example, if the user's singing goal is to sing a particular song perfectly, the feedback unit will suggest areas for improvement specific to that song. The feedback unit can also provide specific areas for improvement based on the user's singing goals. Furthermore, the feedback unit can consider the user's singing goals and suggest the most suitable areas for improvement. This allows the user to practice effectively towards their goals by suggesting areas for improvement based on their singing goals. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's singing goal data into an AI, which can then suggest specific areas for improvement.

[0089] The feedback unit can estimate the user's emotions and adjust the timing of feedback based on the estimated emotions. For example, if the user is nervous, the feedback unit may delay the timing of feedback. Conversely, if the user is relaxed, the feedback unit may speed up the timing of feedback. Furthermore, if the user is in a hurry, the feedback unit may provide feedback quickly. This allows for more effective instruction by providing feedback at a timing appropriate to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the feedback unit may be performed using AI, or not using AI. For example, the feedback unit can input the user's voice data into a generative AI, which can estimate the user's emotions and adjust the timing of feedback.

[0090] The feedback unit can adjust the feedback content by taking into account the user's singing environment. For example, the feedback unit can adjust the feedback content by considering the type of microphone the user is using. The feedback unit can also provide optimal feedback by considering the user's singing environment. Furthermore, the feedback unit can comprehensively analyze the microphone type and singing environment and adjust the feedback content. This allows for more accurate feedback by considering the singing environment. Some or all of the above processing in the feedback unit may be performed using AI, for example, or without AI. For example, the feedback unit can input the user's singing environment data into the AI, which can then adjust the feedback content.

[0091] The feedback unit can provide different feedback formats depending on the user's singing style. For example, if the user's singing style is pop, the feedback unit provides audio feedback. Alternatively, if the user's singing style is classical, the feedback unit can provide text feedback. Furthermore, the feedback unit can select and provide the most appropriate feedback format based on the user's singing style. This allows for more appropriate instruction by providing feedback formats tailored to the singing style. Some or all of the above processing in the feedback unit may be performed using AI, or without AI. For example, the feedback unit can input the user's singing style data into an AI, which can then provide the most appropriate feedback format.

[0092] The visual feedback unit can estimate the user's emotions and adjust the display method of the visual feedback based on the estimated user emotions. For example, if the user is nervous, the visual feedback unit can provide a simple and highly visible display method. If the user is relaxed, the visual feedback unit can also provide a display method that includes detailed information. If the user is in a hurry, the visual feedback unit can also provide a display method that gets straight to the point. By providing a display method that matches the user's emotions, it becomes easier for the user to understand. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's voice data into the generative AI, which can estimate the user's emotions and adjust the display method.

[0093] The visual feedback unit can provide personalized advice by referring to the user's past visual feedback history. For example, the visual feedback unit can point out specific areas for improvement based on the user's past visual feedback history. The visual feedback unit can also refer to the user's past visual feedback history and provide ongoing areas for improvement. Furthermore, the visual feedback unit can analyze the user's past visual feedback history and provide optimal advice. This makes it easier to provide personalized advice by referring to past visual feedback history. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's past visual feedback history data into AI, and the AI ​​can provide personalized advice.

[0094] The visual feedback unit can indicate specific areas for improvement based on the user's singing goals. For example, if the user's singing goal is to sing a particular song perfectly, the visual feedback unit will indicate areas for improvement specific to that song. The visual feedback unit can also provide specific areas for improvement based on the user's singing goals. Furthermore, the visual feedback unit can consider the user's singing goals and indicate the most suitable areas for improvement. This allows the user to practice effectively towards their goals by indicating areas for improvement based on their singing goals. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's singing goal data into AI, which can then indicate specific areas for improvement.

[0095] The visual feedback unit can estimate the user's emotions and adjust the timing of visual feedback based on the estimated emotions. For example, if the user is nervous, the visual feedback unit can delay the timing of visual feedback. Conversely, if the user is relaxed, the visual feedback unit can also advance the timing of visual feedback. Furthermore, if the user is in a hurry, the visual feedback unit can provide visual feedback quickly. This enables more effective instruction by providing visual feedback at a timing appropriate to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the visual feedback unit may be performed using AI, or not using AI. For example, the visual feedback unit can input user voice data into a generative AI, which can estimate the user's emotions and adjust the timing of visual feedback.

[0096] The visual feedback unit can correct the visual feedback content by taking into account the user's singing environment. For example, the visual feedback unit can provide optimal visual feedback by considering the user's singing environment. The visual feedback unit can also analyze the acoustic characteristics of the room and correct the visual feedback content based on those characteristics. Furthermore, the visual feedback unit can comprehensively analyze the user's singing environment and the acoustic characteristics of the room to provide optimal visual feedback. This allows for more accurate visual feedback by taking the singing environment into consideration. Some or all of the above processing in the visual feedback unit may be performed using AI, for example, or without AI. For example, the visual feedback unit can input the user's singing environment data into the AI, which can then correct the visual feedback content.

[0097] The visual feedback unit can provide different visual feedback formats depending on the user's singing style. For example, if the user's singing style is pop, the visual feedback unit can provide graph-based visual feedback. Alternatively, if the user's singing style is classical, the visual feedback unit can provide video-based visual feedback. Furthermore, the visual feedback unit can select and provide the most appropriate visual feedback format based on the user's singing style. This allows for more appropriate instruction by providing visual feedback formats tailored to the singing style. Some or all of the above processing in the visual feedback unit may be performed using AI, or without AI. For example, the visual feedback unit can input the user's singing style data into an AI, which can then provide the most appropriate visual feedback format.

[0098] The generation unit can estimate the user's emotions and adjust the accuracy of the correct version of the pitch it generates based on the estimated user emotions. For example, if the user is tense, the generation unit may loosen the accuracy of the correct version of the pitch it generates. Conversely, if the user is relaxed, the generation unit may increase the accuracy of the correct version of the pitch it generates. Furthermore, if the user is excited, the generation unit may adjust the accuracy of the correct version of the pitch it generates. This allows for more appropriate feedback by generating the correct version of the pitch with an accuracy that matches the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's voice data into the generation AI, which can estimate the user's emotions and adjust the accuracy of the correct version of the pitch it generates.

[0099] The generation unit can refer to the user's past singing data and reflect individual areas for improvement. For example, the generation unit can reflect specific pitch deviations based on the user's past singing data. The generation unit can also analyze the user's past rhythm patterns and reflect areas for improvement in rhythm. Furthermore, the generation unit can refer to the user's past vocalization techniques and reflect areas for improvement in vocalization techniques. This makes it easier to reflect individual areas for improvement by referring to past data. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's past singing data into AI, and the AI ​​can reflect individual areas for improvement.

[0100] The generation unit can generate the optimal pitch by considering the user's voice quality and vocal cord condition. For example, the generation unit analyzes the user's voice quality and generates the optimal pitch. The generation unit can also consider the user's vocal cord condition and generate a pitch that does not strain the vocal cords. Furthermore, the generation unit can comprehensively analyze the user's voice quality and vocal cord condition to generate the optimal pitch. This enables effective practice by generating a pitch that matches the user's voice quality and vocal cord condition. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's voice quality data into AI, and the AI ​​can generate the optimal pitch.

[0101] The generation unit can estimate the user's emotions and adjust the display method of the correct version of the generated pitch based on the estimated user emotions. For example, if the user is nervous, the generation unit can provide a simple and highly visible display method. If the user is relaxed, the generation unit can also provide a display method that includes detailed information. If the user is in a hurry, the generation unit can also provide a display method that gets straight to the point. By providing a display method that matches the user's emotions, it becomes easier for the user to understand. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input the user's voice data into the generation AI, which can estimate the user's emotions and adjust the display method.

[0102] The generation unit can correct the generation results by taking into account the user's singing environment. For example, the generation unit can consider the user's singing environment and provide the optimal generation result. The generation unit can also analyze the acoustic characteristics of the room and correct the generation result based on those characteristics. Furthermore, the generation unit can comprehensively analyze the user's singing environment and the acoustic characteristics of the room to provide the optimal generation result. This allows for the provision of more accurate generation results by considering the singing environment. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the user's singing environment data into the AI, and the AI ​​can correct the generation result.

[0103] The generation unit can apply different generation algorithms depending on the user's singing style. For example, if the user's singing style is pop, the generation unit will apply a generation algorithm suitable for pop music. Similarly, if the user's singing style is classical, the generation unit can apply a generation algorithm suitable for classical music. Furthermore, the generation unit can select and apply the optimal generation algorithm according to the user's singing style. This allows for more appropriate feedback by applying a generation algorithm tailored to the singing style. Some or all of the above-described processes in the generation unit may be performed using AI, or without AI. For example, the generation unit can input the user's singing style data into an AI, which can then apply the optimal generation algorithm.

[0104] The evaluation unit can estimate the user's emotions and adjust the evaluation criteria based on the estimated emotions. For example, if the user is tense, the evaluation unit may loosen the evaluation criteria. Conversely, if the user is relaxed, the evaluation unit may tighten the evaluation criteria. Furthermore, if the user is excited, the evaluation unit may adjust the evaluation criteria. This allows for more appropriate feedback to be provided by evaluating based on criteria that match the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the evaluation unit may be performed using AI, or not using AI. For example, the evaluation unit can input the user's voice data into a generative AI, which can estimate the user's emotions and adjust the evaluation criteria.

[0105] The evaluation unit can provide individual evaluations by referring to the user's past evaluation data. For example, the evaluation unit can provide specific evaluation points based on the user's past evaluation data. The evaluation unit can also refer to the user's past evaluation data and provide continuous evaluation points. Furthermore, the evaluation unit can analyze the user's past evaluation data and provide optimal evaluation points. This makes it easier to provide individual evaluations by referring to past evaluation data. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's past evaluation data into AI, and the AI ​​can provide individual evaluations.

[0106] The evaluation unit can provide specific evaluation points based on the user's singing goals. For example, if the user's singing goal is to sing a particular song perfectly, the evaluation unit will provide evaluation points specific to that song. The evaluation unit can also provide specific evaluation points based on the user's singing goals. Furthermore, the evaluation unit can consider the user's singing goals and provide optimal evaluation points. This allows the user to practice effectively towards their goals by providing evaluation points based on their singing goals. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's singing goal data into AI, which can then provide specific evaluation points.

[0107] The evaluation unit can estimate the user's emotions and adjust the timing of the evaluation based on the estimated emotions. For example, if the user is nervous, the evaluation unit may delay the evaluation. Conversely, if the user is relaxed, the evaluation unit may speed up the evaluation. Furthermore, if the user is in a hurry, the evaluation unit may provide the evaluation quickly. This allows for more effective instruction by providing evaluations at a timing appropriate to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the evaluation unit may be performed using AI, or not using AI. For example, the evaluation unit can input the user's voice data into a generative AI, which can estimate the user's emotions and adjust the timing of the evaluation.

[0108] The evaluation unit can adjust the evaluation results by taking into account the user's singing environment. For example, the evaluation unit can provide an optimal evaluation by considering the user's singing environment. The evaluation unit can also analyze the acoustic characteristics of the room and adjust the evaluation results based on those characteristics. Furthermore, the evaluation unit can comprehensively analyze the user's singing environment and the acoustic characteristics of the room to provide an optimal evaluation. This allows for a more accurate evaluation by considering the singing environment. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can input the user's singing environment data into the AI, which can then adjust the evaluation results.

[0109] The evaluation unit can provide different evaluation formats depending on the user's singing style. For example, if the user's singing style is pop, the evaluation unit can provide a score-based evaluation. Alternatively, if the user's singing style is classical, the evaluation unit can provide a comment-based evaluation. Furthermore, the evaluation unit can select and provide the most appropriate evaluation format based on the user's singing style. This allows for more appropriate guidance by providing evaluation formats tailored to the singing style. Some or all of the above processing in the evaluation unit may be performed using AI, or without AI. For example, the evaluation unit can input the user's singing style data into an AI, which can then provide the most appropriate evaluation format.

[0110] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0111] An AI voice training system can have the functionality to save users' singing data to the cloud and share it with other users. For example, a user can upload their singing data to the cloud, and other users can download that data for reference. Users can also provide comments and feedback on other users' singing data. Furthermore, a ranking function can be provided on the cloud where users can compete with each other. This allows users to interact with other users and learn from each other.

[0112] The AI ​​voice training system can automatically generate personalized training plans based on the user's singing data. For example, it can identify the user's weaknesses in pitch and rhythm and provide a training menu accordingly. It can also adjust the training plan according to the user's progress. Furthermore, it can provide a training plan tailored to the user's goals. This allows users to receive the most suitable training for themselves.

[0113] The AI ​​voice training system can estimate the user's emotions and adjust the training content based on those emotions. For example, if the user is tired, it can provide lighter training. If the user is highly motivated, it can provide more challenging training. Furthermore, if the user is stressed, it can provide relaxing training. By providing training tailored to the user's emotions, it enables more effective practice.

[0114] An AI voice training system can feature a virtual coach that provides real-time instruction based on the user's singing data. For example, the virtual coach can evaluate the user's singing in real time and point out specific areas for improvement. The virtual coach can also offer words of encouragement. Furthermore, the virtual coach can adjust the training menu according to the user's progress. This allows the user to receive real-time guidance.

[0115] An AI voice training system can feature a function that allows users to perform with a virtual band based on their singing data. For example, a user can play their own singing data along with a virtual band and perform together. The virtual band can also adjust its performance to match the user's singing. Furthermore, the virtual band can provide feedback on the user's singing. This allows users to practice while experiencing the feeling of performing with a band.

[0116] The AI ​​voice training system can estimate the user's emotions and adjust the training difficulty based on those emotions. For example, if the user is nervous, the training difficulty can be lowered. Conversely, if the user is relaxed, the difficulty can be increased. Furthermore, if the user is excited, the difficulty can also be adjusted. This allows for more effective practice by providing training at a difficulty level that matches the user's emotions.

[0117] An AI voice training system can feature a virtual audience that reacts in real time based on the user's singing data. For example, the virtual audience can applaud or cheer in response to the user's singing. The virtual audience can also provide feedback on the user's singing. Furthermore, the virtual audience can change its reactions according to the user's progress. This allows the user to practice while feeling the audience's reactions.

[0118] The AI ​​voice training system can estimate the user's emotions and adjust the training feedback based on those emotions. For example, if the user is feeling down, it can provide encouraging feedback. If the user is confident, it can provide detailed feedback. Furthermore, if the user is anxious, it can provide concise feedback. This allows for more effective instruction by providing feedback tailored to the user's emotions.

[0119] An AI voice training system can have the functionality to simulate performance on a virtual stage based on the user's singing data. For example, the user can sing on a virtual stage and have their performance evaluated. It can also simulate movements and facial expressions on the virtual stage. Furthermore, it can provide feedback on the performance on the virtual stage. This allows users to practice while simulating their performance on a real stage.

[0120] The AI ​​voice training system can estimate the user's emotions and adjust the training pace based on those emotions. For example, if the user is anxious, the training pace can be slowed down. Conversely, if the user is relaxed, the training pace can be increased. Furthermore, if the user is focused, the training pace can be adjusted accordingly. This allows for more effective practice by providing training at a pace that matches the user's emotions.

[0121] The following briefly describes the processing flow for example form 2.

[0122] Step 1: The analysis unit analyzes the user's singing voice in real time. The analysis unit points out specific areas for improvement, such as pitch, rhythm, and vocal technique. Step 2: The feedback unit provides audio feedback based on the results analyzed by the analysis unit. For example, the feedback unit displays pitch deviations graphically. Step 3: The visual feedback unit provides visual feedback based on the audio feedback provided by the feedback unit. For example, the visual feedback unit might explain the correct pronunciation technique with a video. Step 4: The generation unit automatically generates a version with correct pitch based on the results analyzed by the analysis unit. The generation unit generates a version with correct pitch using, for example, a pitch correction algorithm. Step 5: The evaluation unit provides periodic evaluations and reports based on the correct version of the pitch generated by the generation unit. The evaluation unit performs weekly and monthly evaluations, for example, to visualize user growth.

[0123] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0124] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0125] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0126] Each of the multiple elements described above, including the analysis unit, feedback unit, visual feedback unit, generation unit, and evaluation unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the analysis unit acquires the user's singing voice using the microphone 38B of the smart device 14 and analyzes it in real time by the specific processing unit 290 of the data processing unit 12. The feedback unit generates audio feedback using the specific processing unit 290 of the data processing unit 12 and provides it through the speaker 40B of the smart device 14. The visual feedback unit provides visual feedback using the display 40A of the smart device 14. The generation unit automatically generates a version with correct pitch using the specific processing unit 290 of the data processing unit 12. The evaluation unit generates periodic evaluations and reports using the specific processing unit 290 of the data processing unit 12 and provides them to the user through the display 40A of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0127] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0128] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0129] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0130] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0131] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0132] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0133] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0134] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0135] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0136] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0137] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0138] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0139] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0140] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0141] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0142] Each of the multiple elements described above, including the analysis unit, feedback unit, visual feedback unit, generation unit, and evaluation unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the analysis unit acquires the user's singing voice using the microphone 238 of the smart glasses 214 and analyzes it in real time by the specific processing unit 290 of the data processing unit 12. The feedback unit generates audio feedback, for example, by the specific processing unit 290 of the data processing unit 12 and provides it through the speaker 240 of the smart glasses 214. The visual feedback unit provides visual feedback, for example, using the display of the smart glasses 214. The generation unit automatically generates a version with correct pitch, for example, by the specific processing unit 290 of the data processing unit 12. The evaluation unit generates periodic evaluations and reports, for example, by the specific processing unit 290 of the data processing unit 12 and provides them to the user through the display of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.

[0143] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0144] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0145] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0146] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0147] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0148] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0149] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0150] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0151] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0152] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0153] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0154] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0155] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0156] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0157] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0158] Each of the multiple elements described above, including the analysis unit, feedback unit, visual feedback unit, generation unit, and evaluation unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the analysis unit acquires the user's singing voice using the microphone 238 of the headset terminal 314 and analyzes it in real time by the specific processing unit 290 of the data processing unit 12. The feedback unit generates audio feedback using the specific processing unit 290 of the data processing unit 12 and provides it through the speaker 240 of the headset terminal 314. The visual feedback unit provides visual feedback using the display 343 of the headset terminal 314. The generation unit automatically generates a version with correct pitch using the specific processing unit 290 of the data processing unit 12. The evaluation unit generates periodic evaluations and reports using the specific processing unit 290 of the data processing unit 12 and provides them to the user through the display 343 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0159] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0160] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0161] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0162] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0163] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0164] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0165] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0166] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0167] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0168] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0169] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0170] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0171] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0172] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0173] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0174] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0175] Each of the multiple elements described above, including the analysis unit, feedback unit, visual feedback unit, generation unit, and evaluation unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the analysis unit acquires the user's singing voice using the microphone 238 of the robot 414 and analyzes it in real time by the specific processing unit 290 of the data processing unit 12. The feedback unit generates audio feedback by, for example, the specific processing unit 290 of the data processing unit 12 and provides it through the speaker 240 of the robot 414. The visual feedback unit provides visual feedback using, for example, the display of the robot 414. The generation unit automatically generates a version with correct pitch by, for example, the specific processing unit 290 of the data processing unit 12. The evaluation unit generates periodic evaluations and reports by, for example, the specific processing unit 290 of the data processing unit 12 and provides them to the user through the display of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.

[0176] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0177] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0178] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0179] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0180] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0181] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0182] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0183] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0184] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0185] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0186] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0187] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0188] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0189] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0190] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0191] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0192] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0193] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0194] (Note 1) The analysis department analyzes the singing voice in real time, A feedback unit provides voice feedback based on the results of the analysis performed by the aforementioned analysis unit, A visual feedback unit provides visual feedback based on the audio feedback provided by the aforementioned feedback unit, A generation unit that automatically generates a version with correct pitch based on the results of analysis by the aforementioned analysis unit, The system includes an evaluation unit that provides periodic evaluations and reports based on the correct version of the pitch generated by the generation unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit is I will point out specific areas for improvement regarding pitch, rhythm, and vocal technique. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned feedback unit is The pitch deviation is shown in a graph. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned visual feedback unit is This video explains the correct way to sing. The system described in Appendix 1, characterized by the features described herein. (Note 5) The generating unit is Automatically generates a version with correct pitch. The system described in Appendix 1, characterized by the features described herein. (Note 6) The evaluation unit, We provide regular evaluations and reports. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned analysis unit is It estimates the user's emotions and adjusts the accuracy of the analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned analysis unit is Referencing the user's past singing data identifies individual areas for improvement. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned analysis unit is We propose the optimal vocalization method, taking into account the user's voice quality and vocal cord condition. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit is It estimates the user's emotions and adjusts how the analysis results are displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit is The analysis results are corrected to take into account the user's singing environment. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit is Different analysis algorithms are applied depending on the user's singing style. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned feedback unit is It estimates the user's emotions and adjusts the content of the feedback based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned feedback unit is We provide personalized advice by referring to the user's past feedback history. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned feedback unit is Based on the user's singing goals, specific areas for improvement will be suggested. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned feedback unit is It estimates the user's emotions and adjusts the timing of feedback based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned feedback unit is The feedback content will be adjusted to take into account the user's singing environment. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned feedback unit is Provides different feedback formats depending on the user's singing style. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned visual feedback unit is It estimates the user's emotions and adjusts how visual feedback is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned visual feedback unit is Referencing the user's past visual feedback history provides personalized advice. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned visual feedback unit is Based on the user's singing goals, specific areas for improvement are indicated. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned visual feedback unit is It estimates the user's emotions and adjusts the timing of visual feedback based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned visual feedback unit is The visual feedback content is adjusted to take into account the user's singing environment. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned visual feedback unit is Provides different visual feedback formats depending on the user's singing style. The system described in Appendix 1, characterized by the features described herein. (Note 25) The generating unit is It estimates the user's emotions and adjusts the accuracy of the correct version of the pitch generated based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The generating unit is Referencing the user's past singing data will allow for individual improvements to be reflected. The system described in Appendix 1, characterized by the features described herein. (Note 27) The generating unit is It generates the optimal pitch by taking into account the user's voice quality and vocal cord condition. The system described in Appendix 1, characterized by the features described herein. (Note 28) The generating unit is It estimates the user's emotions and adjusts how the correct version of the pitch generated based on those emotions is displayed. The system described in Appendix 1, characterized by the features described herein. (Note 29) The generating unit is The generated results are corrected to take into account the user's singing environment. The system described in Appendix 1, characterized by the features described herein. (Note 30) The generating unit is Apply different generation algorithms depending on the user's singing style. The system described in Appendix 1, characterized by the features described herein. (Note 31) The evaluation unit, It estimates the user's emotions and adjusts the evaluation criteria based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The evaluation unit, Provides individualized ratings by referencing the user's past rating data. The system described in Appendix 1, characterized by the features described herein. (Note 33) The evaluation unit, Based on the user's singing goals, specific evaluation points are displayed. The system described in Appendix 1, characterized by the features described herein. (Note 34) The evaluation unit, It estimates the user's emotions and adjusts the timing of evaluations based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The evaluation unit, The evaluation results will be adjusted to take into account the user's singing environment. The system described in Appendix 1, characterized by the features described herein. (Note 36) The evaluation unit, It provides different evaluation formats depending on the user's singing style. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0195] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. The analysis department analyzes the singing voice in real time, A feedback unit provides voice feedback based on the results of the analysis performed by the aforementioned analysis unit, A visual feedback unit provides visual feedback based on the audio feedback provided by the aforementioned feedback unit, A generation unit that automatically generates a version with correct pitch based on the results of analysis by the aforementioned analysis unit, The system includes an evaluation unit that provides periodic evaluations and reports based on the correct version of the pitch generated by the generation unit. A system characterized by the following features.

2. The aforementioned analysis unit is I will point out specific areas for improvement regarding pitch, rhythm, and vocal technique. The system according to feature 1.

3. The aforementioned feedback unit is The pitch deviation is shown in a graph. The system according to feature 1.

4. The aforementioned visual feedback unit is This video explains the correct way to sing. The system according to feature 1.

5. The generating unit is Automatically generates a version with correct pitch. The system according to feature 1.

6. The evaluation unit, We provide regular evaluations and reports. The system according to feature 1.

7. The aforementioned analysis unit is It estimates the user's emotions and adjusts the accuracy of the analysis based on the estimated user emotions. The system according to feature 1.

8. The aforementioned analysis unit is Referencing the user's past singing data identifies individual areas for improvement. The system according to feature 1.

9. The aforementioned analysis unit is We propose the optimal vocalization method, taking into account the user's voice quality and vocal cord condition. The system according to feature 1.

10. The aforementioned analysis unit is It estimates the user's emotions and adjusts how the analysis results are displayed based on those estimated emotions. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A