System
A system that analyzes user input audio and video data to provide specific advice on pitch and emotional expression helps individuals improve their musical skills independently.
Patent Information
- Application Number
- JP2024121554
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Individuals practicing musical instruments at home lack clear guidelines for improvement, making it difficult to enhance their skills effectively without professional instruction.
A system that allows users to input audio and video data, which is analyzed by a server to identify areas for improvement, and generates specific advice for correcting pitch deviations and enhancing emotional expression and expressiveness.
Enables users to practice music efficiently by providing objective feedback on pitch, rhythm, and emotional expression, facilitating effective skill development.
Smart Images

Figure 2026019806000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, with the increasing popularity of personal music distribution and karaoke, an increasing number of people are starting to practice musical instruments at home. However, most people do not know how to practice to improve, and so they have to rely on professional voice training or instrument classes. Furthermore, when practicing independently, there is a problem in that it is difficult to obtain clear guidelines on what needs to be improved, limiting the effectiveness of practice. This invention aims to solve these problems and provide specific support means for individuals to practice music efficiently. [Means for solving the problem]
[0005] The present invention proposes to solve the above problems by the following means.
[0006] By providing a system including a means for a user to input audio or video, a server that receives and analyzes the input audio or video data, a server that identifies areas for improvement based on the analysis results and generates specific advice, and a means for providing the generated advice to the user, it becomes possible for individuals to effectively practice music.
[0007] Specifically, by proposing a system that includes a means for analyzing the pitch of input audio data, a means for detecting deviations in the analyzed pitch, and a means for generating specific advice for correcting the deviations in pitch, it is possible to practice pitch with high accuracy.
[0008] Furthermore, we propose a system that includes a means for analyzing facial expressions and gestures in input video data, a means for evaluating the appropriateness of emotional expression for the analyzed facial expressions and gestures, and a means for generating specific advice for improving emotional expression, thereby supporting the improvement of expressiveness.
[0009] "User" refers to an individual or group who uses this system to input audio or video and receive analysis results and advice.
[0010] "Audio" refers to a series of sound data generated by a user singing or playing an instrument.
[0011] "Video" refers to video data that captures the user's facial expressions and movements while playing or singing.
[0012] "Input means" refers to a device or interface that allows a user to provide audio or video to a system, such as a microphone or camera.
[0013] "Server" means a computer system that provides the computational resources to analyze input audio or video data, identify areas for improvement, and generate recommendations.
[0014] "Analysis" refers to the process of analyzing input audio or video data and evaluating technical characteristics such as pitch, rhythm, and expressiveness.
[0015] "Points for improvement" refers to specific areas or items that should be improved based on the analysis results in order to improve the user's performance.
[0016] "Advice" refers to information that describes specific methods and practice for how to improve identified areas for improvement.
[0017] "Providing means" refers to an interface or device for conveying the generated advice to the user, such as a smartphone display or audio output.
[0018] "Pitch" refers to a musical element that evaluates the frequency or pitch of consecutive sounds within a certain period of time.
[0019] "Deviation" refers to the results of an analysis that identifies deviations from ideal or standard pitch or rhythm.
[0020] "Facial expression" refers to the emotions and moods conveyed by a user's facial movements and expressions.
[0021] "Gestures" refer to hand and body movements that users make while playing music.
[0022] "Evaluation" refers to the process of determining technical or presentational adequacy based on analyzed data. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] The embodiments of the present invention will be specifically described below.
[0045] System Overview
[0046] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[0047] Audio and video capture
[0048] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0049] Data transmission and analysis
[0050] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0051] Identifying areas for improvement and generating recommendations
[0052] The server identifies areas for improvement based on the analysis results. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how.
[0053] Providing advice
[0054] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a little higher," or video advice such as, "Use a metronome to practice maintaining a steady tempo."
[0055] Specific examples
[0056] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[0057] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[0058] The processing flow will be explained below.
[0059] Step 1:
[0060] The user launches a dedicated application and starts recording or recording.
[0061] Pressing the "Start Recording" button on the application screen will begin audio and video input.
[0062] Step 2:
[0063] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[0064] The recorded data is temporarily stored in a memory area on the terminal.
[0065] Step 3:
[0066] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[0067] After completing the data transmission, the terminal waits for a response from the server.
[0068] Step 4:
[0069] The server receives the received audio and video data and passes it to the analysis module.
[0070] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[0071] Step 5:
[0072] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[0073] Specific algorithms are used to analyze emotional expressions and performance movements.
[0074] Step 6:
[0075] The server lists specific areas for improvement based on the analysis of pitch, rhythm, expressiveness, etc.
[0076] This improvement highlights areas of user performance that require special attention.
[0077] Step 7:
[0078] The server generates specific advice for the identified improvements.
[0079] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0080] Step 8:
[0081] The generated advice is encoded in text or video format and sent to the terminal.
[0082] The server performs the encoding process and sends the data to the device in the optimal format.
[0083] Step 9:
[0084] The advice received by the terminal is displayed on the user interface.
[0085] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[0086] Step 10:
[0087] The user checks the displayed advice and practices independently based on the specific practice method.
[0088] By following the advice and practicing repeatedly, you can improve your skills.
[0089] Example 1
[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0091] In the past, when practicing music or improving performance, it was difficult for users to objectively evaluate their own performance and identify specific areas for improvement. Furthermore, due to limited opportunities to receive professional instruction, it was difficult to find an efficient practice method.
[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0093] In this invention, the server includes: means for a user to input audio or video; means for temporarily storing the input audio or video data; means for transmitting the stored data to the server; means for receiving and analyzing the transmitted audio or video data; means for performing spectral analysis on the audio data to evaluate pitch and rhythm; means for analyzing facial expressions and gestures on the video data; means for identifying pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness based on the analysis results; means including a generative AI model for generating specific advice for identified areas for improvement; and means for providing the generated advice to the user. This allows the user to objectively evaluate their own performance and identify specific areas for improvement.
[0094] "User" refers to an individual who uses the system to input audio or video and receive the resulting advice.
[0095] "Means for inputting audio or video" refers to a device such as a microphone or camera that is used to input the user's singing voice, playing, movements, etc. into the system.
[0096] The "means for temporarily storing audio or video data" refers to a storage device or memory for temporarily saving captured audio or video data.
[0097] The "means for transmitting data to a server" refers to communication techniques and protocols for transmitting temporarily stored data to a remote server over a network such as the Internet.
[0098] "Means for receiving and analyzing data" refers to software or algorithms that allow the server to receive transmitted audio or video data and analyze that data.
[0099] "Spectral analysis" is an analytical technique for extracting frequency components from audio data and evaluating pitch and rhythm.
[0100] The "means for evaluating pitch and rhythm" is a technology for evaluating the accuracy of pitch and the degree of rhythmic agreement from the results of spectrum analysis of audio data.
[0101] "Means for analyzing facial expressions and gestures" refers to image analysis technology for recognizing and analyzing a user's facial expressions and body movements based on video data.
[0102] "Means for identifying areas for improvement based on analysis results" refers to technology that identifies areas that need improvement based on discrepancies or deficiencies detected from analyzed audio and video data.
[0103] A "generative AI model that generates specific advice for identified improvement points" is an artificial intelligence model that automatically generates specific advice, such as how users should practice, for detected improvement points.
[0104] The "means for providing the generated advice to the user" is a technology for encoding the generated advice in text or video format, transmitting it to the user's terminal, and displaying it.
[0105] MODE FOR CARRYING OUT THE INVENTION
[0106] The embodiments of the present invention will be specifically described below.
[0107] System Overview
[0108] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[0109] Audio and video capture
[0110] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0111] Data transmission and temporary storage
[0112] The audio and video data is temporarily stored on the device's memory and then sent to a server over the Internet. The device uploads the data to the server using Wi-Fi or mobile networks.
[0113] Data analysis
[0114] The server then passes the received data to an analysis module that evaluates technical characteristics such as pitch, rhythm, and expressiveness. This analysis includes spectral analysis of the audio data and facial and gesture analysis of the video data. The server then uses algorithms to detect specific patterns and deviations in the analyzed data.
[0115] Identifying areas for improvement and generating recommendations
[0116] The server identifies areas for improvement based on the analysis results, such as pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. For each identified area for improvement, the server uses a generative AI model to generate advice, including specific practice methods and points to note.
[0117] Providing advice
[0118] The generated advice is encoded in text or video format and sent to the device. The device then displays the received advice on its user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a little higher" or video advice such as "Use a metronome to practice maintaining a steady tempo" are displayed on the smartphone screen.
[0119] Specific examples
[0120] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[0121] Prompt Sentence Examples
[0122] "Based on a recording of my singing using a karaoke app, please analyze any inconsistencies in pitch or rhythm and suggest specific practice methods."
[0123] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[0124] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0125] Step 1:
[0126] User inputs audio and video data
[0127] Users start a dedicated application on their smartphone or PC and press the "Start Recording" button on the screen to input audio and video data. As users sing or play an instrument, the device's microphone and camera capture this data in real time.
[0128] Input: Real-time audio and video data of singing and performance
[0129] Output: Audio and video data temporarily stored in a storage device
[0130] Step 2:
[0131] The device temporarily stores data
[0132] The device temporarily stores the captured audio and video data in a storage device (e.g., a smartphone's internal storage). After recording is complete, press the "Start Analysis" button to complete the data storage.
[0133] Input: Real-time audio and video data
[0134] Output: Temporarily saved data (audio files, video files)
[0135] Step 3:
[0136] The device sends data to the server
[0137] The device transmits the stored audio and video data to a server over the internet, using Wi-Fi or mobile networks.
[0138] Specifically, the device automatically uploads data after the user presses the "Start Analysis" button.
[0139] Input: Temporarily saved data (audio files, video files)
[0140] Output: Audio and video data sent to the server
[0141] Step 4:
[0142] The server receives and analyzes the data
[0143] The server passes the received audio and video data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0144] Input: Audio and video data sent to the server
[0145] Output: Analyzed technical features (pitch, rhythm, facial expressions, gestures, etc.)
[0146] Step 5:
[0147] The server identifies areas for improvement and generates specific advice
[0148] Based on the analysis results, the server identifies pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. Based on the identified areas for improvement, the server uses a generative AI model to generate specific advice, including which areas should be improved and how.
[0149] Input: Parsed technical features
[0150] Output: Specific advice generated
[0151] Step 6:
[0152] The server sends the generated advice to the device.
[0153] The generated advice is encoded in text or video format and sent to the device.
[0154] Input: Generated specific advice
[0155] Output: Text or video advice sent to your device
[0156] Step 7:
[0157] The device displays advice to the user
[0158] The device displays the received advice on the user interface, and the user can view the advice as text or video on the screen of their smartphone or PC.
[0159] Specific actions include instructions such as "Your pitch is too low in the chorus, so try singing a bit higher," and video advice such as "Use a metronome to practice maintaining a steady tempo."
[0160] Input: Text or video advice sent to your device
[0161] Output: Advice displayed to the user
[0162] (Application example 1)
[0163] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0164] To improve the safety of autonomous vehicles, it is essential to analyze the driving situation in real time and provide appropriate advice based on the results. However, current systems do not efficiently analyze audio and video data while driving and provide specific and timely advice based on that analysis, which means that safety is not sufficiently improved. Another issue is the lack of technology to generate advice that accurately reflects the driving situation.
[0165] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0166] In this invention, the server includes a means for a user to input audio or video, a means for receiving and analyzing the input audio or video data, a means for generating advice on safe driving from the analysis results, and a means for displaying the generated advice on an in-vehicle display or an audio output device in real time, thereby enabling the provision of specific and appropriate advice on safe driving in accordance with the driving situation in real time.
[0167] "User" means a person who uses the system to input audio and video data.
[0168] A "means" is a device or part of a system used to achieve a specified purpose.
[0169] "Voice data" means a digital representation of the voice input by a user.
[0170] "Video data" refers to a digital representation of a video input by a user.
[0171] A "server" is a computer system that analyzes received audio and video data and generates advice based on the results of that analysis.
[0172] "Analysis" refers to the evaluation and detection of technical characteristics of audio and video data.
[0173] "Advice" is specific instructions or recommendations generated by the server based on the analysis results.
[0174] An "in-vehicle display" is a screen display device installed inside an autonomous vehicle.
[0175] The "audio output device" is a device that conveys the generated advice to the user as audio.
[0176] "Safe driving" refers to an autonomous vehicle operating appropriately and safely in accordance with the surrounding conditions.
[0177] "Real-time" means that data is collected, analyzed, and results are provided without delay.
[0178] "Driving conditions" refers to information about vehicle operation, including the vehicle's current speed, location, surrounding traffic conditions, and the like.
[0179] "Obstacle" refers to an object that is in the vehicle's driving path and may affect driving.
[0180] The embodiments of the present invention will be specifically described below.
[0181] System Overview
[0182] This system allows users to input audio and video data, analyzes the data, and provides specific advice on safe driving. The system consists of an in-vehicle terminal used by the user and a server that performs the analysis.
[0183] Audio and video capture
[0184] The user inputs audio and video into the vehicle's microphone and camera, and the device captures the audio and video data in real time. This is done using the vehicle's camera and microphone. The user can start audio input or video capture using, for example, a voice command.
[0185] Data transmission and analysis
[0186] The captured audio and video data is temporarily stored in the in-vehicle terminal's storage device and then sent to a server via the Internet. The server receives the data and passes it to an analysis module. The audio data is analyzed using speech recognition technology, and the video data is analyzed using computer vision technology. Specific software used includes the speech_recognition library for speech recognition and OpenCV for video analysis.
[0187] Identifying areas for improvement and generating recommendations
[0188] The server identifies areas for improvement in driving based on the analysis results. For example, it detects when there is an obstacle ahead or when the speed exceeds the speed limit. Based on these analysis results, the server generates specific advice. This advice includes specific instructions on what to improve and how to improve it. The optimal advice is automatically generated using a generative AI model.
[0189] Providing advice
[0190] The generated advice is encoded in text or audio format and sent back to the terminal, where it is provided to the user in real time via the in-car display or audio output device. For example, the in-car display might say "Obstacle ahead. Please reduce speed" or a similar instruction might be played back via audio.
[0191] Program processing
[0192] The server converts the received audio data into text using the speech_recognition library.
[0193] Video data acquired by the onboard camera is analyzed using the OpenCV library to detect obstacles and other vehicles.
[0194] Based on the analysis results, the generative AI model generates appropriate driving advice.
[0195] The generated advice is displayed in text format on an in-vehicle display or output in audio format from the car's speakers.
[0196] Specific examples
[0197] As a specific example, consider the case where a user asks a question into the car's microphone while driving, "What's the situation ahead?" This voice data is captured in real time and sent to a server for analysis. The server converts the voice into text and analyzes the video data about the situation ahead. Based on the analysis results, a generative AI model generates advice such as "There is a pedestrian ahead, please slow down." This advice is displayed as text on the in-car display or communicated aloud through a voice output device.
[0198] Prompt Sentence Examples
[0199] Input: In-car voice "What's the situation ahead?", external camera footage
[0200] Output: "Slow down, pedestrians ahead."
[0201] As described above, this system can improve the safety of self-driving vehicles by analyzing audio and video data in real time while the user is driving and providing appropriate driving advice.
[0202] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0203] Step 1:
[0204] Users input audio and video into the microphone and camera inside the car.
[0205] Input: User's voice command "What's the situation ahead?" and video data from inside and outside the vehicle.
[0206] Output: Audio and video data temporarily stored on the device.
[0207] How it works: The user asks a question into the car's microphone, which is recorded in real time. At the same time, the car's camera captures images of the interior and exterior of the car.
[0208] Step 2:
[0209] The audio and video data acquired by the terminal is temporarily stored in a storage device.
[0210] Input: Raw audio and video data captured by the device.
[0211] Output: Temporarily stored audio and video data.
[0212] Specific operation: Save the data in the storage device of the in-vehicle terminal (for example, the hard disk or flash memory of the in-vehicle computer).
[0213] Step 3:
[0214] The device transmits the stored data to a server via the Internet.
[0215] Input: Temporarily stored audio and video data.
[0216] Output: Audio and video data sent to the server.
[0217] Specific operation: The saved data is sent to the server using an HTTP request, etc. An internet connection is required for this.
[0218] Step 4:
[0219] The server analyzes the voice data and converts it into text.
[0220] Input: The transmitted audio data.
[0221] Output: Text data converted from audio data.
[0222] Specific operation: The server uses the speech_recognition library to analyze the audio data and perform speech recognition. For example, the audio "What's the situation ahead?" is converted into the text "What's the situation ahead?"
[0223] Step 5:
[0224] The server analyzes the video data and detects obstacles and other vehicles.
[0225] Input: Transmitted video data.
[0226] Output: Obstacle and vehicle position information based on the analyzed video data.
[0227] Specific operation: The server uses the OpenCV library to analyze the video data and, for example, detect the presence of a pedestrian ahead.
[0228] Step 6:
[0229] The server integrates the results of voice recognition and video analysis and generates specific advice using a generative AI model.
[0230] Input: Text data of speech recognition results and location information of video analysis results.
[0231] Output: Specific advice on safe driving.
[0232] Specific operation: The server uses a generative AI model to automatically generate advice such as "Please slow down as there is a pedestrian ahead."
[0233] Step 7:
[0234] The server transmits the generated advice to the vehicle-mounted terminal.
[0235] Input: The generated advice.
[0236] Output: Advice data sent to the in-vehicle terminal.
[0237] Specific operation: The server sends the generated advice to the in-vehicle terminal using an HTTP request or the like.
[0238] Step 8:
[0239] The advice received by the terminal is displayed on the in-vehicle display or played back as audio from an audio output device.
[0240] Input: Advice data sent by the server.
[0241] Output: Text advice displayed on the in-car display or audio advice played from an audio output device.
[0242] Specific action: The in-car display will display the text "Please slow down as there is a pedestrian ahead," or a similar advice will be played aloud from the speaker.
[0243] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0244] The embodiments of the present invention will be specifically described below.
[0245] System Overview
[0246] This system allows users to input audio and video data, analyzes that data, and provides specific advice. By combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[0247] Audio and video capture
[0248] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0249] Data transmission and analysis
[0250] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0251] Use of emotion engine
[0252] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of pitch and rhythm analysis and evaluated. For example, if the user is playing in an unstable emotional state, it can determine whether this is affecting the pitch or rhythm.
[0253] Identifying areas for improvement and generating recommendations
[0254] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on the user's emotions, showing how to improve expressiveness in accordance with the user's state.
[0255] Providing advice
[0256] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0257] Specific examples
[0258] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can practice using this advice as a reference.
[0259] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills and expressiveness.
[0260] The processing flow will be explained below.
[0261] Step 1:
[0262] The user launches a dedicated application and starts recording or recording.
[0263] Press the "Start Recording" button on the application screen to begin inputting audio and video.
[0264] Step 2:
[0265] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[0266] The recorded data is temporarily stored in a memory area on the terminal.
[0267] Step 3:
[0268] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[0269] After completing the data transmission, the terminal waits for a response from the server.
[0270] Step 4:
[0271] The server receives the received audio and video data and passes it to the analysis module.
[0272] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[0273] Step 5:
[0274] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[0275] Specific algorithms are used to analyze emotional expressions and performance movements.
[0276] Step 6:
[0277] The server's emotion engine analyzes the user's emotions in real time from the input audio and video data.
[0278] The analyzed emotional information is then evaluated in combination with the results of pitch and rhythm analysis. For example, if the user is playing in an unstable emotional state, it will be determined that this is affecting the pitch and rhythm.
[0279] Step 7:
[0280] The server lists specific areas for improvement based on the analysis results of pitch, rhythm, expressiveness, etc., as well as emotional information.
[0281] This improvement highlights areas of user performance that require special attention.
[0282] Step 8:
[0283] The server generates specific advice based on the identified improvements and emotions.
[0284] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo." It also generates emotion-based advice such as, "Take deep breaths to relax when playing," depending on your emotional state.
[0285] Step 9:
[0286] The generated advice is encoded in text or video format and sent to the terminal.
[0287] The server performs the encoding process and sends the data to the device in the optimal format.
[0288] Step 10:
[0289] The advice received by the terminal is displayed on the user interface.
[0290] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[0291] Step 11:
[0292] The user checks the displayed advice and practices independently based on the specific practice method.
[0293] By following the advice and practicing repeatedly, you can improve your technique and emotional management skills.
[0294] As a concrete example, consider the case where a user sings a favorite song using a karaoke app. The user launches the app and starts recording. At this time, the device captures and saves the audio and video. The transmitted data is analyzed by the server, which detects pitch discrepancies, rhythmic inconsistencies, emotional instability, and so on. Based on this, the server generates specific advice and sends it to the device. The user can then receive this advice and practice again, effectively improving their skills.
[0295] Example 2
[0296] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0297] Conventional audio and video analysis systems, when analyzing a user's audio and video data, are unable to provide appropriate feedback based on the user's emotional state, in addition to suggesting areas for technical improvement. Therefore, in order for users to practice efficiently, they need support not only in terms of technical aspects but also mental aspects, but this has not been achieved with conventional technologies.
[0298] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0299] In this invention, the server includes means for a user to input audio and video, means for temporarily storing the input audio and video data in a storage device of the terminal, means for transmitting the stored audio and video data to the server, means for analyzing the received audio and video data, means for performing spectrum analysis on the analyzed audio data, means for analyzing facial expressions and gestures of the analyzed video data, means including an emotion engine for analyzing emotions based on the analysis results, means for identifying areas for improvement based on the analysis results and emotion information and generating specific advice, and means for providing the generated advice to the user. This allows the user to receive not only technical improvements but also appropriate feedback according to their emotional state, enabling more effective self-practice.
[0300] "User" means the end user who inputs audio and video data into the system.
[0301] "Audio and video input means" means a device or software that allows a user to input audio and video data into the system.
[0302] A "terminal" is a device that stores audio and video data and transmits it to a server. Examples of such devices include smartphones, tablets, and personal computers.
[0303] "Device storage device" refers to a storage means for temporarily storing audio and video data. This includes memory, storage devices, etc.
[0304] A "server" is a computer system that receives and analyzes audio and video data transmitted from terminals via a network.
[0305] "Means for analyzing" means algorithms or software for extracting and evaluating the technical and emotional characteristics of received audio and video data.
[0306] "Spectral analysis" is a technique for analyzing the frequency components of audio data and evaluating pitch and rhythm.
[0307] "Means for analyzing facial expressions and gestures" refers to algorithms or software for recognizing and evaluating a user's facial expressions and gestures from video data.
[0308] An "emotion engine" is an algorithm or software for analyzing a user's emotions based on audio and video data.
[0309] The "means for identifying improvements" is an algorithm or software for extracting technical and emotional improvements based on the analytical results and emotional information.
[0310] The "means for generating specific advice" is an algorithm or software for generating specific feedback and practice methods for identified areas for improvement.
[0311] "Means for providing" refers to a device or software for providing the generated advice to the user in text or video format.
[0312] MODE FOR CARRYING OUT THE INVENTION
[0313] This invention is a system that allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. This system consists of a user terminal and a server.
[0314] Audio and video capture
[0315] The user inputs sound or plays an instrument into the device. At this time, the device captures the sound and video in real time using the microphone and camera of the smartphone or PC. The user starts the dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0316] Data transmission and analysis
[0317] The audio and video data is temporarily stored in the device's memory and then transmitted over the Internet to a server. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0318] Use of emotion engine
[0319] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. This analysis uses, for example, the tone and pitch of the audio and changes in facial expressions in the video. The analyzed emotional information is then combined with the results of the analysis of pitch and rhythm and evaluated. For example, if the user is playing in an unstable emotional state, it can be determined that this is affecting the pitch or rhythm.
[0320] Identifying areas for improvement and generating advice
[0321] The server identifies technical and emotional areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on emotions, showing how to improve expressiveness according to the user's state.
[0322] Providing advice
[0323] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0324] Specific examples
[0325] As a concrete example, let us consider the case where a user sings using a karaoke app. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and this data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, which analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can refer to this advice to practice efficiently.
[0326] Prompt Sentence Examples
[0327] Analyzing pitch discrepancies: "Analyze the audio data sung by the user and identify the parts where pitch discrepancies occur."
[0328] To analyze rhythmic discrepancies: "Analyze the rhythmic data played by the user and identify where there are rhythmic discrepancies."
[0329] For sentiment analysis: "Analyze the emotions from the user's audio and video data and present the results."
[0330] To generate overall advice: "Analyze the user's performance data and emotional state to identify areas for improvement and generate specific advice."
[0331] This invention allows users to practice more effectively by providing a practice method that takes into account not only their technical skills but also their emotional state.
[0332] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0333] Step 1:
[0334] The user starts a dedicated application and prepares to input audio and video. For example, they start a karaoke application and select a favorite song. The application displays a "Start Recording" button, and when the user presses that button, recording begins. The input data is audio data and video data.
[0335] Input: User's singing voice and video
[0336] Output: Audio data (WAV files, etc.), video data (MP4 files, etc.)
[0337] Step 2:
[0338] The device temporarily stores the recorded and captured audio and video data in a storage device. During this storage process, the data format is automatically set. For example, audio data is stored as a WAV file and video data is stored as an MP4 file in the smartphone's memory.
[0339] Input: Recorded and filmed audio and video data
[0340] Output: Audio and video files saved on a storage device
[0341] Step 3:
[0342] The device then transmits the stored audio and video data to a server over the Internet, using secure encryption protocols such as HTTPS.
[0343] Input: Audio and video files stored on a storage device
[0344] Output: Audio and video data sent to the server
[0345] Step 4:
[0346] The server passes the received audio and video data to the analysis module, which then analyzes the data. Specifically, the audio data undergoes spectral analysis using FFT (Fast Fourier Transform), and the video data undergoes analysis using facial expression and gesture recognition algorithms (e.g., OpenCV).
[0347] Input: Received audio and video data
[0348] Output: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[0349] Step 5:
[0350] The server uses an emotion engine to analyze the user's emotions based on the analysis results, determining the user's emotions (e.g., joy, sadness, surprise) based on the tone and pitch of the audio data and facial expressions in the video data.
[0351] Input: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[0352] Output: Emotion analysis result (e.g., user is calm, nervous)
[0353] Step 6:
[0354] The server identifies areas for improvement based on the analysis results and emotional information and generates specific advice. For example, if the pitch is low, the server generates advice such as, "Your pitch is too low in the chorus, so try singing a bit higher." It also generates emotionally appropriate advice for a nervous user, such as, "Take a deep breath to relax and play."
[0355] Input: Analysis results and emotion information
[0356] Output: Specific advice (text or video format)
[0357] Step 7:
[0358] The server then encodes the generated advice and sends it back to the terminal, again securely using an encryption protocol.
[0359] Input: Specific advice
[0360] Output: Advice data sent to the terminal
[0361] Step 8:
[0362] The device displays the received advice on its user interface. The advice can be in the form of text such as "Your pitch is too low in the chorus, try singing a bit higher," or it can be in the form of a video clip.
[0363] Input: Advice data sent to the terminal
[0364] Output: Specific advice displayed in the user interface
[0365] The specific operations and inputs and outputs at each step have been described above.
[0366] (Application example 2)
[0367] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0368] Conventional audio and video analysis systems generate feedback based on the results of technical analysis of music. However, because they do not take into account the user's emotional state, many users are dissatisfied with technical instruction alone. In particular, when emotions significantly affect a user's playing or singing, the feedback is incomplete, resulting in reduced practice efficiency. To address these issues, the present invention aims to provide customized advice that takes into account the user's emotional state by analyzing the user's emotional state in real time.
[0369] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input audio or video, means for receiving and analyzing the input audio or video data, means for identifying areas for improvement based on the analysis results and generating specific advice, means for providing the generated advice to the user, means for analyzing the user's emotions, and means for generating individually customized advice based on the analyzed emotional information. This not only improves the user's technique, but also enables the user to receive feedback according to their emotional state, resulting in more effective practice.
[0370] "User" refers to an individual or organization that uses the system to input audio or video data.
[0371] "Voice data" refers to voice or music data that users input into the system.
[0372] "Video data" refers to video and image data that users input into the system.
[0373] "Analysis" refers to the process of evaluating and analyzing the technical characteristics and emotional information of audio and video data.
[0374] "Server" refers to the computer system that receives and processes data for analysis and generates feedback.
[0375] "Areas for improvement" refers to points identified from the analyzed data that the user can use to improve their technical or expressive abilities.
[0376] "Advice" refers to feedback provided based on the analysis results, including specific ways to improve or practice.
[0377] "Emotion analysis" refers to the process of assessing a user's emotional state from audio and video data.
[0378] "Customized advice" refers to advice provided based on the user's individual state, based on analyzed emotional information.
[0379] We will now explain a specific system for implementing this invention. This system allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[0380] System Overview
[0381] The system is outlined as follows: Users input audio and video using the microphone and camera on their smartphone or PC. The input data is temporarily stored in the device's memory and then sent to a server via the Internet. The server passes the received data to an analysis module, which evaluates the technical features. It then uses an emotion engine to analyze the user's emotional state and generates specific advice based on the analysis results.
[0382] Audio and video capture
[0383] Users can input audio or play an instrument into the device, and the microphone and camera on their smartphone or PC will capture the audio and video data in real time. Users simply launch the dedicated application and press the "Start Recording" button on the screen to begin recording.
[0384] Data transmission and analysis
[0385] The captured audio and video data is temporarily stored on the device's memory and then transmitted over the Internet to a server using an HTTP library such as "requests." The server then analyzes the received data using software such as "MusicAnalyzer" or "EmotionEngine" to evaluate technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0386] Use of emotion engine
[0387] The server is equipped with an "Emotion Engine" that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of technical analysis and evaluated. For example, if a user is performing in an unstable emotional state, it can be determined that this is affecting pitch and rhythm deviations.
[0388] Identifying areas for improvement and generating recommendations
[0389] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state on these. For identified areas for improvement, the server uses a "generative AI model" to generate specific advice. For example, the server inputs prompts such as "Please suggest practice methods for when my emotional state is unstable" or "Please tell me specific ways to improve when I'm out of tune" into the AI model, and obtains results.
[0390] Providing advice
[0391] The generated advice is encoded in text or video format and sent back to the device. The device receives the "prompt sentence" and displays it on the user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a bit higher" may be displayed on the smartphone screen. Other advice may include "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0392] Specific examples
[0393] Let's consider a scenario where a user sings using an interactive music learning app on a smartphone. First, the user launches the app, selects a favorite song, and sings. The smartphone's microphone records the user's singing, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet and analyzed by "MusicAnalyzer" and "EmotionEngine." Based on the analysis results, the server detects pitch discrepancies and rhythmic inconsistencies and generates specific advice. This advice is sent to the smartphone and displayed on the screen. The user can use this advice as a reference for practicing.
[0394] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0395] Step 1:
[0396] The user launches a dedicated application on their smartphone or PC and presses the start button for recording. As the user inputs sound or plays an instrument, the microphone and camera on the smartphone or PC captures the audio and video data in real time. The input audio and video data is temporarily stored in the device's memory device.
[0397] Input: User's audio and video data
[0398] Output: Audio and video data stored in the device's storage device
[0399] Step 2:
[0400] The device sends the stored audio and video data to the server by uploading the data using an HTTP request and passing it to the server.
[0401] Input: Audio and video data stored on the device
[0402] Output: Audio and video data sent to the server
[0403] Step 3:
[0404] The server analyzes the received audio and video data. First, it uses MusicAnalyzer to evaluate technical characteristics (pitch, rhythm, and expressiveness). Spectral analysis is performed on the audio data to detect deviations in pitch and rhythm. Facial expressions and gestures are analyzed on the video data.
[0405] Input: Audio and video data sent to the server
[0406] Output: Evaluation results of technical features (pitch, rhythm, expressiveness, facial expressions, gestures)
[0407] Step 4:
[0408] The server then uses the "Emotion Engine" to perform emotion analysis, assessing the user's emotional state from audio and video data and combining this with the evaluation of the technical features.
[0409] Input: Technical characteristics evaluation results, audio and video data
[0410] Output: Integrated evaluation results of emotional state and technical features
[0411] Step 5:
[0412] The server identifies areas for improvement based on the analysis results. It comprehensively evaluates technical features and emotional information and generates specific advice. Using a "generative AI model," it inputs prompts into the model to generate individually customized advice. For example, it uses prompts such as, "Please suggest practice methods for when my emotional state is unstable."
[0413] Input: Integrated evaluation results, prompts to the generative AI model
[0414] Output: Personalized, specific advice
[0415] Step 6:
[0416] The server encodes the generated advice in text or video format and sends it back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0417] Input: Specific, personalized advice
[0418] Output: Feedback displayed on the device (text or video format)
[0419] Through these steps, users can not only improve their technique, but also receive feedback based on their emotional state, enabling more effective practice.
[0420] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0421] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0422] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0423] [Second embodiment]
[0424] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0425] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0426] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0427] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0428] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0429] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0430] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0431] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0432] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0433] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0434] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0435] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0436] The embodiments of the present invention will be specifically described below.
[0437] System Overview
[0438] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[0439] Audio and video capture
[0440] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0441] Data transmission and analysis
[0442] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0443] Identifying areas for improvement and generating recommendations
[0444] The server identifies areas for improvement based on the analysis results. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how.
[0445] Providing advice
[0446] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a little higher," or video advice such as, "Use a metronome to practice maintaining a steady tempo."
[0447] Specific examples
[0448] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[0449] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[0450] The processing flow will be explained below.
[0451] Step 1:
[0452] The user launches a dedicated application and starts recording or recording.
[0453] Pressing the "Start Recording" button on the application screen will begin audio and video input.
[0454] Step 2:
[0455] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[0456] The recorded data is temporarily stored in a memory area on the terminal.
[0457] Step 3:
[0458] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[0459] After completing the data transmission, the terminal waits for a response from the server.
[0460] Step 4:
[0461] The server receives the received audio and video data and passes it to the analysis module.
[0462] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[0463] Step 5:
[0464] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[0465] Specific algorithms are used to analyze emotional expressions and performance movements.
[0466] Step 6:
[0467] The server lists specific areas for improvement based on the analysis of pitch, rhythm, expressiveness, etc.
[0468] This improvement highlights areas of user performance that require special attention.
[0469] Step 7:
[0470] The server generates specific advice for the identified improvements.
[0471] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0472] Step 8:
[0473] The generated advice is encoded in text or video format and sent to the terminal.
[0474] The server performs the encoding process and sends the data to the device in the optimal format.
[0475] Step 9:
[0476] The advice received by the terminal is displayed on the user interface.
[0477] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[0478] Step 10:
[0479] The user checks the displayed advice and practices independently based on the specific practice method.
[0480] By following the advice and practicing repeatedly, you can improve your skills.
[0481] Example 1
[0482] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0483] In the past, when practicing music or improving performance, it was difficult for users to objectively evaluate their own performance and identify specific areas for improvement. Furthermore, due to limited opportunities to receive professional instruction, it was difficult to find an efficient practice method.
[0484] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0485] In this invention, the server includes: means for a user to input audio or video; means for temporarily storing the input audio or video data; means for transmitting the stored data to the server; means for receiving and analyzing the transmitted audio or video data; means for performing spectral analysis on the audio data to evaluate pitch and rhythm; means for analyzing facial expressions and gestures on the video data; means for identifying pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness based on the analysis results; means including a generative AI model for generating specific advice for identified areas for improvement; and means for providing the generated advice to the user. This allows the user to objectively evaluate their own performance and identify specific areas for improvement.
[0486] "User" refers to an individual who uses the system to input audio or video and receive the resulting advice.
[0487] "Means for inputting audio or video" refers to a device such as a microphone or camera that is used to input the user's singing voice, playing, movements, etc. into the system.
[0488] The "means for temporarily storing audio or video data" refers to a storage device or memory for temporarily saving captured audio or video data.
[0489] The "means for transmitting data to a server" refers to communication techniques and protocols for transmitting temporarily stored data to a remote server over a network such as the Internet.
[0490] "Means for receiving and analyzing data" refers to software or algorithms that allow the server to receive transmitted audio or video data and analyze that data.
[0491] "Spectral analysis" is an analytical technique for extracting frequency components from audio data and evaluating pitch and rhythm.
[0492] The "means for evaluating pitch and rhythm" is a technology for evaluating the accuracy of pitch and the degree of rhythmic agreement from the results of spectrum analysis of audio data.
[0493] "Means for analyzing facial expressions and gestures" refers to image analysis technology for recognizing and analyzing a user's facial expressions and body movements based on video data.
[0494] "Means for identifying areas for improvement based on analysis results" refers to technology that identifies areas that need improvement based on discrepancies or deficiencies detected from analyzed audio and video data.
[0495] A "generative AI model that generates specific advice for identified improvement points" is an artificial intelligence model that automatically generates specific advice, such as how users should practice, for detected improvement points.
[0496] The "means for providing the generated advice to the user" is a technology for encoding the generated advice in text or video format, transmitting it to the user's terminal, and displaying it.
[0497] MODE FOR CARRYING OUT THE INVENTION
[0498] The embodiments of the present invention will be specifically described below.
[0499] System Overview
[0500] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[0501] Audio and video capture
[0502] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0503] Data transmission and temporary storage
[0504] The audio and video data is temporarily stored on the device's memory and then sent to a server over the Internet. The device uploads the data to the server using Wi-Fi or mobile networks.
[0505] Data analysis
[0506] The server then passes the received data to an analysis module that evaluates technical characteristics such as pitch, rhythm, and expressiveness. This analysis includes spectral analysis of the audio data and facial and gesture analysis of the video data. The server then uses algorithms to detect specific patterns and deviations in the analyzed data.
[0507] Identifying areas for improvement and generating recommendations
[0508] The server identifies areas for improvement based on the analysis results, such as pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. For each identified area for improvement, the server uses a generative AI model to generate advice, including specific practice methods and points to note.
[0509] Providing advice
[0510] The generated advice is encoded in text or video format and sent to the device. The device then displays the received advice on its user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a little higher" or video advice such as "Use a metronome to practice maintaining a steady tempo" are displayed on the smartphone screen.
[0511] Specific examples
[0512] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[0513] Prompt Sentence Examples
[0514] "Based on a recording of my singing using a karaoke app, please analyze any inconsistencies in pitch or rhythm and suggest specific practice methods."
[0515] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[0516] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0517] Step 1:
[0518] User inputs audio and video data
[0519] Users start a dedicated application on their smartphone or PC and press the "Start Recording" button on the screen to input audio and video data. As users sing or play an instrument, the device's microphone and camera capture this data in real time.
[0520] Input: Real-time audio and video data of singing and performance
[0521] Output: Audio and video data temporarily stored in a storage device
[0522] Step 2:
[0523] The device temporarily stores data
[0524] The device temporarily stores the captured audio and video data in a storage device (e.g., a smartphone's internal storage). After recording is complete, press the "Start Analysis" button to complete the data storage.
[0525] Input: Real-time audio and video data
[0526] Output: Temporarily saved data (audio files, video files)
[0527] Step 3:
[0528] The device sends data to the server
[0529] The device transmits the stored audio and video data to a server over the internet, using Wi-Fi or mobile networks.
[0530] Specifically, the device automatically uploads data after the user presses the "Start Analysis" button.
[0531] Input: Temporarily saved data (audio files, video files)
[0532] Output: Audio and video data sent to the server
[0533] Step 4:
[0534] The server receives and analyzes the data
[0535] The server passes the received audio and video data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0536] Input: Audio and video data sent to the server
[0537] Output: Analyzed technical features (pitch, rhythm, facial expressions, gestures, etc.)
[0538] Step 5:
[0539] The server identifies areas for improvement and generates specific advice
[0540] Based on the analysis results, the server identifies pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. Based on the identified areas for improvement, the server uses a generative AI model to generate specific advice, including which areas should be improved and how.
[0541] Input: Parsed technical features
[0542] Output: Specific advice generated
[0543] Step 6:
[0544] The server sends the generated advice to the device.
[0545] The generated advice is encoded in text or video format and sent to the device.
[0546] Input: Generated specific advice
[0547] Output: Text or video advice sent to your device
[0548] Step 7:
[0549] The device displays advice to the user
[0550] The device displays the received advice on the user interface, and the user can view the advice as text or video on the screen of their smartphone or PC.
[0551] Specific actions include instructions such as "Your pitch is too low in the chorus, so try singing a bit higher," and video advice such as "Use a metronome to practice maintaining a steady tempo."
[0552] Input: Text or video advice sent to your device
[0553] Output: Advice displayed to the user
[0554] (Application example 1)
[0555] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0556] To improve the safety of autonomous vehicles, it is essential to analyze the driving situation in real time and provide appropriate advice based on the results. However, current systems do not efficiently analyze audio and video data while driving and provide specific and timely advice based on that analysis, which means that safety is not sufficiently improved. Another issue is the lack of technology to generate advice that accurately reflects the driving situation.
[0557] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0558] In this invention, the server includes a means for a user to input audio or video, a means for receiving and analyzing the input audio or video data, a means for generating advice on safe driving from the analysis results, and a means for displaying the generated advice on an in-vehicle display or an audio output device in real time, thereby enabling the provision of specific and appropriate advice on safe driving in accordance with the driving situation in real time.
[0559] "User" means a person who uses the system to input audio and video data.
[0560] A "means" is a device or part of a system used to achieve a specified purpose.
[0561] "Voice data" means a digital representation of the voice input by a user.
[0562] "Video data" refers to a digital representation of a video input by a user.
[0563] A "server" is a computer system that analyzes received audio and video data and generates advice based on the results of that analysis.
[0564] "Analysis" refers to the evaluation and detection of technical characteristics of audio and video data.
[0565] "Advice" is specific instructions or recommendations generated by the server based on the analysis results.
[0566] An "in-vehicle display" is a screen display device installed inside an autonomous vehicle.
[0567] The "audio output device" is a device that conveys the generated advice to the user as audio.
[0568] "Safe driving" refers to an autonomous vehicle operating appropriately and safely in accordance with the surrounding conditions.
[0569] "Real-time" means that data is collected, analyzed, and results are provided without delay.
[0570] "Driving conditions" refers to information about vehicle operation, including the vehicle's current speed, location, surrounding traffic conditions, and the like.
[0571] "Obstacle" refers to an object that is in the vehicle's driving path and may affect driving.
[0572] The embodiments of the present invention will be specifically described below.
[0573] System Overview
[0574] This system allows users to input audio and video data, analyzes the data, and provides specific advice on safe driving. The system consists of an in-vehicle terminal used by the user and a server that performs the analysis.
[0575] Audio and video capture
[0576] The user inputs audio and video into the vehicle's microphone and camera, and the device captures the audio and video data in real time. This is done using the vehicle's camera and microphone. The user can start audio input or video capture using, for example, a voice command.
[0577] Data transmission and analysis
[0578] The captured audio and video data is temporarily stored in the in-vehicle terminal's storage device and then sent to a server via the Internet. The server receives the data and passes it to an analysis module. The audio data is analyzed using speech recognition technology, and the video data is analyzed using computer vision technology. Specific software used includes the speech_recognition library for speech recognition and OpenCV for video analysis.
[0579] Identifying areas for improvement and generating recommendations
[0580] The server identifies areas for improvement in driving based on the analysis results. For example, it detects when there is an obstacle ahead or when the speed exceeds the speed limit. Based on these analysis results, the server generates specific advice. This advice includes specific instructions on what to improve and how to improve it. The optimal advice is automatically generated using a generative AI model.
[0581] Providing advice
[0582] The generated advice is encoded in text or audio format and sent back to the terminal, where it is provided to the user in real time via the in-car display or audio output device. For example, the in-car display might say "Obstacle ahead. Please reduce speed" or a similar instruction might be played back via audio.
[0583] Program processing
[0584] The server converts the received audio data into text using the speech_recognition library.
[0585] Video data acquired by the onboard camera is analyzed using the OpenCV library to detect obstacles and other vehicles.
[0586] Based on the analysis results, the generative AI model generates appropriate driving advice.
[0587] The generated advice is displayed in text format on an in-vehicle display or output in audio format from the car's speakers.
[0588] Specific examples
[0589] As a specific example, consider the case where a user asks a question into the car's microphone while driving, "What's the situation ahead?" This voice data is captured in real time and sent to a server for analysis. The server converts the voice into text and analyzes the video data about the situation ahead. Based on the analysis results, a generative AI model generates advice such as "There is a pedestrian ahead, please slow down." This advice is displayed as text on the in-car display or communicated aloud through a voice output device.
[0590] Prompt Sentence Examples
[0591] Input: In-car voice "What's the situation ahead?", external camera footage
[0592] Output: "Slow down, pedestrians ahead."
[0593] As described above, this system can improve the safety of self-driving vehicles by analyzing audio and video data in real time while the user is driving and providing appropriate driving advice.
[0594] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0595] Step 1:
[0596] Users input audio and video into the microphone and camera inside the car.
[0597] Input: User's voice command "What's the situation ahead?" and video data from inside and outside the vehicle.
[0598] Output: Audio and video data temporarily stored on the device.
[0599] How it works: The user asks a question into the car's microphone, which is recorded in real time. At the same time, the car's camera captures images of the interior and exterior of the car.
[0600] Step 2:
[0601] The audio and video data acquired by the terminal is temporarily stored in a storage device.
[0602] Input: Raw audio and video data captured by the device.
[0603] Output: Temporarily stored audio and video data.
[0604] Specific operation: Save the data in the storage device of the in-vehicle terminal (for example, the hard disk or flash memory of the in-vehicle computer).
[0605] Step 3:
[0606] The device transmits the stored data to a server via the Internet.
[0607] Input: Temporarily stored audio and video data.
[0608] Output: Audio and video data sent to the server.
[0609] Specific operation: The saved data is sent to the server using an HTTP request, etc. An internet connection is required for this.
[0610] Step 4:
[0611] The server analyzes the voice data and converts it into text.
[0612] Input: The transmitted audio data.
[0613] Output: Text data converted from audio data.
[0614] Specific operation: The server uses the speech_recognition library to analyze the audio data and perform speech recognition. For example, the audio "What's the situation ahead?" is converted into the text "What's the situation ahead?"
[0615] Step 5:
[0616] The server analyzes the video data and detects obstacles and other vehicles.
[0617] Input: Transmitted video data.
[0618] Output: Obstacle and vehicle position information based on the analyzed video data.
[0619] Specific operation: The server uses the OpenCV library to analyze the video data and, for example, detect the presence of a pedestrian ahead.
[0620] Step 6:
[0621] The server integrates the results of voice recognition and video analysis and generates specific advice using a generative AI model.
[0622] Input: Text data of speech recognition results and location information of video analysis results.
[0623] Output: Specific advice on safe driving.
[0624] Specific operation: The server uses a generative AI model to automatically generate advice such as "Please slow down as there is a pedestrian ahead."
[0625] Step 7:
[0626] The server transmits the generated advice to the vehicle-mounted terminal.
[0627] Input: The generated advice.
[0628] Output: Advice data sent to the in-vehicle terminal.
[0629] Specific operation: The server sends the generated advice to the in-vehicle terminal using an HTTP request or the like.
[0630] Step 8:
[0631] The advice received by the terminal is displayed on the in-vehicle display or played back as audio from an audio output device.
[0632] Input: Advice data sent by the server.
[0633] Output: Text advice displayed on the in-car display or audio advice played from an audio output device.
[0634] Specific action: The in-car display will display the text "Please slow down as there is a pedestrian ahead," or a similar advice will be played aloud from the speaker.
[0635] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0636] The embodiments of the present invention will be specifically described below.
[0637] System Overview
[0638] This system allows users to input audio and video data, analyzes that data, and provides specific advice. By combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[0639] Audio and video capture
[0640] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0641] Data transmission and analysis
[0642] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0643] Use of emotion engine
[0644] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of pitch and rhythm analysis and evaluated. For example, if the user is playing in an unstable emotional state, it can determine whether this is affecting the pitch or rhythm.
[0645] Identifying areas for improvement and generating recommendations
[0646] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on the user's emotions, showing how to improve expressiveness in accordance with the user's state.
[0647] Providing advice
[0648] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0649] Specific examples
[0650] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can practice using this advice as a reference.
[0651] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills and expressiveness.
[0652] The processing flow will be explained below.
[0653] Step 1:
[0654] The user launches a dedicated application and starts recording or recording.
[0655] Press the "Start Recording" button on the application screen to begin inputting audio and video.
[0656] Step 2:
[0657] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[0658] The recorded data is temporarily stored in a memory area on the terminal.
[0659] Step 3:
[0660] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[0661] After completing the data transmission, the terminal waits for a response from the server.
[0662] Step 4:
[0663] The server receives the received audio and video data and passes it to the analysis module.
[0664] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[0665] Step 5:
[0666] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[0667] Specific algorithms are used to analyze emotional expressions and performance movements.
[0668] Step 6:
[0669] The server's emotion engine analyzes the user's emotions in real time from the input audio and video data.
[0670] The analyzed emotional information is then evaluated in combination with the results of pitch and rhythm analysis. For example, if the user is playing in an unstable emotional state, it will be determined that this is affecting the pitch and rhythm.
[0671] Step 7:
[0672] The server lists specific areas for improvement based on the analysis results of pitch, rhythm, expressiveness, etc., as well as emotional information.
[0673] This improvement highlights areas of user performance that require special attention.
[0674] Step 8:
[0675] The server generates specific advice based on the identified improvements and emotions.
[0676] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo." It also generates emotion-based advice such as, "Take deep breaths to relax when playing," depending on your emotional state.
[0677] Step 9:
[0678] The generated advice is encoded in text or video format and sent to the terminal.
[0679] The server performs the encoding process and sends the data to the device in the optimal format.
[0680] Step 10:
[0681] The advice received by the terminal is displayed on the user interface.
[0682] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[0683] Step 11:
[0684] The user checks the displayed advice and practices independently based on the specific practice method.
[0685] By following the advice and practicing repeatedly, you can improve your technique and emotional management skills.
[0686] As a concrete example, consider the case where a user sings a favorite song using a karaoke app. The user launches the app and starts recording. At this time, the device captures and saves the audio and video. The transmitted data is analyzed by the server, which detects pitch discrepancies, rhythmic inconsistencies, emotional instability, and so on. Based on this, the server generates specific advice and sends it to the device. The user can then receive this advice and practice again, effectively improving their skills.
[0687] Example 2
[0688] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0689] Conventional audio and video analysis systems, when analyzing a user's audio and video data, are unable to provide appropriate feedback based on the user's emotional state, in addition to suggesting areas for technical improvement. Therefore, in order for users to practice efficiently, they need support not only in terms of technical aspects but also mental aspects, but this has not been achieved with conventional technologies.
[0690] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0691] In this invention, the server includes means for a user to input audio and video, means for temporarily storing the input audio and video data in a storage device of the terminal, means for transmitting the stored audio and video data to the server, means for analyzing the received audio and video data, means for performing spectrum analysis on the analyzed audio data, means for analyzing facial expressions and gestures of the analyzed video data, means including an emotion engine for analyzing emotions based on the analysis results, means for identifying areas for improvement based on the analysis results and emotion information and generating specific advice, and means for providing the generated advice to the user. This allows the user to receive not only technical improvements but also appropriate feedback according to their emotional state, enabling more effective self-practice.
[0692] "User" means the end user who inputs audio and video data into the system.
[0693] "Audio and video input means" means a device or software that allows a user to input audio and video data into the system.
[0694] A "terminal" is a device that stores audio and video data and transmits it to a server. Examples of such devices include smartphones, tablets, and personal computers.
[0695] "Device storage device" refers to a storage means for temporarily storing audio and video data. This includes memory, storage devices, etc.
[0696] A "server" is a computer system that receives and analyzes audio and video data transmitted from terminals via a network.
[0697] "Means for analyzing" means algorithms or software for extracting and evaluating the technical and emotional characteristics of received audio and video data.
[0698] "Spectral analysis" is a technique for analyzing the frequency components of audio data and evaluating pitch and rhythm.
[0699] "Means for analyzing facial expressions and gestures" refers to algorithms or software for recognizing and evaluating a user's facial expressions and gestures from video data.
[0700] An "emotion engine" is an algorithm or software for analyzing a user's emotions based on audio and video data.
[0701] The "means for identifying improvements" is an algorithm or software for extracting technical and emotional improvements based on the analytical results and emotional information.
[0702] The "means for generating specific advice" is an algorithm or software for generating specific feedback and practice methods for identified areas for improvement.
[0703] "Means for providing" refers to a device or software for providing the generated advice to the user in text or video format.
[0704] MODE FOR CARRYING OUT THE INVENTION
[0705] This invention is a system that allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. This system consists of a user terminal and a server.
[0706] Audio and video capture
[0707] The user inputs sound or plays an instrument into the device. At this time, the device captures the sound and video in real time using the microphone and camera of the smartphone or PC. The user starts the dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0708] Data transmission and analysis
[0709] The audio and video data is temporarily stored in the device's memory and then transmitted over the Internet to a server. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0710] Use of emotion engine
[0711] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. This analysis uses, for example, the tone and pitch of the audio and changes in facial expressions in the video. The analyzed emotional information is then combined with the results of the analysis of pitch and rhythm and evaluated. For example, if the user is playing in an unstable emotional state, it can be determined that this is affecting the pitch or rhythm.
[0712] Identifying areas for improvement and generating advice
[0713] The server identifies technical and emotional areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on emotions, showing how to improve expressiveness according to the user's state.
[0714] Providing advice
[0715] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0716] Specific examples
[0717] As a concrete example, let us consider the case where a user sings using a karaoke app. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and this data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, which analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can refer to this advice to practice efficiently.
[0718] Prompt Sentence Examples
[0719] Analyzing pitch discrepancies: "Analyze the audio data sung by the user and identify the parts where pitch discrepancies occur."
[0720] To analyze rhythmic discrepancies: "Analyze the rhythmic data played by the user and identify where there are rhythmic discrepancies."
[0721] For sentiment analysis: "Analyze the emotions from the user's audio and video data and present the results."
[0722] To generate overall advice: "Analyze the user's performance data and emotional state to identify areas for improvement and generate specific advice."
[0723] This invention allows users to practice more effectively by providing a practice method that takes into account not only their technical skills but also their emotional state.
[0724] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0725] Step 1:
[0726] The user starts a dedicated application and prepares to input audio and video. For example, they start a karaoke application and select a favorite song. The application displays a "Start Recording" button, and when the user presses that button, recording begins. The input data is audio data and video data.
[0727] Input: User's singing voice and video
[0728] Output: Audio data (WAV files, etc.), video data (MP4 files, etc.)
[0729] Step 2:
[0730] The device temporarily stores the recorded and captured audio and video data in a storage device. During this storage process, the data format is automatically set. For example, audio data is stored as a WAV file and video data is stored as an MP4 file in the smartphone's memory.
[0731] Input: Recorded and filmed audio and video data
[0732] Output: Audio and video files saved on a storage device
[0733] Step 3:
[0734] The device then transmits the stored audio and video data to a server over the Internet, using secure encryption protocols such as HTTPS.
[0735] Input: Audio and video files stored on a storage device
[0736] Output: Audio and video data sent to the server
[0737] Step 4:
[0738] The server passes the received audio and video data to the analysis module, which then analyzes the data. Specifically, the audio data undergoes spectral analysis using FFT (Fast Fourier Transform), and the video data undergoes analysis using facial expression and gesture recognition algorithms (e.g., OpenCV).
[0739] Input: Received audio and video data
[0740] Output: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[0741] Step 5:
[0742] The server uses an emotion engine to analyze the user's emotions based on the analysis results, determining the user's emotions (e.g., joy, sadness, surprise) based on the tone and pitch of the audio data and facial expressions in the video data.
[0743] Input: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[0744] Output: Emotion analysis result (e.g., user is calm, nervous)
[0745] Step 6:
[0746] The server identifies areas for improvement based on the analysis results and emotional information and generates specific advice. For example, if the pitch is low, the server generates advice such as, "Your pitch is too low in the chorus, so try singing a bit higher." It also generates emotionally appropriate advice for a nervous user, such as, "Take a deep breath to relax and play."
[0747] Input: Analysis results and emotion information
[0748] Output: Specific advice (text or video format)
[0749] Step 7:
[0750] The server then encodes the generated advice and sends it back to the terminal, again securely using an encryption protocol.
[0751] Input: Specific advice
[0752] Output: Advice data sent to the terminal
[0753] Step 8:
[0754] The device displays the received advice on its user interface. The advice can be in the form of text such as "Your pitch is too low in the chorus, try singing a bit higher," or it can be in the form of a video clip.
[0755] Input: Advice data sent to the terminal
[0756] Output: Specific advice displayed in the user interface
[0757] The specific operations and inputs and outputs at each step have been described above.
[0758] (Application example 2)
[0759] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0760] Conventional audio and video analysis systems generate feedback based on the results of technical analysis of music. However, because they do not take into account the user's emotional state, many users are dissatisfied with technical instruction alone. In particular, when emotions significantly affect a user's playing or singing, the feedback is incomplete, resulting in reduced practice efficiency. To address these issues, the present invention aims to provide customized advice that takes into account the user's emotional state by analyzing the user's emotional state in real time.
[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input audio or video, means for receiving and analyzing the input audio or video data, means for identifying areas for improvement based on the analysis results and generating specific advice, means for providing the generated advice to the user, means for analyzing the user's emotions, and means for generating individually customized advice based on the analyzed emotional information. This not only improves the user's technique, but also enables the user to receive feedback according to their emotional state, resulting in more effective practice.
[0762] "User" refers to an individual or organization that uses the system to input audio or video data.
[0763] "Voice data" refers to voice or music data that users input into the system.
[0764] "Video data" refers to video and image data that users input into the system.
[0765] "Analysis" refers to the process of evaluating and analyzing the technical characteristics and emotional information of audio and video data.
[0766] "Server" refers to the computer system that receives and processes data for analysis and generates feedback.
[0767] "Areas for improvement" refers to points identified from the analyzed data that the user can use to improve their technical or expressive abilities.
[0768] "Advice" refers to feedback provided based on the analysis results, including specific ways to improve or practice.
[0769] "Emotion analysis" refers to the process of assessing a user's emotional state from audio and video data.
[0770] "Customized advice" refers to advice provided based on the user's individual state, based on analyzed emotional information.
[0771] We will now explain a specific system for implementing this invention. This system allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[0772] System Overview
[0773] The system is outlined as follows: Users input audio and video using the microphone and camera on their smartphone or PC. The input data is temporarily stored in the device's memory and then sent to a server via the Internet. The server passes the received data to an analysis module, which evaluates the technical features. It then uses an emotion engine to analyze the user's emotional state and generates specific advice based on the analysis results.
[0774] Audio and video capture
[0775] Users can input audio or play an instrument into the device, and the microphone and camera on their smartphone or PC will capture the audio and video data in real time. Users simply launch the dedicated application and press the "Start Recording" button on the screen to begin recording.
[0776] Data transmission and analysis
[0777] The captured audio and video data is temporarily stored on the device's memory and then transmitted over the Internet to a server using an HTTP library such as "requests." The server then analyzes the received data using software such as "MusicAnalyzer" or "EmotionEngine" to evaluate technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0778] Use of emotion engine
[0779] The server is equipped with an "Emotion Engine" that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of technical analysis and evaluated. For example, if a user is performing in an unstable emotional state, it can be determined that this is affecting pitch and rhythm deviations.
[0780] Identifying areas for improvement and generating recommendations
[0781] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state on these. For identified areas for improvement, the server uses a "generative AI model" to generate specific advice. For example, the server inputs prompts such as "Please suggest practice methods for when my emotional state is unstable" or "Please tell me specific ways to improve when I'm out of tune" into the AI model, and obtains results.
[0782] Providing advice
[0783] The generated advice is encoded in text or video format and sent back to the device. The device receives the "prompt sentence" and displays it on the user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a bit higher" may be displayed on the smartphone screen. Other advice may include "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0784] Specific examples
[0785] Let's consider a scenario where a user sings using an interactive music learning app on a smartphone. First, the user launches the app, selects a favorite song, and sings. The smartphone's microphone records the user's singing, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet and analyzed by "MusicAnalyzer" and "EmotionEngine." Based on the analysis results, the server detects pitch discrepancies and rhythmic inconsistencies and generates specific advice. This advice is sent to the smartphone and displayed on the screen. The user can use this advice as a reference for practicing.
[0786] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0787] Step 1:
[0788] The user launches a dedicated application on their smartphone or PC and presses the start button for recording. As the user inputs sound or plays an instrument, the microphone and camera on the smartphone or PC captures the audio and video data in real time. The input audio and video data is temporarily stored in the device's memory device.
[0789] Input: User's audio and video data
[0790] Output: Audio and video data stored in the device's storage device
[0791] Step 2:
[0792] The device sends the stored audio and video data to the server by uploading the data using an HTTP request and passing it to the server.
[0793] Input: Audio and video data stored on the device
[0794] Output: Audio and video data sent to the server
[0795] Step 3:
[0796] The server analyzes the received audio and video data. First, it uses MusicAnalyzer to evaluate technical characteristics (pitch, rhythm, and expressiveness). Spectral analysis is performed on the audio data to detect deviations in pitch and rhythm. Facial expressions and gestures are analyzed on the video data.
[0797] Input: Audio and video data sent to the server
[0798] Output: Evaluation results of technical features (pitch, rhythm, expressiveness, facial expressions, gestures)
[0799] Step 4:
[0800] The server then uses the "Emotion Engine" to perform emotion analysis, assessing the user's emotional state from audio and video data and combining this with the evaluation of the technical features.
[0801] Input: Technical characteristics evaluation results, audio and video data
[0802] Output: Integrated evaluation results of emotional state and technical features
[0803] Step 5:
[0804] The server identifies areas for improvement based on the analysis results. It comprehensively evaluates technical features and emotional information and generates specific advice. Using a "generative AI model," it inputs prompts into the model to generate individually customized advice. For example, it uses prompts such as, "Please suggest practice methods for when my emotional state is unstable."
[0805] Input: Integrated evaluation results, prompts to the generative AI model
[0806] Output: Personalized, specific advice
[0807] Step 6:
[0808] The server encodes the generated advice in text or video format and sends it back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0809] Input: Specific, personalized advice
[0810] Output: Feedback displayed on the device (text or video format)
[0811] Through these steps, users can not only improve their technique, but also receive feedback based on their emotional state, enabling more effective practice.
[0812] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0813] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0814] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0815] [Third embodiment]
[0816] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0817] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0818] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0819] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0820] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0821] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0822] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0823] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0824] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0825] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0826] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0827] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0828] The embodiments of the present invention will be specifically described below.
[0829] System Overview
[0830] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[0831] Audio and video capture
[0832] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0833] Data transmission and analysis
[0834] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0835] Identifying areas for improvement and generating recommendations
[0836] The server identifies areas for improvement based on the analysis results. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how.
[0837] Providing advice
[0838] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a little higher," or video advice such as, "Use a metronome to practice maintaining a steady tempo."
[0839] Specific examples
[0840] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[0841] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[0842] The processing flow will be explained below.
[0843] Step 1:
[0844] The user launches a dedicated application and starts recording or recording.
[0845] Pressing the "Start Recording" button on the application screen will begin audio and video input.
[0846] Step 2:
[0847] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[0848] The recorded data is temporarily stored in a memory area on the terminal.
[0849] Step 3:
[0850] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[0851] After completing the data transmission, the terminal waits for a response from the server.
[0852] Step 4:
[0853] The server receives the received audio and video data and passes it to the analysis module.
[0854] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[0855] Step 5:
[0856] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[0857] Specific algorithms are used to analyze emotional expressions and performance movements.
[0858] Step 6:
[0859] The server lists specific areas for improvement based on the analysis of pitch, rhythm, expressiveness, etc.
[0860] This improvement highlights areas of user performance that require special attention.
[0861] Step 7:
[0862] The server generates specific advice for the identified improvements.
[0863] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[0864] Step 8:
[0865] The generated advice is encoded in text or video format and sent to the terminal.
[0866] The server performs the encoding process and sends the data to the device in the optimal format.
[0867] Step 9:
[0868] The advice received by the terminal is displayed on the user interface.
[0869] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[0870] Step 10:
[0871] The user checks the displayed advice and practices independently based on the specific practice method.
[0872] By following the advice and practicing repeatedly, you can improve your skills.
[0873] Example 1
[0874] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0875] In the past, when practicing music or improving performance, it was difficult for users to objectively evaluate their own performance and identify specific areas for improvement. Furthermore, due to limited opportunities to receive professional instruction, it was difficult to find an efficient practice method.
[0876] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0877] In this invention, the server includes: means for a user to input audio or video; means for temporarily storing the input audio or video data; means for transmitting the stored data to the server; means for receiving and analyzing the transmitted audio or video data; means for performing spectral analysis on the audio data to evaluate pitch and rhythm; means for analyzing facial expressions and gestures on the video data; means for identifying pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness based on the analysis results; means including a generative AI model for generating specific advice for identified areas for improvement; and means for providing the generated advice to the user. This allows the user to objectively evaluate their own performance and identify specific areas for improvement.
[0878] "User" refers to an individual who uses the system to input audio or video and receive the resulting advice.
[0879] "Means for inputting audio or video" refers to a device such as a microphone or camera that is used to input the user's singing voice, playing, movements, etc. into the system.
[0880] The "means for temporarily storing audio or video data" refers to a storage device or memory for temporarily saving captured audio or video data.
[0881] The "means for transmitting data to a server" refers to communication techniques and protocols for transmitting temporarily stored data to a remote server over a network such as the Internet.
[0882] "Means for receiving and analyzing data" refers to software or algorithms that allow the server to receive transmitted audio or video data and analyze that data.
[0883] "Spectral analysis" is an analytical technique for extracting frequency components from audio data and evaluating pitch and rhythm.
[0884] The "means for evaluating pitch and rhythm" is a technology for evaluating the accuracy of pitch and the degree of rhythmic agreement from the results of spectrum analysis of audio data.
[0885] "Means for analyzing facial expressions and gestures" refers to image analysis technology for recognizing and analyzing a user's facial expressions and body movements based on video data.
[0886] "Means for identifying areas for improvement based on analysis results" refers to technology that identifies areas that need improvement based on discrepancies or deficiencies detected from analyzed audio and video data.
[0887] A "generative AI model that generates specific advice for identified improvement points" is an artificial intelligence model that automatically generates specific advice, such as how users should practice, for detected improvement points.
[0888] The "means for providing the generated advice to the user" is a technology for encoding the generated advice in text or video format, transmitting it to the user's terminal, and displaying it.
[0889] MODE FOR CARRYING OUT THE INVENTION
[0890] The embodiments of the present invention will be specifically described below.
[0891] System Overview
[0892] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[0893] Audio and video capture
[0894] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[0895] Data transmission and temporary storage
[0896] The audio and video data is temporarily stored on the device's memory and then sent to a server over the Internet. The device uploads the data to the server using Wi-Fi or mobile networks.
[0897] Data analysis
[0898] The server then passes the received data to an analysis module that evaluates technical characteristics such as pitch, rhythm, and expressiveness. This analysis includes spectral analysis of the audio data and facial and gesture analysis of the video data. The server then uses algorithms to detect specific patterns and deviations in the analyzed data.
[0899] Identifying areas for improvement and generating recommendations
[0900] The server identifies areas for improvement based on the analysis results, such as pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. For each identified area for improvement, the server uses a generative AI model to generate advice, including specific practice methods and points to note.
[0901] Providing advice
[0902] The generated advice is encoded in text or video format and sent to the device. The device then displays the received advice on its user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a little higher" or video advice such as "Use a metronome to practice maintaining a steady tempo" are displayed on the smartphone screen.
[0903] Specific examples
[0904] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[0905] Prompt Sentence Examples
[0906] "Based on a recording of my singing using a karaoke app, please analyze any inconsistencies in pitch or rhythm and suggest specific practice methods."
[0907] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[0908] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0909] Step 1:
[0910] User inputs audio and video data
[0911] Users start a dedicated application on their smartphone or PC and press the "Start Recording" button on the screen to input audio and video data. As users sing or play an instrument, the device's microphone and camera capture this data in real time.
[0912] Input: Real-time audio and video data of singing and performance
[0913] Output: Audio and video data temporarily stored in a storage device
[0914] Step 2:
[0915] The device temporarily stores data
[0916] The device temporarily stores the captured audio and video data in a storage device (e.g., a smartphone's internal storage). After recording is complete, press the "Start Analysis" button to complete the data storage.
[0917] Input: Real-time audio and video data
[0918] Output: Temporarily saved data (audio files, video files)
[0919] Step 3:
[0920] The device sends data to the server
[0921] The device transmits the stored audio and video data to a server over the internet, using Wi-Fi or mobile networks.
[0922] Specifically, the device automatically uploads data after the user presses the "Start Analysis" button.
[0923] Input: Temporarily saved data (audio files, video files)
[0924] Output: Audio and video data sent to the server
[0925] Step 4:
[0926] The server receives and analyzes the data
[0927] The server passes the received audio and video data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[0928] Input: Audio and video data sent to the server
[0929] Output: Analyzed technical features (pitch, rhythm, facial expressions, gestures, etc.)
[0930] Step 5:
[0931] The server identifies areas for improvement and generates specific advice
[0932] Based on the analysis results, the server identifies pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. Based on the identified areas for improvement, the server uses a generative AI model to generate specific advice, including which areas should be improved and how.
[0933] Input: Parsed technical features
[0934] Output: Specific advice generated
[0935] Step 6:
[0936] The server sends the generated advice to the device.
[0937] The generated advice is encoded in text or video format and sent to the device.
[0938] Input: Generated specific advice
[0939] Output: Text or video advice sent to your device
[0940] Step 7:
[0941] The device displays advice to the user
[0942] The device displays the received advice on the user interface, and the user can view the advice as text or video on the screen of their smartphone or PC.
[0943] Specific actions include instructions such as "Your pitch is too low in the chorus, so try singing a bit higher," and video advice such as "Use a metronome to practice maintaining a steady tempo."
[0944] Input: Text or video advice sent to your device
[0945] Output: Advice displayed to the user
[0946] (Application example 1)
[0947] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0948] To improve the safety of autonomous vehicles, it is essential to analyze the driving situation in real time and provide appropriate advice based on the results. However, current systems do not efficiently analyze audio and video data while driving and provide specific and timely advice based on that analysis, which means that safety is not sufficiently improved. Another issue is the lack of technology to generate advice that accurately reflects the driving situation.
[0949] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0950] In this invention, the server includes a means for a user to input audio or video, a means for receiving and analyzing the input audio or video data, a means for generating advice on safe driving from the analysis results, and a means for displaying the generated advice on an in-vehicle display or an audio output device in real time, thereby enabling the provision of specific and appropriate advice on safe driving in accordance with the driving situation in real time.
[0951] "User" means a person who uses the system to input audio and video data.
[0952] A "means" is a device or part of a system used to achieve a specified purpose.
[0953] "Voice data" means a digital representation of the voice input by a user.
[0954] "Video data" refers to a digital representation of a video input by a user.
[0955] A "server" is a computer system that analyzes received audio and video data and generates advice based on the results of that analysis.
[0956] "Analysis" refers to the evaluation and detection of technical characteristics of audio and video data.
[0957] "Advice" is specific instructions or recommendations generated by the server based on the analysis results.
[0958] An "in-vehicle display" is a screen display device installed inside an autonomous vehicle.
[0959] The "audio output device" is a device that conveys the generated advice to the user as audio.
[0960] "Safe driving" refers to an autonomous vehicle operating appropriately and safely in accordance with the surrounding conditions.
[0961] "Real-time" means that data is collected, analyzed, and results are provided without delay.
[0962] "Driving conditions" refers to information about vehicle operation, including the vehicle's current speed, location, surrounding traffic conditions, and the like.
[0963] "Obstacle" refers to an object that is in the vehicle's driving path and may affect driving.
[0964] The embodiments of the present invention will be specifically described below.
[0965] System Overview
[0966] This system allows users to input audio and video data, analyzes the data, and provides specific advice on safe driving. The system consists of an in-vehicle terminal used by the user and a server that performs the analysis.
[0967] Audio and video capture
[0968] The user inputs audio and video into the vehicle's microphone and camera, and the device captures the audio and video data in real time. This is done using the vehicle's camera and microphone. The user can start audio input or video capture using, for example, a voice command.
[0969] Data transmission and analysis
[0970] The captured audio and video data is temporarily stored in the in-vehicle terminal's storage device and then sent to a server via the Internet. The server receives the data and passes it to an analysis module. The audio data is analyzed using speech recognition technology, and the video data is analyzed using computer vision technology. Specific software used includes the speech_recognition library for speech recognition and OpenCV for video analysis.
[0971] Identifying areas for improvement and generating recommendations
[0972] The server identifies areas for improvement in driving based on the analysis results. For example, it detects when there is an obstacle ahead or when the speed exceeds the speed limit. Based on these analysis results, the server generates specific advice. This advice includes specific instructions on what to improve and how to improve it. The optimal advice is automatically generated using a generative AI model.
[0973] Providing advice
[0974] The generated advice is encoded in text or audio format and sent back to the terminal, where it is provided to the user in real time via the in-car display or audio output device. For example, the in-car display might say "Obstacle ahead. Please reduce speed" or a similar instruction might be played back via audio.
[0975] Program processing
[0976] The server converts the received audio data into text using the speech_recognition library.
[0977] Video data acquired by the onboard camera is analyzed using the OpenCV library to detect obstacles and other vehicles.
[0978] Based on the analysis results, the generative AI model generates appropriate driving advice.
[0979] The generated advice is displayed in text format on an in-vehicle display or output in audio format from the car's speakers.
[0980] Specific examples
[0981] As a specific example, consider the case where a user asks a question into the car's microphone while driving, "What's the situation ahead?" This voice data is captured in real time and sent to a server for analysis. The server converts the voice into text and analyzes the video data about the situation ahead. Based on the analysis results, a generative AI model generates advice such as "There is a pedestrian ahead, please slow down." This advice is displayed as text on the in-car display or communicated aloud through a voice output device.
[0982] Prompt Sentence Examples
[0983] Input: In-car voice "What's the situation ahead?", external camera footage
[0984] Output: "Slow down, pedestrians ahead."
[0985] As described above, this system can improve the safety of self-driving vehicles by analyzing audio and video data in real time while the user is driving and providing appropriate driving advice.
[0986] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0987] Step 1:
[0988] Users input audio and video into the microphone and camera inside the car.
[0989] Input: User's voice command "What's the situation ahead?" and video data from inside and outside the vehicle.
[0990] Output: Audio and video data temporarily stored on the device.
[0991] How it works: The user asks a question into the car's microphone, which is recorded in real time. At the same time, the car's camera captures images of the interior and exterior of the car.
[0992] Step 2:
[0993] The audio and video data acquired by the terminal is temporarily stored in a storage device.
[0994] Input: Raw audio and video data captured by the device.
[0995] Output: Temporarily stored audio and video data.
[0996] Specific operation: Save the data in the storage device of the in-vehicle terminal (for example, the hard disk or flash memory of the in-vehicle computer).
[0997] Step 3:
[0998] The device transmits the stored data to a server via the Internet.
[0999] Input: Temporarily stored audio and video data.
[1000] Output: Audio and video data sent to the server.
[1001] Specific operation: The saved data is sent to the server using an HTTP request, etc. An internet connection is required for this.
[1002] Step 4:
[1003] The server analyzes the voice data and converts it into text.
[1004] Input: The transmitted audio data.
[1005] Output: Text data converted from audio data.
[1006] Specific operation: The server uses the speech_recognition library to analyze the audio data and perform speech recognition. For example, the audio "What's the situation ahead?" is converted into the text "What's the situation ahead?"
[1007] Step 5:
[1008] The server analyzes the video data and detects obstacles and other vehicles.
[1009] Input: Transmitted video data.
[1010] Output: Obstacle and vehicle position information based on the analyzed video data.
[1011] Specific operation: The server uses the OpenCV library to analyze the video data and, for example, detect the presence of a pedestrian ahead.
[1012] Step 6:
[1013] The server integrates the results of voice recognition and video analysis and generates specific advice using a generative AI model.
[1014] Input: Text data of speech recognition results and location information of video analysis results.
[1015] Output: Specific advice on safe driving.
[1016] Specific operation: The server uses a generative AI model to automatically generate advice such as "Please slow down as there is a pedestrian ahead."
[1017] Step 7:
[1018] The server transmits the generated advice to the vehicle-mounted terminal.
[1019] Input: The generated advice.
[1020] Output: Advice data sent to the in-vehicle terminal.
[1021] Specific operation: The server sends the generated advice to the in-vehicle terminal using an HTTP request or the like.
[1022] Step 8:
[1023] The advice received by the terminal is displayed on the in-vehicle display or played back as audio from an audio output device.
[1024] Input: Advice data sent by the server.
[1025] Output: Text advice displayed on the in-car display or audio advice played from an audio output device.
[1026] Specific action: The in-car display will display the text "Please slow down as there is a pedestrian ahead," or a similar advice will be played aloud from the speaker.
[1027] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1028] The embodiments of the present invention will be specifically described below.
[1029] System Overview
[1030] This system allows users to input audio and video data, analyzes that data, and provides specific advice. By combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[1031] Audio and video capture
[1032] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[1033] Data transmission and analysis
[1034] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1035] Use of emotion engine
[1036] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of pitch and rhythm analysis and evaluated. For example, if the user is playing in an unstable emotional state, it can determine whether this is affecting the pitch or rhythm.
[1037] Identifying areas for improvement and generating recommendations
[1038] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on the user's emotions, showing how to improve expressiveness in accordance with the user's state.
[1039] Providing advice
[1040] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1041] Specific examples
[1042] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can practice using this advice as a reference.
[1043] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills and expressiveness.
[1044] The processing flow will be explained below.
[1045] Step 1:
[1046] The user launches a dedicated application and starts recording or recording.
[1047] Press the "Start Recording" button on the application screen to begin inputting audio and video.
[1048] Step 2:
[1049] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[1050] The recorded data is temporarily stored in a memory area on the terminal.
[1051] Step 3:
[1052] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[1053] After completing the data transmission, the terminal waits for a response from the server.
[1054] Step 4:
[1055] The server receives the received audio and video data and passes it to the analysis module.
[1056] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[1057] Step 5:
[1058] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[1059] Specific algorithms are used to analyze emotional expressions and performance movements.
[1060] Step 6:
[1061] The server's emotion engine analyzes the user's emotions in real time from the input audio and video data.
[1062] The analyzed emotional information is then evaluated in combination with the results of pitch and rhythm analysis. For example, if the user is playing in an unstable emotional state, it will be determined that this is affecting the pitch and rhythm.
[1063] Step 7:
[1064] The server lists specific areas for improvement based on the analysis results of pitch, rhythm, expressiveness, etc., as well as emotional information.
[1065] This improvement highlights areas of user performance that require special attention.
[1066] Step 8:
[1067] The server generates specific advice based on the identified improvements and emotions.
[1068] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo." It also generates emotion-based advice such as, "Take deep breaths to relax when playing," depending on your emotional state.
[1069] Step 9:
[1070] The generated advice is encoded in text or video format and sent to the terminal.
[1071] The server performs the encoding process and sends the data to the device in the optimal format.
[1072] Step 10:
[1073] The advice received by the terminal is displayed on the user interface.
[1074] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[1075] Step 11:
[1076] The user checks the displayed advice and practices independently based on the specific practice method.
[1077] By following the advice and practicing repeatedly, you can improve your technique and emotional management skills.
[1078] As a concrete example, consider the case where a user sings a favorite song using a karaoke app. The user launches the app and starts recording. At this time, the device captures and saves the audio and video. The transmitted data is analyzed by the server, which detects pitch discrepancies, rhythmic inconsistencies, emotional instability, and so on. Based on this, the server generates specific advice and sends it to the device. The user can then receive this advice and practice again, effectively improving their skills.
[1079] Example 2
[1080] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1081] Conventional audio and video analysis systems, when analyzing a user's audio and video data, are unable to provide appropriate feedback based on the user's emotional state, in addition to suggesting areas for technical improvement. Therefore, in order for users to practice efficiently, they need support not only in terms of technical aspects but also mental aspects, but this has not been achieved with conventional technologies.
[1082] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1083] In this invention, the server includes means for a user to input audio and video, means for temporarily storing the input audio and video data in a storage device of the terminal, means for transmitting the stored audio and video data to the server, means for analyzing the received audio and video data, means for performing spectrum analysis on the analyzed audio data, means for analyzing facial expressions and gestures of the analyzed video data, means including an emotion engine for analyzing emotions based on the analysis results, means for identifying areas for improvement based on the analysis results and emotion information and generating specific advice, and means for providing the generated advice to the user. This allows the user to receive not only technical improvements but also appropriate feedback according to their emotional state, enabling more effective self-practice.
[1084] "User" means the end user who inputs audio and video data into the system.
[1085] "Audio and video input means" means a device or software that allows a user to input audio and video data into the system.
[1086] A "terminal" is a device that stores audio and video data and transmits it to a server. Examples of such devices include smartphones, tablets, and personal computers.
[1087] "Device storage device" refers to a storage means for temporarily storing audio and video data. This includes memory, storage devices, etc.
[1088] A "server" is a computer system that receives and analyzes audio and video data transmitted from terminals via a network.
[1089] "Means for analyzing" means algorithms or software for extracting and evaluating the technical and emotional characteristics of received audio and video data.
[1090] "Spectral analysis" is a technique for analyzing the frequency components of audio data and evaluating pitch and rhythm.
[1091] "Means for analyzing facial expressions and gestures" refers to algorithms or software for recognizing and evaluating a user's facial expressions and gestures from video data.
[1092] An "emotion engine" is an algorithm or software for analyzing a user's emotions based on audio and video data.
[1093] The "means for identifying improvements" is an algorithm or software for extracting technical and emotional improvements based on the analytical results and emotional information.
[1094] The "means for generating specific advice" is an algorithm or software for generating specific feedback and practice methods for identified areas for improvement.
[1095] "Means for providing" refers to a device or software for providing the generated advice to the user in text or video format.
[1096] MODE FOR CARRYING OUT THE INVENTION
[1097] This invention is a system that allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. This system consists of a user terminal and a server.
[1098] Audio and video capture
[1099] The user inputs sound or plays an instrument into the device. At this time, the device captures the sound and video in real time using the microphone and camera of the smartphone or PC. The user starts the dedicated application and presses the "Start Recording" button on the screen to begin recording.
[1100] Data transmission and analysis
[1101] The audio and video data is temporarily stored in the device's memory and then transmitted over the Internet to a server. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1102] Use of emotion engine
[1103] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. This analysis uses, for example, the tone and pitch of the audio and changes in facial expressions in the video. The analyzed emotional information is then combined with the results of the analysis of pitch and rhythm and evaluated. For example, if the user is playing in an unstable emotional state, it can be determined that this is affecting the pitch or rhythm.
[1104] Identifying areas for improvement and generating advice
[1105] The server identifies technical and emotional areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on emotions, showing how to improve expressiveness according to the user's state.
[1106] Providing advice
[1107] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1108] Specific examples
[1109] As a concrete example, let us consider the case where a user sings using a karaoke app. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and this data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, which analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can refer to this advice to practice efficiently.
[1110] Prompt Sentence Examples
[1111] Analyzing pitch discrepancies: "Analyze the audio data sung by the user and identify the parts where pitch discrepancies occur."
[1112] To analyze rhythmic discrepancies: "Analyze the rhythmic data played by the user and identify where there are rhythmic discrepancies."
[1113] For sentiment analysis: "Analyze the emotions from the user's audio and video data and present the results."
[1114] To generate overall advice: "Analyze the user's performance data and emotional state to identify areas for improvement and generate specific advice."
[1115] This invention allows users to practice more effectively by providing a practice method that takes into account not only their technical skills but also their emotional state.
[1116] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1117] Step 1:
[1118] The user starts a dedicated application and prepares to input audio and video. For example, they start a karaoke application and select a favorite song. The application displays a "Start Recording" button, and when the user presses that button, recording begins. The input data is audio data and video data.
[1119] Input: User's singing voice and video
[1120] Output: Audio data (WAV files, etc.), video data (MP4 files, etc.)
[1121] Step 2:
[1122] The device temporarily stores the recorded and captured audio and video data in a storage device. During this storage process, the data format is automatically set. For example, audio data is stored as a WAV file and video data is stored as an MP4 file in the smartphone's memory.
[1123] Input: Recorded and filmed audio and video data
[1124] Output: Audio and video files saved on a storage device
[1125] Step 3:
[1126] The device then transmits the stored audio and video data to a server over the Internet, using secure encryption protocols such as HTTPS.
[1127] Input: Audio and video files stored on a storage device
[1128] Output: Audio and video data sent to the server
[1129] Step 4:
[1130] The server passes the received audio and video data to the analysis module, which then analyzes the data. Specifically, the audio data undergoes spectral analysis using FFT (Fast Fourier Transform), and the video data undergoes analysis using facial expression and gesture recognition algorithms (e.g., OpenCV).
[1131] Input: Received audio and video data
[1132] Output: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[1133] Step 5:
[1134] The server uses an emotion engine to analyze the user's emotions based on the analysis results, determining the user's emotions (e.g., joy, sadness, surprise) based on the tone and pitch of the audio data and facial expressions in the video data.
[1135] Input: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[1136] Output: Emotion analysis result (e.g., user is calm, nervous)
[1137] Step 6:
[1138] The server identifies areas for improvement based on the analysis results and emotional information and generates specific advice. For example, if the pitch is low, the server generates advice such as, "Your pitch is too low in the chorus, so try singing a bit higher." It also generates emotionally appropriate advice for a nervous user, such as, "Take a deep breath to relax and play."
[1139] Input: Analysis results and emotion information
[1140] Output: Specific advice (text or video format)
[1141] Step 7:
[1142] The server then encodes the generated advice and sends it back to the terminal, again securely using an encryption protocol.
[1143] Input: Specific advice
[1144] Output: Advice data sent to the terminal
[1145] Step 8:
[1146] The device displays the received advice on its user interface. The advice can be in the form of text such as "Your pitch is too low in the chorus, try singing a bit higher," or it can be in the form of a video clip.
[1147] Input: Advice data sent to the terminal
[1148] Output: Specific advice displayed in the user interface
[1149] The specific operations and inputs and outputs at each step have been described above.
[1150] (Application example 2)
[1151] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1152] Conventional audio and video analysis systems generate feedback based on the results of technical analysis of music. However, because they do not take into account the user's emotional state, many users are dissatisfied with technical instruction alone. In particular, when emotions significantly affect a user's playing or singing, the feedback is incomplete, resulting in reduced practice efficiency. To address these issues, the present invention aims to provide customized advice that takes into account the user's emotional state by analyzing the user's emotional state in real time.
[1153] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input audio or video, means for receiving and analyzing the input audio or video data, means for identifying areas for improvement based on the analysis results and generating specific advice, means for providing the generated advice to the user, means for analyzing the user's emotions, and means for generating individually customized advice based on the analyzed emotional information. This not only improves the user's technique, but also enables the user to receive feedback according to their emotional state, resulting in more effective practice.
[1154] "User" refers to an individual or organization that uses the system to input audio or video data.
[1155] "Voice data" refers to voice or music data that users input into the system.
[1156] "Video data" refers to video and image data that users input into the system.
[1157] "Analysis" refers to the process of evaluating and analyzing the technical characteristics and emotional information of audio and video data.
[1158] "Server" refers to the computer system that receives and processes data for analysis and generates feedback.
[1159] "Areas for improvement" refers to points identified from the analyzed data that the user can use to improve their technical or expressive abilities.
[1160] "Advice" refers to feedback provided based on the analysis results, including specific ways to improve or practice.
[1161] "Emotion analysis" refers to the process of assessing a user's emotional state from audio and video data.
[1162] "Customized advice" refers to advice provided based on the user's individual state, based on analyzed emotional information.
[1163] We will now explain a specific system for implementing this invention. This system allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[1164] System Overview
[1165] The system is outlined as follows: Users input audio and video using the microphone and camera on their smartphone or PC. The input data is temporarily stored in the device's memory and then sent to a server via the Internet. The server passes the received data to an analysis module, which evaluates the technical features. It then uses an emotion engine to analyze the user's emotional state and generates specific advice based on the analysis results.
[1166] Audio and video capture
[1167] Users can input audio or play an instrument into the device, and the microphone and camera on their smartphone or PC will capture the audio and video data in real time. Users simply launch the dedicated application and press the "Start Recording" button on the screen to begin recording.
[1168] Data transmission and analysis
[1169] The captured audio and video data is temporarily stored on the device's memory and then transmitted over the Internet to a server using an HTTP library such as "requests." The server then analyzes the received data using software such as "MusicAnalyzer" or "EmotionEngine" to evaluate technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1170] Use of emotion engine
[1171] The server is equipped with an "Emotion Engine" that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of technical analysis and evaluated. For example, if a user is performing in an unstable emotional state, it can be determined that this is affecting pitch and rhythm deviations.
[1172] Identifying areas for improvement and generating recommendations
[1173] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state on these. For identified areas for improvement, the server uses a "generative AI model" to generate specific advice. For example, the server inputs prompts such as "Please suggest practice methods for when my emotional state is unstable" or "Please tell me specific ways to improve when I'm out of tune" into the AI model, and obtains results.
[1174] Providing advice
[1175] The generated advice is encoded in text or video format and sent back to the device. The device receives the "prompt sentence" and displays it on the user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a bit higher" may be displayed on the smartphone screen. Other advice may include "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1176] Specific examples
[1177] Let's consider a scenario where a user sings using an interactive music learning app on a smartphone. First, the user launches the app, selects a favorite song, and sings. The smartphone's microphone records the user's singing, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet and analyzed by "MusicAnalyzer" and "EmotionEngine." Based on the analysis results, the server detects pitch discrepancies and rhythmic inconsistencies and generates specific advice. This advice is sent to the smartphone and displayed on the screen. The user can use this advice as a reference for practicing.
[1178] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1179] Step 1:
[1180] The user launches a dedicated application on their smartphone or PC and presses the start button for recording. As the user inputs sound or plays an instrument, the microphone and camera on the smartphone or PC captures the audio and video data in real time. The input audio and video data is temporarily stored in the device's memory device.
[1181] Input: User's audio and video data
[1182] Output: Audio and video data stored in the device's storage device
[1183] Step 2:
[1184] The device sends the stored audio and video data to the server by uploading the data using an HTTP request and passing it to the server.
[1185] Input: Audio and video data stored on the device
[1186] Output: Audio and video data sent to the server
[1187] Step 3:
[1188] The server analyzes the received audio and video data. First, it uses MusicAnalyzer to evaluate technical characteristics (pitch, rhythm, and expressiveness). Spectral analysis is performed on the audio data to detect deviations in pitch and rhythm. Facial expressions and gestures are analyzed on the video data.
[1189] Input: Audio and video data sent to the server
[1190] Output: Evaluation results of technical features (pitch, rhythm, expressiveness, facial expressions, gestures)
[1191] Step 4:
[1192] The server then uses the "Emotion Engine" to perform emotion analysis, assessing the user's emotional state from audio and video data and combining this with the evaluation of the technical features.
[1193] Input: Technical characteristics evaluation results, audio and video data
[1194] Output: Integrated evaluation results of emotional state and technical features
[1195] Step 5:
[1196] The server identifies areas for improvement based on the analysis results. It comprehensively evaluates technical features and emotional information and generates specific advice. Using a "generative AI model," it inputs prompts into the model to generate individually customized advice. For example, it uses prompts such as, "Please suggest practice methods for when my emotional state is unstable."
[1197] Input: Integrated evaluation results, prompts to the generative AI model
[1198] Output: Personalized, specific advice
[1199] Step 6:
[1200] The server encodes the generated advice in text or video format and sends it back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1201] Input: Specific, personalized advice
[1202] Output: Feedback displayed on the device (text or video format)
[1203] Through these steps, users can not only improve their technique, but also receive feedback based on their emotional state, enabling more effective practice.
[1204] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1205] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1206] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1207] [Fourth embodiment]
[1208] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1209] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1210] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1211] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1212] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1213] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1214] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1215] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1216] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1217] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1218] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1219] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1220] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1221] The embodiments of the present invention will be specifically described below.
[1222] System Overview
[1223] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[1224] Audio and video capture
[1225] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[1226] Data transmission and analysis
[1227] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1228] Identifying areas for improvement and generating recommendations
[1229] The server identifies areas for improvement based on the analysis results. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how.
[1230] Providing advice
[1231] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a little higher," or video advice such as, "Use a metronome to practice maintaining a steady tempo."
[1232] Specific examples
[1233] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[1234] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[1235] The processing flow will be explained below.
[1236] Step 1:
[1237] The user launches a dedicated application and starts recording or recording.
[1238] Pressing the "Start Recording" button on the application screen will begin audio and video input.
[1239] Step 2:
[1240] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[1241] The recorded data is temporarily stored in a memory area on the terminal.
[1242] Step 3:
[1243] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[1244] After completing the data transmission, the terminal waits for a response from the server.
[1245] Step 4:
[1246] The server receives the received audio and video data and passes it to the analysis module.
[1247] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[1248] Step 5:
[1249] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[1250] Specific algorithms are used to analyze emotional expressions and performance movements.
[1251] Step 6:
[1252] The server lists specific areas for improvement based on the analysis of pitch, rhythm, expressiveness, etc.
[1253] This improvement highlights areas of user performance that require special attention.
[1254] Step 7:
[1255] The server generates specific advice for the identified improvements.
[1256] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1257] Step 8:
[1258] The generated advice is encoded in text or video format and sent to the terminal.
[1259] The server performs the encoding process and sends the data to the device in the optimal format.
[1260] Step 9:
[1261] The advice received by the terminal is displayed on the user interface.
[1262] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[1263] Step 10:
[1264] The user checks the displayed advice and practices independently based on the specific practice method.
[1265] By following the advice and practicing repeatedly, you can improve your skills.
[1266] Example 1
[1267] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1268] In the past, when practicing music or improving performance, it was difficult for users to objectively evaluate their own performance and identify specific areas for improvement. Furthermore, due to limited opportunities to receive professional instruction, it was difficult to find an efficient practice method.
[1269] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1270] In this invention, the server includes: means for a user to input audio or video; means for temporarily storing the input audio or video data; means for transmitting the stored data to the server; means for receiving and analyzing the transmitted audio or video data; means for performing spectral analysis on the audio data to evaluate pitch and rhythm; means for analyzing facial expressions and gestures on the video data; means for identifying pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness based on the analysis results; means including a generative AI model for generating specific advice for identified areas for improvement; and means for providing the generated advice to the user. This allows the user to objectively evaluate their own performance and identify specific areas for improvement.
[1271] "User" refers to an individual who uses the system to input audio or video and receive the resulting advice.
[1272] "Means for inputting audio or video" refers to a device such as a microphone or camera that is used to input the user's singing voice, playing, movements, etc. into the system.
[1273] The "means for temporarily storing audio or video data" refers to a storage device or memory for temporarily saving captured audio or video data.
[1274] The "means for transmitting data to a server" refers to communication techniques and protocols for transmitting temporarily stored data to a remote server over a network such as the Internet.
[1275] "Means for receiving and analyzing data" refers to software or algorithms that allow the server to receive transmitted audio or video data and analyze that data.
[1276] "Spectral analysis" is an analytical technique for extracting frequency components from audio data and evaluating pitch and rhythm.
[1277] The "means for evaluating pitch and rhythm" is a technology for evaluating the accuracy of pitch and the degree of rhythmic agreement from the results of spectrum analysis of audio data.
[1278] "Means for analyzing facial expressions and gestures" refers to image analysis technology for recognizing and analyzing a user's facial expressions and body movements based on video data.
[1279] "Means for identifying areas for improvement based on analysis results" refers to technology that identifies areas that need improvement based on discrepancies or deficiencies detected from analyzed audio and video data.
[1280] A "generative AI model that generates specific advice for identified improvement points" is an artificial intelligence model that automatically generates specific advice, such as how users should practice, for detected improvement points.
[1281] The "means for providing the generated advice to the user" is a technology for encoding the generated advice in text or video format, transmitting it to the user's terminal, and displaying it.
[1282] MODE FOR CARRYING OUT THE INVENTION
[1283] The embodiments of the present invention will be specifically described below.
[1284] System Overview
[1285] This system allows users to input audio and video data, analyzes the data, and provides specific advice. The system consists of a user terminal and a server.
[1286] Audio and video capture
[1287] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[1288] Data transmission and temporary storage
[1289] The audio and video data is temporarily stored on the device's memory and then sent to a server over the Internet. The device uploads the data to the server using Wi-Fi or mobile networks.
[1290] Data analysis
[1291] The server then passes the received data to an analysis module that evaluates technical characteristics such as pitch, rhythm, and expressiveness. This analysis includes spectral analysis of the audio data and facial and gesture analysis of the video data. The server then uses algorithms to detect specific patterns and deviations in the analyzed data.
[1292] Identifying areas for improvement and generating recommendations
[1293] The server identifies areas for improvement based on the analysis results, such as pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. For each identified area for improvement, the server uses a generative AI model to generate advice, including specific practice methods and points to note.
[1294] Providing advice
[1295] The generated advice is encoded in text or video format and sent to the device. The device then displays the received advice on its user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a little higher" or video advice such as "Use a metronome to practice maintaining a steady tempo" are displayed on the smartphone screen.
[1296] Specific examples
[1297] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. The user can practice using this advice as a reference.
[1298] Prompt Sentence Examples
[1299] "Based on a recording of my singing using a karaoke app, please analyze any inconsistencies in pitch or rhythm and suggest specific practice methods."
[1300] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills.
[1301] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1302] Step 1:
[1303] User inputs audio and video data
[1304] Users start a dedicated application on their smartphone or PC and press the "Start Recording" button on the screen to input audio and video data. As users sing or play an instrument, the device's microphone and camera capture this data in real time.
[1305] Input: Real-time audio and video data of singing and performance
[1306] Output: Audio and video data temporarily stored in a storage device
[1307] Step 2:
[1308] The device temporarily stores data
[1309] The device temporarily stores the captured audio and video data in a storage device (e.g., a smartphone's internal storage). After recording is complete, press the "Start Analysis" button to complete the data storage.
[1310] Input: Real-time audio and video data
[1311] Output: Temporarily saved data (audio files, video files)
[1312] Step 3:
[1313] The device sends data to the server
[1314] The device transmits the stored audio and video data to a server over the internet, using Wi-Fi or mobile networks.
[1315] Specifically, the device automatically uploads data after the user presses the "Start Analysis" button.
[1316] Input: Temporarily saved data (audio files, video files)
[1317] Output: Audio and video data sent to the server
[1318] Step 4:
[1319] The server receives and analyzes the data
[1320] The server passes the received audio and video data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1321] Input: Audio and video data sent to the server
[1322] Output: Analyzed technical features (pitch, rhythm, facial expressions, gestures, etc.)
[1323] Step 5:
[1324] The server identifies areas for improvement and generates specific advice
[1325] Based on the analysis results, the server identifies pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness. Based on the identified areas for improvement, the server uses a generative AI model to generate specific advice, including which areas should be improved and how.
[1326] Input: Parsed technical features
[1327] Output: Specific advice generated
[1328] Step 6:
[1329] The server sends the generated advice to the device.
[1330] The generated advice is encoded in text or video format and sent to the device.
[1331] Input: Generated specific advice
[1332] Output: Text or video advice sent to your device
[1333] Step 7:
[1334] The device displays advice to the user
[1335] The device displays the received advice on the user interface, and the user can view the advice as text or video on the screen of their smartphone or PC.
[1336] Specific actions include instructions such as "Your pitch is too low in the chorus, so try singing a bit higher," and video advice such as "Use a metronome to practice maintaining a steady tempo."
[1337] Input: Text or video advice sent to your device
[1338] Output: Advice displayed to the user
[1339] (Application example 1)
[1340] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1341] To improve the safety of autonomous vehicles, it is essential to analyze the driving situation in real time and provide appropriate advice based on the results. However, current systems do not efficiently analyze audio and video data while driving and provide specific and timely advice based on that analysis, which means that safety is not sufficiently improved. Another issue is the lack of technology to generate advice that accurately reflects the driving situation.
[1342] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1343] In this invention, the server includes a means for a user to input audio or video, a means for receiving and analyzing the input audio or video data, a means for generating advice on safe driving from the analysis results, and a means for displaying the generated advice on an in-vehicle display or an audio output device in real time, thereby enabling the provision of specific and appropriate advice on safe driving in accordance with the driving situation in real time.
[1344] "User" means a person who uses the system to input audio and video data.
[1345] A "means" is a device or part of a system used to achieve a specified purpose.
[1346] "Voice data" means a digital representation of the voice input by a user.
[1347] "Video data" refers to a digital representation of a video input by a user.
[1348] A "server" is a computer system that analyzes received audio and video data and generates advice based on the results of that analysis.
[1349] "Analysis" refers to the evaluation and detection of technical characteristics of audio and video data.
[1350] "Advice" is specific instructions or recommendations generated by the server based on the analysis results.
[1351] An "in-vehicle display" is a screen display device installed inside an autonomous vehicle.
[1352] The "audio output device" is a device that conveys the generated advice to the user as audio.
[1353] "Safe driving" refers to an autonomous vehicle operating appropriately and safely in accordance with the surrounding conditions.
[1354] "Real-time" means that data is collected, analyzed, and results are provided without delay.
[1355] "Driving conditions" refers to information about vehicle operation, including the vehicle's current speed, location, surrounding traffic conditions, and the like.
[1356] "Obstacle" refers to an object that is in the vehicle's driving path and may affect driving.
[1357] The embodiments of the present invention will be specifically described below.
[1358] System Overview
[1359] This system allows users to input audio and video data, analyzes the data, and provides specific advice on safe driving. The system consists of an in-vehicle terminal used by the user and a server that performs the analysis.
[1360] Audio and video capture
[1361] The user inputs audio and video into the vehicle's microphone and camera, and the device captures the audio and video data in real time. This is done using the vehicle's camera and microphone. The user can start audio input or video capture using, for example, a voice command.
[1362] Data transmission and analysis
[1363] The captured audio and video data is temporarily stored in the in-vehicle terminal's storage device and then sent to a server via the Internet. The server receives the data and passes it to an analysis module. The audio data is analyzed using speech recognition technology, and the video data is analyzed using computer vision technology. Specific software used includes the speech_recognition library for speech recognition and OpenCV for video analysis.
[1364] Identifying areas for improvement and generating recommendations
[1365] The server identifies areas for improvement in driving based on the analysis results. For example, it detects when there is an obstacle ahead or when the speed exceeds the speed limit. Based on these analysis results, the server generates specific advice. This advice includes specific instructions on what to improve and how to improve it. The optimal advice is automatically generated using a generative AI model.
[1366] Providing advice
[1367] The generated advice is encoded in text or audio format and sent back to the terminal, where it is provided to the user in real time via the in-car display or audio output device. For example, the in-car display might say "Obstacle ahead. Please reduce speed" or a similar instruction might be played back via audio.
[1368] Program processing
[1369] The server converts the received audio data into text using the speech_recognition library.
[1370] Video data acquired by the onboard camera is analyzed using the OpenCV library to detect obstacles and other vehicles.
[1371] Based on the analysis results, the generative AI model generates appropriate driving advice.
[1372] The generated advice is displayed in text format on an in-vehicle display or output in audio format from the car's speakers.
[1373] Specific examples
[1374] As a specific example, consider the case where a user asks a question into the car's microphone while driving, "What's the situation ahead?" This voice data is captured in real time and sent to a server for analysis. The server converts the voice into text and analyzes the video data about the situation ahead. Based on the analysis results, a generative AI model generates advice such as "There is a pedestrian ahead, please slow down." This advice is displayed as text on the in-car display or communicated aloud through a voice output device.
[1375] Prompt Sentence Examples
[1376] Input: In-car voice "What's the situation ahead?", external camera footage
[1377] Output: "Slow down, pedestrians ahead."
[1378] As described above, this system can improve the safety of self-driving vehicles by analyzing audio and video data in real time while the user is driving and providing appropriate driving advice.
[1379] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1380] Step 1:
[1381] Users input audio and video into the microphone and camera inside the car.
[1382] Input: User's voice command "What's the situation ahead?" and video data from inside and outside the vehicle.
[1383] Output: Audio and video data temporarily stored on the device.
[1384] How it works: The user asks a question into the car's microphone, which is recorded in real time. At the same time, the car's camera captures images of the interior and exterior of the car.
[1385] Step 2:
[1386] The audio and video data acquired by the terminal is temporarily stored in a storage device.
[1387] Input: Raw audio and video data captured by the device.
[1388] Output: Temporarily stored audio and video data.
[1389] Specific operation: Save the data in the storage device of the in-vehicle terminal (for example, the hard disk or flash memory of the in-vehicle computer).
[1390] Step 3:
[1391] The device transmits the stored data to a server via the Internet.
[1392] Input: Temporarily stored audio and video data.
[1393] Output: Audio and video data sent to the server.
[1394] Specific operation: The saved data is sent to the server using an HTTP request, etc. An internet connection is required for this.
[1395] Step 4:
[1396] The server analyzes the voice data and converts it into text.
[1397] Input: The transmitted audio data.
[1398] Output: Text data converted from audio data.
[1399] Specific operation: The server uses the speech_recognition library to analyze the audio data and perform speech recognition. For example, the audio "What's the situation ahead?" is converted into the text "What's the situation ahead?"
[1400] Step 5:
[1401] The server analyzes the video data and detects obstacles and other vehicles.
[1402] Input: Transmitted video data.
[1403] Output: Obstacle and vehicle position information based on the analyzed video data.
[1404] Specific operation: The server uses the OpenCV library to analyze the video data and, for example, detect the presence of a pedestrian ahead.
[1405] Step 6:
[1406] The server integrates the results of voice recognition and video analysis and generates specific advice using a generative AI model.
[1407] Input: Text data of speech recognition results and location information of video analysis results.
[1408] Output: Specific advice on safe driving.
[1409] Specific operation: The server uses a generative AI model to automatically generate advice such as "Please slow down as there is a pedestrian ahead."
[1410] Step 7:
[1411] The server transmits the generated advice to the vehicle-mounted terminal.
[1412] Input: The generated advice.
[1413] Output: Advice data sent to the in-vehicle terminal.
[1414] Specific operation: The server sends the generated advice to the in-vehicle terminal using an HTTP request or the like.
[1415] Step 8:
[1416] The advice received by the terminal is displayed on the in-vehicle display or played back as audio from an audio output device.
[1417] Input: Advice data sent by the server.
[1418] Output: Text advice displayed on the in-car display or audio advice played from an audio output device.
[1419] Specific action: The in-car display will display the text "Please slow down as there is a pedestrian ahead," or a similar advice will be played aloud from the speaker.
[1420] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1421] The embodiments of the present invention will be specifically described below.
[1422] System Overview
[1423] This system allows users to input audio and video data, analyzes that data, and provides specific advice. By combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[1424] Audio and video capture
[1425] When a user speaks into the device or plays an instrument, the device captures the audio and video in real time. This is done using the microphone and camera of a smartphone or PC. The user simply launches a dedicated application and presses the "Start Recording" button on the screen to begin recording.
[1426] Data transmission and analysis
[1427] The audio and video data is temporarily stored in the device's memory and then transmitted to a server via the Internet. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1428] Use of emotion engine
[1429] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of pitch and rhythm analysis and evaluated. For example, if the user is playing in an unstable emotional state, it can determine whether this is affecting the pitch or rhythm.
[1430] Identifying areas for improvement and generating recommendations
[1431] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on the user's emotions, showing how to improve expressiveness in accordance with the user's state.
[1432] Providing advice
[1433] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1434] Specific examples
[1435] As an example of a specific embodiment, a case where a user sings using a karaoke app will be described. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, and the server analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can practice using this advice as a reference.
[1436] The above is a specific embodiment for carrying out the present invention. This system allows users to practice on their own efficiently and improve their musical skills and expressiveness.
[1437] The processing flow will be explained below.
[1438] Step 1:
[1439] The user launches a dedicated application and starts recording or recording.
[1440] Press the "Start Recording" button on the application screen to begin inputting audio and video.
[1441] Step 2:
[1442] The device uses its built-in microphone and camera to capture the user's voice and video in real time and generate recorded data.
[1443] The recorded data is temporarily stored in a memory area on the terminal.
[1444] Step 3:
[1445] The device compresses the recorded audio and video data and transmits it to a server using a secure communication protocol.
[1446] After completing the data transmission, the terminal waits for a response from the server.
[1447] Step 4:
[1448] The server receives the received audio and video data and passes it to the analysis module.
[1449] The analysis module performs spectral analysis on the audio data to identify discrepancies in pitch and rhythm.
[1450] Step 5:
[1451] The server analyzes the video data and evaluates the user's facial expressions and gestures.
[1452] Specific algorithms are used to analyze emotional expressions and performance movements.
[1453] Step 6:
[1454] The server's emotion engine analyzes the user's emotions in real time from the input audio and video data.
[1455] The analyzed emotional information is then evaluated in combination with the results of pitch and rhythm analysis. For example, if the user is playing in an unstable emotional state, it will be determined that this is affecting the pitch and rhythm.
[1456] Step 7:
[1457] The server lists specific areas for improvement based on the analysis results of pitch, rhythm, expressiveness, etc., as well as emotional information.
[1458] This improvement highlights areas of user performance that require special attention.
[1459] Step 8:
[1460] The server generates specific advice based on the identified improvements and emotions.
[1461] For example, it generates advice such as, "Your pitch is too low, so try singing a bit higher," or, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo." It also generates emotion-based advice such as, "Take deep breaths to relax when playing," depending on your emotional state.
[1462] Step 9:
[1463] The generated advice is encoded in text or video format and sent to the terminal.
[1464] The server performs the encoding process and sends the data to the device in the optimal format.
[1465] Step 10:
[1466] The advice received by the terminal is displayed on the user interface.
[1467] For example, text advice will be displayed on the smartphone screen, with relevant video links provided where necessary.
[1468] Step 11:
[1469] The user checks the displayed advice and practices independently based on the specific practice method.
[1470] By following the advice and practicing repeatedly, you can improve your technique and emotional management skills.
[1471] As a concrete example, consider the case where a user sings a favorite song using a karaoke app. The user launches the app and starts recording. At this time, the device captures and saves the audio and video. The transmitted data is analyzed by the server, which detects pitch discrepancies, rhythmic inconsistencies, emotional instability, and so on. Based on this, the server generates specific advice and sends it to the device. The user can then receive this advice and practice again, effectively improving their skills.
[1472] Example 2
[1473] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1474] Conventional audio and video analysis systems, when analyzing a user's audio and video data, are unable to provide appropriate feedback based on the user's emotional state, in addition to suggesting areas for technical improvement. Therefore, in order for users to practice efficiently, they need support not only in terms of technical aspects but also mental aspects, but this has not been achieved with conventional technologies.
[1475] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1476] In this invention, the server includes means for a user to input audio and video, means for temporarily storing the input audio and video data in a storage device of the terminal, means for transmitting the stored audio and video data to the server, means for analyzing the received audio and video data, means for performing spectrum analysis on the analyzed audio data, means for analyzing facial expressions and gestures of the analyzed video data, means including an emotion engine for analyzing emotions based on the analysis results, means for identifying areas for improvement based on the analysis results and emotion information and generating specific advice, and means for providing the generated advice to the user. This allows the user to receive not only technical improvements but also appropriate feedback according to their emotional state, enabling more effective self-practice.
[1477] "User" means the end user who inputs audio and video data into the system.
[1478] "Audio and video input means" means a device or software that allows a user to input audio and video data into the system.
[1479] A "terminal" is a device that stores audio and video data and transmits it to a server. Examples of such devices include smartphones, tablets, and personal computers.
[1480] "Device storage device" refers to a storage means for temporarily storing audio and video data. This includes memory, storage devices, etc.
[1481] A "server" is a computer system that receives and analyzes audio and video data transmitted from terminals via a network.
[1482] "Means for analyzing" means algorithms or software for extracting and evaluating the technical and emotional characteristics of received audio and video data.
[1483] "Spectral analysis" is a technique for analyzing the frequency components of audio data and evaluating pitch and rhythm.
[1484] "Means for analyzing facial expressions and gestures" refers to algorithms or software for recognizing and evaluating a user's facial expressions and gestures from video data.
[1485] An "emotion engine" is an algorithm or software for analyzing a user's emotions based on audio and video data.
[1486] The "means for identifying improvements" is an algorithm or software for extracting technical and emotional improvements based on the analytical results and emotional information.
[1487] The "means for generating specific advice" is an algorithm or software for generating specific feedback and practice methods for identified areas for improvement.
[1488] "Means for providing" refers to a device or software for providing the generated advice to the user in text or video format.
[1489] MODE FOR CARRYING OUT THE INVENTION
[1490] This invention is a system that allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. This system consists of a user terminal and a server.
[1491] Audio and video capture
[1492] The user inputs sound or plays an instrument into the device. At this time, the device captures the sound and video in real time using the microphone and camera of the smartphone or PC. The user starts the dedicated application and presses the "Start Recording" button on the screen to begin recording.
[1493] Data transmission and analysis
[1494] The audio and video data is temporarily stored in the device's memory and then transmitted over the Internet to a server. The server then passes the data to an analysis module, which evaluates technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1495] Use of emotion engine
[1496] The server is equipped with an emotion engine that analyzes the user's emotions in real time based on the input audio and video data. This analysis uses, for example, the tone and pitch of the audio and changes in facial expressions in the video. The analyzed emotional information is then combined with the results of the analysis of pitch and rhythm and evaluated. For example, if the user is playing in an unstable emotional state, it can be determined that this is affecting the pitch or rhythm.
[1497] Identifying areas for improvement and generating advice
[1498] The server identifies technical and emotional areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state. The server then generates specific advice for the identified areas for improvement. This advice includes specific practice methods and points to note, such as which areas should be improved and how. Furthermore, it generates advice based on emotions, showing how to improve expressiveness according to the user's state.
[1499] Providing advice
[1500] The generated advice is encoded in text or video format and sent back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or video advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1501] Specific examples
[1502] As a concrete example, let us consider the case where a user sings using a karaoke app. First, the user launches the karaoke app on their smartphone, selects a favorite song, and sings. The smartphone's microphone records the user's singing voice, and this data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet, which analyzes the received data. Based on the analysis results, the server detects pitch discrepancies and rhythm discrepancies and generates specific advice. The generated advice is sent to the smartphone and displayed on the smartphone screen. Furthermore, if the user's emotional state is unstable, appropriate advice is also provided. The user can refer to this advice to practice efficiently.
[1503] Prompt Sentence Examples
[1504] Analyzing pitch discrepancies: "Analyze the audio data sung by the user and identify the parts where pitch discrepancies occur."
[1505] To analyze rhythmic discrepancies: "Analyze the rhythmic data played by the user and identify where there are rhythmic discrepancies."
[1506] For sentiment analysis: "Analyze the emotions from the user's audio and video data and present the results."
[1507] To generate overall advice: "Analyze the user's performance data and emotional state to identify areas for improvement and generate specific advice."
[1508] This invention allows users to practice more effectively by providing a practice method that takes into account not only their technical skills but also their emotional state.
[1509] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1510] Step 1:
[1511] The user starts a dedicated application and prepares to input audio and video. For example, they start a karaoke application and select a favorite song. The application displays a "Start Recording" button, and when the user presses that button, recording begins. The input data is audio data and video data.
[1512] Input: User's singing voice and video
[1513] Output: Audio data (WAV files, etc.), video data (MP4 files, etc.)
[1514] Step 2:
[1515] The device temporarily stores the recorded and captured audio and video data in a storage device. During this storage process, the data format is automatically set. For example, audio data is stored as a WAV file and video data is stored as an MP4 file in the smartphone's memory.
[1516] Input: Recorded and filmed audio and video data
[1517] Output: Audio and video files saved on a storage device
[1518] Step 3:
[1519] The device then transmits the stored audio and video data to a server over the Internet, using secure encryption protocols such as HTTPS.
[1520] Input: Audio and video files stored on a storage device
[1521] Output: Audio and video data sent to the server
[1522] Step 4:
[1523] The server passes the received audio and video data to the analysis module, which then analyzes the data. Specifically, the audio data undergoes spectral analysis using FFT (Fast Fourier Transform), and the video data undergoes analysis using facial expression and gesture recognition algorithms (e.g., OpenCV).
[1524] Input: Received audio and video data
[1525] Output: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[1526] Step 5:
[1527] The server uses an emotion engine to analyze the user's emotions based on the analysis results, determining the user's emotions (e.g., joy, sadness, surprise) based on the tone and pitch of the audio data and facial expressions in the video data.
[1528] Input: Analysis results of pitch, rhythm, facial expressions, gestures, etc.
[1529] Output: Emotion analysis result (e.g., user is calm, nervous)
[1530] Step 6:
[1531] The server identifies areas for improvement based on the analysis results and emotional information and generates specific advice. For example, if the pitch is low, the server generates advice such as, "Your pitch is too low in the chorus, so try singing a bit higher." It also generates emotionally appropriate advice for a nervous user, such as, "Take a deep breath to relax and play."
[1532] Input: Analysis results and emotion information
[1533] Output: Specific advice (text or video format)
[1534] Step 7:
[1535] The server then encodes the generated advice and sends it back to the terminal, again securely using an encryption protocol.
[1536] Input: Specific advice
[1537] Output: Advice data sent to the terminal
[1538] Step 8:
[1539] The device displays the received advice on its user interface. The advice can be in the form of text such as "Your pitch is too low in the chorus, try singing a bit higher," or it can be in the form of a video clip.
[1540] Input: Advice data sent to the terminal
[1541] Output: Specific advice displayed in the user interface
[1542] The specific operations and inputs and outputs at each step have been described above.
[1543] (Application example 2)
[1544] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1545] Conventional audio and video analysis systems generate feedback based on the results of technical analysis of music. However, because they do not take into account the user's emotional state, many users are dissatisfied with technical instruction alone. In particular, when emotions significantly affect a user's playing or singing, the feedback is incomplete, resulting in reduced practice efficiency. To address these issues, the present invention aims to provide customized advice that takes into account the user's emotional state by analyzing the user's emotional state in real time.
[1546] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input audio or video, means for receiving and analyzing the input audio or video data, means for identifying areas for improvement based on the analysis results and generating specific advice, means for providing the generated advice to the user, means for analyzing the user's emotions, and means for generating individually customized advice based on the analyzed emotional information. This not only improves the user's technique, but also enables the user to receive feedback according to their emotional state, resulting in more effective practice.
[1547] "User" refers to an individual or organization that uses the system to input audio or video data.
[1548] "Voice data" refers to voice or music data that users input into the system.
[1549] "Video data" refers to video and image data that users input into the system.
[1550] "Analysis" refers to the process of evaluating and analyzing the technical characteristics and emotional information of audio and video data.
[1551] "Server" refers to the computer system that receives and processes data for analysis and generates feedback.
[1552] "Areas for improvement" refers to points identified from the analyzed data that the user can use to improve their technical or expressive abilities.
[1553] "Advice" refers to feedback provided based on the analysis results, including specific ways to improve or practice.
[1554] "Emotion analysis" refers to the process of assessing a user's emotional state from audio and video data.
[1555] "Customized advice" refers to advice provided based on the user's individual state, based on analyzed emotional information.
[1556] We will now explain a specific system for implementing this invention. This system allows users to input audio and video data, analyzes that data, and provides specific advice. Furthermore, by combining it with an emotion engine, it is possible to analyze the user's emotions and provide more appropriate feedback. The system consists of a user terminal and a server.
[1557] System Overview
[1558] The system is outlined as follows: Users input audio and video using the microphone and camera on their smartphone or PC. The input data is temporarily stored in the device's memory and then sent to a server via the Internet. The server passes the received data to an analysis module, which evaluates the technical features. It then uses an emotion engine to analyze the user's emotional state and generates specific advice based on the analysis results.
[1559] Audio and video capture
[1560] Users can input audio or play an instrument into the device, and the microphone and camera on their smartphone or PC will capture the audio and video data in real time. Users simply launch the dedicated application and press the "Start Recording" button on the screen to begin recording.
[1561] Data transmission and analysis
[1562] The captured audio and video data is temporarily stored on the device's memory and then transmitted over the Internet to a server using an HTTP library such as "requests." The server then analyzes the received data using software such as "MusicAnalyzer" or "EmotionEngine" to evaluate technical characteristics such as pitch, rhythm, and expressiveness. Spectral analysis is performed on the audio data, and facial expressions and gestures are analyzed on the video data.
[1563] Use of emotion engine
[1564] The server is equipped with an "Emotion Engine" that analyzes the user's emotions in real time based on the input audio and video data. The analyzed emotional information is then combined with the results of technical analysis and evaluated. For example, if a user is performing in an unstable emotional state, it can be determined that this is affecting pitch and rhythm deviations.
[1565] Identifying areas for improvement and generating recommendations
[1566] The server identifies areas for improvement based on the analysis results and emotional information. For example, it detects pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness, while also taking into account the impact of the user's emotional state on these. For identified areas for improvement, the server uses a "generative AI model" to generate specific advice. For example, the server inputs prompts such as "Please suggest practice methods for when my emotional state is unstable" or "Please tell me specific ways to improve when I'm out of tune" into the AI model, and obtains results.
[1567] Providing advice
[1568] The generated advice is encoded in text or video format and sent back to the device. The device receives the "prompt sentence" and displays it on the user interface. For example, specific instructions such as "Your pitch is too low in the chorus, so try singing a bit higher" may be displayed on the smartphone screen. Other advice may include "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1569] Specific examples
[1570] Let's consider a scenario where a user sings using an interactive music learning app on a smartphone. First, the user launches the app, selects a favorite song, and sings. The smartphone's microphone records the user's singing, and the data is temporarily saved on the smartphone. This data is then uploaded to a server via the Internet and analyzed by "MusicAnalyzer" and "EmotionEngine." Based on the analysis results, the server detects pitch discrepancies and rhythmic inconsistencies and generates specific advice. This advice is sent to the smartphone and displayed on the screen. The user can use this advice as a reference for practicing.
[1571] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1572] Step 1:
[1573] The user launches a dedicated application on their smartphone or PC and presses the start button for recording. As the user inputs sound or plays an instrument, the microphone and camera on the smartphone or PC captures the audio and video data in real time. The input audio and video data is temporarily stored in the device's memory device.
[1574] Input: User's audio and video data
[1575] Output: Audio and video data stored in the device's storage device
[1576] Step 2:
[1577] The device sends the stored audio and video data to the server by uploading the data using an HTTP request and passing it to the server.
[1578] Input: Audio and video data stored on the device
[1579] Output: Audio and video data sent to the server
[1580] Step 3:
[1581] The server analyzes the received audio and video data. First, it uses MusicAnalyzer to evaluate technical characteristics (pitch, rhythm, and expressiveness). Spectral analysis is performed on the audio data to detect deviations in pitch and rhythm. Facial expressions and gestures are analyzed on the video data.
[1582] Input: Audio and video data sent to the server
[1583] Output: Evaluation results of technical features (pitch, rhythm, expressiveness, facial expressions, gestures)
[1584] Step 4:
[1585] The server then uses the "Emotion Engine" to perform emotion analysis, assessing the user's emotional state from audio and video data and combining this with the evaluation of the technical features.
[1586] Input: Technical characteristics evaluation results, audio and video data
[1587] Output: Integrated evaluation results of emotional state and technical features
[1588] Step 5:
[1589] The server identifies areas for improvement based on the analysis results. It comprehensively evaluates technical features and emotional information and generates specific advice. Using a "generative AI model," it inputs prompts into the model to generate individually customized advice. For example, it uses prompts such as, "Please suggest practice methods for when my emotional state is unstable."
[1590] Input: Integrated evaluation results, prompts to the generative AI model
[1591] Output: Personalized, specific advice
[1592] Step 6:
[1593] The server encodes the generated advice in text or video format and sends it back to the device. The device then displays the received advice on its user interface. For example, the smartphone screen might show specific instructions such as, "Your pitch is too low in the chorus, so try singing a bit higher," or advice such as, "Your rhythm is too fast, so use a metronome to practice maintaining a steady tempo."
[1594] Input: Specific, personalized advice
[1595] Output: Feedback displayed on the device (text or video format)
[1596] Through these steps, users can not only improve their technique, but also receive feedback based on their emotional state, enabling more effective practice.
[1597] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1598] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1599] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1600] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1601] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1602] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1603] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1604] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1605] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1606] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1607] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1608] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1609] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1610] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1611] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1612] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1613] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1614] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1615] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1616] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1617] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1618] The following is further disclosed regarding the above embodiment.
[1619] (Claim 1)
[1620] a means for a user to input audio or video;
[1621] a server that receives and analyzes input audio or video data;
[1622] a server that identifies areas for improvement based on the analysis results and generates specific advice;
[1623] The system includes a means for providing the generated advice to a user.
[1624] (Claim 2)
[1625] means for analyzing the pitch of input audio data;
[1626] a means for detecting deviations in the analyzed pitch;
[1627] 2. The system of claim 1, further comprising means for generating specific advice for correcting pitch deviations.
[1628] (Claim 3)
[1629] A means for analyzing facial expressions and gestures of input video data;
[1630] A means for evaluating the appropriateness of emotional expression of the analyzed facial expressions and gestures;
[1631] 10. The system of claim 1, further comprising means for generating specific advice for improving emotional expression.
[1632] "Example 1"
[1633] (Claim 1)
[1634] a means for a user to input audio or video;
[1635] means for temporarily storing input audio or video data;
[1636] means for transmitting the stored data to a server;
[1637] means for receiving and analyzing the transmitted audio or video data;
[1638] means for performing a spectral analysis on the audio data to evaluate pitch and rhythm;
[1639] means for analyzing facial expressions and gestures from video data;
[1640] A means to identify pitch discrepancies, rhythmic inconsistencies, and lack of expressiveness based on the analysis results;
[1641] a generative AI model that generates specific recommendations for the identified improvements;
[1642] The system includes a means for providing the generated advice to a user.
[1643] (Claim 2)
[1644] means for analyzing the pitch of input audio data;
[1645] a means for detecting deviations in the analyzed pitch;
[1646] 2. The system of claim 1, further comprising means for generating specific advice for correcting pitch deviations.
[1647] (Claim 3)
[1648] A means for analyzing facial expressions and gestures of input video data;
[1649] A means for evaluating the appropriateness of emotional expression of the analyzed facial expressions and gestures;
[1650] 10. The system of claim 1, further comprising means for generating specific advice for improving emotional expression.
[1651] "Application Example 1"
[1652] (Claim 1)
[1653] a means for a user to input audio or video;
[1654] a server that receives and analyzes input audio or video data;
[1655] a server that identifies areas for improvement based on the analysis results and generates specific advice;
[1656] a means for providing the generated advice to a user;
[1657] A means for generating advice on safe driving from the analysis results;
[1658] The system includes a means for displaying the generated safe driving advice in real time on an in-vehicle display or audio output device.
[1659] (Claim 2)
[1660] means for analyzing the pitch of input audio data;
[1661] a means for detecting deviations in the analyzed pitch;
[1662] A means of generating specific advice for correcting pitch discrepancies;
[1663] 2. The system according to claim 1, further comprising: means for analyzing information about a driving situation from the voice data and generating advice for safe driving.
[1664] (Claim 3)
[1665] A means for analyzing facial expressions and gestures of input video data;
[1666] A means for evaluating the appropriateness of emotional expression of the analyzed facial expressions and gestures;
[1667] means for generating specific advice for improving emotional expression;
[1668] The system according to claim 1, further comprising means for detecting other vehicles and obstacles from the video data and generating advice for safe driving.
[1669] "Example 2: Combining Emotion Engines"
[1670] (Claim 1)
[1671] a means for a user to input audio and video;
[1672] means for temporarily storing input audio and video data in a storage device of the terminal;
[1673] means for transmitting the stored audio and video data to a server;
[1674] a server that analyzes the received audio and video data;
[1675] means for performing a spectral analysis on the analyzed audio data;
[1676] A means for analyzing facial expressions and gestures from the analyzed video data;
[1677] a server including an emotion engine that analyzes emotions based on the analysis results;
[1678] A means for identifying areas for improvement and generating specific advice based on the analysis results and sentiment information;
[1679] The system includes a means for providing the generated advice to a user.
[1680] (Claim 2)
[1681] means for analyzing the pitch of input audio data;
[1682] a means for detecting deviations in the analyzed pitch;
[1683] 2. The system of claim 1, further comprising means for generating specific advice for correcting pitch deviations.
[1684] (Claim 3)
[1685] A means for analyzing facial expressions and gestures of input video data;
[1686] A means for evaluating the appropriateness of emotional expression of the analyzed facial expressions and gestures;
[1687] 10. The system of claim 1, further comprising means for generating specific advice for improving emotional expression.
[1688] "Application example 2 when combining emotion engines"
[1689] (Claim 1)
[1690] a means for a user to input audio or video;
[1691] a server that receives and analyzes input audio or video data;
[1692] a server that identifies areas for improvement based on the analysis results and generates specific advice;
[1693] a means for providing the generated advice to a user;
[1694] A means for analyzing user emotions;
[1695] The system includes a means for generating individually customized advice based on the analyzed emotional information.
[1696] (Claim 2)
[1697] means for analyzing the pitch of input audio data;
[1698] a means for detecting deviations in the analyzed pitch;
[1699] 2. The system of claim 1, further comprising means for generating specific advice for correcting pitch deviations.
[1700] (Claim 3)
[1701] A means for analyzing facial expressions and gestures of input video data;
[1702] A means for evaluating the appropriateness of emotional expression of the analyzed facial expressions and gestures;
[1703] 10. The system of claim 1, further comprising means for generating specific advice for improving emotional expression. [Explanation of symbols]
[1704] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for a user to input audio or video; a server that receives and analyzes input audio or video data; a server that identifies areas for improvement based on the analysis results and generates specific advice; The system includes a means for providing the generated advice to a user.
2. means for analyzing the pitch of input audio data; a means for detecting deviations in the analyzed pitch; 2. The system according to claim 1, further comprising means for generating specific advice for correcting pitch deviation.
3. A means for analyzing facial expressions and gestures of input video data; A means for evaluating the appropriateness of emotional expression of the analyzed facial expressions and gestures; 10. The system of claim 1, further comprising means for generating specific advice for improving emotional expression.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A