System

The system addresses the challenge of inadequate emotion recognition in conversational AI by analyzing voice input for content, pitch, and pauses to generate empathetic suggestions, enhancing user interaction with improved emotional understanding.

JP2026024035APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126356
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Conventional conversational AI systems struggle to accurately recognize emotions from users' voice inputs, particularly subtle characteristics such as emotional intonation and pauses, leading to inadequate empathetic responses.

Method used

A system that acquires user voice input, analyzes it for content, pitch, intonation, and pauses, recognizes emotions, and generates empathetic suggestions based on these analyses, using a generative AI model to provide appropriate advice.

Benefits of technology

The system improves emotion recognition accuracy, enabling it to provide empathetic and appropriate suggestions tailored to the user's emotional state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026024035000001_ABST
    Figure 2026024035000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring a voice input of a user; means for analyzing the voice input and extracting information on a content, a pitch, an intonation, and a pause of the voice; means for recognizing an emotion of the user based on a result of the analysis; means for generating a proposal with sympathy based on the recognized emotion; and means for providing the generated proposal to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional conversational AI systems have difficulty in properly recognizing emotions from users' voice input, resulting in the inability to provide appropriate and empathetic suggestions to users. In particular, they are unable to capture subtle characteristics of speech, such as emotional intonation and pauses, which reduces the accuracy of the support and advice users seek. [Means for solving the problem]

[0005] The present invention provides a means for acquiring a user's voice input and analyzing information on the content of the voice, voice pitch, intonation, and pauses. It also includes a means for recognizing the user's emotions based on the analysis results, and a means for generating empathetic suggestions based on the recognition results. The means described in the claims realize a system that can provide appropriate suggestions in a sympathetic manner while being sensitive to the user's emotions. This improves the accuracy of user emotion recognition, making it possible to provide appropriate and empathetic advice.

[0006] "Voice input" refers to the voice information uttered by the user and includes the process of inputting it into the system as digital data.

[0007] "Analysis" refers to the process of analyzing the acquired audio information in detail and extracting different elements such as the content of the audio, pitch, intonation, and pauses.

[0008] "Emotion recognition" refers to the process of identifying a user's emotional state based on voice features extracted through analysis.

[0009] "Suggestion generation" refers to the process of generating appropriate advice and support for users based on recognized emotions.

[0010] "Providing" refers to the process of communicating the generated proposal to the user and delivering it to the user.

[0011] "Empathy" refers to empathizing with the user's emotions and understanding and showing those emotions.

[0012] "Speech content" refers to the linguistic information and speech content contained in the speech input.

[0013] "Voice pitch" refers to information about the high or low pitch of a voice, obtained by analyzing the frequency components of the voice input.

[0014] "Intonation" refers to information obtained by analyzing changes in the strength and momentum of speech.

[0015] "Pauses" refers to information obtained by analyzing pauses, intervals, and silent parts of speech. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] System Overview

[0038] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user. The following describes the system's components and processing details in detail.

[0039] System Components

[0040] 1. Voice input acquisition means (terminal):

[0041] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0042] 2. Voice analysis means (server):

[0043] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0044] 3. Emotion recognition means (server):

[0045] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[0046] 4. Proposal generation means (server):

[0047] The server generates optimal suggestions for the user based on the recognized emotions, which are expressed in a manner that is empathetic and easy to accept.

[0048] 5. Proposal providing method (terminal):

[0049] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0050] Example

[0051] 1. Acquiring voice input

[0052] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0053] The device records this voice and sends it to the server as voice data.

[0054] 2. Audio analysis

[0055] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0056] 3. Emotion recognition

[0057] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[0058] 4. Proposal generation

[0059] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0060] 5. Providing suggestions

[0061] The server sends the generated proposal to the terminal.

[0062] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0063] This allows users to receive appropriate advice while having their own emotions and state recognized. The system is able to provide reliable suggestions while empathizing with the user's emotions.

[0064] The processing flow will be explained below.

[0065] Step 1:

[0066] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[0067] Step 2:

[0068] The terminal records the user's voice through a microphone and generates the voice data.

[0069] Step 3:

[0070] The device transmits the recorded audio data to a server via the Internet.

[0071] Step 4:

[0072] The server receives the voice data transmitted from the terminal.

[0073] Step 5:

[0074] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[0075] Step 6:

[0076] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[0077] Step 7:

[0078] Based on the emotion recognition results, the server uses a proposal generation engine to generate optimal proposals for the user, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0079] Step 8:

[0080] The server sends the generated proposal to the terminal.

[0081] Step 9:

[0082] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0083] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[0084] Example 1

[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0086] Conventional speech recognition systems simply convert the user's speech into text data, making it difficult to understand the emotions and nuances behind it. Furthermore, they lacked the means to provide appropriate suggestions based on the user's emotions, making it difficult to show empathy for the user. For this reason, there was a need to develop a system that could understand the user's actual emotions and state and provide appropriate advice based on that.

[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0088] In this invention, the server includes means for acquiring voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions (including means for using a generative AI model), and means for providing the generated suggestions to the user. This makes it possible to appropriately recognize emotions from the user's voice and provide advice that is sensitive to those emotions.

[0089] A "means for obtaining voice input" is a hardware or software feature that digitally records a user's voice and transmits it to other system components.

[0090] The "means for analyzing voice input" is a function for analyzing acquired voice data and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses.

[0091] "Means for recognizing emotions" refers to algorithms or engines that identify the user's emotional state based on the results of voice analysis.

[0092] The "means for generating suggestions" refers to a function for generating appropriate suggestions incorporating empathy based on the recognized emotions, including the use of a generative AI model.

[0093] A "generative AI model" is an artificial intelligence model trained with a large dataset to generate natural language text based on specific input (prompts).

[0094] A "prompt" is text input to a generative AI model, and is an instruction statement that controls the content and format of the generated output.

[0095] System Overview

[0096] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user.

[0097] System Components

[0098] 1. Voice input acquisition means (terminal):

[0099] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0100] As a specific example, a microphone installed in a smartphone or a personal computer is used.

[0101] 2. Voice analysis means (server):

[0102] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0103] A common voice recognition engine used is the "voice recognition API."

[0104] 3. Emotion recognition means (server):

[0105] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[0106] A common emotion recognition engine used is the "Emotion Analysis API."

[0107] 4. Proposal generation means (server):

[0108] The server generates optimal suggestions for the user based on the recognized emotions, using a generative AI model to express the suggestions in a way that is empathetic and easy to accept.

[0109] A "natural language generation model" is used as a generative AI model, such as the "GPT (generative pre-trained transformer)" model.

[0110] Example prompt: "Based on the user's voice analysis, we've determined that the user is feeling tired and stressed. Based on this, generate suggestions to help the user relax. Specifically, messages recommending deep breathing or taking a short break would be appropriate."

[0111] 5. Proposal providing method (terminal):

[0112] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0113] A specific example is a method of notifying by voice using a "speech synthesis engine" on the terminal, for example, using a "text-to-speech engine."

[0114] Example

[0115] 1. Acquiring voice input:

[0116] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0117] The device records this voice and sends it to the server as voice data.

[0118] 2. Audio analysis:

[0119] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0120] 3. Emotion recognition:

[0121] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[0122] 4. Proposal generation:

[0123] The server generates suggestions based on the emotion recognition results. For example, it generates suggestions such as, "You seem tired lately. Why don't you try taking a deep breath to relax?". It uses a generative AI model for generation.

[0124] 5. Providing suggestions:

[0125] The server sends the generated proposal to the terminal.

[0126] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0127] In this way, users can receive appropriate advice based on their own emotions and state of mind, enabling the system to provide highly reliable suggestions while showing empathy and understanding of the user's emotions.

[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0129] Step 1: Getting voice input

[0130] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0131] The device uses a built-in microphone to capture audio and stores it as digital audio data.

[0132] Input: User's voice

[0133] Output: Digital audio data

[0134] What happens: The device application starts the audio stream, buffers the user's voice, and when the recording is finished, a digital audio file is created.

[0135] Step 2: Sending audio data

[0136] The terminal transmits the acquired voice data to the server.

[0137] Input: Digital audio data

[0138] Output: Audio data sent to the server

[0139] Specific operation: The audio data is encoded and sent as a POST request to the server's API endpoint using the HTTPS protocol.

[0140] Step 3: Audio analysis

[0141] The server analyzes the received audio data.

[0142] Input: Audio data sent to the server

[0143] Output: Text data and speech characteristics (voice pitch, intonation, pauses, etc.)

[0144] Specific operation: The server uses a speech recognition engine (e.g., speech recognition API) to convert the voice data into text. It also uses an acoustic analysis module to extract voice characteristics such as pitch, intonation, and pauses. For example, it calls the function "speech_to_text(audio_data)" to obtain the analysis results.

[0145] Step 4: Emotion Recognition

[0146] The server recognizes emotions based on the results of voice analysis.

[0147] Input: Text data and speech features

[0148] Output: Emotion recognition result (e.g., "Tired" or "Stressed")

[0149] Specific behavior: The server uses an emotion recognition engine (e.g., emotion analysis API) to determine the emotional state from text and voice features. For example, if the user's voice is high-pitched and flat, it will recognize that the user is tired and stressed.

[0150] Step 5: Proposal Generation

[0151] The server generates suggestions based on the recognized emotions.

[0152] Input: Emotion recognition results (e.g., "Tired" or "Stressed")

[0153] Output: Generated suggestion message (e.g., "You seem tired lately. Why don't you try taking some deep breaths to relax?")

[0154] Specific operation: Using a generative AI model (e.g., GPT model), generate a suggested message by inputting a prompt sentence. Example of a prompt sentence: "The results of the user's voice analysis indicate that the user is feeling tired and stressed. Based on this state, generate a suggestion that will help the user relax. Specifically, a message recommending deep breathing or a short break would be appropriate."

[0155] Step 6: Submit a proposal

[0156] The server sends the generated proposal to the terminal.

[0157] Input: The generated proposal message

[0158] Output: Proposal message sent to the terminal

[0159] Specific operation: The server sends the generated proposal message to the terminal using the HTTPS protocol.

[0160] The device notifies the user of the suggestion in voice or text format.

[0161] Input: Proposal message sent to the terminal

[0162] Output: Suggestion message provided to the user

[0163] Specific operation: The application on the device receives the suggestion message and notifies it by voice using a speech synthesis engine (e.g., a text-to-speech engine), or displays it on the screen in text format.

[0164] (Application example 1)

[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0166] Conventional food delivery applications do not suggest optimal products or services based on the user's emotions or state, and this has limited the ability to improve user satisfaction and experience. For example, it is not possible to suggest relaxing meals to a tired user, so it has been a challenge to provide appropriate products and services that accurately grasp the user's needs.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0168] In this invention, the server includes means for acquiring a voice input from a user, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, means for providing the generated suggestions to the user, and means for recommending appropriate products and services based on the user's emotions. This makes it possible to suggest optimal products and services according to the user's emotions and state.

[0169] A "user" is an individual who uses the system.

[0170] "Voice input" refers to voice information that a user utters to the system.

[0171] "Speech content" refers to the linguistic information in the speech input, that is, the textual information uttered by the user.

[0172] "Voice pitch" is a feature related to the frequency of speech and indicates the tone or pitch of a user's voice.

[0173] "Intonation" refers to the intonation or change in speech, the ups and downs of the voice that indicate the user's emotions or intentions.

[0174] A "pause" refers to a temporary silence or pause in the speech input, which indicates a break or emphasis in what the user is saying.

[0175] The "analysis results" are data obtained after analyzing the voice input and extracting information about the voice content, pitch, intonation, and pauses.

[0176] "Means for recognizing emotions" refers to methods and technologies for identifying a user's emotional state based on the analysis results.

[0177] The "means for generating suggestions" refers to a method or technology for generating appropriate suggestions for a user based on the recognized emotions.

[0178] "Means for providing" refers to the methods and techniques for communicating the generated suggestions to users.

[0179] "Means for recommending products and services" refers to methods and techniques for suggesting appropriate products and services based on the user's emotions.

[0180] The present invention provides a system for analyzing a user's voice input and recognizing emotions to provide appropriate suggestions. The system includes a means for acquiring a user's voice input, a voice analysis means, an emotion recognition means, a suggestion generation means, and a suggestion providing means. The system also includes a means for recommending appropriate products and services based on the user's emotions.

[0181] System configuration:

[0182] Hardware:

[0183] Smartphone or robot: Built-in microphone for voice input

[0184] Server: for data analysis and emotion recognition

[0185] software:

[0186] 1. SpeechRecognition: A speech recognition library that converts voice data into text.

[0187] 2. Transformers (pipeline): Emotion recognition library. Analyzes input text and identifies emotions.

[0188] How it works:

[0189] 1. Voice input acquisition:

[0190] The user speaks into the microphone on their smartphone or the robot, saying, "I'm tired today, so I want to be soothed." The device then digitally records this voice and sends it to the server.

[0191] 2. Audio analysis:

[0192] The server analyzes the received voice data using the SpeechRecognition library, extracting information about the voice content (text data), voice pitch, intonation, and pauses.

[0193] 3. Emotion recognition:

[0194] The server uses Transformers (pipeline) to recognize emotions based on the results of voice analysis, such as "tired" or "stressed."

[0195] 4. Proposal generation:

[0196] The server generates appropriate suggestions based on the recognized emotion, such as "You seem tired lately. How about a relaxing herbal tea or a relaxing restaurant?"

[0197] 5. Providing suggestions:

[0198] The server sends the generated suggestions to the device, which then notifies the user of the suggestions, which may be provided in voice or text format.

[0199] Examples:

[0200] Example 1:

[0201] The user says, "I'm tired today and I want to be soothed."

[0202] The server analyzes this voice and recognizes the emotion "tiredness."

[0203] The server suggests "relaxing herbal teas and restaurants," and the device notifies the user of the suggestions via voice.

[0204] Example prompt sentence:

[0205] User: "I'm feeling great today. What's your recommended meal?"

[0206] App: "That sounds great! How about some pizza or steak to give you some energy?"

[0207] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0208] Step 1:

[0209] The user inputs voice data. The user speaks to the smartphone or robot about their condition and wishes. For example, they might say, "I'm tired today, so I want to be soothed." This voice data becomes the input.

[0210] Step 2:

[0211] The device acquires the voice data and converts it into a digital format. The device records the voice through the microphone and saves it as a digital audio file. The device then sends this data to the server as is. The input is the recorded voice data, and the output is digital voice data.

[0212] Step 3:

[0213] The server receives the voice data and analyzes it using the SpeechRecognition library. This analysis extracts information about the voice content (converts it to text), voice pitch, intonation, and pauses. For example, text data such as "I'm tired today, so I want to be soothed" and voice feature information are output. The input is digital voice data, and the output is text data and voice feature information.

[0214] Step 4:

[0215] The server uses Transformers (pipeline) to perform emotion recognition based on the results of speech analysis. Specifically, it inputs text data and speech feature information into an emotion recognition model to identify emotions such as "tired" or "stressed." The input is the results of speech analysis, and the output is emotional information.

[0216] Step 5:

[0217] The server generates suggestions based on the recognized emotions. Based on the emotion recognition results, for example, if you feel "tired," it will suggest products or services that will help you relax. Specifically, it will generate suggestions such as, "How about some relaxing herbal tea or a relaxing restaurant?" The input is emotional information, and the output is a suggested sentence.

[0218] Step 6:

[0219] The server sends the generated proposal to the terminal. The generated proposal is converted into text or audio format and sent to the terminal. The input is the proposal, and the output is the transmitted proposal data.

[0220] Step 7:

[0221] The device notifies the user of the suggestion. The device presents the received suggestion data to the user in the form of voice or text display. For example, the device might say to the user, "How about some relaxing herbal tea or a relaxing restaurant?" The input is the received suggestion data, and the output is voice or displayed text information.

[0222] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0223] System Overview

[0224] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an emotion recognition engine that recognizes the emotion. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. The following describes the system's components and processing details in detail.

[0225] System Components

[0226] 1. Voice input acquisition means (terminal):

[0227] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0228] 2. Voice analysis means (server):

[0229] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0230] 3. Emotion recognition means (server including emotion recognition engine):

[0231] The server uses an emotion recognition engine based on the results of voice analysis to identify the user's emotional state. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[0232] 4. Proposal generation means (server):

[0233] The server uses a proposal generation engine based on the recognized emotions to generate optimal proposals for the user, which are expressed in a manner that is empathetic to the user's emotions and easy to accept.

[0234] 5. Proposal providing method (terminal):

[0235] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0236] Example

[0237] 1. Acquiring voice input

[0238] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0239] The device records this voice and sends it to the server as voice data.

[0240] 2. Audio analysis

[0241] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0242] 3. Emotion recognition

[0243] Based on the voice analysis results, the server uses an emotion recognition engine to determine that the user is "tired." The emotion recognition engine uses a machine learning algorithm to identify multiple emotional states based on information such as voice pitch, intonation, and pauses. It also detects changes in emotion by comparing the data with past voice data, allowing for more accurate identification of the user's emotional state.

[0244] 4. Proposal generation

[0245] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you take a deep breath to relax?" The suggestion generation engine is designed to empathize with the user's emotions and express them in an acceptable way.

[0246] 5. Providing suggestions

[0247] The server sends the generated proposal to the terminal.

[0248] The device will notify the user of suggestions via voice or text. For example, the device might say to the user, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[0249] About program processing

[0250] The program for this system operates in the following steps: When a user speaks into the device, the device records the voice and sends the voice data to the server. The server analyzes the voice data and identifies the emotion using an emotion recognition engine. Based on the recognized emotion, a proposal generation engine generates an appropriate proposal and sends it to the device. The device notifies the user of the proposal in voice or text format. Through this series of steps, the user is able to recognize their own emotions and state and receive appropriate advice based on that.

[0251] Specific examples

[0252] 1. Acquiring voice input

[0253] User: Say, "I've been really busy lately and I'm tired every day."

[0254] Terminal: Records the user's voice and sends the voice data to the server.

[0255] 2. Audio analysis

[0256] Server: Analyzes the voice data and generates text data such as "I've been really busy lately and I'm tired every day." It also extracts information about the pitch, intonation, and pauses of the voice.

[0257] 3. Emotion recognition

[0258] Server: Using an emotion recognition engine, the server recognizes the user's emotion as "tired." It uses a machine learning algorithm to analyze the intonation and pauses of the voice and compare them with past data to identify changes in emotion.

[0259] 4. Proposal generation

[0260] Server: Based on the emotion recognition results, it generates a suggestion such as, "Why not take a deep breath to relax and refresh yourself?"

[0261] 5. Providing suggestions

[0262] Server: Sends proposals to the device.

[0263] Device: Speak a suggestion to the user: "Why not take a deep breath to relax and refresh yourself?"

[0264] In this way, this system can analyze the user's voice, recognize their emotions, and generate and provide appropriate suggestions, thereby providing support that is sensitive to the user's emotions.

[0265] The processing flow will be explained below.

[0266] Step 1:

[0267] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[0268] Step 2:

[0269] The terminal records the user's voice through a microphone and generates the voice data.

[0270] Step 3:

[0271] The device transmits the recorded audio data to a server via the Internet.

[0272] Step 4:

[0273] The server receives the voice data transmitted from the terminal.

[0274] Step 5:

[0275] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[0276] Step 6:

[0277] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[0278] Step 7:

[0279] The server uses an emotion recognition engine and machine learning algorithms to recognize emotions, utilizing voice pitch, intonation, and pauses to identify multiple emotional states (e.g., fatigue, stress) of the user.

[0280] Step 8:

[0281] The server identifies changes in the user's emotional state by comparing with past voice data to detect changes in emotion.

[0282] Step 9:

[0283] The server uses a suggestion generation engine to generate optimal suggestions based on the recognized emotions and their changes, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0284] Step 10:

[0285] The server sends the generated proposal to the terminal.

[0286] Step 11:

[0287] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0288] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[0289] Example 2

[0290] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0291] While existing systems can analyze a user's voice input and recognize their emotions, they face the challenge of effectively generating and providing appropriate suggestions based on those emotions. Another problem is that they have not yet fully achieved the ability to identify emotions or states that the user is not aware of and to make suggestions based on those emotions or states.

[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0293] In this invention, the server includes means for digitally storing and transmitting the user's voice input, means for analyzing the voice input and extracting information on the voice content, pitch, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, and means for providing the generated suggestions to the user in voice or text format. This makes it possible to grasp the user's emotions in detail and provide accurate and empathetic suggestions. Furthermore, a series of processes can be efficiently realized, including identifying emotions and states that the user is not aware of based on past data from the voice input and making suggestions based on them.

[0294] "User" refers to an individual or end user who uses the System.

[0295] "Voice input" refers to voice information that a user speaks into a terminal.

[0296] "Device" means a hardware device for recording and storing audio in digital form, such as a smartphone or personal computer.

[0297] "Server" refers to a computer system for receiving voice data and performing analysis and emotion recognition processing.

[0298] A "voice recognition engine" refers to software that converts voice data into text data and extracts elements such as voice pitch, intonation, and pauses.

[0299] "Emotion recognition engine" refers to software that identifies a user's emotional state from the results of voice analysis. It uses machine learning algorithms to recognize emotions.

[0300] "Suggestion Generation Engine" means software for generating suggestions to a user based on recognized emotions.

[0301] "Voice or text format" refers to the presentation method of the suggestions provided to the user, and means speech output using a speech synthesis engine or text message output.

[0302] A "speech synthesis engine" refers to software that converts text data into speech and outputs it in an easy-to-listen format.

[0303] MODE FOR CARRYING OUT THE INVENTION

[0304] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an engine that recognizes the user's emotions. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. Specific embodiments of this system are described below.

[0305] The system mainly includes the following components:

[0306] 1. Voice input acquisition means (terminal): The terminal acquires voice input from the user through a microphone. This voice information is stored in digital format. Terminals include smartphones and PCs.

[0307] 2. Voice analysis means (server): The server receives the voice data sent from the device and analyzes the data using a voice recognition engine (e.g., Google Speech-to-Text API). Through this analysis, elements such as the content of the voice, voice pitch, intonation, and pauses are extracted.

[0308] 3. Emotion recognition means (server including emotion recognition engine): The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state based on the results of voice analysis. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[0309] 4. Proposal generation means (server including a proposal generation engine): The server uses a proposal generation engine (e.g., OpenAI GPT-4) based on the recognized emotions to generate optimal proposals for the user. These proposals are expressed in a way that is empathetic to the user's emotions and easy to accept.

[0310] 5. Proposal providing means (terminal): The terminal receives the proposal sent from the server and provides it to the user. The proposal is notified to the user in voice (using a TTS engine) or text format.

[0311] Specific examples

[0312] 1. Acquiring voice input:

[0313] User: "I've been really busy lately and I can't seem to get rid of my fatigue."

[0314] Device: This audio is recorded using a microphone and stored digitally.

[0315] 2. Sending audio data:

[0316] On the device: The recorded audio data is sent to the server using the HTTPS protocol.

[0317] 3. Audio analysis:

[0318] Server: Receives the voice data and analyzes it using a voice recognition engine. The voice content is converted into text data, and information such as voice pitch, intonation, and pauses is extracted.

[0319] 4. Emotion recognition:

[0320] Server: Inputs the results of voice analysis into an emotion recognition engine to recognize the user's emotion. For example, the analysis results identify the emotion as "tired."

[0321] 5. Proposal generation:

[0322] Server: Using the suggestion generation engine based on the emotion recognition results, generate a suggestion such as "Why not take a deep breath to relax and refresh yourself?"

[0323] 6. Providing suggestions:

[0324] Server: Sends the generated proposals to the device.

[0325] On the device: Providing suggestions to the user in the form of voice or text, such as "Why not try taking a few deep breaths to relax and refresh yourself?"

[0326] Through this series of processes, the system is able to grasp the user's emotions in detail and provide appropriate, empathetic suggestions. As a concrete example, in response to the prompt "I've been so busy lately, I can't shake off my fatigue," the system generates the suggestion "Why don't you take a deep breath and refresh yourself?" In this way, by analyzing emotions from the user's voice input and providing suggestions based on that, the system achieves support that is sensitive to the user's emotions.

[0327] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0328] Program processing flow

[0329] (Step 1: Acquiring voice input)

[0330] The user speaks into the terminal.

[0331] Example: Say, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0332] The device captures the user's voice through a microphone and stores it in digital form as voice data.

[0333] Input: Audio generated by the user's speech

[0334] Output: Digital audio data

[0335] What it does: The microphone picks up audio and records it in an internal buffer.

[0336] (Step 2: Sending audio data)

[0337] The terminal transmits the stored voice data to the server.

[0338] Input: Digitally stored audio data

[0339] Output: Audio data sent to the server

[0340] What it does: Data is sent securely to the server using the HTTPS protocol.

[0341] (Step 3: Audio analysis)

[0342] The server analyzes the voice data received from the terminal.

[0343] Example: Audio data of "I've been so busy lately, I can't get rid of my fatigue"

[0344] The server uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the speech into text, and also extracts information such as voice pitch, intonation, and pauses.

[0345] Input: Audio data sent to the server

[0346] Output: Text data "I've been so busy lately, I can't get rid of my fatigue" and information on voice pitch, intonation, pauses, etc.

[0347] Specific operation: Sends audio data to the API and receives the analysis results in return.

[0348] (Step 4: Emotion Recognition)

[0349] The server uses an emotion recognition engine to identify the user's emotions based on the results of the voice analysis.

[0350] Example: Recognizing the emotion "I'm tired."

[0351] The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the emotional state from the text data and voice elements of the voice analysis results.

[0352] Input: Text data of the voice analysis results and voice pitch, intonation, and pause information

[0353] Output: Emotional state, such as "tired"

[0354] Specific operation: The analysis results are input into the emotion recognition engine and the returned emotional state data is received.

[0355] (Step 5: Proposal Generation)

[0356] The server generates suggestions based on the recognized emotions.

[0357] Example: Suggestion: "Why don't you take a deep breath to relax and refresh yourself?"

[0358] The server uses a proposal generation engine (e.g., OpenAI GPT-4) to create optimal proposals for the user in a way that empathizes with their emotional state.

[0359] Input: Recognized emotional state "tired"

[0360] Output: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[0361] Specific operation: Input the emotional state into the proposal generation engine and obtain the generated proposals.

[0362] (Step 6: Submit your proposal)

[0363] The server transmits the generated proposal to the terminal.

[0364] Input: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[0365] Output: The proposal is sent to the terminal

[0366] Specific operation: Proposal data is sent to the terminal using the HTTPS protocol.

[0367] (Step 7: Provide a proposal)

[0368] The terminal receives the suggestions and provides them to the user in voice or text format.

[0369] Example: A voice notification saying, "Why not take a deep breath to relax and refresh yourself?"

[0370] The device may use a text-to-speech engine (TTS) to communicate the suggestions to the user in audible form or display them in text form.

[0371] Input: Proposal sent by the server

[0372] Output: Suggestions that are communicated to the user in audio or text format

[0373] What it does: Inputs the proposed data into a speech synthesis engine to generate speech output that can be played through a speaker or displayed as text on a screen.

[0374] (Application example 2)

[0375] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0376] Conventional systems that recognize a user's emotions through voice input and make general suggestions based on the results are unable to recommend content that meets the user's specific needs, making it difficult to provide appropriate support based on emotions. In particular, there is a need for a system that can automatically recommend content for relaxation and stress reduction when a user is feeling fatigued or stressed.

[0377] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0378] In this invention, the server includes means for acquiring a user's voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating suggestions that incorporate empathy, means for providing the generated suggestions to the user, means for recommending appropriate content based on the emotion recognition results, means including a recommendation engine for selecting specific content for the emotional state input by voice, and means for notifying the user of the recommended content. This makes it possible to recommend specific content according to the user's emotions, thereby better meeting user needs.

[0379] "Voice input"

[0380] It is voice data that captures what a user says in digital form.

[0381] "Speech recognition engine"

[0382] is a software system that analyzes captured audio data and extracts the content and characteristics of the audio.

[0383] Emotion recognition engine

[0384] It is a system that uses machine learning algorithms to identify a user's emotions based on data analyzed by a voice recognition engine.

[0385] "Proposal generation engine"

[0386] It is a system that generates empathetic suggestions for users based on recognized emotions.

[0387] Recommendation engine

[0388] It is a system that selects and recommends appropriate content according to the user's emotional state.

[0389] "content"

[0390] It refers to information media that can be viewed and played in digital format, such as videos, music, documentaries, and movies.

[0391] "Notification means"

[0392] is an interface for informing users of generated suggestions and recommended content, and has the ability to notify them in voice or text format.

[0393] "Voice pitch"

[0394] is information that indicates frequency changes in audio data, and is also known as the pitch of speech.

[0395] "intonation"

[0396] This is information that indicates the intonation and emotional strength contained in the voice data.

[0397] "Information between"

[0398] is information that indicates the timing of speech and pauses in audio data.

[0399] "Specific content"

[0400] It is a digital media of visual and audio content that best suits the user's emotional state.

[0401] System Overview

[0402] The present invention provides a system for recognizing a user's emotions through voice input, recommending appropriate content based on the results, and providing the content to the user. The system includes a voice input unit, a voice analysis unit, an emotion recognition unit, a suggestion generation unit, a recommendation engine, and a notification unit.

[0403] Hardware and software used

[0404] Hardware:

[0405] Smartphone: Equipped with a microphone to capture voice input, a speaker and a display to announce generated suggestions and recommendations.

[0406] software:

[0407] Speech recognition engine: Analyzes voice data and extracts the content and characteristics of the voice.

[0408] Emotion recognition engine: Based on data analyzed by the voice recognition engine, a machine learning algorithm is used to identify the user's emotions.

[0409] Suggestion generation engine: Generates empathetic suggestions based on recognized emotions.

[0410] Recommendation engine: Selects and recommends appropriate content based on emotion recognition results.

[0411] Notification Method: The user is notified of generated suggestions and recommended content via text or audio.

[0412] Specific examples of processing

[0413] Example 1: A user says to their smartphone, "I've been so busy lately and I can't seem to get rid of my fatigue."

[0414] The server analyzes the voice data and generates text data such as, "I've been really busy lately and I can't get rid of my fatigue."

[0415] The emotion recognition engine identifies "fatigue" as an emotion from the data of the voice recognition engine.

[0416] The suggestion generation engine generates a suggestion such as, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[0417] The recommendation engine selects relaxing video content (e.g., videos of natural scenery or relaxing music) according to the emotion.

[0418] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[0419] Example 2: A user says, "I'm feeling very stressed and irritable today."

[0420] The server analyzes the voice data and generates text data such as "I'm feeling very stressed and irritated today."

[0421] The emotion recognition engine identifies "anger" as an emotion from the data of the voice recognition engine.

[0422] The suggestion generation engine generates a suggestion such as, "You seem stressed. Take a deep breath and refresh yourself."

[0423] The recommendation engine selects video content that helps reduce stress (e.g., relaxation videos or comedy movies) based on emotions.

[0424] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[0425] Prompt Sentence Examples

[0426] You are an emotion recognition engine. Analyze the following user voice inputs and identify the appropriate emotion:

[0427] Dictation: "I'm feeling very stressed and irritable today."

[0428] This allows users to easily receive specific content that corresponds to their emotions, enabling them to receive appropriate support.

[0429] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0430] Step 1:

[0431] The user provides voice input. The user speaks into the smartphone's microphone, expressing their mood and emotions. This voice input is captured in digital form. The input data is voice data, which forms the basis for subsequent processing.

[0432] Step 2:

[0433] The terminal sends the acquired voice data to the server. The input data is voice data and is transferred to the server. At this stage, it is confirmed that the voice data has reached the server.

[0434] Step 3:

[0435] The server analyzes the voice data using a voice recognition engine. Specifically, it converts the voice data into text data and extracts information about the pitch, intonation, and pauses of the voice. The output data is the analysis result, and includes text data and voice feature information.

[0436] Step 4:

[0437] The server uses an emotion recognition engine to identify emotions from the analysis results. The input data is text data and voice feature information, and a machine learning algorithm identifies the user's emotions. The output data is emotional information, such as fatigue or stress.

[0438] Step 5:

[0439] The server uses a proposal generation engine to generate proposals based on emotions and incorporating empathy. The input data is emotional information, and appropriate proposals are generated based on this. The output data is a proposal message, which includes content that shows empathy to the user.

[0440] Step 6:

[0441] The server uses a recommendation engine to recommend content according to emotions. The input data is emotion information, and appropriate content is selected based on this. The output data is content recommendation information, including specific video and audio content.

[0442] Step 7:

[0443] The server sends the generated suggestion message and recommended content to the terminal. The input data is the suggestion message and content recommendation information, which are transferred to the terminal. The output data is notification information on the terminal.

[0444] Step 8:

[0445] The device notifies the user of the suggestion message and the recommended content. The input data is notification information, which is communicated to the user via the smartphone's display or voice output. Specific actions include a text message being displayed on the screen or a voice announcement of the suggestion.

[0446] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0448] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0449] [Second embodiment]

[0450] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0451] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0452] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0453] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0454] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0456] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0457] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0458] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0459] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0460] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0461] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0462] System Overview

[0463] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user. The following describes the system's components and processing details in detail.

[0464] System Components

[0465] 1. Voice input acquisition means (terminal):

[0466] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0467] 2. Voice analysis means (server):

[0468] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0469] 3. Emotion recognition means (server):

[0470] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[0471] 4. Proposal generation means (server):

[0472] The server generates optimal suggestions for the user based on the recognized emotions, which are expressed in a manner that is empathetic and easy to accept.

[0473] 5. Proposal providing method (terminal):

[0474] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0475] Example

[0476] 1. Acquiring voice input

[0477] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0478] The device records this voice and sends it to the server as voice data.

[0479] 2. Audio analysis

[0480] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0481] 3. Emotion recognition

[0482] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[0483] 4. Proposal generation

[0484] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0485] 5. Providing suggestions

[0486] The server sends the generated proposal to the terminal.

[0487] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0488] This allows users to receive appropriate advice while having their own emotions and state recognized. The system is able to provide reliable suggestions while empathizing with the user's emotions.

[0489] The processing flow will be explained below.

[0490] Step 1:

[0491] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[0492] Step 2:

[0493] The terminal records the user's voice through a microphone and generates the voice data.

[0494] Step 3:

[0495] The device transmits the recorded audio data to a server via the Internet.

[0496] Step 4:

[0497] The server receives the voice data transmitted from the terminal.

[0498] Step 5:

[0499] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[0500] Step 6:

[0501] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[0502] Step 7:

[0503] Based on the emotion recognition results, the server uses a proposal generation engine to generate optimal proposals for the user, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0504] Step 8:

[0505] The server sends the generated proposal to the terminal.

[0506] Step 9:

[0507] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0508] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[0509] Example 1

[0510] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0511] Conventional speech recognition systems simply convert the user's speech into text data, making it difficult to understand the emotions and nuances behind it. Furthermore, they lacked the means to provide appropriate suggestions based on the user's emotions, making it difficult to show empathy for the user. For this reason, there was a need to develop a system that could understand the user's actual emotions and state and provide appropriate advice based on that.

[0512] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0513] In this invention, the server includes means for acquiring voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions (including means for using a generative AI model), and means for providing the generated suggestions to the user. This makes it possible to appropriately recognize emotions from the user's voice and provide advice that is sensitive to those emotions.

[0514] A "means for obtaining voice input" is a hardware or software feature that digitally records a user's voice and transmits it to other system components.

[0515] The "means for analyzing voice input" is a function for analyzing acquired voice data and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses.

[0516] "Means for recognizing emotions" refers to algorithms or engines that identify the user's emotional state based on the results of voice analysis.

[0517] The "means for generating suggestions" refers to a function for generating appropriate suggestions incorporating empathy based on the recognized emotions, including the use of a generative AI model.

[0518] A "generative AI model" is an artificial intelligence model trained with a large dataset to generate natural language text based on specific input (prompts).

[0519] A "prompt" is text input to a generative AI model, and is an instruction statement that controls the content and format of the generated output.

[0520] System Overview

[0521] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user.

[0522] System Components

[0523] 1. Voice input acquisition means (terminal):

[0524] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0525] As a specific example, a microphone installed in a smartphone or a personal computer is used.

[0526] 2. Voice analysis means (server):

[0527] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0528] A common voice recognition engine used is the "voice recognition API."

[0529] 3. Emotion recognition means (server):

[0530] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[0531] A common emotion recognition engine used is the "Emotion Analysis API."

[0532] 4. Proposal generation means (server):

[0533] The server generates optimal suggestions for the user based on the recognized emotions, using a generative AI model to express the suggestions in a way that is empathetic and easy to accept.

[0534] A "natural language generation model" is used as a generative AI model, such as the "GPT (generative pre-trained transformer)" model.

[0535] Example prompt: "Based on the user's voice analysis, we've determined that the user is feeling tired and stressed. Based on this, generate suggestions to help the user relax. Specifically, messages recommending deep breathing or taking a short break would be appropriate."

[0536] 5. Proposal providing method (terminal):

[0537] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0538] A specific example is a method of notifying by voice using a "speech synthesis engine" on the terminal, for example, using a "text-to-speech engine."

[0539] Example

[0540] 1. Acquiring voice input:

[0541] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0542] The device records this voice and sends it to the server as voice data.

[0543] 2. Audio analysis:

[0544] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0545] 3. Emotion recognition:

[0546] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[0547] 4. Proposal generation:

[0548] The server generates suggestions based on the emotion recognition results. For example, it generates suggestions such as, "You seem tired lately. Why don't you try taking a deep breath to relax?". It uses a generative AI model for generation.

[0549] 5. Providing suggestions:

[0550] The server sends the generated proposal to the terminal.

[0551] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0552] In this way, users can receive appropriate advice based on their own emotions and state of mind, enabling the system to provide highly reliable suggestions while showing empathy and understanding of the user's emotions.

[0553] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0554] Step 1: Getting voice input

[0555] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0556] The device uses a built-in microphone to capture audio and stores it as digital audio data.

[0557] Input: User's voice

[0558] Output: Digital audio data

[0559] What happens: The device application starts the audio stream, buffers the user's voice, and when the recording is finished, a digital audio file is created.

[0560] Step 2: Sending audio data

[0561] The terminal transmits the acquired voice data to the server.

[0562] Input: Digital audio data

[0563] Output: Audio data sent to the server

[0564] Specific operation: The audio data is encoded and sent as a POST request to the server's API endpoint using the HTTPS protocol.

[0565] Step 3: Audio analysis

[0566] The server analyzes the received audio data.

[0567] Input: Audio data sent to the server

[0568] Output: Text data and speech characteristics (voice pitch, intonation, pauses, etc.)

[0569] Specific operation: The server uses a speech recognition engine (e.g., speech recognition API) to convert the voice data into text. It also uses an acoustic analysis module to extract voice characteristics such as pitch, intonation, and pauses. For example, it calls the function "speech_to_text(audio_data)" to obtain the analysis results.

[0570] Step 4: Emotion Recognition

[0571] The server recognizes emotions based on the results of voice analysis.

[0572] Input: Text data and speech features

[0573] Output: Emotion recognition result (e.g., "Tired" or "Stressed")

[0574] Specific behavior: The server uses an emotion recognition engine (e.g., emotion analysis API) to determine the emotional state from text and voice features. For example, if the user's voice is high-pitched and flat, it will recognize that the user is tired and stressed.

[0575] Step 5: Proposal Generation

[0576] The server generates suggestions based on the recognized emotions.

[0577] Input: Emotion recognition results (e.g., "Tired" or "Stressed")

[0578] Output: Generated suggestion message (e.g., "You seem tired lately. Why don't you try taking some deep breaths to relax?")

[0579] Specific operation: Using a generative AI model (e.g., GPT model), generate a suggested message by inputting a prompt sentence. Example of a prompt sentence: "The results of the user's voice analysis indicate that the user is feeling tired and stressed. Based on this state, generate a suggestion that will help the user relax. Specifically, a message recommending deep breathing or a short break would be appropriate."

[0580] Step 6: Submit a proposal

[0581] The server sends the generated proposal to the terminal.

[0582] Input: The generated proposal message

[0583] Output: Proposal message sent to the terminal

[0584] Specific operation: The server sends the generated proposal message to the terminal using the HTTPS protocol.

[0585] The device notifies the user of the suggestion in voice or text format.

[0586] Input: Proposal message sent to the terminal

[0587] Output: Suggestion message provided to the user

[0588] Specific operation: The application on the device receives the suggestion message and notifies it by voice using a speech synthesis engine (e.g., a text-to-speech engine), or displays it on the screen in text format.

[0589] (Application example 1)

[0590] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0591] Conventional food delivery applications do not suggest optimal products or services based on the user's emotions or state, and this has limited the ability to improve user satisfaction and experience. For example, it is not possible to suggest relaxing meals to a tired user, so it has been a challenge to provide appropriate products and services that accurately grasp the user's needs.

[0592] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0593] In this invention, the server includes means for acquiring a voice input from a user, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, means for providing the generated suggestions to the user, and means for recommending appropriate products and services based on the user's emotions. This makes it possible to suggest optimal products and services according to the user's emotions and state.

[0594] A "user" is an individual who uses the system.

[0595] "Voice input" refers to voice information that a user utters to the system.

[0596] "Speech content" refers to the linguistic information in the speech input, that is, the textual information uttered by the user.

[0597] "Voice pitch" is a feature related to the frequency of speech and indicates the tone or pitch of a user's voice.

[0598] "Intonation" refers to the intonation or change in speech, the ups and downs of the voice that indicate the user's emotions or intentions.

[0599] A "pause" refers to a temporary silence or pause in the speech input, which indicates a break or emphasis in what the user is saying.

[0600] The "analysis results" are data obtained after analyzing the voice input and extracting information about the voice content, pitch, intonation, and pauses.

[0601] "Means for recognizing emotions" refers to methods and technologies for identifying a user's emotional state based on the analysis results.

[0602] The "means for generating suggestions" refers to a method or technology for generating appropriate suggestions for a user based on the recognized emotions.

[0603] "Means for providing" refers to the methods and techniques for communicating the generated suggestions to users.

[0604] "Means for recommending products and services" refers to methods and techniques for suggesting appropriate products and services based on the user's emotions.

[0605] The present invention provides a system for analyzing a user's voice input and recognizing emotions to provide appropriate suggestions. The system includes a means for acquiring a user's voice input, a voice analysis means, an emotion recognition means, a suggestion generation means, and a suggestion providing means. The system also includes a means for recommending appropriate products and services based on the user's emotions.

[0606] System configuration:

[0607] Hardware:

[0608] Smartphone or robot: Built-in microphone for voice input

[0609] Server: for data analysis and emotion recognition

[0610] software:

[0611] 1. SpeechRecognition: A speech recognition library that converts voice data into text.

[0612] 2. Transformers (pipeline): Emotion recognition library. Analyzes input text and identifies emotions.

[0613] How it works:

[0614] 1. Voice input acquisition:

[0615] The user speaks into the microphone on their smartphone or the robot, saying, "I'm tired today, so I want to be soothed." The device then digitally records this voice and sends it to the server.

[0616] 2. Audio analysis:

[0617] The server analyzes the received voice data using the SpeechRecognition library, extracting information about the voice content (text data), voice pitch, intonation, and pauses.

[0618] 3. Emotion recognition:

[0619] The server uses Transformers (pipeline) to recognize emotions based on the results of voice analysis, such as "tired" or "stressed."

[0620] 4. Proposal generation:

[0621] The server generates appropriate suggestions based on the recognized emotion, such as "You seem tired lately. How about a relaxing herbal tea or a relaxing restaurant?"

[0622] 5. Providing suggestions:

[0623] The server sends the generated suggestions to the device, which then notifies the user of the suggestions, which may be provided in voice or text format.

[0624] Examples:

[0625] Example 1:

[0626] The user says, "I'm tired today and I want to be soothed."

[0627] The server analyzes this voice and recognizes the emotion "tiredness."

[0628] The server suggests "relaxing herbal teas and restaurants," and the device notifies the user of the suggestions via voice.

[0629] Example prompt sentence:

[0630] User: "I'm feeling great today. What's your recommended meal?"

[0631] App: "That sounds great! How about some pizza or steak to give you some energy?"

[0632] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0633] Step 1:

[0634] The user inputs voice data. The user speaks to the smartphone or robot about their condition and wishes. For example, they might say, "I'm tired today, so I want to be soothed." This voice data becomes the input.

[0635] Step 2:

[0636] The device acquires the voice data and converts it into a digital format. The device records the voice through the microphone and saves it as a digital audio file. The device then sends this data to the server as is. The input is the recorded voice data, and the output is digital voice data.

[0637] Step 3:

[0638] The server receives the voice data and analyzes it using the SpeechRecognition library. This analysis extracts information about the voice content (converts it to text), voice pitch, intonation, and pauses. For example, text data such as "I'm tired today, so I want to be soothed" and voice feature information are output. The input is digital voice data, and the output is text data and voice feature information.

[0639] Step 4:

[0640] The server uses Transformers (pipeline) to perform emotion recognition based on the results of speech analysis. Specifically, it inputs text data and speech feature information into an emotion recognition model to identify emotions such as "tired" or "stressed." The input is the results of speech analysis, and the output is emotional information.

[0641] Step 5:

[0642] The server generates suggestions based on the recognized emotions. Based on the emotion recognition results, for example, if you feel "tired," it will suggest products or services that will help you relax. Specifically, it will generate suggestions such as, "How about some relaxing herbal tea or a relaxing restaurant?" The input is emotional information, and the output is a suggested sentence.

[0643] Step 6:

[0644] The server sends the generated proposal to the terminal. The generated proposal is converted into text or audio format and sent to the terminal. The input is the proposal, and the output is the transmitted proposal data.

[0645] Step 7:

[0646] The device notifies the user of the suggestion. The device presents the received suggestion data to the user in the form of voice or text display. For example, the device might say to the user, "How about some relaxing herbal tea or a relaxing restaurant?" The input is the received suggestion data, and the output is voice or displayed text information.

[0647] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0648] System Overview

[0649] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an emotion recognition engine that recognizes the emotion. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. The following describes the system's components and processing details in detail.

[0650] System Components

[0651] 1. Voice input acquisition means (terminal):

[0652] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0653] 2. Voice analysis means (server):

[0654] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0655] 3. Emotion recognition means (server including emotion recognition engine):

[0656] The server uses an emotion recognition engine based on the results of voice analysis to identify the user's emotional state. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[0657] 4. Proposal generation means (server):

[0658] The server uses a proposal generation engine based on the recognized emotions to generate optimal proposals for the user, which are expressed in a manner that is empathetic to the user's emotions and easy to accept.

[0659] 5. Proposal providing method (terminal):

[0660] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0661] Example

[0662] 1. Acquiring voice input

[0663] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0664] The device records this voice and sends it to the server as voice data.

[0665] 2. Audio analysis

[0666] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0667] 3. Emotion recognition

[0668] Based on the voice analysis results, the server uses an emotion recognition engine to determine that the user is "tired." The emotion recognition engine uses a machine learning algorithm to identify multiple emotional states based on information such as voice pitch, intonation, and pauses. It also detects changes in emotion by comparing the data with past voice data, allowing for more accurate identification of the user's emotional state.

[0669] 4. Proposal generation

[0670] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you take a deep breath to relax?" The suggestion generation engine is designed to empathize with the user's emotions and express them in an acceptable way.

[0671] 5. Providing suggestions

[0672] The server sends the generated proposal to the terminal.

[0673] The device will notify the user of suggestions via voice or text. For example, the device might say to the user, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[0674] About program processing

[0675] The program for this system operates in the following steps: When a user speaks into the device, the device records the voice and sends the voice data to the server. The server analyzes the voice data and identifies the emotion using an emotion recognition engine. Based on the recognized emotion, a proposal generation engine generates an appropriate proposal and sends it to the device. The device notifies the user of the proposal in voice or text format. Through this series of steps, the user is able to recognize their own emotions and state and receive appropriate advice based on that.

[0676] Specific examples

[0677] 1. Acquiring voice input

[0678] User: Say, "I've been really busy lately and I'm tired every day."

[0679] Terminal: Records the user's voice and sends the voice data to the server.

[0680] 2. Audio analysis

[0681] Server: Analyzes the voice data and generates text data such as "I've been really busy lately and I'm tired every day." It also extracts information about the pitch, intonation, and pauses of the voice.

[0682] 3. Emotion recognition

[0683] Server: Using an emotion recognition engine, the server recognizes the user's emotion as "tired." It uses a machine learning algorithm to analyze the intonation and pauses of the voice and compare them with past data to identify changes in emotion.

[0684] 4. Proposal generation

[0685] Server: Based on the emotion recognition results, it generates a suggestion such as, "Why not take a deep breath to relax and refresh yourself?"

[0686] 5. Providing suggestions

[0687] Server: Sends proposals to the device.

[0688] Device: Speak a suggestion to the user: "Why not take a deep breath to relax and refresh yourself?"

[0689] In this way, this system can analyze the user's voice, recognize their emotions, and generate and provide appropriate suggestions, thereby providing support that is sensitive to the user's emotions.

[0690] The processing flow will be explained below.

[0691] Step 1:

[0692] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[0693] Step 2:

[0694] The terminal records the user's voice through a microphone and generates the voice data.

[0695] Step 3:

[0696] The device transmits the recorded audio data to a server via the Internet.

[0697] Step 4:

[0698] The server receives the voice data transmitted from the terminal.

[0699] Step 5:

[0700] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[0701] Step 6:

[0702] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[0703] Step 7:

[0704] The server uses an emotion recognition engine and machine learning algorithms to recognize emotions, utilizing voice pitch, intonation, and pauses to identify multiple emotional states (e.g., fatigue, stress) of the user.

[0705] Step 8:

[0706] The server identifies changes in the user's emotional state by comparing with past voice data to detect changes in emotion.

[0707] Step 9:

[0708] The server uses a suggestion generation engine to generate optimal suggestions based on the recognized emotions and their changes, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0709] Step 10:

[0710] The server sends the generated proposal to the terminal.

[0711] Step 11:

[0712] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0713] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[0714] Example 2

[0715] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0716] While existing systems can analyze a user's voice input and recognize their emotions, they face the challenge of effectively generating and providing appropriate suggestions based on those emotions. Another problem is that they have not yet fully achieved the ability to identify emotions or states that the user is not aware of and to make suggestions based on those emotions or states.

[0717] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0718] In this invention, the server includes means for digitally storing and transmitting the user's voice input, means for analyzing the voice input and extracting information on the voice content, pitch, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, and means for providing the generated suggestions to the user in voice or text format. This makes it possible to grasp the user's emotions in detail and provide accurate and empathetic suggestions. Furthermore, a series of processes can be efficiently realized, including identifying emotions and states that the user is not aware of based on past data from the voice input and making suggestions based on them.

[0719] "User" refers to an individual or end user who uses the System.

[0720] "Voice input" refers to voice information that a user speaks into a terminal.

[0721] "Device" means a hardware device for recording and storing audio in digital form, such as a smartphone or personal computer.

[0722] "Server" refers to a computer system for receiving voice data and performing analysis and emotion recognition processing.

[0723] A "voice recognition engine" refers to software that converts voice data into text data and extracts elements such as voice pitch, intonation, and pauses.

[0724] "Emotion recognition engine" refers to software that identifies a user's emotional state from the results of voice analysis. It uses machine learning algorithms to recognize emotions.

[0725] "Suggestion Generation Engine" means software for generating suggestions to a user based on recognized emotions.

[0726] "Voice or text format" refers to the presentation method of the suggestions provided to the user, and means speech output using a speech synthesis engine or text message output.

[0727] A "speech synthesis engine" refers to software that converts text data into speech and outputs it in an easy-to-listen format.

[0728] MODE FOR CARRYING OUT THE INVENTION

[0729] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an engine that recognizes the user's emotions. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. Specific embodiments of this system are described below.

[0730] The system mainly includes the following components:

[0731] 1. Voice input acquisition means (terminal): The terminal acquires voice input from the user through a microphone. This voice information is stored in digital format. Terminals include smartphones and PCs.

[0732] 2. Voice analysis means (server): The server receives the voice data sent from the device and analyzes the data using a voice recognition engine (e.g., Google Speech-to-Text API). Through this analysis, elements such as the content of the voice, voice pitch, intonation, and pauses are extracted.

[0733] 3. Emotion recognition means (server including emotion recognition engine): The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state based on the results of voice analysis. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[0734] 4. Proposal generation means (server including a proposal generation engine): The server uses a proposal generation engine (e.g., OpenAI GPT-4) based on the recognized emotions to generate optimal proposals for the user. These proposals are expressed in a way that is empathetic to the user's emotions and easy to accept.

[0735] 5. Proposal providing means (terminal): The terminal receives the proposal sent from the server and provides it to the user. The proposal is notified to the user in voice (using a TTS engine) or text format.

[0736] Specific examples

[0737] 1. Acquiring voice input:

[0738] User: "I've been really busy lately and I can't seem to get rid of my fatigue."

[0739] Device: This audio is recorded using a microphone and stored digitally.

[0740] 2. Sending audio data:

[0741] On the device: The recorded audio data is sent to the server using the HTTPS protocol.

[0742] 3. Audio analysis:

[0743] Server: Receives the voice data and analyzes it using a voice recognition engine. The voice content is converted into text data, and information such as voice pitch, intonation, and pauses is extracted.

[0744] 4. Emotion recognition:

[0745] Server: Inputs the results of voice analysis into an emotion recognition engine to recognize the user's emotion. For example, the analysis results identify the emotion as "tired."

[0746] 5. Proposal generation:

[0747] Server: Using the suggestion generation engine based on the emotion recognition results, generate a suggestion such as "Why not take a deep breath to relax and refresh yourself?"

[0748] 6. Providing suggestions:

[0749] Server: Sends the generated proposals to the device.

[0750] On the device: Providing suggestions to the user in the form of voice or text, such as "Why not try taking a few deep breaths to relax and refresh yourself?"

[0751] Through this series of processes, the system is able to grasp the user's emotions in detail and provide appropriate, empathetic suggestions. As a concrete example, in response to the prompt "I've been so busy lately, I can't shake off my fatigue," the system generates the suggestion "Why don't you take a deep breath and refresh yourself?" In this way, by analyzing emotions from the user's voice input and providing suggestions based on that, the system achieves support that is sensitive to the user's emotions.

[0752] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0753] Program processing flow

[0754] (Step 1: Acquiring voice input)

[0755] The user speaks into the terminal.

[0756] Example: Say, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0757] The device captures the user's voice through a microphone and stores it in digital form as voice data.

[0758] Input: Audio generated by the user's speech

[0759] Output: Digital audio data

[0760] What it does: The microphone picks up audio and records it in an internal buffer.

[0761] (Step 2: Sending audio data)

[0762] The terminal transmits the stored voice data to the server.

[0763] Input: Digitally stored audio data

[0764] Output: Audio data sent to the server

[0765] What it does: Data is sent securely to the server using the HTTPS protocol.

[0766] (Step 3: Audio analysis)

[0767] The server analyzes the voice data received from the terminal.

[0768] Example: Audio data of "I've been so busy lately, I can't get rid of my fatigue"

[0769] The server uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the speech into text, and also extracts information such as voice pitch, intonation, and pauses.

[0770] Input: Audio data sent to the server

[0771] Output: Text data "I've been so busy lately, I can't get rid of my fatigue" and information on voice pitch, intonation, pauses, etc.

[0772] Specific operation: Sends audio data to the API and receives the analysis results in return.

[0773] (Step 4: Emotion Recognition)

[0774] The server uses an emotion recognition engine to identify the user's emotions based on the results of the voice analysis.

[0775] Example: Recognizing the emotion "I'm tired."

[0776] The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the emotional state from the text data and voice elements of the voice analysis results.

[0777] Input: Text data of the voice analysis results and voice pitch, intonation, and pause information

[0778] Output: Emotional state, such as "tired"

[0779] Specific operation: The analysis results are input into the emotion recognition engine and the returned emotional state data is received.

[0780] (Step 5: Proposal Generation)

[0781] The server generates suggestions based on the recognized emotions.

[0782] Example: Suggestion: "Why don't you take a deep breath to relax and refresh yourself?"

[0783] The server uses a proposal generation engine (e.g., OpenAI GPT-4) to create optimal proposals for the user in a way that empathizes with their emotional state.

[0784] Input: Recognized emotional state "tired"

[0785] Output: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[0786] Specific operation: Input the emotional state into the proposal generation engine and obtain the generated proposals.

[0787] (Step 6: Submit your proposal)

[0788] The server transmits the generated proposal to the terminal.

[0789] Input: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[0790] Output: The proposal is sent to the terminal

[0791] Specific operation: Proposal data is sent to the terminal using the HTTPS protocol.

[0792] (Step 7: Provide a proposal)

[0793] The terminal receives the suggestions and provides them to the user in voice or text format.

[0794] Example: A voice notification saying, "Why not take a deep breath to relax and refresh yourself?"

[0795] The device may use a text-to-speech engine (TTS) to communicate the suggestions to the user in audible form or display them in text form.

[0796] Input: Proposal sent by the server

[0797] Output: Suggestions that are communicated to the user in audio or text format

[0798] What it does: Inputs the proposed data into a speech synthesis engine to generate speech output that can be played through a speaker or displayed as text on a screen.

[0799] (Application example 2)

[0800] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0801] Conventional systems that recognize a user's emotions through voice input and make general suggestions based on the results are unable to recommend content that meets the user's specific needs, making it difficult to provide appropriate support based on emotions. In particular, there is a need for a system that can automatically recommend content for relaxation and stress reduction when a user is feeling fatigued or stressed.

[0802] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0803] In this invention, the server includes means for acquiring a user's voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating suggestions that incorporate empathy, means for providing the generated suggestions to the user, means for recommending appropriate content based on the emotion recognition results, means including a recommendation engine for selecting specific content for the emotional state input by voice, and means for notifying the user of the recommended content. This makes it possible to recommend specific content according to the user's emotions, thereby better meeting user needs.

[0804] "Voice input"

[0805] It is voice data that captures what a user says in digital form.

[0806] "Speech recognition engine"

[0807] is a software system that analyzes captured audio data and extracts the content and characteristics of the audio.

[0808] Emotion recognition engine

[0809] It is a system that uses machine learning algorithms to identify a user's emotions based on data analyzed by a voice recognition engine.

[0810] "Proposal generation engine"

[0811] It is a system that generates empathetic suggestions for users based on recognized emotions.

[0812] Recommendation engine

[0813] It is a system that selects and recommends appropriate content according to the user's emotional state.

[0814] "content"

[0815] It refers to information media that can be viewed and played in digital format, such as videos, music, documentaries, and movies.

[0816] "Notification means"

[0817] is an interface for informing users of generated suggestions and recommended content, and has the ability to notify them in voice or text format.

[0818] "Voice pitch"

[0819] is information that indicates frequency changes in audio data, and is also known as the pitch of speech.

[0820] "intonation"

[0821] This is information that indicates the intonation and emotional strength contained in the voice data.

[0822] "Information between"

[0823] is information that indicates the timing of speech and pauses in audio data.

[0824] "Specific content"

[0825] It is a digital media of visual and audio content that best suits the user's emotional state.

[0826] System Overview

[0827] The present invention provides a system for recognizing a user's emotions through voice input, recommending appropriate content based on the results, and providing the content to the user. The system includes a voice input unit, a voice analysis unit, an emotion recognition unit, a suggestion generation unit, a recommendation engine, and a notification unit.

[0828] Hardware and software used

[0829] Hardware:

[0830] Smartphone: Equipped with a microphone to capture voice input, a speaker and a display to announce generated suggestions and recommendations.

[0831] software:

[0832] Speech recognition engine: Analyzes voice data and extracts the content and characteristics of the voice.

[0833] Emotion recognition engine: Based on data analyzed by the voice recognition engine, a machine learning algorithm is used to identify the user's emotions.

[0834] Suggestion generation engine: Generates empathetic suggestions based on recognized emotions.

[0835] Recommendation engine: Selects and recommends appropriate content based on emotion recognition results.

[0836] Notification Method: The user is notified of generated suggestions and recommended content via text or audio.

[0837] Specific examples of processing

[0838] Example 1: A user says to their smartphone, "I've been so busy lately and I can't seem to get rid of my fatigue."

[0839] The server analyzes the voice data and generates text data such as, "I've been really busy lately and I can't get rid of my fatigue."

[0840] The emotion recognition engine identifies "fatigue" as an emotion from the data of the voice recognition engine.

[0841] The suggestion generation engine generates a suggestion such as, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[0842] The recommendation engine selects relaxing video content (e.g., videos of natural scenery or relaxing music) according to the emotion.

[0843] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[0844] Example 2: A user says, "I'm feeling very stressed and irritable today."

[0845] The server analyzes the voice data and generates text data such as "I'm feeling very stressed and irritated today."

[0846] The emotion recognition engine identifies "anger" as an emotion from the data of the voice recognition engine.

[0847] The suggestion generation engine generates a suggestion such as, "You seem stressed. Take a deep breath and refresh yourself."

[0848] The recommendation engine selects video content that helps reduce stress (e.g., relaxation videos or comedy movies) based on emotions.

[0849] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[0850] Prompt Sentence Examples

[0851] You are an emotion recognition engine. Analyze the following user voice inputs and identify the appropriate emotion:

[0852] Dictation: "I'm feeling very stressed and irritable today."

[0853] This allows users to easily receive specific content that corresponds to their emotions, enabling them to receive appropriate support.

[0854] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0855] Step 1:

[0856] The user provides voice input. The user speaks into the smartphone's microphone, expressing their mood and emotions. This voice input is captured in digital form. The input data is voice data, which forms the basis for subsequent processing.

[0857] Step 2:

[0858] The terminal sends the acquired voice data to the server. The input data is voice data and is transferred to the server. At this stage, it is confirmed that the voice data has reached the server.

[0859] Step 3:

[0860] The server analyzes the voice data using a voice recognition engine. Specifically, it converts the voice data into text data and extracts information about the pitch, intonation, and pauses of the voice. The output data is the analysis result, and includes text data and voice feature information.

[0861] Step 4:

[0862] The server uses an emotion recognition engine to identify emotions from the analysis results. The input data is text data and voice feature information, and a machine learning algorithm identifies the user's emotions. The output data is emotional information, such as fatigue or stress.

[0863] Step 5:

[0864] The server uses a proposal generation engine to generate proposals based on emotions and incorporating empathy. The input data is emotional information, and appropriate proposals are generated based on this. The output data is a proposal message, which includes content that shows empathy to the user.

[0865] Step 6:

[0866] The server uses a recommendation engine to recommend content according to emotions. The input data is emotion information, and appropriate content is selected based on this. The output data is content recommendation information, including specific video and audio content.

[0867] Step 7:

[0868] The server sends the generated suggestion message and recommended content to the terminal. The input data is the suggestion message and content recommendation information, which are transferred to the terminal. The output data is notification information on the terminal.

[0869] Step 8:

[0870] The device notifies the user of the suggestion message and the recommended content. The input data is notification information, which is communicated to the user via the smartphone's display or voice output. Specific actions include a text message being displayed on the screen or a voice announcement of the suggestion.

[0871] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0872] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0873] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0874] [Third embodiment]

[0875] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0876] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0877] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0878] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0879] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0880] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0881] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0882] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0883] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0884] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0885] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0886] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0887] System Overview

[0888] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user. The following describes the system's components and processing details in detail.

[0889] System Components

[0890] 1. Voice input acquisition means (terminal):

[0891] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0892] 2. Voice analysis means (server):

[0893] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0894] 3. Emotion recognition means (server):

[0895] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[0896] 4. Proposal generation means (server):

[0897] The server generates optimal suggestions for the user based on the recognized emotions, which are expressed in a manner that is empathetic and easy to accept.

[0898] 5. Proposal providing method (terminal):

[0899] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0900] Example

[0901] 1. Acquiring voice input

[0902] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0903] The device records this voice and sends it to the server as voice data.

[0904] 2. Audio analysis

[0905] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0906] 3. Emotion recognition

[0907] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[0908] 4. Proposal generation

[0909] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0910] 5. Providing suggestions

[0911] The server sends the generated proposal to the terminal.

[0912] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0913] This allows users to receive appropriate advice while having their own emotions and state recognized. The system is able to provide reliable suggestions while empathizing with the user's emotions.

[0914] The processing flow will be explained below.

[0915] Step 1:

[0916] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[0917] Step 2:

[0918] The terminal records the user's voice through a microphone and generates the voice data.

[0919] Step 3:

[0920] The device transmits the recorded audio data to a server via the Internet.

[0921] Step 4:

[0922] The server receives the voice data transmitted from the terminal.

[0923] Step 5:

[0924] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[0925] Step 6:

[0926] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[0927] Step 7:

[0928] Based on the emotion recognition results, the server uses a proposal generation engine to generate optimal proposals for the user, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0929] Step 8:

[0930] The server sends the generated proposal to the terminal.

[0931] Step 9:

[0932] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0933] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[0934] Example 1

[0935] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0936] Conventional speech recognition systems simply convert the user's speech into text data, making it difficult to understand the emotions and nuances behind it. Furthermore, they lacked the means to provide appropriate suggestions based on the user's emotions, making it difficult to show empathy for the user. For this reason, there was a need to develop a system that could understand the user's actual emotions and state and provide appropriate advice based on that.

[0937] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0938] In this invention, the server includes means for acquiring voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions (including means for using a generative AI model), and means for providing the generated suggestions to the user. This makes it possible to appropriately recognize emotions from the user's voice and provide advice that is sensitive to those emotions.

[0939] A "means for obtaining voice input" is a hardware or software feature that digitally records a user's voice and transmits it to other system components.

[0940] The "means for analyzing voice input" is a function for analyzing acquired voice data and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses.

[0941] "Means for recognizing emotions" refers to algorithms or engines that identify the user's emotional state based on the results of voice analysis.

[0942] The "means for generating suggestions" refers to a function for generating appropriate suggestions incorporating empathy based on the recognized emotions, including the use of a generative AI model.

[0943] A "generative AI model" is an artificial intelligence model trained with a large dataset to generate natural language text based on specific input (prompts).

[0944] A "prompt" is text input to a generative AI model, and is an instruction statement that controls the content and format of the generated output.

[0945] System Overview

[0946] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user.

[0947] System Components

[0948] 1. Voice input acquisition means (terminal):

[0949] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[0950] As a specific example, a microphone installed in a smartphone or a personal computer is used.

[0951] 2. Voice analysis means (server):

[0952] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[0953] A common voice recognition engine used is the "voice recognition API."

[0954] 3. Emotion recognition means (server):

[0955] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[0956] A common emotion recognition engine used is the "Emotion Analysis API."

[0957] 4. Proposal generation means (server):

[0958] The server generates optimal suggestions for the user based on the recognized emotions, using a generative AI model to express the suggestions in a way that is empathetic and easy to accept.

[0959] A "natural language generation model" is used as a generative AI model, such as the "GPT (generative pre-trained transformer)" model.

[0960] Example prompt: "Based on the user's voice analysis, we've determined that the user is feeling tired and stressed. Based on this, generate suggestions to help the user relax. Specifically, messages recommending deep breathing or taking a short break would be appropriate."

[0961] 5. Proposal providing method (terminal):

[0962] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[0963] A specific example is a method of notifying by voice using a "speech synthesis engine" on the terminal, for example, using a "text-to-speech engine."

[0964] Example

[0965] 1. Acquiring voice input:

[0966] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0967] The device records this voice and sends it to the server as voice data.

[0968] 2. Audio analysis:

[0969] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[0970] 3. Emotion recognition:

[0971] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[0972] 4. Proposal generation:

[0973] The server generates suggestions based on the emotion recognition results. For example, it generates suggestions such as, "You seem tired lately. Why don't you try taking a deep breath to relax?". It uses a generative AI model for generation.

[0974] 5. Providing suggestions:

[0975] The server sends the generated proposal to the terminal.

[0976] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[0977] In this way, users can receive appropriate advice based on their own emotions and state of mind, enabling the system to provide highly reliable suggestions while showing empathy and understanding of the user's emotions.

[0978] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0979] Step 1: Getting voice input

[0980] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[0981] The device uses a built-in microphone to capture audio and stores it as digital audio data.

[0982] Input: User's voice

[0983] Output: Digital audio data

[0984] What happens: The device application starts the audio stream, buffers the user's voice, and when the recording is finished, a digital audio file is created.

[0985] Step 2: Sending audio data

[0986] The terminal transmits the acquired voice data to the server.

[0987] Input: Digital audio data

[0988] Output: Audio data sent to the server

[0989] Specific operation: The audio data is encoded and sent as a POST request to the server's API endpoint using the HTTPS protocol.

[0990] Step 3: Audio analysis

[0991] The server analyzes the received audio data.

[0992] Input: Audio data sent to the server

[0993] Output: Text data and speech characteristics (voice pitch, intonation, pauses, etc.)

[0994] Specific operation: The server uses a speech recognition engine (e.g., speech recognition API) to convert the voice data into text. It also uses an acoustic analysis module to extract voice characteristics such as pitch, intonation, and pauses. For example, it calls the function "speech_to_text(audio_data)" to obtain the analysis results.

[0995] Step 4: Emotion Recognition

[0996] The server recognizes emotions based on the results of voice analysis.

[0997] Input: Text data and speech features

[0998] Output: Emotion recognition result (e.g., "Tired" or "Stressed")

[0999] Specific behavior: The server uses an emotion recognition engine (e.g., emotion analysis API) to determine the emotional state from text and voice features. For example, if the user's voice is high-pitched and flat, it will recognize that the user is tired and stressed.

[1000] Step 5: Proposal Generation

[1001] The server generates suggestions based on the recognized emotions.

[1002] Input: Emotion recognition results (e.g., "Tired" or "Stressed")

[1003] Output: Generated suggestion message (e.g., "You seem tired lately. Why don't you try taking some deep breaths to relax?")

[1004] Specific operation: Using a generative AI model (e.g., GPT model), generate a suggested message by inputting a prompt sentence. Example of a prompt sentence: "The results of the user's voice analysis indicate that the user is feeling tired and stressed. Based on this state, generate a suggestion that will help the user relax. Specifically, a message recommending deep breathing or a short break would be appropriate."

[1005] Step 6: Submit a proposal

[1006] The server sends the generated proposal to the terminal.

[1007] Input: The generated proposal message

[1008] Output: Proposal message sent to the terminal

[1009] Specific operation: The server sends the generated proposal message to the terminal using the HTTPS protocol.

[1010] The device notifies the user of the suggestion in voice or text format.

[1011] Input: Proposal message sent to the terminal

[1012] Output: Suggestion message provided to the user

[1013] Specific operation: The application on the device receives the suggestion message and notifies it by voice using a speech synthesis engine (e.g., a text-to-speech engine), or displays it on the screen in text format.

[1014] (Application example 1)

[1015] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1016] Conventional food delivery applications do not suggest optimal products or services based on the user's emotions or state, and this has limited the ability to improve user satisfaction and experience. For example, it is not possible to suggest relaxing meals to a tired user, so it has been a challenge to provide appropriate products and services that accurately grasp the user's needs.

[1017] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1018] In this invention, the server includes means for acquiring a voice input from a user, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, means for providing the generated suggestions to the user, and means for recommending appropriate products and services based on the user's emotions. This makes it possible to suggest optimal products and services according to the user's emotions and state.

[1019] A "user" is an individual who uses the system.

[1020] "Voice input" refers to voice information that a user utters to the system.

[1021] "Speech content" refers to the linguistic information in the speech input, that is, the textual information uttered by the user.

[1022] "Voice pitch" is a feature related to the frequency of speech and indicates the tone or pitch of a user's voice.

[1023] "Intonation" refers to the intonation or change in speech, the ups and downs of the voice that indicate the user's emotions or intentions.

[1024] A "pause" refers to a temporary silence or pause in the speech input, which indicates a break or emphasis in what the user is saying.

[1025] The "analysis results" are data obtained after analyzing the voice input and extracting information about the voice content, pitch, intonation, and pauses.

[1026] "Means for recognizing emotions" refers to methods and technologies for identifying a user's emotional state based on the analysis results.

[1027] The "means for generating suggestions" refers to a method or technology for generating appropriate suggestions for a user based on the recognized emotions.

[1028] "Means for providing" refers to the methods and techniques for communicating the generated suggestions to users.

[1029] "Means for recommending products and services" refers to methods and techniques for suggesting appropriate products and services based on the user's emotions.

[1030] The present invention provides a system for analyzing a user's voice input and recognizing emotions to provide appropriate suggestions. The system includes a means for acquiring a user's voice input, a voice analysis means, an emotion recognition means, a suggestion generation means, and a suggestion providing means. The system also includes a means for recommending appropriate products and services based on the user's emotions.

[1031] System configuration:

[1032] Hardware:

[1033] Smartphone or robot: Built-in microphone for voice input

[1034] Server: for data analysis and emotion recognition

[1035] software:

[1036] 1. SpeechRecognition: A speech recognition library that converts voice data into text.

[1037] 2. Transformers (pipeline): Emotion recognition library. Analyzes input text and identifies emotions.

[1038] How it works:

[1039] 1. Voice input acquisition:

[1040] The user speaks into the microphone on their smartphone or the robot, saying, "I'm tired today, so I want to be soothed." The device then digitally records this voice and sends it to the server.

[1041] 2. Audio analysis:

[1042] The server analyzes the received voice data using the SpeechRecognition library, extracting information about the voice content (text data), voice pitch, intonation, and pauses.

[1043] 3. Emotion recognition:

[1044] The server uses Transformers (pipeline) to recognize emotions based on the results of voice analysis, such as "tired" or "stressed."

[1045] 4. Proposal generation:

[1046] The server generates appropriate suggestions based on the recognized emotion, such as "You seem tired lately. How about a relaxing herbal tea or a relaxing restaurant?"

[1047] 5. Providing suggestions:

[1048] The server sends the generated suggestions to the device, which then notifies the user of the suggestions, which may be provided in voice or text format.

[1049] Examples:

[1050] Example 1:

[1051] The user says, "I'm tired today and I want to be soothed."

[1052] The server analyzes this voice and recognizes the emotion "tiredness."

[1053] The server suggests "relaxing herbal teas and restaurants," and the device notifies the user of the suggestions via voice.

[1054] Example prompt sentence:

[1055] User: "I'm feeling great today. What's your recommended meal?"

[1056] App: "That sounds great! How about some pizza or steak to give you some energy?"

[1057] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1058] Step 1:

[1059] The user inputs voice data. The user speaks to the smartphone or robot about their condition and wishes. For example, they might say, "I'm tired today, so I want to be soothed." This voice data becomes the input.

[1060] Step 2:

[1061] The device acquires the voice data and converts it into a digital format. The device records the voice through the microphone and saves it as a digital audio file. The device then sends this data to the server as is. The input is the recorded voice data, and the output is digital voice data.

[1062] Step 3:

[1063] The server receives the voice data and analyzes it using the SpeechRecognition library. This analysis extracts information about the voice content (converts it to text), voice pitch, intonation, and pauses. For example, text data such as "I'm tired today, so I want to be soothed" and voice feature information are output. The input is digital voice data, and the output is text data and voice feature information.

[1064] Step 4:

[1065] The server uses Transformers (pipeline) to perform emotion recognition based on the results of speech analysis. Specifically, it inputs text data and speech feature information into an emotion recognition model to identify emotions such as "tired" or "stressed." The input is the results of speech analysis, and the output is emotional information.

[1066] Step 5:

[1067] The server generates suggestions based on the recognized emotions. Based on the emotion recognition results, for example, if you feel "tired," it will suggest products or services that will help you relax. Specifically, it will generate suggestions such as, "How about some relaxing herbal tea or a relaxing restaurant?" The input is emotional information, and the output is a suggested sentence.

[1068] Step 6:

[1069] The server sends the generated proposal to the terminal. The generated proposal is converted into text or audio format and sent to the terminal. The input is the proposal, and the output is the transmitted proposal data.

[1070] Step 7:

[1071] The device notifies the user of the suggestion. The device presents the received suggestion data to the user in the form of voice or text display. For example, the device might say to the user, "How about some relaxing herbal tea or a relaxing restaurant?" The input is the received suggestion data, and the output is voice or displayed text information.

[1072] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1073] System Overview

[1074] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an emotion recognition engine that recognizes the emotion. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. The following describes the system's components and processing details in detail.

[1075] System Components

[1076] 1. Voice input acquisition means (terminal):

[1077] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[1078] 2. Voice analysis means (server):

[1079] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[1080] 3. Emotion recognition means (server including emotion recognition engine):

[1081] The server uses an emotion recognition engine based on the results of voice analysis to identify the user's emotional state. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[1082] 4. Proposal generation means (server):

[1083] The server uses a proposal generation engine based on the recognized emotions to generate optimal proposals for the user, which are expressed in a manner that is empathetic to the user's emotions and easy to accept.

[1084] 5. Proposal providing method (terminal):

[1085] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[1086] Example

[1087] 1. Acquiring voice input

[1088] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1089] The device records this voice and sends it to the server as voice data.

[1090] 2. Audio analysis

[1091] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[1092] 3. Emotion recognition

[1093] Based on the voice analysis results, the server uses an emotion recognition engine to determine that the user is "tired." The emotion recognition engine uses a machine learning algorithm to identify multiple emotional states based on information such as voice pitch, intonation, and pauses. It also detects changes in emotion by comparing the data with past voice data, allowing for more accurate identification of the user's emotional state.

[1094] 4. Proposal generation

[1095] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you take a deep breath to relax?" The suggestion generation engine is designed to empathize with the user's emotions and express them in an acceptable way.

[1096] 5. Providing suggestions

[1097] The server sends the generated proposal to the terminal.

[1098] The device will notify the user of suggestions via voice or text. For example, the device might say to the user, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[1099] About program processing

[1100] The program for this system operates in the following steps: When a user speaks into the device, the device records the voice and sends the voice data to the server. The server analyzes the voice data and identifies the emotion using an emotion recognition engine. Based on the recognized emotion, a proposal generation engine generates an appropriate proposal and sends it to the device. The device notifies the user of the proposal in voice or text format. Through this series of steps, the user is able to recognize their own emotions and state and receive appropriate advice based on that.

[1101] Specific examples

[1102] 1. Acquiring voice input

[1103] User: Say, "I've been really busy lately and I'm tired every day."

[1104] Terminal: Records the user's voice and sends the voice data to the server.

[1105] 2. Audio analysis

[1106] Server: Analyzes the voice data and generates text data such as "I've been really busy lately and I'm tired every day." It also extracts information about the pitch, intonation, and pauses of the voice.

[1107] 3. Emotion recognition

[1108] Server: Using an emotion recognition engine, the server recognizes the user's emotion as "tired." It uses a machine learning algorithm to analyze the intonation and pauses of the voice and compare them with past data to identify changes in emotion.

[1109] 4. Proposal generation

[1110] Server: Based on the emotion recognition results, it generates a suggestion such as, "Why not take a deep breath to relax and refresh yourself?"

[1111] 5. Providing suggestions

[1112] Server: Sends proposals to the device.

[1113] Device: Speak a suggestion to the user: "Why not take a deep breath to relax and refresh yourself?"

[1114] In this way, this system can analyze the user's voice, recognize their emotions, and generate and provide appropriate suggestions, thereby providing support that is sensitive to the user's emotions.

[1115] The processing flow will be explained below.

[1116] Step 1:

[1117] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[1118] Step 2:

[1119] The terminal records the user's voice through a microphone and generates the voice data.

[1120] Step 3:

[1121] The device transmits the recorded audio data to a server via the Internet.

[1122] Step 4:

[1123] The server receives the voice data transmitted from the terminal.

[1124] Step 5:

[1125] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[1126] Step 6:

[1127] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[1128] Step 7:

[1129] The server uses an emotion recognition engine and machine learning algorithms to recognize emotions, utilizing voice pitch, intonation, and pauses to identify multiple emotional states (e.g., fatigue, stress) of the user.

[1130] Step 8:

[1131] The server identifies changes in the user's emotional state by comparing with past voice data to detect changes in emotion.

[1132] Step 9:

[1133] The server uses a suggestion generation engine to generate optimal suggestions based on the recognized emotions and their changes, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1134] Step 10:

[1135] The server sends the generated proposal to the terminal.

[1136] Step 11:

[1137] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1138] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[1139] Example 2

[1140] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1141] While existing systems can analyze a user's voice input and recognize their emotions, they face the challenge of effectively generating and providing appropriate suggestions based on those emotions. Another problem is that they have not yet fully achieved the ability to identify emotions or states that the user is not aware of and to make suggestions based on those emotions or states.

[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1143] In this invention, the server includes means for digitally storing and transmitting the user's voice input, means for analyzing the voice input and extracting information on the voice content, pitch, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, and means for providing the generated suggestions to the user in voice or text format. This makes it possible to grasp the user's emotions in detail and provide accurate and empathetic suggestions. Furthermore, a series of processes can be efficiently realized, including identifying emotions and states that the user is not aware of based on past data from the voice input and making suggestions based on them.

[1144] "User" refers to an individual or end user who uses the System.

[1145] "Voice input" refers to voice information that a user speaks into a terminal.

[1146] "Device" means a hardware device for recording and storing audio in digital form, such as a smartphone or personal computer.

[1147] "Server" refers to a computer system for receiving voice data and performing analysis and emotion recognition processing.

[1148] A "voice recognition engine" refers to software that converts voice data into text data and extracts elements such as voice pitch, intonation, and pauses.

[1149] "Emotion recognition engine" refers to software that identifies a user's emotional state from the results of voice analysis. It uses machine learning algorithms to recognize emotions.

[1150] "Suggestion Generation Engine" means software for generating suggestions to a user based on recognized emotions.

[1151] "Voice or text format" refers to the presentation method of the suggestions provided to the user, and means speech output using a speech synthesis engine or text message output.

[1152] A "speech synthesis engine" refers to software that converts text data into speech and outputs it in an easy-to-listen format.

[1153] MODE FOR CARRYING OUT THE INVENTION

[1154] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an engine that recognizes the user's emotions. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. Specific embodiments of this system are described below.

[1155] The system mainly includes the following components:

[1156] 1. Voice input acquisition means (terminal): The terminal acquires voice input from the user through a microphone. This voice information is stored in digital format. Terminals include smartphones and PCs.

[1157] 2. Voice analysis means (server): The server receives the voice data sent from the device and analyzes the data using a voice recognition engine (e.g., Google Speech-to-Text API). Through this analysis, elements such as the content of the voice, voice pitch, intonation, and pauses are extracted.

[1158] 3. Emotion recognition means (server including emotion recognition engine): The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state based on the results of voice analysis. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[1159] 4. Proposal generation means (server including a proposal generation engine): The server uses a proposal generation engine (e.g., OpenAI GPT-4) based on the recognized emotions to generate optimal proposals for the user. These proposals are expressed in a way that is empathetic to the user's emotions and easy to accept.

[1160] 5. Proposal providing means (terminal): The terminal receives the proposal sent from the server and provides it to the user. The proposal is notified to the user in voice (using a TTS engine) or text format.

[1161] Specific examples

[1162] 1. Acquiring voice input:

[1163] User: "I've been really busy lately and I can't seem to get rid of my fatigue."

[1164] Device: This audio is recorded using a microphone and stored digitally.

[1165] 2. Sending audio data:

[1166] On the device: The recorded audio data is sent to the server using the HTTPS protocol.

[1167] 3. Audio analysis:

[1168] Server: Receives the voice data and analyzes it using a voice recognition engine. The voice content is converted into text data, and information such as voice pitch, intonation, and pauses is extracted.

[1169] 4. Emotion recognition:

[1170] Server: Inputs the results of voice analysis into an emotion recognition engine to recognize the user's emotion. For example, the analysis results identify the emotion as "tired."

[1171] 5. Proposal generation:

[1172] Server: Using the suggestion generation engine based on the emotion recognition results, generate a suggestion such as "Why not take a deep breath to relax and refresh yourself?"

[1173] 6. Providing suggestions:

[1174] Server: Sends the generated proposals to the device.

[1175] On the device: Providing suggestions to the user in the form of voice or text, such as "Why not try taking a few deep breaths to relax and refresh yourself?"

[1176] Through this series of processes, the system is able to grasp the user's emotions in detail and provide appropriate, empathetic suggestions. As a concrete example, in response to the prompt "I've been so busy lately, I can't shake off my fatigue," the system generates the suggestion "Why don't you take a deep breath and refresh yourself?" In this way, by analyzing emotions from the user's voice input and providing suggestions based on that, the system achieves support that is sensitive to the user's emotions.

[1177] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1178] Program processing flow

[1179] (Step 1: Acquiring voice input)

[1180] The user speaks into the terminal.

[1181] Example: Say, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1182] The device captures the user's voice through a microphone and stores it in digital form as voice data.

[1183] Input: Audio generated by the user's speech

[1184] Output: Digital audio data

[1185] What it does: The microphone picks up audio and records it in an internal buffer.

[1186] (Step 2: Sending audio data)

[1187] The terminal transmits the stored voice data to the server.

[1188] Input: Digitally stored audio data

[1189] Output: Audio data sent to the server

[1190] What it does: Data is sent securely to the server using the HTTPS protocol.

[1191] (Step 3: Audio analysis)

[1192] The server analyzes the voice data received from the terminal.

[1193] Example: Audio data of "I've been so busy lately, I can't get rid of my fatigue"

[1194] The server uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the speech into text, and also extracts information such as voice pitch, intonation, and pauses.

[1195] Input: Audio data sent to the server

[1196] Output: Text data "I've been so busy lately, I can't get rid of my fatigue" and information on voice pitch, intonation, pauses, etc.

[1197] Specific operation: Sends audio data to the API and receives the analysis results in return.

[1198] (Step 4: Emotion Recognition)

[1199] The server uses an emotion recognition engine to identify the user's emotions based on the results of the voice analysis.

[1200] Example: Recognizing the emotion "I'm tired."

[1201] The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the emotional state from the text data and voice elements of the voice analysis results.

[1202] Input: Text data of the voice analysis results and voice pitch, intonation, and pause information

[1203] Output: Emotional state, such as "tired"

[1204] Specific operation: The analysis results are input into the emotion recognition engine and the returned emotional state data is received.

[1205] (Step 5: Proposal Generation)

[1206] The server generates suggestions based on the recognized emotions.

[1207] Example: Suggestion: "Why don't you take a deep breath to relax and refresh yourself?"

[1208] The server uses a proposal generation engine (e.g., OpenAI GPT-4) to create optimal proposals for the user in a way that empathizes with their emotional state.

[1209] Input: Recognized emotional state "tired"

[1210] Output: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[1211] Specific operation: Input the emotional state into the proposal generation engine and obtain the generated proposals.

[1212] (Step 6: Submit your proposal)

[1213] The server transmits the generated proposal to the terminal.

[1214] Input: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[1215] Output: The proposal is sent to the terminal

[1216] Specific operation: Proposal data is sent to the terminal using the HTTPS protocol.

[1217] (Step 7: Provide a proposal)

[1218] The terminal receives the suggestions and provides them to the user in voice or text format.

[1219] Example: A voice notification saying, "Why not take a deep breath to relax and refresh yourself?"

[1220] The device may use a text-to-speech engine (TTS) to communicate the suggestions to the user in audible form or display them in text form.

[1221] Input: Proposal sent by the server

[1222] Output: Suggestions that are communicated to the user in audio or text format

[1223] What it does: Inputs the proposed data into a speech synthesis engine to generate speech output that can be played through a speaker or displayed as text on a screen.

[1224] (Application example 2)

[1225] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1226] Conventional systems that recognize a user's emotions through voice input and make general suggestions based on the results are unable to recommend content that meets the user's specific needs, making it difficult to provide appropriate support based on emotions. In particular, there is a need for a system that can automatically recommend content for relaxation and stress reduction when a user is feeling fatigued or stressed.

[1227] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1228] In this invention, the server includes means for acquiring a user's voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating suggestions that incorporate empathy, means for providing the generated suggestions to the user, means for recommending appropriate content based on the emotion recognition results, means including a recommendation engine for selecting specific content for the emotional state input by voice, and means for notifying the user of the recommended content. This makes it possible to recommend specific content according to the user's emotions, thereby better meeting user needs.

[1229] "Voice input"

[1230] It is voice data that captures what a user says in digital form.

[1231] "Speech recognition engine"

[1232] is a software system that analyzes captured audio data and extracts the content and characteristics of the audio.

[1233] Emotion recognition engine

[1234] It is a system that uses machine learning algorithms to identify a user's emotions based on data analyzed by a voice recognition engine.

[1235] "Proposal generation engine"

[1236] It is a system that generates empathetic suggestions for users based on recognized emotions.

[1237] Recommendation engine

[1238] It is a system that selects and recommends appropriate content according to the user's emotional state.

[1239] "content"

[1240] It refers to information media that can be viewed and played in digital format, such as videos, music, documentaries, and movies.

[1241] "Notification means"

[1242] is an interface for informing users of generated suggestions and recommended content, and has the ability to notify them in voice or text format.

[1243] "Voice pitch"

[1244] is information that indicates frequency changes in audio data, and is also known as the pitch of speech.

[1245] "intonation"

[1246] This is information that indicates the intonation and emotional strength contained in the voice data.

[1247] "Information between"

[1248] is information that indicates the timing of speech and pauses in audio data.

[1249] "Specific content"

[1250] It is a digital media of visual and audio content that best suits the user's emotional state.

[1251] System Overview

[1252] The present invention provides a system for recognizing a user's emotions through voice input, recommending appropriate content based on the results, and providing the content to the user. The system includes a voice input unit, a voice analysis unit, an emotion recognition unit, a suggestion generation unit, a recommendation engine, and a notification unit.

[1253] Hardware and software used

[1254] Hardware:

[1255] Smartphone: Equipped with a microphone to capture voice input, a speaker and a display to announce generated suggestions and recommendations.

[1256] software:

[1257] Speech recognition engine: Analyzes voice data and extracts the content and characteristics of the voice.

[1258] Emotion recognition engine: Based on data analyzed by the voice recognition engine, a machine learning algorithm is used to identify the user's emotions.

[1259] Suggestion generation engine: Generates empathetic suggestions based on recognized emotions.

[1260] Recommendation engine: Selects and recommends appropriate content based on emotion recognition results.

[1261] Notification Method: The user is notified of generated suggestions and recommended content via text or audio.

[1262] Specific examples of processing

[1263] Example 1: A user says to their smartphone, "I've been so busy lately and I can't seem to get rid of my fatigue."

[1264] The server analyzes the voice data and generates text data such as, "I've been really busy lately and I can't get rid of my fatigue."

[1265] The emotion recognition engine identifies "fatigue" as an emotion from the data of the voice recognition engine.

[1266] The suggestion generation engine generates a suggestion such as, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[1267] The recommendation engine selects relaxing video content (e.g., videos of natural scenery or relaxing music) according to the emotion.

[1268] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[1269] Example 2: A user says, "I'm feeling very stressed and irritable today."

[1270] The server analyzes the voice data and generates text data such as "I'm feeling very stressed and irritated today."

[1271] The emotion recognition engine identifies "anger" as an emotion from the data of the voice recognition engine.

[1272] The suggestion generation engine generates a suggestion such as, "You seem stressed. Take a deep breath and refresh yourself."

[1273] The recommendation engine selects video content that helps reduce stress (e.g., relaxation videos or comedy movies) based on emotions.

[1274] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[1275] Prompt Sentence Examples

[1276] You are an emotion recognition engine. Analyze the following user voice inputs and identify the appropriate emotion:

[1277] Dictation: "I'm feeling very stressed and irritable today."

[1278] This allows users to easily receive specific content that corresponds to their emotions, enabling them to receive appropriate support.

[1279] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1280] Step 1:

[1281] The user provides voice input. The user speaks into the smartphone's microphone, expressing their mood and emotions. This voice input is captured in digital form. The input data is voice data, which forms the basis for subsequent processing.

[1282] Step 2:

[1283] The terminal sends the acquired voice data to the server. The input data is voice data and is transferred to the server. At this stage, it is confirmed that the voice data has reached the server.

[1284] Step 3:

[1285] The server analyzes the voice data using a voice recognition engine. Specifically, it converts the voice data into text data and extracts information about the pitch, intonation, and pauses of the voice. The output data is the analysis result, and includes text data and voice feature information.

[1286] Step 4:

[1287] The server uses an emotion recognition engine to identify emotions from the analysis results. The input data is text data and voice feature information, and a machine learning algorithm identifies the user's emotions. The output data is emotional information, such as fatigue or stress.

[1288] Step 5:

[1289] The server uses a proposal generation engine to generate proposals based on emotions and incorporating empathy. The input data is emotional information, and appropriate proposals are generated based on this. The output data is a proposal message, which includes content that shows empathy to the user.

[1290] Step 6:

[1291] The server uses a recommendation engine to recommend content according to emotions. The input data is emotion information, and appropriate content is selected based on this. The output data is content recommendation information, including specific video and audio content.

[1292] Step 7:

[1293] The server sends the generated suggestion message and recommended content to the terminal. The input data is the suggestion message and content recommendation information, which are transferred to the terminal. The output data is notification information on the terminal.

[1294] Step 8:

[1295] The device notifies the user of the suggestion message and the recommended content. The input data is notification information, which is communicated to the user via the smartphone's display or voice output. Specific actions include a text message being displayed on the screen or a voice announcement of the suggestion.

[1296] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1297] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1298] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1299] [Fourth embodiment]

[1300] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1301] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1302] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1303] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1304] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1305] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1306] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1307] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1308] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1309] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1310] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1311] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1312] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1313] System Overview

[1314] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user. The following describes the system's components and processing details in detail.

[1315] System Components

[1316] 1. Voice input acquisition means (terminal):

[1317] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[1318] 2. Voice analysis means (server):

[1319] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[1320] 3. Emotion recognition means (server):

[1321] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[1322] 4. Proposal generation means (server):

[1323] The server generates optimal suggestions for the user based on the recognized emotions, which are expressed in a manner that is empathetic and easy to accept.

[1324] 5. Proposal providing method (terminal):

[1325] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[1326] Example

[1327] 1. Acquiring voice input

[1328] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1329] The device records this voice and sends it to the server as voice data.

[1330] 2. Audio analysis

[1331] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[1332] 3. Emotion recognition

[1333] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[1334] 4. Proposal generation

[1335] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1336] 5. Providing suggestions

[1337] The server sends the generated proposal to the terminal.

[1338] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1339] This allows users to receive appropriate advice while having their own emotions and state recognized. The system is able to provide reliable suggestions while empathizing with the user's emotions.

[1340] The processing flow will be explained below.

[1341] Step 1:

[1342] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[1343] Step 2:

[1344] The terminal records the user's voice through a microphone and generates the voice data.

[1345] Step 3:

[1346] The device transmits the recorded audio data to a server via the Internet.

[1347] Step 4:

[1348] The server receives the voice data transmitted from the terminal.

[1349] Step 5:

[1350] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[1351] Step 6:

[1352] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[1353] Step 7:

[1354] Based on the emotion recognition results, the server uses a proposal generation engine to generate optimal proposals for the user, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1355] Step 8:

[1356] The server sends the generated proposal to the terminal.

[1357] Step 9:

[1358] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1359] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[1360] Example 1

[1361] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1362] Conventional speech recognition systems simply convert the user's speech into text data, making it difficult to understand the emotions and nuances behind it. Furthermore, they lacked the means to provide appropriate suggestions based on the user's emotions, making it difficult to show empathy for the user. For this reason, there was a need to develop a system that could understand the user's actual emotions and state and provide appropriate advice based on that.

[1363] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1364] In this invention, the server includes means for acquiring voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions (including means for using a generative AI model), and means for providing the generated suggestions to the user. This makes it possible to appropriately recognize emotions from the user's voice and provide advice that is sensitive to those emotions.

[1365] A "means for obtaining voice input" is a hardware or software feature that digitally records a user's voice and transmits it to other system components.

[1366] The "means for analyzing voice input" is a function for analyzing acquired voice data and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses.

[1367] "Means for recognizing emotions" refers to algorithms or engines that identify the user's emotional state based on the results of voice analysis.

[1368] The "means for generating suggestions" refers to a function for generating appropriate suggestions incorporating empathy based on the recognized emotions, including the use of a generative AI model.

[1369] A "generative AI model" is an artificial intelligence model trained with a large dataset to generate natural language text based on specific input (prompts).

[1370] A "prompt" is text input to a generative AI model, and is an instruction statement that controls the content and format of the generated output.

[1371] System Overview

[1372] The present invention is a system that analyzes a user's voice input, recognizes their emotions, and generates and provides appropriate suggestions. This system operates through interactions between a terminal, a server, and the user.

[1373] System Components

[1374] 1. Voice input acquisition means (terminal):

[1375] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[1376] As a specific example, a microphone installed in a smartphone or a personal computer is used.

[1377] 2. Voice analysis means (server):

[1378] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[1379] A common voice recognition engine used is the "voice recognition API."

[1380] 3. Emotion recognition means (server):

[1381] The server uses an emotion recognition engine to identify the user's emotional state based on the results of voice analysis. For example, if the voice is high-pitched, it is judged to be excited, and if it is low-pitched, it is judged to be calm. It also understands emotional nuances from intonation and pauses.

[1382] A common emotion recognition engine used is the "Emotion Analysis API."

[1383] 4. Proposal generation means (server):

[1384] The server generates optimal suggestions for the user based on the recognized emotions, using a generative AI model to express the suggestions in a way that is empathetic and easy to accept.

[1385] A "natural language generation model" is used as a generative AI model, such as the "GPT (generative pre-trained transformer)" model.

[1386] Example prompt: "Based on the user's voice analysis, we've determined that the user is feeling tired and stressed. Based on this, generate suggestions to help the user relax. Specifically, messages recommending deep breathing or taking a short break would be appropriate."

[1387] 5. Proposal providing method (terminal):

[1388] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[1389] A specific example is a method of notifying by voice using a "speech synthesis engine" on the terminal, for example, using a "text-to-speech engine."

[1390] Example

[1391] 1. Acquiring voice input:

[1392] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1393] The device records this voice and sends it to the server as voice data.

[1394] 2. Audio analysis:

[1395] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[1396] 3. Emotion recognition:

[1397] The server uses its emotion recognition engine to determine from the results of voice analysis that the user is "tired." It also determines that the user is "stressed" because of the lack of intonation.

[1398] 4. Proposal generation:

[1399] The server generates suggestions based on the emotion recognition results. For example, it generates suggestions such as, "You seem tired lately. Why don't you try taking a deep breath to relax?". It uses a generative AI model for generation.

[1400] 5. Providing suggestions:

[1401] The server sends the generated proposal to the terminal.

[1402] The device will notify the user of the suggestion by voice or text. The device might say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1403] In this way, users can receive appropriate advice based on their own emotions and state of mind, enabling the system to provide highly reliable suggestions while showing empathy and understanding of the user's emotions.

[1404] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1405] Step 1: Getting voice input

[1406] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1407] The device uses a built-in microphone to capture audio and stores it as digital audio data.

[1408] Input: User's voice

[1409] Output: Digital audio data

[1410] What happens: The device application starts the audio stream, buffers the user's voice, and when the recording is finished, a digital audio file is created.

[1411] Step 2: Sending audio data

[1412] The terminal transmits the acquired voice data to the server.

[1413] Input: Digital audio data

[1414] Output: Audio data sent to the server

[1415] Specific operation: The audio data is encoded and sent as a POST request to the server's API endpoint using the HTTPS protocol.

[1416] Step 3: Audio analysis

[1417] The server analyzes the received audio data.

[1418] Input: Audio data sent to the server

[1419] Output: Text data and speech characteristics (voice pitch, intonation, pauses, etc.)

[1420] Specific operation: The server uses a speech recognition engine (e.g., speech recognition API) to convert the voice data into text. It also uses an acoustic analysis module to extract voice characteristics such as pitch, intonation, and pauses. For example, it calls the function "speech_to_text(audio_data)" to obtain the analysis results.

[1421] Step 4: Emotion Recognition

[1422] The server recognizes emotions based on the results of voice analysis.

[1423] Input: Text data and speech features

[1424] Output: Emotion recognition result (e.g., "Tired" or "Stressed")

[1425] Specific behavior: The server uses an emotion recognition engine (e.g., emotion analysis API) to determine the emotional state from text and voice features. For example, if the user's voice is high-pitched and flat, it will recognize that the user is tired and stressed.

[1426] Step 5: Proposal Generation

[1427] The server generates suggestions based on the recognized emotions.

[1428] Input: Emotion recognition results (e.g., "Tired" or "Stressed")

[1429] Output: Generated suggestion message (e.g., "You seem tired lately. Why don't you try taking some deep breaths to relax?")

[1430] Specific operation: Using a generative AI model (e.g., GPT model), generate a suggested message by inputting a prompt sentence. Example of a prompt sentence: "The results of the user's voice analysis indicate that the user is feeling tired and stressed. Based on this state, generate a suggestion that will help the user relax. Specifically, a message recommending deep breathing or a short break would be appropriate."

[1431] Step 6: Submit a proposal

[1432] The server sends the generated proposal to the terminal.

[1433] Input: The generated proposal message

[1434] Output: Proposal message sent to the terminal

[1435] Specific operation: The server sends the generated proposal message to the terminal using the HTTPS protocol.

[1436] The device notifies the user of the suggestion in voice or text format.

[1437] Input: Proposal message sent to the terminal

[1438] Output: Suggestion message provided to the user

[1439] Specific operation: The application on the device receives the suggestion message and notifies it by voice using a speech synthesis engine (e.g., a text-to-speech engine), or displays it on the screen in text format.

[1440] (Application example 1)

[1441] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1442] Conventional food delivery applications do not suggest optimal products or services based on the user's emotions or state, and this has limited the ability to improve user satisfaction and experience. For example, it is not possible to suggest relaxing meals to a tired user, so it has been a challenge to provide appropriate products and services that accurately grasp the user's needs.

[1443] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1444] In this invention, the server includes means for acquiring a voice input from a user, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, means for providing the generated suggestions to the user, and means for recommending appropriate products and services based on the user's emotions. This makes it possible to suggest optimal products and services according to the user's emotions and state.

[1445] A "user" is an individual who uses the system.

[1446] "Voice input" refers to voice information that a user utters to the system.

[1447] "Speech content" refers to the linguistic information in the speech input, that is, the textual information uttered by the user.

[1448] "Voice pitch" is a feature related to the frequency of speech and indicates the tone or pitch of a user's voice.

[1449] "Intonation" refers to the intonation or change in speech, the ups and downs of the voice that indicate the user's emotions or intentions.

[1450] A "pause" refers to a temporary silence or pause in the speech input, which indicates a break or emphasis in what the user is saying.

[1451] The "analysis results" are data obtained after analyzing the voice input and extracting information about the voice content, pitch, intonation, and pauses.

[1452] "Means for recognizing emotions" refers to methods and technologies for identifying a user's emotional state based on the analysis results.

[1453] The "means for generating suggestions" refers to a method or technology for generating appropriate suggestions for a user based on the recognized emotions.

[1454] "Means for providing" refers to the methods and techniques for communicating the generated suggestions to users.

[1455] "Means for recommending products and services" refers to methods and techniques for suggesting appropriate products and services based on the user's emotions.

[1456] The present invention provides a system for analyzing a user's voice input and recognizing emotions to provide appropriate suggestions. The system includes a means for acquiring a user's voice input, a voice analysis means, an emotion recognition means, a suggestion generation means, and a suggestion providing means. The system also includes a means for recommending appropriate products and services based on the user's emotions.

[1457] System configuration:

[1458] Hardware:

[1459] Smartphone or robot: Built-in microphone for voice input

[1460] Server: for data analysis and emotion recognition

[1461] software:

[1462] 1. SpeechRecognition: A speech recognition library that converts voice data into text.

[1463] 2. Transformers (pipeline): Emotion recognition library. Analyzes input text and identifies emotions.

[1464] How it works:

[1465] 1. Voice input acquisition:

[1466] The user speaks into the microphone on their smartphone or the robot, saying, "I'm tired today, so I want to be soothed." The device then digitally records this voice and sends it to the server.

[1467] 2. Audio analysis:

[1468] The server analyzes the received voice data using the SpeechRecognition library, extracting information about the voice content (text data), voice pitch, intonation, and pauses.

[1469] 3. Emotion recognition:

[1470] The server uses Transformers (pipeline) to recognize emotions based on the results of voice analysis, such as "tired" or "stressed."

[1471] 4. Proposal generation:

[1472] The server generates appropriate suggestions based on the recognized emotion, such as "You seem tired lately. How about a relaxing herbal tea or a relaxing restaurant?"

[1473] 5. Providing suggestions:

[1474] The server sends the generated suggestions to the device, which then notifies the user of the suggestions, which may be provided in voice or text format.

[1475] Examples:

[1476] Example 1:

[1477] The user says, "I'm tired today and I want to be soothed."

[1478] The server analyzes this voice and recognizes the emotion "tiredness."

[1479] The server suggests "relaxing herbal teas and restaurants," and the device notifies the user of the suggestions via voice.

[1480] Example prompt sentence:

[1481] User: "I'm feeling great today. What's your recommended meal?"

[1482] App: "That sounds great! How about some pizza or steak to give you some energy?"

[1483] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1484] Step 1:

[1485] The user inputs voice data. The user speaks to the smartphone or robot about their condition and wishes. For example, they might say, "I'm tired today, so I want to be soothed." This voice data becomes the input.

[1486] Step 2:

[1487] The device acquires the voice data and converts it into a digital format. The device records the voice through the microphone and saves it as a digital audio file. The device then sends this data to the server as is. The input is the recorded voice data, and the output is digital voice data.

[1488] Step 3:

[1489] The server receives the voice data and analyzes it using the SpeechRecognition library. This analysis extracts information about the voice content (converts it to text), voice pitch, intonation, and pauses. For example, text data such as "I'm tired today, so I want to be soothed" and voice feature information are output. The input is digital voice data, and the output is text data and voice feature information.

[1490] Step 4:

[1491] The server uses Transformers (pipeline) to perform emotion recognition based on the results of speech analysis. Specifically, it inputs text data and speech feature information into an emotion recognition model to identify emotions such as "tired" or "stressed." The input is the results of speech analysis, and the output is emotional information.

[1492] Step 5:

[1493] The server generates suggestions based on the recognized emotions. Based on the emotion recognition results, for example, if you feel "tired," it will suggest products or services that will help you relax. Specifically, it will generate suggestions such as, "How about some relaxing herbal tea or a relaxing restaurant?" The input is emotional information, and the output is a suggested sentence.

[1494] Step 6:

[1495] The server sends the generated proposal to the terminal. The generated proposal is converted into text or audio format and sent to the terminal. The input is the proposal, and the output is the transmitted proposal data.

[1496] Step 7:

[1497] The device notifies the user of the suggestion. The device presents the received suggestion data to the user in the form of voice or text display. For example, the device might say to the user, "How about some relaxing herbal tea or a relaxing restaurant?" The input is the received suggestion data, and the output is voice or displayed text information.

[1498] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1499] System Overview

[1500] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an emotion recognition engine that recognizes the emotion. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. The following describes the system's components and processing details in detail.

[1501] System Components

[1502] 1. Voice input acquisition means (terminal):

[1503] The device receives voice input from the user through a microphone, and this voice information is stored in digital form.

[1504] 2. Voice analysis means (server):

[1505] The server receives the voice data sent from the device and analyzes it using a voice recognition engine, extracting elements such as the content of the voice, pitch, intonation, and pauses.

[1506] 3. Emotion recognition means (server including emotion recognition engine):

[1507] The server uses an emotion recognition engine based on the results of voice analysis to identify the user's emotional state. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[1508] 4. Proposal generation means (server):

[1509] The server uses a proposal generation engine based on the recognized emotions to generate optimal proposals for the user, which are expressed in a manner that is empathetic to the user's emotions and easy to accept.

[1510] 5. Proposal providing method (terminal):

[1511] The suggestion sent from the server is received by the terminal and provided to the user. The suggestion is notified to the user in voice or text format.

[1512] Example

[1513] 1. Acquiring voice input

[1514] The user speaks to the terminal, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1515] The device records this voice and sends it to the server as voice data.

[1516] 2. Audio analysis

[1517] The server analyzes the received voice data using a speech recognition engine, generating text data such as "I've been so busy lately, I can't get rid of my fatigue," and extracting information such as the pitch, intonation, and pauses of the voice.

[1518] 3. Emotion recognition

[1519] Based on the voice analysis results, the server uses an emotion recognition engine to determine that the user is "tired." The emotion recognition engine uses a machine learning algorithm to identify multiple emotional states based on information such as voice pitch, intonation, and pauses. It also detects changes in emotion by comparing the data with past voice data, allowing for more accurate identification of the user's emotional state.

[1520] 4. Proposal generation

[1521] The server generates suggestions based on the emotion recognition results, such as "You seem tired lately. Why don't you take a deep breath to relax?" The suggestion generation engine is designed to empathize with the user's emotions and express them in an acceptable way.

[1522] 5. Providing suggestions

[1523] The server sends the generated proposal to the terminal.

[1524] The device will notify the user of suggestions via voice or text. For example, the device might say to the user, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[1525] About program processing

[1526] The program for this system operates in the following steps: When a user speaks into the device, the device records the voice and sends the voice data to the server. The server analyzes the voice data and identifies the emotion using an emotion recognition engine. Based on the recognized emotion, a proposal generation engine generates an appropriate proposal and sends it to the device. The device notifies the user of the proposal in voice or text format. Through this series of steps, the user is able to recognize their own emotions and state and receive appropriate advice based on that.

[1527] Specific examples

[1528] 1. Acquiring voice input

[1529] User: Say, "I've been really busy lately and I'm tired every day."

[1530] Terminal: Records the user's voice and sends the voice data to the server.

[1531] 2. Audio analysis

[1532] Server: Analyzes the voice data and generates text data such as "I've been really busy lately and I'm tired every day." It also extracts information about the pitch, intonation, and pauses of the voice.

[1533] 3. Emotion recognition

[1534] Server: Using an emotion recognition engine, the server recognizes the user's emotion as "tired." It uses a machine learning algorithm to analyze the intonation and pauses of the voice and compare them with past data to identify changes in emotion.

[1535] 4. Proposal generation

[1536] Server: Based on the emotion recognition results, it generates a suggestion such as, "Why not take a deep breath to relax and refresh yourself?"

[1537] 5. Providing suggestions

[1538] Server: Sends proposals to the device.

[1539] Device: Speak a suggestion to the user: "Why not take a deep breath to relax and refresh yourself?"

[1540] In this way, this system can analyze the user's voice, recognize their emotions, and generate and provide appropriate suggestions, thereby providing support that is sensitive to the user's emotions.

[1541] The processing flow will be explained below.

[1542] Step 1:

[1543] The user speaks to the device. For example, the user might say, "I've been really busy lately and I can't get rid of my fatigue."

[1544] Step 2:

[1545] The terminal records the user's voice through a microphone and generates the voice data.

[1546] Step 3:

[1547] The device transmits the recorded audio data to a server via the Internet.

[1548] Step 4:

[1549] The server receives the voice data transmitted from the terminal.

[1550] Step 5:

[1551] The server uses a speech recognition engine to convert the received voice data into text data, while simultaneously extracting elements such as the content of the voice, voice pitch, intonation, and pauses.

[1552] Step 6:

[1553] The server uses an emotion recognition engine to identify the user's emotions based on the results of voice analysis (text data, voice pitch, intonation, and pauses). For example, if the user's voice is low and flat, the server may determine that the user is "tired."

[1554] Step 7:

[1555] The server uses an emotion recognition engine and machine learning algorithms to recognize emotions, utilizing voice pitch, intonation, and pauses to identify multiple emotional states (e.g., fatigue, stress) of the user.

[1556] Step 8:

[1557] The server identifies changes in the user's emotional state by comparing with past voice data to detect changes in emotion.

[1558] Step 9:

[1559] The server uses a suggestion generation engine to generate optimal suggestions based on the recognized emotions and their changes, such as "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1560] Step 10:

[1561] The server sends the generated proposal to the terminal.

[1562] Step 11:

[1563] The device notifies the user of the received suggestion in voice or text format. For example, the device may say to the user, "You seem tired lately. Why don't you try taking a deep breath to relax?"

[1564] In this way, the system performs a series of processes to analyze the user's voice, recognize emotions, and generate and provide appropriate suggestions.

[1565] Example 2

[1566] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1567] While existing systems can analyze a user's voice input and recognize their emotions, they face the challenge of effectively generating and providing appropriate suggestions based on those emotions. Another problem is that they have not yet fully achieved the ability to identify emotions or states that the user is not aware of and to make suggestions based on those emotions or states.

[1568] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1569] In this invention, the server includes means for digitally storing and transmitting the user's voice input, means for analyzing the voice input and extracting information on the voice content, pitch, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating empathetic suggestions based on the recognized emotions, and means for providing the generated suggestions to the user in voice or text format. This makes it possible to grasp the user's emotions in detail and provide accurate and empathetic suggestions. Furthermore, a series of processes can be efficiently realized, including identifying emotions and states that the user is not aware of based on past data from the voice input and making suggestions based on them.

[1570] "User" refers to an individual or end user who uses the System.

[1571] "Voice input" refers to voice information that a user speaks into a terminal.

[1572] "Device" means a hardware device for recording and storing audio in digital form, such as a smartphone or personal computer.

[1573] "Server" refers to a computer system for receiving voice data and performing analysis and emotion recognition processing.

[1574] A "voice recognition engine" refers to software that converts voice data into text data and extracts elements such as voice pitch, intonation, and pauses.

[1575] "Emotion recognition engine" refers to software that identifies a user's emotional state from the results of voice analysis. It uses machine learning algorithms to recognize emotions.

[1576] "Suggestion Generation Engine" means software for generating suggestions to a user based on recognized emotions.

[1577] "Voice or text format" refers to the presentation method of the suggestions provided to the user, and means speech output using a speech synthesis engine or text message output.

[1578] A "speech synthesis engine" refers to software that converts text data into speech and outputs it in an easy-to-listen format.

[1579] MODE FOR CARRYING OUT THE INVENTION

[1580] The present invention is a system that analyzes a user's voice input and generates and provides appropriate suggestions using an engine that recognizes the user's emotions. This system combines a terminal, a server, and an emotion recognition engine, and operates through interaction with the user. Specific embodiments of this system are described below.

[1581] The system mainly includes the following components:

[1582] 1. Voice input acquisition means (terminal): The terminal acquires voice input from the user through a microphone. This voice information is stored in digital format. Terminals include smartphones and PCs.

[1583] 2. Voice analysis means (server): The server receives the voice data sent from the device and analyzes the data using a voice recognition engine (e.g., Google Speech-to-Text API). Through this analysis, elements such as the content of the voice, voice pitch, intonation, and pauses are extracted.

[1584] 3. Emotion recognition means (server including emotion recognition engine): The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state based on the results of voice analysis. The emotion recognition engine uses a machine learning algorithm to recognize emotions and identifies multiple emotional states using information on the content of the voice, voice pitch, intonation, and pauses. It also detects changes in emotions by comparing with past voice data.

[1585] 4. Proposal generation means (server including a proposal generation engine): The server uses a proposal generation engine (e.g., OpenAI GPT-4) based on the recognized emotions to generate optimal proposals for the user. These proposals are expressed in a way that is empathetic to the user's emotions and easy to accept.

[1586] 5. Proposal providing means (terminal): The terminal receives the proposal sent from the server and provides it to the user. The proposal is notified to the user in voice (using a TTS engine) or text format.

[1587] Specific examples

[1588] 1. Acquiring voice input:

[1589] User: "I've been really busy lately and I can't seem to get rid of my fatigue."

[1590] Device: This audio is recorded using a microphone and stored digitally.

[1591] 2. Sending audio data:

[1592] On the device: The recorded audio data is sent to the server using the HTTPS protocol.

[1593] 3. Audio analysis:

[1594] Server: Receives the voice data and analyzes it using a voice recognition engine. The voice content is converted into text data, and information such as voice pitch, intonation, and pauses is extracted.

[1595] 4. Emotion recognition:

[1596] Server: Inputs the results of voice analysis into an emotion recognition engine to recognize the user's emotion. For example, the analysis results identify the emotion as "tired."

[1597] 5. Proposal generation:

[1598] Server: Using the suggestion generation engine based on the emotion recognition results, generate a suggestion such as "Why not take a deep breath to relax and refresh yourself?"

[1599] 6. Providing suggestions:

[1600] Server: Sends the generated proposals to the device.

[1601] On the device: Providing suggestions to the user in the form of voice or text, such as "Why not try taking a few deep breaths to relax and refresh yourself?"

[1602] Through this series of processes, the system is able to grasp the user's emotions in detail and provide appropriate, empathetic suggestions. As a concrete example, in response to the prompt "I've been so busy lately, I can't shake off my fatigue," the system generates the suggestion "Why don't you take a deep breath and refresh yourself?" In this way, by analyzing emotions from the user's voice input and providing suggestions based on that, the system achieves support that is sensitive to the user's emotions.

[1603] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1604] Program processing flow

[1605] (Step 1: Acquiring voice input)

[1606] The user speaks into the terminal.

[1607] Example: Say, "I've been really busy lately and I can't seem to get rid of my fatigue."

[1608] The device captures the user's voice through a microphone and stores it in digital form as voice data.

[1609] Input: Audio generated by the user's speech

[1610] Output: Digital audio data

[1611] What it does: The microphone picks up audio and records it in an internal buffer.

[1612] (Step 2: Sending audio data)

[1613] The terminal transmits the stored voice data to the server.

[1614] Input: Digitally stored audio data

[1615] Output: Audio data sent to the server

[1616] What it does: Data is sent securely to the server using the HTTPS protocol.

[1617] (Step 3: Audio analysis)

[1618] The server analyzes the voice data received from the terminal.

[1619] Example: Audio data of "I've been so busy lately, I can't get rid of my fatigue"

[1620] The server uses a speech recognition engine (e.g., Google Speech-to-Text API) to convert the speech into text, and also extracts information such as voice pitch, intonation, and pauses.

[1621] Input: Audio data sent to the server

[1622] Output: Text data "I've been so busy lately, I can't get rid of my fatigue" and information on voice pitch, intonation, pauses, etc.

[1623] Specific operation: Sends audio data to the API and receives the analysis results in return.

[1624] (Step 4: Emotion Recognition)

[1625] The server uses an emotion recognition engine to identify the user's emotions based on the results of the voice analysis.

[1626] Example: Recognizing the emotion "I'm tired."

[1627] The server uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to identify the emotional state from the text data and voice elements of the voice analysis results.

[1628] Input: Text data of the voice analysis results and voice pitch, intonation, and pause information

[1629] Output: Emotional state, such as "tired"

[1630] Specific operation: The analysis results are input into the emotion recognition engine and the returned emotional state data is received.

[1631] (Step 5: Proposal Generation)

[1632] The server generates suggestions based on the recognized emotions.

[1633] Example: Suggestion: "Why don't you take a deep breath to relax and refresh yourself?"

[1634] The server uses a proposal generation engine (e.g., OpenAI GPT-4) to create optimal proposals for the user in a way that empathizes with their emotional state.

[1635] Input: Recognized emotional state "tired"

[1636] Output: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[1637] Specific operation: Input the emotional state into the proposal generation engine and obtain the generated proposals.

[1638] (Step 6: Submit your proposal)

[1639] The server transmits the generated proposal to the terminal.

[1640] Input: Suggestion: "Why not take a deep breath to relax and refresh yourself?"

[1641] Output: The proposal is sent to the terminal

[1642] Specific operation: Proposal data is sent to the terminal using the HTTPS protocol.

[1643] (Step 7: Provide a proposal)

[1644] The terminal receives the suggestions and provides them to the user in voice or text format.

[1645] Example: A voice notification saying, "Why not take a deep breath to relax and refresh yourself?"

[1646] The device may use a text-to-speech engine (TTS) to communicate the suggestions to the user in audible form or display them in text form.

[1647] Input: Proposal sent by the server

[1648] Output: Suggestions that are communicated to the user in audio or text format

[1649] What it does: Inputs the proposed data into a speech synthesis engine to generate speech output that can be played through a speaker or displayed as text on a screen.

[1650] (Application example 2)

[1651] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1652] Conventional systems that recognize a user's emotions through voice input and make general suggestions based on the results are unable to recommend content that meets the user's specific needs, making it difficult to provide appropriate support based on emotions. In particular, there is a need for a system that can automatically recommend content for relaxation and stress reduction when a user is feeling fatigued or stressed.

[1653] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1654] In this invention, the server includes means for acquiring a user's voice input, means for analyzing the voice input and extracting information on the content of the voice, the pitch of the voice, intonation, and pauses, means for recognizing the user's emotions based on the analysis results, means for generating suggestions that incorporate empathy, means for providing the generated suggestions to the user, means for recommending appropriate content based on the emotion recognition results, means including a recommendation engine for selecting specific content for the emotional state input by voice, and means for notifying the user of the recommended content. This makes it possible to recommend specific content according to the user's emotions, thereby better meeting user needs.

[1655] "Voice input"

[1656] It is voice data that captures what a user says in digital form.

[1657] "Speech recognition engine"

[1658] is a software system that analyzes captured audio data and extracts the content and characteristics of the audio.

[1659] Emotion recognition engine

[1660] It is a system that uses machine learning algorithms to identify a user's emotions based on data analyzed by a voice recognition engine.

[1661] "Proposal generation engine"

[1662] It is a system that generates empathetic suggestions for users based on recognized emotions.

[1663] Recommendation engine

[1664] It is a system that selects and recommends appropriate content according to the user's emotional state.

[1665] "content"

[1666] It refers to information media that can be viewed and played in digital format, such as videos, music, documentaries, and movies.

[1667] "Notification means"

[1668] is an interface for informing users of generated suggestions and recommended content, and has the ability to notify them in voice or text format.

[1669] "Voice pitch"

[1670] is information that indicates frequency changes in audio data, and is also known as the pitch of speech.

[1671] "intonation"

[1672] This is information that indicates the intonation and emotional strength contained in the voice data.

[1673] "Information between"

[1674] is information that indicates the timing of speech and pauses in audio data.

[1675] "Specific content"

[1676] It is a digital media of visual and audio content that best suits the user's emotional state.

[1677] System Overview

[1678] The present invention provides a system for recognizing a user's emotions through voice input, recommending appropriate content based on the results, and providing the content to the user. The system includes a voice input unit, a voice analysis unit, an emotion recognition unit, a suggestion generation unit, a recommendation engine, and a notification unit.

[1679] Hardware and software used

[1680] Hardware:

[1681] Smartphone: Equipped with a microphone to capture voice input, a speaker and a display to announce generated suggestions and recommendations.

[1682] software:

[1683] Speech recognition engine: Analyzes voice data and extracts the content and characteristics of the voice.

[1684] Emotion recognition engine: Based on data analyzed by the voice recognition engine, a machine learning algorithm is used to identify the user's emotions.

[1685] Suggestion generation engine: Generates empathetic suggestions based on recognized emotions.

[1686] Recommendation engine: Selects and recommends appropriate content based on emotion recognition results.

[1687] Notification Method: The user is notified of generated suggestions and recommended content via text or audio.

[1688] Specific examples of processing

[1689] Example 1: A user says to their smartphone, "I've been so busy lately and I can't seem to get rid of my fatigue."

[1690] The server analyzes the voice data and generates text data such as, "I've been really busy lately and I can't get rid of my fatigue."

[1691] The emotion recognition engine identifies "fatigue" as an emotion from the data of the voice recognition engine.

[1692] The suggestion generation engine generates a suggestion such as, "You seem tired lately. Why don't you try taking some deep breaths to relax?"

[1693] The recommendation engine selects relaxing video content (e.g., videos of natural scenery or relaxing music) according to the emotion.

[1694] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[1695] Example 2: A user says, "I'm feeling very stressed and irritable today."

[1696] The server analyzes the voice data and generates text data such as "I'm feeling very stressed and irritated today."

[1697] The emotion recognition engine identifies "anger" as an emotion from the data of the voice recognition engine.

[1698] The suggestion generation engine generates a suggestion such as, "You seem stressed. Take a deep breath and refresh yourself."

[1699] The recommendation engine selects video content that helps reduce stress (e.g., relaxation videos or comedy movies) based on emotions.

[1700] The notification means notifies the generated proposal and recommended content by displaying it on the display of the smartphone or by voice.

[1701] Prompt Sentence Examples

[1702] You are an emotion recognition engine. Analyze the following user voice inputs and identify the appropriate emotion:

[1703] Dictation: "I'm feeling very stressed and irritable today."

[1704] This allows users to easily receive specific content that corresponds to their emotions, enabling them to receive appropriate support.

[1705] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1706] Step 1:

[1707] The user provides voice input. The user speaks into the smartphone's microphone, expressing their mood and emotions. This voice input is captured in digital form. The input data is voice data, which forms the basis for subsequent processing.

[1708] Step 2:

[1709] The terminal sends the acquired voice data to the server. The input data is voice data and is transferred to the server. At this stage, it is confirmed that the voice data has reached the server.

[1710] Step 3:

[1711] The server analyzes the voice data using a voice recognition engine. Specifically, it converts the voice data into text data and extracts information about the pitch, intonation, and pauses of the voice. The output data is the analysis result, and includes text data and voice feature information.

[1712] Step 4:

[1713] The server uses an emotion recognition engine to identify emotions from the analysis results. The input data is text data and voice feature information, and a machine learning algorithm identifies the user's emotions. The output data is emotional information, such as fatigue or stress.

[1714] Step 5:

[1715] The server uses a proposal generation engine to generate proposals based on emotions and incorporating empathy. The input data is emotional information, and appropriate proposals are generated based on this. The output data is a proposal message, which includes content that shows empathy to the user.

[1716] Step 6:

[1717] The server uses a recommendation engine to recommend content according to emotions. The input data is emotion information, and appropriate content is selected based on this. The output data is content recommendation information, including specific video and audio content.

[1718] Step 7:

[1719] The server sends the generated suggestion message and recommended content to the terminal. The input data is the suggestion message and content recommendation information, which are transferred to the terminal. The output data is notification information on the terminal.

[1720] Step 8:

[1721] The device notifies the user of the suggestion message and the recommended content. The input data is notification information, which is communicated to the user via the smartphone's display or voice output. Specific actions include a text message being displayed on the screen or a voice announcement of the suggestion.

[1722] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1723] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1724] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1725] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1726] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1727] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1728] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1729] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1730] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1731] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1732] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1733] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1734] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1735] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1736] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1737] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1738] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1739] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1740] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1741] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1742] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1743] The following is further disclosed regarding the above embodiment.

[1744] (Claim 1)

[1745] means for obtaining a user's voice input;

[1746] means for analyzing the voice input and extracting information on voice content, pitch, intonation, and pauses;

[1747] means for recognizing the emotion of the user based on the analysis result;

[1748] means for generating empathetic suggestions based on the recognized emotions;

[1749] means for providing the generated suggestions to a user;

[1750] A system including:

[1751] (Claim 2)

[1752] The method further includes a means for identifying an emotion or state that the user is not aware of from the analysis results.

[1753] 10. The system of claim 1.

[1754] (Claim 3)

[1755] Further includes a means for making suggestions in a manner that shows empathy while being considerate of the user's feelings.

[1756] 10. The system of claim 1.

[1757] "Example 1"

[1758] (Claim 1)

[1759] means for obtaining a user's voice input;

[1760] means for analyzing the voice input and extracting information on voice content, pitch, intonation, and pauses;

[1761] means for recognizing the emotion of the user based on the analysis result;

[1762] A means for generating empathetic suggestions based on the recognized emotions (including using a generative AI model);

[1763] means for providing the generated suggestions to a user;

[1764] A system including:

[1765] (Claim 2)

[1766] The system according to claim 1, further comprising means for identifying an emotion or state that the user is not aware of from the analysis results.

[1767] (Claim 3)

[1768] 10. The system of claim 1, further comprising: means for generating the suggestions using prompt sentences for a generative AI model.

[1769] "Application Example 1"

[1770] (Claim 1)

[1771] means for obtaining a user's voice input;

[1772] means for analyzing the voice input and extracting information on voice content, pitch, intonation, and pauses;

[1773] means for recognizing the emotion of the user based on the analysis result;

[1774] means for generating empathetic suggestions based on the recognized emotions;

[1775] means for providing the generated suggestions to a user;

[1776] A means for recommending appropriate products and services based on the user's emotions;

[1777] A system including:

[1778] (Claim 2)

[1779] The system according to claim 1, further comprising means for identifying an emotion or state that the user is not aware of from the analysis results.

[1780] (Claim 3)

[1781] 2. The system according to claim 1, further comprising means for suggesting products and services in a manner that is sensitive to the user's feelings and shows empathy.

[1782] "Example 2: Combining Emotion Engines"

[1783] (Claim 1)

[1784] means for obtaining a user's voice input;

[1785] means for digitally storing and transmitting said voice input to a server;

[1786] means for analyzing the voice input and extracting information on voice content, pitch, intonation, and pauses;

[1787] means for recognizing the emotion of the user based on the analysis result;

[1788] means for generating empathetic suggestions based on the recognized emotions;

[1789] means for providing said generated suggestions to a user in audio or text format;

[1790] A system including:

[1791] (Claim 2)

[1792] The method further includes a means for identifying an emotion or state that the user is not aware of from the analysis results.

[1793] 10. The system of claim 1.

[1794] (Claim 3)

[1795] The system further includes a means for providing the generated suggestions in a voice format by a voice synthesis engine while being sensitive to the user's emotions.

[1796] 10. The system of claim 1.

[1797] "Application example 2 when combining emotion engines"

[1798] (Claim 1)

[1799] means for obtaining a user's voice input;

[1800] means for analyzing the voice input and extracting information on voice content, pitch, intonation, and pauses;

[1801] means for recognizing the emotion of the user based on the analysis result;

[1802] means for generating empathetic suggestions based on the recognized emotions;

[1803] means for providing the generated suggestions to a user;

[1804] A means of recommending appropriate content based on emotion recognition results,

[1805] means including a recommendation engine for selecting specific content for a voice-input emotional state;

[1806] means for notifying a user of the recommended content;

[1807] A system including:

[1808] (Claim 2)

[1809] The method further includes a means for identifying an emotion or state that the user is not aware of from the analysis results.

[1810] 10. The system of claim 1.

[1811] (Claim 3)

[1812] Further includes a means for making suggestions in a manner that shows empathy while being considerate of the user's feelings.

[1813] 10. The system of claim 1. [Explanation of symbols]

[1814] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for obtaining a user's voice input; means for analyzing the voice input and extracting information on voice content, pitch, intonation, and pauses; means for recognizing the emotion of the user based on the analysis result; means for generating empathetic suggestions based on the recognized emotions; means for providing the generated suggestions to a user; A system including:

2. The method further includes a means for identifying an emotion or state that the user is not aware of from the analysis results. The system of claim 1 .

3. Further includes a means for making suggestions in a manner that shows empathy while being considerate of the user's feelings. The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A