System

A system for individuals with speech disorders converts voice input to text, uses generative AI for personalized responses, and provides voice feedback, addressing communication challenges and enhancing rehabilitation at home.

JP2026021077APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122759
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

People with speech disorders due to illness or accidents face challenges in communicating effectively, leading to social isolation and psychological stress, exacerbated by the scarcity of speech-language-hearing therapists and the lack of effective rehabilitation methods at home.

Method used

A system that acquires user voice input, converts it into text, analyzes it using generative AI for appropriate responses, converts the response into voice, and provides feedback, while saving conversation data for personalized rehabilitation and communication experiences.

Benefits of technology

Enables effective rehabilitation and communication at home by providing personalized responses based on user habits and interests, improving the quality of life for individuals with speech disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021077000001_ABST
    Figure 2026021077000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for acquiring a voice input of a user, a means for converting the acquired voice into a text, a means for analyzing the text by using a generation AI and generating an appropriate response, a means for converting the generated response into a voice, a means for giving voice feedback to the user, and a means for storing conversation date with the user and using it in the next use.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] People with speech disorders due to illness or accidents often have difficulty communicating, leading to social isolation and psychological stress. This problem is exacerbated by the limited number of speech-language-hearing therapists, making it difficult for them to receive the rehabilitation they need. There is a need for a method that allows people with speech disorders to undergo simple and effective rehabilitation at home and maintain a positive attitude. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following means: a means for acquiring a user's voice input, a means for converting the acquired voice into text, a means for analyzing the text using a generation AI and generating an appropriate response, a means for converting the generated response into voice, a means for providing voice feedback to the user, and a means for saving conversation data with the user for future use. This allows people with speech disorders to utilize the generation AI at home for rehabilitation and to enjoy conversations. In particular, the generation AI learns the user's past conversation data and provides personalized responses, achieving a more effective rehabilitation and communication experience. Furthermore, the system analyzes the user's utterances in real time and generates appropriate responses based on the conversation context, providing the user with a natural conversation experience.

[0006] "Voice input" refers to the process of acquiring a voice signal from a user speaking.

[0007] "Speech recognition" is a technology that converts acquired voice input into text data.

[0008] "Generative AI" is an artificial intelligence technology that uses natural language processing to analyze text and automatically generate appropriate responses.

[0009] "Speech synthesis" is a technology that converts text data into a voice signal.

[0010] "Feedback" refers to the process of providing generated responses to the user in real time.

[0011] "Conversation data" refers to data that records the content of a conversation with a user.

[0012] "Personalized responses" are responses that are optimized for a specific user based on the user's past conversation data, habits, and interests.

[0013] "Rehabilitation" refers to training to help people who have suffered language disorders due to illness or accidents recover and improve their language abilities.

[0014] "Context" refers to the background and context of a user's statement, as well as the sequence of statements.

[0015] "Real-time" means immediate processing or response without delay. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. The system is composed of a means for acquiring the user's voice input, converting it into text, generating an appropriate response using a generative AI, and converting the response back into speech to provide feedback to the user. The system also stores conversation data with the user and makes it available for the next use, providing a personalized rehabilitation and communication experience.

[0038] The system program operates while the server and terminals exchange data with each other. Below, we will explain the flow of the system's processing and each step.

[0039] Acquiring and converting voice input

[0040] Voice input begins when a user launches an application and speaks into the microphone. The device captures the voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text, "The weather is nice today."

[0041] Text analysis and response generation

[0042] The converted text is sent from the device to the server. The server receives this text and analyzes it using a generation AI. The generation AI understands the context of the user's remarks and generates an appropriate response. For example, it generates a response such as, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0043] Audio Feedback

[0044] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[0045] Data storage and training

[0046] The server securely stores all conversation data, including the user's utterances, responses, and conversation timestamps, and this data is used in the next conversation. The generative AI uses this data to learn the user's habits and past conversations, and provides personalized responses in the future.

[0047] Specific examples

[0048] For example, if a user launches the application at the same time every day, the generative AI can suggest topics related to that time of day. Also, if the user has previously talked about their favorite movies, the AI ​​can bring up new movies in the next conversation. This allows communication to be tailored to a specific user's interests and lifestyle, improving the effectiveness of rehabilitation.

[0049] In this way, the present invention helps users maintain a positive attitude while undergoing effective rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly and stress-free communication environment.

[0050] The processing flow will be explained below.

[0051] Step 1:

[0052] The user launches the application on the device. Upon launch, the main interface is displayed and the device is ready for voice input.

[0053] Step 2:

[0054] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the user's voice through the microphone.

[0055] Step 3:

[0056] The device calls the speech recognition API and converts the user's speech into text in real time. In this case, the speech "The weather is nice today" is recognized as the text "The weather is nice today."

[0057] Step 4:

[0058] The terminal sends the converted text to the server, where it is queued for analysis.

[0059] Step 5:

[0060] The server passes the received text to the generation AI, which analyzes the content of the text. The generation AI understands the context of the user's remarks and generates the most appropriate response. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0061] Step 6:

[0062] The server generates a response text and sends it to the terminal, where it is prepared for speech output.

[0063] Step 7:

[0064] The device passes the received response text to the speech synthesis API, which converts the text into speech. In this case, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into speech.

[0065] Step 8:

[0066] The device plays the audio output from the speech synthesis API, and the user hears the response, "What a beautiful day. It would be nice to go for a walk on a day like this."

[0067] Step 9:

[0068] The server securely stores all conversation data with the user, including the content of comments, responses, timestamps, etc.

[0069] Step 10:

[0070] The server uses the stored data to train the generative AI to provide optimized responses for the next conversation based on the user's habits and preferences. This learning process results in a more personalized rehabilitation experience.

[0071] Through the above specific steps, the user can undergo rehabilitation through a natural conversation experience.

[0072] Example 1

[0073] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0074] To effectively rehabilitate users with speech disorders at home and improve their communication skills, a system that generates personalized responses and provides a seamless conversation experience is needed. However, current technology does not adequately provide the means to analyze users' utterances in real time and generate appropriate responses. Furthermore, there are only a limited number of systems that utilize past conversation data to provide responses tailored to the user. This results in issues such as users not being able to fully benefit from rehabilitation and a decline in the quality of their communication.

[0075] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0076] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generative AI model and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time, means for acquiring voice input in real time and performing preprocessing, means for converting voice into text using a voice recognition API, means for transmitting the text data to the server using a secure communication protocol, means for inputting a prompt sentence into the generative AI model for analysis and generating a response, means for transmitting the response data to the terminal using a secure communication protocol, means for converting the response text into voice using a voice synthesis API, and means for learning the user's habits and past conversation content based on the saved conversation data. This enables the user to undergo effective rehabilitation at home and improve the quality of communication by receiving personalized responses.

[0077] The "means for acquiring voice input" refers to a device or method for acquiring the voice uttered by the user as digital voice data through an input device such as a microphone.

[0078] "Means for converting captured speech to text" refers to a device or method that uses a speech recognition API or software to convert captured speech data into corresponding text data.

[0079] A "generative AI model" is an artificial intelligence that uses natural language processing techniques to analyze text and generate appropriate responses.

[0080] A "means for analyzing text and generating appropriate responses" is a device or method for analyzing input text data using a generative AI model and automatically generating appropriate responses.

[0081] A "means for converting generated responses into speech" is a device or method that uses a speech synthesis API or software to convert generated text responses into speech data.

[0082] The "means for providing audio feedback to the user" refers to a device or method for playing back the generated audio data to the user through an output device such as a speaker of the terminal.

[0083] The "means for storing conversation data with a user" refers to a device or method for recording and storing the user's utterances, responses, and related data in digital form.

[0084] A "next use method" is a device or method for referencing saved conversation data in a next session to generate a response that takes past interactions into account.

[0085] The "means for acquiring and pre-processing speech input in real time" refers to a device or method for instantly acquiring speech input from a user and performing pre-processing such as noise removal and data shaping.

[0086] A "means for converting speech to text using a speech recognition API" is a device or method for utilizing a speech recognition service to instantly convert captured speech data into corresponding text data.

[0087] "Means for transmitting text data to a server using a secure communication protocol" means a device or method for transmitting text data from a terminal to a server using an encrypted protocol (e.g., HTTPS) to ensure the security of the communication.

[0088] "Means for inputting a prompt sentence into a generative AI model, analyzing it, and generating a response" refers to a device or method for providing an appropriate input sentence (prompt sentence) to a generative AI model and generating a response based on the analysis results.

[0089] "Means for transmitting response data to a terminal using a secure communication protocol" refers to a device or method for transmitting generated response data from a server to a terminal using an encrypted protocol (e.g., HTTPS) to ensure the security of communications.

[0090] "Means for converting response text into speech using a speech synthesis API" refers to a device or method for converting generated response text into speech data using a speech synthesis service.

[0091] "Means for learning user habits and past conversation content based on saved conversation data" refers to a device or method for generating more appropriate responses by analyzing saved conversation history and learning the user's speech patterns and past conversation content.

[0092] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. This system is composed of a series of means to acquire the user's voice input, convert it into text, analyze and generate a response using a generative AI model, and convert the response into speech and provide feedback to the user. Furthermore, conversation data with the user is saved and used the next time the system is used, providing a personalized rehabilitation and communication experience.

[0093] Hardware and software used

[0094] This system uses the following hardware and software:

[0095] Hardware: microphone, speaker, device (smartphone, tablet, PC, etc.)

[0096] Software: Speech recognition API (general name example: speech recognition service), generative AI model (general name example: natural language generation engine), speech synthesis API (general name example: speech synthesis service), secure communication protocol (e.g., HTTPS)

[0097] Acquiring and converting voice input

[0098] When a user launches an application on their device, voice input begins by speaking into the microphone. The device captures the voice through a built-in or external microphone and preprocesses the captured voice data in real time. Once preprocessed, the voice data is converted into text data using a speech recognition API.

[0099] Text analysis and response generation

[0100] The converted text data is sent from the device to the server using a secure communication protocol (HTTPS). The server inputs the received text data into the generative AI model for analysis. The generative AI model understands the context of the user's remarks and generates an appropriate response. For example, if the user says, "I'm tired today," the generative AI model will generate a response such as, "Please get plenty of rest. I think you'll feel better tomorrow."

[0101] Audio Feedback

[0102] The generated response text is sent from the server to the device using a secure communication protocol. The device inputs the received response text into a speech synthesis API and converts it into speech. This converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[0103] Data storage and training

[0104] The server securely stores all conversation data (user utterances, responses, timestamps, etc.). The stored data is used in the next conversation. The generative AI model uses the stored data to learn the user's habits and past conversations, and can provide personalized responses in future conversations. For example, by revisiting a movie the user previously mentioned, the model can create a continuous conversation and increase familiarity with the user.

[0105] Specific examples

[0106] When a user says, "Good morning, what shall we talk about today?", the voice is picked up through the microphone. The device's speech recognition API converts this voice into text, and the text data "Good morning, what shall we talk about today?" is sent to the server. The server's generative AI model analyzes it and generates a response: "Good morning! How about we talk about a new movie today?" This response is converted into speech by the speech synthesis API and played through the device's speaker.

[0107] Prompt Sentence Examples

[0108] Below are some example prompts to input to a generative AI model:

[0109] If the user says "I'm tired today," the prompt is:

[0110] Text format: "User said 'I am tired today'. Please generate an encouraging response."

[0111] In this way, the present invention provides users with effective rehabilitation at home and a seamless, personalized communication experience.

[0112] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0113] Step 1:

[0114] The user launches an application on the device and speaks into the microphone.

[0115] Input: User's voice data

[0116] How it works: The device captures audio in real time through the microphone.

[0117] Output: Audio data

[0118] Step 2:

[0119] Preprocessing is performed on the voice data acquired by the terminal.

[0120] Input: Acquired audio data

[0121] What it does: Performs preprocessing such as noise removal and volume normalization.

[0122] Output: Preprocessed audio data

[0123] Step 3:

[0124] The device uses a speech recognition API to convert the preprocessed voice data into text data.

[0125] Input: Preprocessed audio data

[0126] What it does: Calls a speech recognition API and converts the audio data into text data, for example, using the Google Cloud Speech-to-Text API.

[0127] Output: Text data

[0128] Step 4:

[0129] The terminal transmits the converted text data to the server using a secure communication protocol (HTTPS).

[0130] Input: Text data

[0131] How it works: Uses HTTPS to securely transmit text data.

[0132] Output: Text data received by the server

[0133] Step 5:

[0134] The text data received by the server is input into the generative AI model, where it is analyzed and an appropriate response is generated.

[0135] Input: Text data

[0136] How it works: A prompt is input to a generative AI model (e.g., a natural language generation engine), which then analyzes it and generates an appropriate response.

[0137] Output: Response text data

[0138] Step 6:

[0139] The server transmits the generated response text to the terminal using a secure communication protocol.

[0140] Input: Response text data

[0141] What it does: Uses HTTPS to securely transmit response text.

[0142] Output: Response text data received by the device

[0143] Step 7:

[0144] The device converts the received response text data into voice data using a speech synthesis API.

[0145] Input: Response text data

[0146] Behavior: Calls a speech synthesis API (e.g., a speech synthesis service) to convert text data into speech data.

[0147] Output: Audio data

[0148] Step 8:

[0149] The terminal plays the converted voice data through a speaker and provides feedback to the user.

[0150] Input: Audio data

[0151] What it does: Plays a sound through the speaker, providing feedback to the user.

[0152] Output: Audio feedback the user hears

[0153] Step 9:

[0154] The server securely stores all conversation data and uses it for the next conversation.

[0155] Input: Conversation data (user utterances, responses, timestamps)

[0156] What it does: Securely stores conversation data and keeps it in a format that can be learned and used by generative AI models.

[0157] Output: Saved conversation data

[0158] Step 10:

[0159] The server learns the user's habits and past conversation content based on the saved conversation data and personalizes the next response.

[0160] Input: Saved conversation data

[0161] How it works: Generative AI models analyze past conversation data, learn user habits and tendencies, and provide personalized responses the next time you talk to them.

[0162] Output: personalized response

[0163] Through these steps, users can undergo effective rehabilitation at home and experience seamless, personalized communication.

[0164] (Application example 1)

[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0166] The goal of this project is to solve the problem that users with speech disorders have difficulty communicating smoothly with employees and customers in physical stores, which causes problems in their daily work. Furthermore, conventional rehabilitation systems were unable to flexibly respond to the individual circumstances of each user and were therefore unable to demonstrate sufficient effectiveness.

[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0168] In this invention, the server includes means for acquiring user voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time the server is used, and means for supporting real-time communication between employees and customers. This enables users with speech disabilities to communicate smoothly in physical stores and interact efficiently with employees and customers.

[0169] The "means for acquiring user's voice input" refers to a device or method for recognizing the voice spoken by the user and capturing it as a digital signal.

[0170] The term "means for converting captured speech to text" refers to a device or method for converting captured speech data into corresponding text data using speech recognition technology.

[0171] "Means for analyzing text and generating appropriate responses using generative AI" refers to devices and methods that use artificial intelligence technology to analyze input text and automatically generate appropriate responses to that text.

[0172] The "means for converting the generated response into speech" refers to a device or method including speech synthesis technology for converting text data into speech data and outputting it as speech.

[0173] A "means for providing audio feedback to a user" is a device or method for communicating the generated audio response to a user through a speaker or headset.

[0174] "Means for saving conversation data with the user and using it the next time" refers to a device or method for saving the dialogue history with the user in a database or the like, and referencing that data during future dialogues to provide personalized responses.

[0175] "Means for supporting real-time communication between employees and customers" refers to devices and methods for generating responses in real time and providing feedback to facilitate smooth dialogue between employees and customers in physical stores.

[0176] This invention is a system that enables users with speech impediments to communicate smoothly with employees and customers in brick-and-mortar stores. This system combines speech recognition, generative AI, and speech synthesis technologies, and is described in detail below.

[0177] Acquiring and converting voice input

[0178] First, when a user starts the application and speaks into the microphone, the device's microphone picks up the audio. The device uses the speech_recognition library to convert this audio data into text. For example, if a user says, "What products do you recommend?", the audio is converted into text data that reads, "What products do you recommend?"

[0179] Text analysis and response generation

[0180] The converted text data is sent from the device to the server, which uses the OpenAI library to analyze the text data and generate an appropriate response. The generation AI understands the context of the user's utterance and generates a response such as "The current recommended product is fresh fruit."

[0181] Audio Feedback

[0182] The generated text response is sent from the server to the device, which then uses the gTTS library to convert the response into speech, which is then played back through the device's speaker and fed back to the user, allowing for seamless communication.

[0183] Data storage and training

[0184] The server securely stores all conversation data, including user utterances, responses, and conversation timestamps, which are referenced the next time the conversation occurs. Using this data, the generative AI can provide personalized responses, making communication more user-friendly.

[0185] Specific examples

[0186] When a user speaks to the device, asking, "What is your recommended product?", the speech is converted into text and sent to the server. The server analyzes the speech and generates a response such as, "Our current recommended product is fresh fruit." This response is then converted into speech and fed back to the user. Using this system, users can communicate smoothly in physical stores.

[0187] Prompt Sentence Examples

[0188] "User Question: What products do you recommend?"

[0189] "Context: A customer and an employee are having a conversation in a store."

[0190] "Answer: Our current recommended product is fresh fruit."

[0191] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0192] Step 1:

[0193] When a user launches an application and speaks into the microphone, the device receives voice input. Specifically, the microphone device detects the user's speech and captures it as digital audio data. This audio data is then sent to the speech_recognition library and converted into text data using Google's speech recognition API. The input is the user's voice, and the output is the converted text data.

[0194] Step 2:

[0195] The acquired text data is sent from the device to the server. The server receives this text data and analyzes it using the OpenAI library. A generative AI model analyzes the text based on the prompt and generates an appropriate response. The input is the text data, and the output is the generated response text. For example, if the input is "What products do you recommend?", the output will be "Our current recommended products are fresh fruit."

[0196] Step 3:

[0197] The generated response text is sent from the server to the device. The device receives this text and converts it into voice data using the gTTS library. Specifically, the text is converted into a digital audio file using speech synthesis technology, and then converted into a format that can be played back through a speaker. The input is the generated response text, and the output is voice data.

[0198] Step 4:

[0199] The voice data is fed back to the user through the device's speaker. The user can listen to this voice and continue the conversation. Specifically, the device's speaker plays the generated voice data and transmits it to the user. The input is the voice data, and the output is the voice that the user hears.

[0200] Step 5:

[0201] The server securely stores all conversation data. The stored data includes the user's utterances, the generated responses, and the conversation timestamp. By referencing this data during the next interaction, the generative AI model can provide a more personalized response. The input is the conversation data with the user, and the output is the stored data.

[0202] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0203] This invention is a system for enabling users with speech disorders to effectively rehabilitate at home and maintain a positive attitude, and aims to achieve more human-like and natural communication by combining it with an emotion engine that recognizes the user's emotions. This system includes a means for acquiring the user's voice input, converting it into text, generating an appropriate response using generative AI, converting the response back into voice and providing feedback to the user, as well as an emotion engine that recognizes the user's emotions.

[0204] The system program operates by exchanging data between the server, terminals, and emotion engine. Below, we will explain the process flow of this system and each step in detail.

[0205] Acquiring and converting voice input

[0206] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text "The weather is nice today."

[0207] Text Analysis and Emotion Recognition

[0208] The converted text is sent from the device to the server. The server receives this text and first analyzes it using an emotion engine. The emotion engine analyzes the user's emotions from the tone of their voice and the content of their speech, and passes that emotion data to the generation AI. For example, if the user is speaking happily, the emotion engine will analyze the emotion as "joy."

[0209] Response Generation

[0210] The generative AI takes into account the text content and emotional data to generate an appropriate response: in this case, "What a lovely day. It would be nice to go for a walk on a day like this," in a bright, joyful tone of voice.

[0211] Audio Feedback

[0212] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless and emotionally relevant conversation.

[0213] Data storage and training

[0214] The server securely stores all conversation data, including the content of statements, responses, emotional data, and timestamps, which are then used in the next conversation. The generative AI and emotion engine use this data to learn the user's habits and emotional history and provide personalized responses in future conversations.

[0215] Specific examples

[0216] For example, if a user has previously spoken under stress, the emotion engine can use that history to generate responses that will reduce stress in the next conversation. If a user says, "I'm tired from work today," the system could provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0217] In this way, the present invention helps users maintain a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to significantly improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotion-sensitive communication.

[0218] The processing flow will be explained below.

[0219] Step 1:

[0220] The user launches an application on the device. When the application launches, the main interface is displayed and the device is ready for voice input.

[0221] Step 2:

[0222] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the voice through the microphone.

[0223] Step 3:

[0224] The device calls the speech recognition API and converts the user's speech into text in real time. Specifically, the speech "The weather is nice today" is converted into the text "The weather is nice today."

[0225] Step 4:

[0226] The terminal sends the converted text to the server, where it is queued for further processing.

[0227] Step 5:

[0228] The server passes the received text to the emotion engine for emotion analysis. The emotion engine analyzes the user's emotion (e.g., joy, sadness, anger, etc.) from the user's tone of voice and the content of the speech.

[0229] Step 6:

[0230] The server receives the emotion data generated by the emotion engine and passes the emotion data along with the text to the generation AI. The generation AI generates an appropriate response based on the user's statement and emotion. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0231] Step 7:

[0232] The server generates a response text and sends it to the device, which then passes the text to the speech synthesis API.

[0233] Step 8:

[0234] The device uses a speech synthesis API to convert the response text into speech, in this case the text "What a lovely day. It would be nice to go for a walk on a day like this" in a bright, happy tone.

[0235] Step 9:

[0236] The device plays the generated speech and provides feedback to the user, who hears the response, "What a lovely day. It would be nice to go for a walk on a day like this."

[0237] Step 10:

[0238] The server securely stores all conversation data with the user, including the content of the conversation, responses, emotional data, and timestamps.

[0239] Step 11:

[0240] The server uses the stored conversation data to train the generative AI and emotion engine so that in future conversations, it can provide personalized responses that take into account the user's habits and emotional history.

[0241] For example, if a user has previously spoken while feeling stressed, the system can generate a response that reduces stress in the next conversation based on that history. If the user says, "I'm tired from work today," the system can provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" In this way, users can enjoy a more personalized and emotionally relevant conversational experience.

[0242] These specific steps allow users to undergo natural and effective rehabilitation, and emotionally-sensitive responses can help users develop a more positive outlook.

[0243] Example 2

[0244] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0245] Conventional dialogue systems using speech recognition systems or generative AI models have difficulty in accurately grasping a user's emotions and reflecting them in responses, making it difficult to provide natural and personalized dialogue, especially for users with speech impediments. Furthermore, there are insufficient means to provide more personalized responses by utilizing a user's past conversation data or emotional history.

[0246] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotion and passing the emotion data to the generative AI model, a means for analyzing the user's emotion in real time and reflecting it in a response, and a means for learning the user's past conversation data and providing a personalized response. This enables natural and personalized dialogue that takes the user's emotion into consideration.

[0247] "User" refers to an individual who uses the system to provide voice input.

[0248] A "server" is a central computing device for performing processing, storing data, and communicating with other components.

[0249] A "terminal" is a device used by a user that has the functions of voice input / output and data transmission / reception.

[0250] The "means for acquiring voice input" is a method for collecting the user's voice through a device such as a microphone.

[0251] "Means for converting voice to text" refers to a method of converting acquired voice data into text information using a voice recognition API.

[0252] "Means of analyzing text using a generative AI model and generating an appropriate response" refers to a method that uses generative AI, an algorithm for providing an appropriate response based on text data.

[0253] The "means of converting the generated response into speech" refers to a method of converting text data into speech data using a speech synthesis API or the like.

[0254] "Means for providing audio feedback to the user" refers to a method of returning the converted audio data to the user via a speaker or the like.

[0255] "Means for saving conversation data and using it the next time" refers to a method for recording the content and emotional data of past conversations and using them in future interactions.

[0256] "Means for recognizing emotions and passing that emotional data to a generative AI model" refers to a method for analyzing emotions from user speech and providing that information to the generative AI.

[0257] "Means for analyzing emotions in real time and reflecting them in responses" refers to a method for analyzing emotions simultaneously with user utterances and immediately incorporating the results into responses.

[0258] This invention is a system that allows users to receive emotional feedback while undergoing speech rehabilitation at home. The system converts the user's voice input into text and generates an appropriate response based on that text and emotional data. The system also converts the generated response into speech and provides feedback to the user. The system also incorporates an emotion recognition engine and has the function of analyzing the user's emotions. The specific processing flow of the system is described below.

[0259] Acquiring and converting voice input

[0260] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and calls a speech recognition API (e.g., a speech recognition API from a major cloud service provider) to convert the voice into text.

[0261] For example, if a user says, "The weather is nice today," the device captures the audio and uses a speech recognition API to convert it into text: "The weather is nice today."

[0262] Text analysis and emotion recognition

[0263] The converted text is sent from the device to a server, which then passes the received text to an emotion recognition engine (e.g., the API of a major emotion analysis service). The emotion recognition engine analyzes the user's emotions from the tone of voice and the content of the speech, and passes the emotional data to a generative AI model.

[0264] For example, if a user is speaking with a happy expression, the emotion recognition engine will analyze it as "joy." The server receives this analysis result and passes it to the generative AI model.

[0265] Response Generation

[0266] Generative AI models (e.g., advanced trained text generation models) generate appropriate responses based on text content and sentiment data.

[0267] In this case, the generated response would be "What a lovely day. It would be nice to go for a walk on a day like this," with a tone that reflects joy based on the emotion data.

[0268] Response transcription and feedback

[0269] The generated response text is sent from the server to the device, which then uses a speech synthesis API (e.g., a service that provides advanced speech synthesis technology) to convert the response text into speech, which is then played back through the device's speaker and provided as feedback to the user.

[0270] This process allows users to enjoy seamless and emotionally relevant interactions.

[0271] Data storage and training

[0272] The server securely stores all conversation data (statements, responses, emotional data, timestamps, etc.) The generative AI model and emotion recognition engine use the stored data to learn the user's habits and emotional history, providing personalized responses for future conversations.

[0273] Examples and prompts

[0274] For example, if a user has previously said, "I'm tired from work today," the emotion recognition engine can detect "stress," and in the next conversation, the generative AI model can generate a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0275] Examples of prompts are:

[0276] User: I'm tired from work today.

[0277] Emotion: Stress

[0278] Prompt: The user says "I'm tired from work today" and is feeling stressed. Generate an appropriate response accordingly.

[0279] In this way, the present invention is a system that supports users in maintaining a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotionally sensitive communication.

[0280] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0281] Step 1: Getting voice input

[0282] When a user starts an application and speaks into the microphone, voice input begins. The device acquires the user's voice data through the microphone. The input is the user's voice, which is acquired by the device as voice data. For example, when a user says, "The weather is nice today," the device collects this voice data.

[0283] Step 2: Speech to text

[0284] The device sends the acquired voice data to a voice recognition API, which converts the voice into text. The input is voice data, and the output is the corresponding text data. For example, voice data such as "The weather is nice today" is converted into text data such as "The weather is nice today."

[0285] Step 3: Sending text data

[0286] The terminal sends the converted text data to the server. The input is text data, which is sent to the server as is. For example, the text "The weather is nice today" is sent.

[0287] Step 4: Emotion Recognition

[0288] The server passes the received text to an emotion recognition engine to analyze the emotion. The input is text data, and the output is emotion data. For example, the emotion recognition engine analyzes the utterance "The weather is nice today" and generates emotion data for "joy."

[0289] Step 5: Input to the generative AI

[0290] The server passes the analyzed text data and emotion data to the generative AI model. The input is text data and emotion data, which are passed to the generative AI model. For example, the text "The weather is nice today" and the emotion data "joy" are input to the generative AI model.

[0291] Step 6: Generate a response

[0292] The generative AI model generates an appropriate response based on text data and emotional data. The input to the generative AI model is text data and emotional data, and the output is a response text. For example, the generated response text might be, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0293] Step 7: Transcribing the response

[0294] The server sends the generated response text to the device, and the device sends the response text to the speech synthesis API to convert it into voice data. The input is the response text, and the output is voice data. For example, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into voice data.

[0295] Step 8: User Feedback

[0296] The device plays the converted voice data from the speaker and provides feedback to the user. The input is voice data, and the output is the voice played back to the user. For example, the user may receive feedback in a bright tone saying, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0297] Step 9: Save your data

[0298] The server securely stores all conversation data. Inputs include utterances, responses, emotional data, and timestamps, and these are stored. For example, the utterance "The weather is nice today," the response "It really is nice weather," the emotional data "Joy," and the associated timestamps are stored.

[0299] Step 10: Data training

[0300] The generative AI model and emotion recognition engine use stored data to learn the user's habits and emotional history and reflect this in the next interaction. The input is stored interaction data, and learning from this improves the next response. For example, if a user previously said, "I'm tired from work today," the next time they make a similar statement, a more personalized response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" will be provided.

[0301] (Application example 2)

[0302] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0303] In modern society, it is not easy for users with speech disorders to undergo effective rehabilitation at home and maintain a positive attitude. Therefore, there is a need to develop systems that can recognize users' emotions and provide appropriate responses. It is also necessary to realize systems that are highly secure and can be used with peace of mind. It is particularly important to develop systems that can detect when a user feels anxiety or stress and respond appropriately.

[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0305] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for recognizing the user's emotions, means for generating a response based on emotion data, and means for saving conversation data with the user and using it the next time the server is used. This enables users with speech disorders to undergo effective rehabilitation at home and communicate naturally in accordance with their emotions.

[0306] "Means for obtaining voice input" refers to a device or software that captures and records the user's speech.

[0307] "Means for converting voice to text" refers to technology that analyzes acquired voice data and converts it into a corresponding text format.

[0308] "Means of using generative AI to analyze text and generate appropriate responses" refers to a process of using artificial intelligence technology to analyze input text and generate appropriate responses.

[0309] "Means for converting the generated response to speech" refers to technology for converting the generated text response to speech.

[0310] "Audio feedback means" refers to a device or software that plays the generated audio to the user.

[0311] "Means for recognizing emotions" refers to technology that analyzes and understands emotions from a user's voice or text.

[0312] "Means for generating a response based on emotional data" refers to a technology that uses recognized emotional data to generate an appropriate response that is in line with the user's emotions.

[0313] "Means for saving conversation data and using it the next time" refers to a system that records conversations and emotional data with the user and saves it for reference the next time the user uses the system.

[0314] This invention relates to a system that receives voice input, converts it into text, generates appropriate responses using generative AI, and provides voice feedback to the user. Furthermore, by recognizing the user's emotions and generating responses based on them, it achieves more natural communication.

[0315] System configuration:

[0316] Hardware

[0317] Device: A device used by a user, such as a smartphone or computer, has a built-in microphone and speaker.

[0318] Server: A device that performs heavy processing, such as a cloud server.

[0319] software

[0320] Speech Recognition API: Installed on the device, it collects the user's voice and converts it into text.

[0321] Generative AI: Runs on a server and analyzes input text to generate appropriate responses, for example using natural language processing models.

[0322] Emotion engine: Software that analyzes text and voice tone to recognize user emotions.

[0323] Text-to-speech API: Software that converts generated text responses into speech.

[0324] Database: Data storage for saving conversation data and emotion data.

[0325] Data processing flow:

[0326] 1. User voice input and text conversion

[0327] When a user speaks into the microphone, the device's speech recognition API captures the speech and converts it into text. For example, if a user says, "I'm tired today," the speech is converted into the text, "I'm tired today."

[0328] 2. Text Analysis and Emotion Recognition

[0329] The text data is sent to a server, which then uses an emotion engine to analyze the user's emotions. For example, when a user says "I'm tired," the emotion engine recognizes the emotion "fatigue" from the tone of voice and choice of words.

[0330] 3. Response Generation

[0331] The server's generation AI generates an appropriate response based on the analyzed emotional data and text. In this case, the generation AI generates the text response, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0332] 4. Audio Feedback

[0333] The generated text response is converted into speech using a speech synthesis API and played back through the device's speaker. The user is told, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0334] 5. Data storage and learning

[0335] All conversation and emotion data is stored in a database on the server, allowing the system to provide personalized responses based on the user's past comments and emotional history in future conversations.

[0336] Specific examples

[0337] Example 1: If a user says, "I'm tired from work today," the system can provide a response such as, "Great work. Would you like some suggestions on how to relax?"

[0338] Example 2: If a user says, "I passed today," the system provides a joyful emotional response such as "Congratulations!"

[0339] Example prompt sentence:

[0340] "I'm tired from work today"

[0341] This invention can be used not only by users with speech impediments but also in everyday communication, and can provide more natural and emotional conversations.

[0342] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0343] Step 1:

[0344] Voice input begins when the user speaks into the microphone. The device uses the microphone to capture the user's voice. Voice data is input and stored in the device as a digital signal. The data captured here is the content of the conversation for rehabilitation purposes.

[0345] Step 2:

[0346] The device converts the acquired voice data into text. This is done by calling a voice recognition API. If the voice input is "I'm tired today," it will be output as text data saying "I'm tired today." This process converts the voice into text.

[0347] Step 3:

[0348] The converted text data is sent to the server. The server first passes this text data to the emotion engine for emotion analysis. For example, the text "I'm tired today" outputs the emotion data "fatigue." This emotion data is used in the next process of the generation AI.

[0349] Step 4:

[0350] The server calls the generation AI using the emotion data and text data received from the emotion engine to generate an appropriate response. The generation AI receives the emotion data "fatigue" and the text "I'm tired today" as input and outputs the text response "Thank you for your hard work. Would you like some suggestions for how to relax?". An appropriate response is generated at this step.

[0351] Step 5:

[0352] The generated text response is sent from the server to the device. The device passes this text data to a speech synthesis API and converts it into voice data. The text "Thank you for your hard work. Would you like some suggestions on how to relax?" is output as voice data and played through the device's speaker. This process provides feedback of the text response to the user as voice.

[0353] Step 6:

[0354] The server stores all conversation and emotion data in a database. The stored data includes voice input, converted text, generated responses, and emotion data. This stored data is used to learn the user's tendencies during future rehabilitation and conversations. This step improves the overall performance of the system.

[0355] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0357] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0358] [Second embodiment]

[0359] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0360] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0361] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0362] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0363] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0364] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0365] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0366] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0367] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0368] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0369] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0370] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0371] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. The system is composed of a means for acquiring the user's voice input, converting it into text, generating an appropriate response using a generative AI, and converting the response back into speech to provide feedback to the user. The system also stores conversation data with the user and makes it available for the next use, providing a personalized rehabilitation and communication experience.

[0372] The system program operates while the server and terminals exchange data with each other. Below, we will explain the flow of the system's processing and each step.

[0373] Acquiring and converting voice input

[0374] Voice input begins when a user launches an application and speaks into the microphone. The device captures the voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text, "The weather is nice today."

[0375] Text analysis and response generation

[0376] The converted text is sent from the device to the server. The server receives this text and analyzes it using a generation AI. The generation AI understands the context of the user's remarks and generates an appropriate response. For example, it generates a response such as, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0377] Audio Feedback

[0378] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[0379] Data storage and training

[0380] The server securely stores all conversation data, including the user's utterances, responses, and conversation timestamps, and this data is used in the next conversation. The generative AI uses this data to learn the user's habits and past conversations, and provides personalized responses in the future.

[0381] Specific examples

[0382] For example, if a user launches the application at the same time every day, the generative AI can suggest topics related to that time of day. Also, if the user has previously talked about their favorite movies, the AI ​​can bring up new movies in the next conversation. This allows communication to be tailored to a specific user's interests and lifestyle, improving the effectiveness of rehabilitation.

[0383] In this way, the present invention helps users maintain a positive attitude while undergoing effective rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly and stress-free communication environment.

[0384] The processing flow will be explained below.

[0385] Step 1:

[0386] The user launches the application on the device. Upon launch, the main interface is displayed and the device is ready for voice input.

[0387] Step 2:

[0388] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the user's voice through the microphone.

[0389] Step 3:

[0390] The device calls the speech recognition API and converts the user's speech into text in real time. In this case, the speech "The weather is nice today" is recognized as the text "The weather is nice today."

[0391] Step 4:

[0392] The terminal sends the converted text to the server, where it is queued for analysis.

[0393] Step 5:

[0394] The server passes the received text to the generation AI, which analyzes the content of the text. The generation AI understands the context of the user's remarks and generates the most appropriate response. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0395] Step 6:

[0396] The server generates a response text and sends it to the terminal, where it is prepared for speech output.

[0397] Step 7:

[0398] The device passes the received response text to the speech synthesis API, which converts the text into speech. In this case, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into speech.

[0399] Step 8:

[0400] The device plays the audio output from the speech synthesis API, and the user hears the response, "What a beautiful day. It would be nice to go for a walk on a day like this."

[0401] Step 9:

[0402] The server securely stores all conversation data with the user, including the content of comments, responses, timestamps, etc.

[0403] Step 10:

[0404] The server uses the stored data to train the generative AI to provide optimized responses for the next conversation based on the user's habits and preferences. This learning process results in a more personalized rehabilitation experience.

[0405] Through the above specific steps, the user can undergo rehabilitation through a natural conversation experience.

[0406] Example 1

[0407] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0408] To effectively rehabilitate users with speech disorders at home and improve their communication skills, a system that generates personalized responses and provides a seamless conversation experience is needed. However, current technology does not adequately provide the means to analyze users' utterances in real time and generate appropriate responses. Furthermore, there are only a limited number of systems that utilize past conversation data to provide responses tailored to the user. This results in issues such as users not being able to fully benefit from rehabilitation and a decline in the quality of their communication.

[0409] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0410] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generative AI model and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time, means for acquiring voice input in real time and performing preprocessing, means for converting voice into text using a voice recognition API, means for transmitting the text data to the server using a secure communication protocol, means for inputting a prompt sentence into the generative AI model for analysis and generating a response, means for transmitting the response data to the terminal using a secure communication protocol, means for converting the response text into voice using a voice synthesis API, and means for learning the user's habits and past conversation content based on the saved conversation data. This enables the user to undergo effective rehabilitation at home and improve the quality of communication by receiving personalized responses.

[0411] The "means for acquiring voice input" refers to a device or method for acquiring the voice uttered by the user as digital voice data through an input device such as a microphone.

[0412] "Means for converting captured speech to text" refers to a device or method that uses a speech recognition API or software to convert captured speech data into corresponding text data.

[0413] A "generative AI model" is an artificial intelligence that uses natural language processing techniques to analyze text and generate appropriate responses.

[0414] A "means for analyzing text and generating appropriate responses" is a device or method for analyzing input text data using a generative AI model and automatically generating appropriate responses.

[0415] A "means for converting generated responses into speech" is a device or method that uses a speech synthesis API or software to convert generated text responses into speech data.

[0416] The "means for providing audio feedback to the user" refers to a device or method for playing back the generated audio data to the user through an output device such as a speaker of the terminal.

[0417] The "means for storing conversation data with a user" refers to a device or method for recording and storing the user's utterances, responses, and related data in digital form.

[0418] A "next use method" is a device or method for referencing saved conversation data in a next session to generate a response that takes past interactions into account.

[0419] The "means for acquiring and pre-processing speech input in real time" refers to a device or method for instantly acquiring speech input from a user and performing pre-processing such as noise removal and data shaping.

[0420] A "means for converting speech to text using a speech recognition API" is a device or method for utilizing a speech recognition service to instantly convert captured speech data into corresponding text data.

[0421] "Means for transmitting text data to a server using a secure communication protocol" means a device or method for transmitting text data from a terminal to a server using an encrypted protocol (e.g., HTTPS) to ensure the security of the communication.

[0422] "Means for inputting a prompt sentence into a generative AI model, analyzing it, and generating a response" refers to a device or method for providing an appropriate input sentence (prompt sentence) to a generative AI model and generating a response based on the analysis results.

[0423] "Means for transmitting response data to a terminal using a secure communication protocol" refers to a device or method for transmitting generated response data from a server to a terminal using an encrypted protocol (e.g., HTTPS) to ensure the security of communications.

[0424] "Means for converting response text into speech using a speech synthesis API" refers to a device or method for converting generated response text into speech data using a speech synthesis service.

[0425] "Means for learning user habits and past conversation content based on saved conversation data" refers to a device or method for generating more appropriate responses by analyzing saved conversation history and learning the user's speech patterns and past conversation content.

[0426] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. This system is composed of a series of means to acquire the user's voice input, convert it into text, analyze and generate a response using a generative AI model, and convert the response into speech and provide feedback to the user. Furthermore, conversation data with the user is saved and used the next time the system is used, providing a personalized rehabilitation and communication experience.

[0427] Hardware and software used

[0428] This system uses the following hardware and software:

[0429] Hardware: microphone, speaker, device (smartphone, tablet, PC, etc.)

[0430] Software: Speech recognition API (general name example: speech recognition service), generative AI model (general name example: natural language generation engine), speech synthesis API (general name example: speech synthesis service), secure communication protocol (e.g., HTTPS)

[0431] Acquiring and converting voice input

[0432] When a user launches an application on their device, voice input begins by speaking into the microphone. The device captures the voice through a built-in or external microphone and preprocesses the captured voice data in real time. Once preprocessed, the voice data is converted into text data using a speech recognition API.

[0433] Text analysis and response generation

[0434] The converted text data is sent from the device to the server using a secure communication protocol (HTTPS). The server inputs the received text data into the generative AI model for analysis. The generative AI model understands the context of the user's remarks and generates an appropriate response. For example, if the user says, "I'm tired today," the generative AI model will generate a response such as, "Please get plenty of rest. I think you'll feel better tomorrow."

[0435] Audio Feedback

[0436] The generated response text is sent from the server to the device using a secure communication protocol. The device inputs the received response text into a speech synthesis API and converts it into speech. This converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[0437] Data storage and training

[0438] The server securely stores all conversation data (user utterances, responses, timestamps, etc.). The stored data is used in the next conversation. The generative AI model uses the stored data to learn the user's habits and past conversations, and can provide personalized responses in future conversations. For example, by revisiting a movie the user previously mentioned, the model can create a continuous conversation and increase familiarity with the user.

[0439] Specific examples

[0440] When a user says, "Good morning, what shall we talk about today?", the voice is picked up through the microphone. The device's speech recognition API converts this voice into text, and the text data "Good morning, what shall we talk about today?" is sent to the server. The server's generative AI model analyzes it and generates a response: "Good morning! How about we talk about a new movie today?" This response is converted into speech by the speech synthesis API and played through the device's speaker.

[0441] Prompt Sentence Examples

[0442] Below are some example prompts to input to a generative AI model:

[0443] If the user says "I'm tired today," the prompt is:

[0444] Text format: "User said 'I am tired today'. Please generate an encouraging response."

[0445] In this way, the present invention provides users with effective rehabilitation at home and a seamless, personalized communication experience.

[0446] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0447] Step 1:

[0448] The user launches an application on the device and speaks into the microphone.

[0449] Input: User's voice data

[0450] How it works: The device captures audio in real time through the microphone.

[0451] Output: Audio data

[0452] Step 2:

[0453] Preprocessing is performed on the voice data acquired by the terminal.

[0454] Input: Acquired audio data

[0455] What it does: Performs preprocessing such as noise removal and volume normalization.

[0456] Output: Preprocessed audio data

[0457] Step 3:

[0458] The device uses a speech recognition API to convert the preprocessed voice data into text data.

[0459] Input: Preprocessed audio data

[0460] What it does: Calls a speech recognition API and converts the audio data into text data, for example, using the Google Cloud Speech-to-Text API.

[0461] Output: Text data

[0462] Step 4:

[0463] The terminal transmits the converted text data to the server using a secure communication protocol (HTTPS).

[0464] Input: Text data

[0465] How it works: Uses HTTPS to securely transmit text data.

[0466] Output: Text data received by the server

[0467] Step 5:

[0468] The text data received by the server is input into the generative AI model, where it is analyzed and an appropriate response is generated.

[0469] Input: Text data

[0470] How it works: A prompt is input to a generative AI model (e.g., a natural language generation engine), which then analyzes it and generates an appropriate response.

[0471] Output: Response text data

[0472] Step 6:

[0473] The server transmits the generated response text to the terminal using a secure communication protocol.

[0474] Input: Response text data

[0475] What it does: Uses HTTPS to securely transmit response text.

[0476] Output: Response text data received by the device

[0477] Step 7:

[0478] The device converts the received response text data into voice data using a speech synthesis API.

[0479] Input: Response text data

[0480] Behavior: Calls a speech synthesis API (e.g., a speech synthesis service) to convert text data into speech data.

[0481] Output: Audio data

[0482] Step 8:

[0483] The terminal plays the converted voice data through a speaker and provides feedback to the user.

[0484] Input: Audio data

[0485] What it does: Plays a sound through the speaker, providing feedback to the user.

[0486] Output: Audio feedback the user hears

[0487] Step 9:

[0488] The server securely stores all conversation data and uses it for the next conversation.

[0489] Input: Conversation data (user utterances, responses, timestamps)

[0490] What it does: Securely stores conversation data and keeps it in a format that can be learned and used by generative AI models.

[0491] Output: Saved conversation data

[0492] Step 10:

[0493] The server learns the user's habits and past conversation content based on the saved conversation data and personalizes the next response.

[0494] Input: Saved conversation data

[0495] How it works: Generative AI models analyze past conversation data, learn user habits and tendencies, and provide personalized responses the next time you talk to them.

[0496] Output: personalized response

[0497] Through these steps, users can undergo effective rehabilitation at home and experience seamless, personalized communication.

[0498] (Application example 1)

[0499] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0500] The goal of this project is to solve the problem that users with speech disorders have difficulty communicating smoothly with employees and customers in physical stores, which causes problems in their daily work. Furthermore, conventional rehabilitation systems were unable to flexibly respond to the individual circumstances of each user and were therefore unable to demonstrate sufficient effectiveness.

[0501] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0502] In this invention, the server includes means for acquiring user voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time the server is used, and means for supporting real-time communication between employees and customers. This enables users with speech disabilities to communicate smoothly in physical stores and interact efficiently with employees and customers.

[0503] The "means for acquiring user's voice input" refers to a device or method for recognizing the voice spoken by the user and capturing it as a digital signal.

[0504] The term "means for converting captured speech to text" refers to a device or method for converting captured speech data into corresponding text data using speech recognition technology.

[0505] "Means for analyzing text and generating appropriate responses using generative AI" refers to devices and methods that use artificial intelligence technology to analyze input text and automatically generate appropriate responses to that text.

[0506] The "means for converting the generated response into speech" refers to a device or method including speech synthesis technology for converting text data into speech data and outputting it as speech.

[0507] A "means for providing audio feedback to a user" is a device or method for communicating the generated audio response to a user through a speaker or headset.

[0508] "Means for saving conversation data with the user and using it the next time" refers to a device or method for saving the dialogue history with the user in a database or the like, and referencing that data during future dialogues to provide personalized responses.

[0509] "Means for supporting real-time communication between employees and customers" refers to devices and methods for generating responses in real time and providing feedback to facilitate smooth dialogue between employees and customers in physical stores.

[0510] This invention is a system that enables users with speech impediments to communicate smoothly with employees and customers in brick-and-mortar stores. This system combines speech recognition, generative AI, and speech synthesis technologies, and is described in detail below.

[0511] Acquiring and converting voice input

[0512] First, when a user starts the application and speaks into the microphone, the device's microphone picks up the audio. The device uses the speech_recognition library to convert this audio data into text. For example, if a user says, "What products do you recommend?", the audio is converted into text data that reads, "What products do you recommend?"

[0513] Text analysis and response generation

[0514] The converted text data is sent from the device to the server, which uses the OpenAI library to analyze the text data and generate an appropriate response. The generation AI understands the context of the user's utterance and generates a response such as "The current recommended product is fresh fruit."

[0515] Audio Feedback

[0516] The generated text response is sent from the server to the device, which then uses the gTTS library to convert the response into speech, which is then played back through the device's speaker and fed back to the user, allowing for seamless communication.

[0517] Data storage and training

[0518] The server securely stores all conversation data, including user utterances, responses, and conversation timestamps, which are referenced the next time the conversation occurs. Using this data, the generative AI can provide personalized responses, making communication more user-friendly.

[0519] Specific examples

[0520] When a user speaks to the device, asking, "What is your recommended product?", the speech is converted into text and sent to the server. The server analyzes the speech and generates a response such as, "Our current recommended product is fresh fruit." This response is then converted into speech and fed back to the user. Using this system, users can communicate smoothly in physical stores.

[0521] Prompt Sentence Examples

[0522] "User Question: What products do you recommend?"

[0523] "Context: A customer and an employee are having a conversation in a store."

[0524] "Answer: Our current recommended product is fresh fruit."

[0525] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0526] Step 1:

[0527] When a user launches an application and speaks into the microphone, the device receives voice input. Specifically, the microphone device detects the user's speech and captures it as digital audio data. This audio data is then sent to the speech_recognition library and converted into text data using Google's speech recognition API. The input is the user's voice, and the output is the converted text data.

[0528] Step 2:

[0529] The acquired text data is sent from the device to the server. The server receives this text data and analyzes it using the OpenAI library. A generative AI model analyzes the text based on the prompt and generates an appropriate response. The input is the text data, and the output is the generated response text. For example, if the input is "What products do you recommend?", the output will be "Our current recommended products are fresh fruit."

[0530] Step 3:

[0531] The generated response text is sent from the server to the device. The device receives this text and converts it into voice data using the gTTS library. Specifically, the text is converted into a digital audio file using speech synthesis technology, and then converted into a format that can be played back through a speaker. The input is the generated response text, and the output is voice data.

[0532] Step 4:

[0533] The voice data is fed back to the user through the device's speaker. The user can listen to this voice and continue the conversation. Specifically, the device's speaker plays the generated voice data and transmits it to the user. The input is the voice data, and the output is the voice that the user hears.

[0534] Step 5:

[0535] The server securely stores all conversation data. The stored data includes the user's utterances, the generated responses, and the conversation timestamp. By referencing this data during the next interaction, the generative AI model can provide a more personalized response. The input is the conversation data with the user, and the output is the stored data.

[0536] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0537] This invention is a system for enabling users with speech disorders to effectively rehabilitate at home and maintain a positive attitude, and aims to achieve more human-like and natural communication by combining it with an emotion engine that recognizes the user's emotions. This system includes a means for acquiring the user's voice input, converting it into text, generating an appropriate response using generative AI, converting the response back into voice and providing feedback to the user, as well as an emotion engine that recognizes the user's emotions.

[0538] The system program operates by exchanging data between the server, terminals, and emotion engine. Below, we will explain the process flow of this system and each step in detail.

[0539] Acquiring and converting voice input

[0540] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text "The weather is nice today."

[0541] Text Analysis and Emotion Recognition

[0542] The converted text is sent from the device to the server. The server receives this text and first analyzes it using an emotion engine. The emotion engine analyzes the user's emotions from the tone of their voice and the content of their speech, and passes that emotion data to the generation AI. For example, if the user is speaking happily, the emotion engine will analyze the emotion as "joy."

[0543] Response Generation

[0544] The generative AI takes into account the text content and emotional data to generate an appropriate response: in this case, "What a lovely day. It would be nice to go for a walk on a day like this," in a bright, joyful tone of voice.

[0545] Audio Feedback

[0546] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless and emotionally relevant conversation.

[0547] Data storage and training

[0548] The server securely stores all conversation data, including the content of statements, responses, emotional data, and timestamps, which are then used in the next conversation. The generative AI and emotion engine use this data to learn the user's habits and emotional history and provide personalized responses in future conversations.

[0549] Specific examples

[0550] For example, if a user has previously spoken under stress, the emotion engine can use that history to generate responses that will reduce stress in the next conversation. If a user says, "I'm tired from work today," the system could provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0551] In this way, the present invention helps users maintain a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to significantly improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotion-sensitive communication.

[0552] The processing flow will be explained below.

[0553] Step 1:

[0554] The user launches an application on the device. When the application launches, the main interface is displayed and the device is ready for voice input.

[0555] Step 2:

[0556] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the voice through the microphone.

[0557] Step 3:

[0558] The device calls the speech recognition API and converts the user's speech into text in real time. Specifically, the speech "The weather is nice today" is converted into the text "The weather is nice today."

[0559] Step 4:

[0560] The terminal sends the converted text to the server, where it is queued for further processing.

[0561] Step 5:

[0562] The server passes the received text to the emotion engine for emotion analysis. The emotion engine analyzes the user's emotion (e.g., joy, sadness, anger, etc.) from the user's tone of voice and the content of the speech.

[0563] Step 6:

[0564] The server receives the emotion data generated by the emotion engine and passes the emotion data along with the text to the generation AI. The generation AI generates an appropriate response based on the user's statement and emotion. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0565] Step 7:

[0566] The server generates a response text and sends it to the device, which then passes the text to the speech synthesis API.

[0567] Step 8:

[0568] The device uses a speech synthesis API to convert the response text into speech, in this case the text "What a lovely day. It would be nice to go for a walk on a day like this" in a bright, happy tone.

[0569] Step 9:

[0570] The device plays the generated speech and provides feedback to the user, who hears the response, "What a lovely day. It would be nice to go for a walk on a day like this."

[0571] Step 10:

[0572] The server securely stores all conversation data with the user, including the content of the conversation, responses, emotional data, and timestamps.

[0573] Step 11:

[0574] The server uses the stored conversation data to train the generative AI and emotion engine so that in future conversations, it can provide personalized responses that take into account the user's habits and emotional history.

[0575] For example, if a user has previously spoken while feeling stressed, the system can generate a response that reduces stress in the next conversation based on that history. If the user says, "I'm tired from work today," the system can provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" In this way, users can enjoy a more personalized and emotionally relevant conversational experience.

[0576] These specific steps allow users to undergo natural and effective rehabilitation, and emotionally-sensitive responses can help users develop a more positive outlook.

[0577] Example 2

[0578] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0579] Conventional dialogue systems using speech recognition systems or generative AI models have difficulty in accurately grasping a user's emotions and reflecting them in responses, making it difficult to provide natural and personalized dialogue, especially for users with speech impediments. Furthermore, there are insufficient means to provide more personalized responses by utilizing a user's past conversation data or emotional history.

[0580] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotion and passing the emotion data to the generative AI model, a means for analyzing the user's emotion in real time and reflecting it in a response, and a means for learning the user's past conversation data and providing a personalized response. This enables natural and personalized dialogue that takes the user's emotion into consideration.

[0581] "User" refers to an individual who uses the system to provide voice input.

[0582] A "server" is a central computing device for performing processing, storing data, and communicating with other components.

[0583] A "terminal" is a device used by a user that has the functions of voice input / output and data transmission / reception.

[0584] The "means for acquiring voice input" is a method for collecting the user's voice through a device such as a microphone.

[0585] "Means for converting voice to text" refers to a method of converting acquired voice data into text information using a voice recognition API.

[0586] "Means of analyzing text using a generative AI model and generating an appropriate response" refers to a method that uses generative AI, an algorithm for providing an appropriate response based on text data.

[0587] The "means of converting the generated response into speech" refers to a method of converting text data into speech data using a speech synthesis API or the like.

[0588] "Means for providing audio feedback to the user" refers to a method of returning the converted audio data to the user via a speaker or the like.

[0589] "Means for saving conversation data and using it the next time" refers to a method for recording the content and emotional data of past conversations and using them in future interactions.

[0590] "Means for recognizing emotions and passing that emotional data to a generative AI model" refers to a method for analyzing emotions from user speech and providing that information to the generative AI.

[0591] "Means for analyzing emotions in real time and reflecting them in responses" refers to a method for analyzing emotions simultaneously with user utterances and immediately incorporating the results into responses.

[0592] This invention is a system that allows users to receive emotional feedback while undergoing speech rehabilitation at home. The system converts the user's voice input into text and generates an appropriate response based on that text and emotional data. The system also converts the generated response into speech and provides feedback to the user. The system also incorporates an emotion recognition engine and has the function of analyzing the user's emotions. The specific processing flow of the system is described below.

[0593] Acquiring and converting voice input

[0594] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and calls a speech recognition API (e.g., a speech recognition API from a major cloud service provider) to convert the voice into text.

[0595] For example, if a user says, "The weather is nice today," the device captures the audio and uses a speech recognition API to convert it into text: "The weather is nice today."

[0596] Text analysis and emotion recognition

[0597] The converted text is sent from the device to a server, which then passes the received text to an emotion recognition engine (e.g., the API of a major emotion analysis service). The emotion recognition engine analyzes the user's emotions from the tone of voice and the content of the speech, and passes the emotional data to a generative AI model.

[0598] For example, if a user is speaking with a happy expression, the emotion recognition engine will analyze it as "joy." The server receives this analysis result and passes it to the generative AI model.

[0599] Response Generation

[0600] Generative AI models (e.g., advanced trained text generation models) generate appropriate responses based on text content and sentiment data.

[0601] In this case, the generated response would be "What a lovely day. It would be nice to go for a walk on a day like this," with a tone that reflects joy based on the emotion data.

[0602] Response transcription and feedback

[0603] The generated response text is sent from the server to the device, which then uses a speech synthesis API (e.g., a service that provides advanced speech synthesis technology) to convert the response text into speech, which is then played back through the device's speaker and provided as feedback to the user.

[0604] This process allows users to enjoy seamless and emotionally relevant interactions.

[0605] Data storage and training

[0606] The server securely stores all conversation data (statements, responses, emotional data, timestamps, etc.) The generative AI model and emotion recognition engine use the stored data to learn the user's habits and emotional history, providing personalized responses for future conversations.

[0607] Examples and prompts

[0608] For example, if a user has previously said, "I'm tired from work today," the emotion recognition engine can detect "stress," and in the next conversation, the generative AI model can generate a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0609] Examples of prompts are:

[0610] User: I'm tired from work today.

[0611] Emotion: Stress

[0612] Prompt: The user says "I'm tired from work today" and is feeling stressed. Generate an appropriate response accordingly.

[0613] In this way, the present invention is a system that supports users in maintaining a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotionally sensitive communication.

[0614] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0615] Step 1: Getting voice input

[0616] When a user starts an application and speaks into the microphone, voice input begins. The device acquires the user's voice data through the microphone. The input is the user's voice, which is acquired by the device as voice data. For example, when a user says, "The weather is nice today," the device collects this voice data.

[0617] Step 2: Speech to text

[0618] The device sends the acquired voice data to a voice recognition API, which converts the voice into text. The input is voice data, and the output is the corresponding text data. For example, voice data such as "The weather is nice today" is converted into text data such as "The weather is nice today."

[0619] Step 3: Sending text data

[0620] The terminal sends the converted text data to the server. The input is text data, which is sent to the server as is. For example, the text "The weather is nice today" is sent.

[0621] Step 4: Emotion Recognition

[0622] The server passes the received text to an emotion recognition engine to analyze the emotion. The input is text data, and the output is emotion data. For example, the emotion recognition engine analyzes the utterance "The weather is nice today" and generates emotion data for "joy."

[0623] Step 5: Input to the generative AI

[0624] The server passes the analyzed text data and emotion data to the generative AI model. The input is text data and emotion data, which are passed to the generative AI model. For example, the text "The weather is nice today" and the emotion data "joy" are input to the generative AI model.

[0625] Step 6: Generate a response

[0626] The generative AI model generates an appropriate response based on text data and emotional data. The input to the generative AI model is text data and emotional data, and the output is a response text. For example, the generated response text might be, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0627] Step 7: Transcribing the response

[0628] The server sends the generated response text to the device, and the device sends the response text to the speech synthesis API to convert it into voice data. The input is the response text, and the output is voice data. For example, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into voice data.

[0629] Step 8: User Feedback

[0630] The device plays the converted voice data from the speaker and provides feedback to the user. The input is voice data, and the output is the voice played back to the user. For example, the user may receive feedback in a bright tone saying, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0631] Step 9: Save your data

[0632] The server securely stores all conversation data. Inputs include utterances, responses, emotional data, and timestamps, and these are stored. For example, the utterance "The weather is nice today," the response "It really is nice weather," the emotional data "Joy," and the associated timestamps are stored.

[0633] Step 10: Data training

[0634] The generative AI model and emotion recognition engine use stored data to learn the user's habits and emotional history and reflect this in the next interaction. The input is stored interaction data, and learning from this improves the next response. For example, if a user previously said, "I'm tired from work today," the next time they make a similar statement, a more personalized response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" will be provided.

[0635] (Application example 2)

[0636] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0637] In modern society, it is not easy for users with speech disorders to undergo effective rehabilitation at home and maintain a positive attitude. Therefore, there is a need to develop systems that can recognize users' emotions and provide appropriate responses. It is also necessary to realize systems that are highly secure and can be used with peace of mind. It is particularly important to develop systems that can detect when a user feels anxiety or stress and respond appropriately.

[0638] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0639] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for recognizing the user's emotions, means for generating a response based on emotion data, and means for saving conversation data with the user and using it the next time the server is used. This enables users with speech disorders to undergo effective rehabilitation at home and communicate naturally in accordance with their emotions.

[0640] "Means for obtaining voice input" refers to a device or software that captures and records the user's speech.

[0641] "Means for converting voice to text" refers to technology that analyzes acquired voice data and converts it into a corresponding text format.

[0642] "Means of using generative AI to analyze text and generate appropriate responses" refers to a process of using artificial intelligence technology to analyze input text and generate appropriate responses.

[0643] "Means for converting the generated response to speech" refers to technology for converting the generated text response to speech.

[0644] "Audio feedback means" refers to a device or software that plays the generated audio to the user.

[0645] "Means for recognizing emotions" refers to technology that analyzes and understands emotions from a user's voice or text.

[0646] "Means for generating a response based on emotional data" refers to a technology that uses recognized emotional data to generate an appropriate response that is in line with the user's emotions.

[0647] "Means for saving conversation data and using it the next time" refers to a system that records conversations and emotional data with the user and saves it for reference the next time the user uses the system.

[0648] This invention relates to a system that receives voice input, converts it into text, generates appropriate responses using generative AI, and provides voice feedback to the user. Furthermore, by recognizing the user's emotions and generating responses based on them, it achieves more natural communication.

[0649] System configuration:

[0650] Hardware

[0651] Device: A device used by a user, such as a smartphone or computer, has a built-in microphone and speaker.

[0652] Server: A device that performs heavy processing, such as a cloud server.

[0653] software

[0654] Speech Recognition API: Installed on the device, it collects the user's voice and converts it into text.

[0655] Generative AI: Runs on a server and analyzes input text to generate appropriate responses, for example using natural language processing models.

[0656] Emotion engine: Software that analyzes text and voice tone to recognize user emotions.

[0657] Text-to-speech API: Software that converts generated text responses into speech.

[0658] Database: Data storage for saving conversation data and emotion data.

[0659] Data processing flow:

[0660] 1. User voice input and text conversion

[0661] When a user speaks into the microphone, the device's speech recognition API captures the speech and converts it into text. For example, if a user says, "I'm tired today," the speech is converted into the text, "I'm tired today."

[0662] 2. Text Analysis and Emotion Recognition

[0663] The text data is sent to a server, which then uses an emotion engine to analyze the user's emotions. For example, when a user says "I'm tired," the emotion engine recognizes the emotion "fatigue" from the tone of voice and choice of words.

[0664] 3. Response Generation

[0665] The server's generation AI generates an appropriate response based on the analyzed emotional data and text. In this case, the generation AI generates the text response, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0666] 4. Audio Feedback

[0667] The generated text response is converted into speech using a speech synthesis API and played back through the device's speaker. The user is told, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0668] 5. Data storage and learning

[0669] All conversation and emotion data is stored in a database on the server, allowing the system to provide personalized responses based on the user's past comments and emotional history in future conversations.

[0670] Specific examples

[0671] Example 1: If a user says, "I'm tired from work today," the system can provide a response such as, "Great work. Would you like some suggestions on how to relax?"

[0672] Example 2: If a user says, "I passed today," the system provides a joyful emotional response such as "Congratulations!"

[0673] Example prompt sentence:

[0674] "I'm tired from work today"

[0675] This invention can be used not only by users with speech impediments but also in everyday communication, and can provide more natural and emotional conversations.

[0676] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0677] Step 1:

[0678] Voice input begins when the user speaks into the microphone. The device uses the microphone to capture the user's voice. Voice data is input and stored in the device as a digital signal. The data captured here is the content of the conversation for rehabilitation purposes.

[0679] Step 2:

[0680] The device converts the acquired voice data into text. This is done by calling a voice recognition API. If the voice input is "I'm tired today," it will be output as text data saying "I'm tired today." This process converts the voice into text.

[0681] Step 3:

[0682] The converted text data is sent to the server. The server first passes this text data to the emotion engine for emotion analysis. For example, the text "I'm tired today" outputs the emotion data "fatigue." This emotion data is used in the next process of the generation AI.

[0683] Step 4:

[0684] The server calls the generation AI using the emotion data and text data received from the emotion engine to generate an appropriate response. The generation AI receives the emotion data "fatigue" and the text "I'm tired today" as input and outputs the text response "Thank you for your hard work. Would you like some suggestions for how to relax?". An appropriate response is generated at this step.

[0685] Step 5:

[0686] The generated text response is sent from the server to the device. The device passes this text data to a speech synthesis API and converts it into voice data. The text "Thank you for your hard work. Would you like some suggestions on how to relax?" is output as voice data and played through the device's speaker. This process provides feedback of the text response to the user as voice.

[0687] Step 6:

[0688] The server stores all conversation and emotion data in a database. The stored data includes voice input, converted text, generated responses, and emotion data. This stored data is used to learn the user's tendencies during future rehabilitation and conversations. This step improves the overall performance of the system.

[0689] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0690] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0691] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0692] [Third embodiment]

[0693] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0694] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0695] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0696] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0697] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0698] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0699] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0700] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0701] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0702] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0703] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0704] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0705] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. The system is composed of a means for acquiring the user's voice input, converting it into text, generating an appropriate response using a generative AI, and converting the response back into speech to provide feedback to the user. The system also stores conversation data with the user and makes it available for the next use, providing a personalized rehabilitation and communication experience.

[0706] The system program operates while the server and terminals exchange data with each other. Below, we will explain the flow of the system's processing and each step.

[0707] Acquiring and converting voice input

[0708] Voice input begins when a user launches an application and speaks into the microphone. The device captures the voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text, "The weather is nice today."

[0709] Text analysis and response generation

[0710] The converted text is sent from the device to the server. The server receives this text and analyzes it using a generation AI. The generation AI understands the context of the user's remarks and generates an appropriate response. For example, it generates a response such as, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0711] Audio Feedback

[0712] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[0713] Data storage and training

[0714] The server securely stores all conversation data, including the user's utterances, responses, and conversation timestamps, and this data is used in the next conversation. The generative AI uses this data to learn the user's habits and past conversations, and provides personalized responses in the future.

[0715] Specific examples

[0716] For example, if a user launches the application at the same time every day, the generative AI can suggest topics related to that time of day. Also, if the user has previously talked about their favorite movies, the AI ​​can bring up new movies in the next conversation. This allows communication to be tailored to a specific user's interests and lifestyle, improving the effectiveness of rehabilitation.

[0717] In this way, the present invention helps users maintain a positive attitude while undergoing effective rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly and stress-free communication environment.

[0718] The processing flow will be explained below.

[0719] Step 1:

[0720] The user launches the application on the device. Upon launch, the main interface is displayed and the device is ready for voice input.

[0721] Step 2:

[0722] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the user's voice through the microphone.

[0723] Step 3:

[0724] The device calls the speech recognition API and converts the user's speech into text in real time. In this case, the speech "The weather is nice today" is recognized as the text "The weather is nice today."

[0725] Step 4:

[0726] The terminal sends the converted text to the server, where it is queued for analysis.

[0727] Step 5:

[0728] The server passes the received text to the generation AI, which analyzes the content of the text. The generation AI understands the context of the user's remarks and generates the most appropriate response. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0729] Step 6:

[0730] The server generates a response text and sends it to the terminal, where it is prepared for speech output.

[0731] Step 7:

[0732] The device passes the received response text to the speech synthesis API, which converts the text into speech. In this case, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into speech.

[0733] Step 8:

[0734] The device plays the audio output from the speech synthesis API, and the user hears the response, "What a beautiful day. It would be nice to go for a walk on a day like this."

[0735] Step 9:

[0736] The server securely stores all conversation data with the user, including the content of comments, responses, timestamps, etc.

[0737] Step 10:

[0738] The server uses the stored data to train the generative AI to provide optimized responses for the next conversation based on the user's habits and preferences. This learning process results in a more personalized rehabilitation experience.

[0739] Through the above specific steps, the user can undergo rehabilitation through a natural conversation experience.

[0740] Example 1

[0741] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0742] To effectively rehabilitate users with speech disorders at home and improve their communication skills, a system that generates personalized responses and provides a seamless conversation experience is needed. However, current technology does not adequately provide the means to analyze users' utterances in real time and generate appropriate responses. Furthermore, there are only a limited number of systems that utilize past conversation data to provide responses tailored to the user. This results in issues such as users not being able to fully benefit from rehabilitation and a decline in the quality of their communication.

[0743] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0744] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generative AI model and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time, means for acquiring voice input in real time and performing preprocessing, means for converting voice into text using a voice recognition API, means for transmitting the text data to the server using a secure communication protocol, means for inputting a prompt sentence into the generative AI model for analysis and generating a response, means for transmitting the response data to the terminal using a secure communication protocol, means for converting the response text into voice using a voice synthesis API, and means for learning the user's habits and past conversation content based on the saved conversation data. This enables the user to undergo effective rehabilitation at home and improve the quality of communication by receiving personalized responses.

[0745] The "means for acquiring voice input" refers to a device or method for acquiring the voice uttered by the user as digital voice data through an input device such as a microphone.

[0746] "Means for converting captured speech to text" refers to a device or method that uses a speech recognition API or software to convert captured speech data into corresponding text data.

[0747] A "generative AI model" is an artificial intelligence that uses natural language processing techniques to analyze text and generate appropriate responses.

[0748] A "means for analyzing text and generating appropriate responses" is a device or method for analyzing input text data using a generative AI model and automatically generating appropriate responses.

[0749] A "means for converting generated responses into speech" is a device or method that uses a speech synthesis API or software to convert generated text responses into speech data.

[0750] The "means for providing audio feedback to the user" refers to a device or method for playing back the generated audio data to the user through an output device such as a speaker of the terminal.

[0751] The "means for storing conversation data with a user" refers to a device or method for recording and storing the user's utterances, responses, and related data in digital form.

[0752] A "next use method" is a device or method for referencing saved conversation data in a next session to generate a response that takes past interactions into account.

[0753] The "means for acquiring and pre-processing speech input in real time" refers to a device or method for instantly acquiring speech input from a user and performing pre-processing such as noise removal and data shaping.

[0754] A "means for converting speech to text using a speech recognition API" is a device or method for utilizing a speech recognition service to instantly convert captured speech data into corresponding text data.

[0755] "Means for transmitting text data to a server using a secure communication protocol" means a device or method for transmitting text data from a terminal to a server using an encrypted protocol (e.g., HTTPS) to ensure the security of the communication.

[0756] "Means for inputting a prompt sentence into a generative AI model, analyzing it, and generating a response" refers to a device or method for providing an appropriate input sentence (prompt sentence) to a generative AI model and generating a response based on the analysis results.

[0757] "Means for transmitting response data to a terminal using a secure communication protocol" refers to a device or method for transmitting generated response data from a server to a terminal using an encrypted protocol (e.g., HTTPS) to ensure the security of communications.

[0758] "Means for converting response text into speech using a speech synthesis API" refers to a device or method for converting generated response text into speech data using a speech synthesis service.

[0759] "Means for learning user habits and past conversation content based on saved conversation data" refers to a device or method for generating more appropriate responses by analyzing saved conversation history and learning the user's speech patterns and past conversation content.

[0760] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. This system is composed of a series of means to acquire the user's voice input, convert it into text, analyze and generate a response using a generative AI model, and convert the response into speech and provide feedback to the user. Furthermore, conversation data with the user is saved and used the next time the system is used, providing a personalized rehabilitation and communication experience.

[0761] Hardware and software used

[0762] This system uses the following hardware and software:

[0763] Hardware: microphone, speaker, device (smartphone, tablet, PC, etc.)

[0764] Software: Speech recognition API (general name example: speech recognition service), generative AI model (general name example: natural language generation engine), speech synthesis API (general name example: speech synthesis service), secure communication protocol (e.g., HTTPS)

[0765] Acquiring and converting voice input

[0766] When a user launches an application on their device, voice input begins by speaking into the microphone. The device captures the voice through a built-in or external microphone and preprocesses the captured voice data in real time. Once preprocessed, the voice data is converted into text data using a speech recognition API.

[0767] Text analysis and response generation

[0768] The converted text data is sent from the device to the server using a secure communication protocol (HTTPS). The server inputs the received text data into the generative AI model for analysis. The generative AI model understands the context of the user's remarks and generates an appropriate response. For example, if the user says, "I'm tired today," the generative AI model will generate a response such as, "Please get plenty of rest. I think you'll feel better tomorrow."

[0769] Audio Feedback

[0770] The generated response text is sent from the server to the device using a secure communication protocol. The device inputs the received response text into a speech synthesis API and converts it into speech. This converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[0771] Data storage and training

[0772] The server securely stores all conversation data (user utterances, responses, timestamps, etc.). The stored data is used in the next conversation. The generative AI model uses the stored data to learn the user's habits and past conversations, and can provide personalized responses in future conversations. For example, by revisiting a movie the user previously mentioned, the model can create a continuous conversation and increase familiarity with the user.

[0773] Specific examples

[0774] When a user says, "Good morning, what shall we talk about today?", the voice is picked up through the microphone. The device's speech recognition API converts this voice into text, and the text data "Good morning, what shall we talk about today?" is sent to the server. The server's generative AI model analyzes it and generates a response: "Good morning! How about we talk about a new movie today?" This response is converted into speech by the speech synthesis API and played through the device's speaker.

[0775] Prompt Sentence Examples

[0776] Below are some example prompts to input to a generative AI model:

[0777] If the user says "I'm tired today," the prompt is:

[0778] Text format: "User said 'I am tired today'. Please generate an encouraging response."

[0779] In this way, the present invention provides users with effective rehabilitation at home and a seamless, personalized communication experience.

[0780] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0781] Step 1:

[0782] The user launches an application on the device and speaks into the microphone.

[0783] Input: User's voice data

[0784] How it works: The device captures audio in real time through the microphone.

[0785] Output: Audio data

[0786] Step 2:

[0787] Preprocessing is performed on the voice data acquired by the terminal.

[0788] Input: Acquired audio data

[0789] What it does: Performs preprocessing such as noise removal and volume normalization.

[0790] Output: Preprocessed audio data

[0791] Step 3:

[0792] The device uses a speech recognition API to convert the preprocessed voice data into text data.

[0793] Input: Preprocessed audio data

[0794] What it does: Calls a speech recognition API and converts the audio data into text data, for example, using the Google Cloud Speech-to-Text API.

[0795] Output: Text data

[0796] Step 4:

[0797] The terminal transmits the converted text data to the server using a secure communication protocol (HTTPS).

[0798] Input: Text data

[0799] How it works: Uses HTTPS to securely transmit text data.

[0800] Output: Text data received by the server

[0801] Step 5:

[0802] The text data received by the server is input into the generative AI model, where it is analyzed and an appropriate response is generated.

[0803] Input: Text data

[0804] How it works: A prompt is input to a generative AI model (e.g., a natural language generation engine), which then analyzes it and generates an appropriate response.

[0805] Output: Response text data

[0806] Step 6:

[0807] The server transmits the generated response text to the terminal using a secure communication protocol.

[0808] Input: Response text data

[0809] What it does: Uses HTTPS to securely transmit response text.

[0810] Output: Response text data received by the device

[0811] Step 7:

[0812] The device converts the received response text data into voice data using a speech synthesis API.

[0813] Input: Response text data

[0814] Behavior: Calls a speech synthesis API (e.g., a speech synthesis service) to convert text data into speech data.

[0815] Output: Audio data

[0816] Step 8:

[0817] The terminal plays the converted voice data through a speaker and provides feedback to the user.

[0818] Input: Audio data

[0819] What it does: Plays a sound through the speaker, providing feedback to the user.

[0820] Output: Audio feedback the user hears

[0821] Step 9:

[0822] The server securely stores all conversation data and uses it for the next conversation.

[0823] Input: Conversation data (user utterances, responses, timestamps)

[0824] What it does: Securely stores conversation data and keeps it in a format that can be learned and used by generative AI models.

[0825] Output: Saved conversation data

[0826] Step 10:

[0827] The server learns the user's habits and past conversation content based on the saved conversation data and personalizes the next response.

[0828] Input: Saved conversation data

[0829] How it works: Generative AI models analyze past conversation data, learn user habits and tendencies, and provide personalized responses the next time you talk to them.

[0830] Output: personalized response

[0831] Through these steps, users can undergo effective rehabilitation at home and experience seamless, personalized communication.

[0832] (Application example 1)

[0833] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0834] The goal of this project is to solve the problem that users with speech disorders have difficulty communicating smoothly with employees and customers in physical stores, which causes problems in their daily work. Furthermore, conventional rehabilitation systems were unable to flexibly respond to the individual circumstances of each user and were therefore unable to demonstrate sufficient effectiveness.

[0835] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0836] In this invention, the server includes means for acquiring user voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time the server is used, and means for supporting real-time communication between employees and customers. This enables users with speech disabilities to communicate smoothly in physical stores and interact efficiently with employees and customers.

[0837] The "means for acquiring user's voice input" refers to a device or method for recognizing the voice spoken by the user and capturing it as a digital signal.

[0838] The term "means for converting captured speech to text" refers to a device or method for converting captured speech data into corresponding text data using speech recognition technology.

[0839] "Means for analyzing text and generating appropriate responses using generative AI" refers to devices and methods that use artificial intelligence technology to analyze input text and automatically generate appropriate responses to that text.

[0840] The "means for converting the generated response into speech" refers to a device or method including speech synthesis technology for converting text data into speech data and outputting it as speech.

[0841] A "means for providing audio feedback to a user" is a device or method for communicating the generated audio response to a user through a speaker or headset.

[0842] "Means for saving conversation data with the user and using it the next time" refers to a device or method for saving the dialogue history with the user in a database or the like, and referencing that data during future dialogues to provide personalized responses.

[0843] "Means for supporting real-time communication between employees and customers" refers to devices and methods for generating responses in real time and providing feedback to facilitate smooth dialogue between employees and customers in physical stores.

[0844] This invention is a system that enables users with speech impediments to communicate smoothly with employees and customers in brick-and-mortar stores. This system combines speech recognition, generative AI, and speech synthesis technologies, and is described in detail below.

[0845] Acquiring and converting voice input

[0846] First, when a user starts the application and speaks into the microphone, the device's microphone picks up the audio. The device uses the speech_recognition library to convert this audio data into text. For example, if a user says, "What products do you recommend?", the audio is converted into text data that reads, "What products do you recommend?"

[0847] Text analysis and response generation

[0848] The converted text data is sent from the device to the server, which uses the OpenAI library to analyze the text data and generate an appropriate response. The generation AI understands the context of the user's utterance and generates a response such as "The current recommended product is fresh fruit."

[0849] Audio Feedback

[0850] The generated text response is sent from the server to the device, which then uses the gTTS library to convert the response into speech, which is then played back through the device's speaker and fed back to the user, allowing for seamless communication.

[0851] Data storage and training

[0852] The server securely stores all conversation data, including user utterances, responses, and conversation timestamps, which are referenced the next time the conversation occurs. Using this data, the generative AI can provide personalized responses, making communication more user-friendly.

[0853] Specific examples

[0854] When a user speaks to the device, asking, "What is your recommended product?", the speech is converted into text and sent to the server. The server analyzes the speech and generates a response such as, "Our current recommended product is fresh fruit." This response is then converted into speech and fed back to the user. Using this system, users can communicate smoothly in physical stores.

[0855] Prompt Sentence Examples

[0856] "User Question: What products do you recommend?"

[0857] "Context: A customer and an employee are having a conversation in a store."

[0858] "Answer: Our current recommended product is fresh fruit."

[0859] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0860] Step 1:

[0861] When a user launches an application and speaks into the microphone, the device receives voice input. Specifically, the microphone device detects the user's speech and captures it as digital audio data. This audio data is then sent to the speech_recognition library and converted into text data using Google's speech recognition API. The input is the user's voice, and the output is the converted text data.

[0862] Step 2:

[0863] The acquired text data is sent from the device to the server. The server receives this text data and analyzes it using the OpenAI library. A generative AI model analyzes the text based on the prompt and generates an appropriate response. The input is the text data, and the output is the generated response text. For example, if the input is "What products do you recommend?", the output will be "Our current recommended products are fresh fruit."

[0864] Step 3:

[0865] The generated response text is sent from the server to the device. The device receives this text and converts it into voice data using the gTTS library. Specifically, the text is converted into a digital audio file using speech synthesis technology, and then converted into a format that can be played back through a speaker. The input is the generated response text, and the output is voice data.

[0866] Step 4:

[0867] The voice data is fed back to the user through the device's speaker. The user can listen to this voice and continue the conversation. Specifically, the device's speaker plays the generated voice data and transmits it to the user. The input is the voice data, and the output is the voice that the user hears.

[0868] Step 5:

[0869] The server securely stores all conversation data. The stored data includes the user's utterances, the generated responses, and the conversation timestamp. By referencing this data during the next interaction, the generative AI model can provide a more personalized response. The input is the conversation data with the user, and the output is the stored data.

[0870] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0871] This invention is a system for enabling users with speech disorders to effectively rehabilitate at home and maintain a positive attitude, and aims to achieve more human-like and natural communication by combining it with an emotion engine that recognizes the user's emotions. This system includes a means for acquiring the user's voice input, converting it into text, generating an appropriate response using generative AI, converting the response back into voice and providing feedback to the user, as well as an emotion engine that recognizes the user's emotions.

[0872] The system program operates by exchanging data between the server, terminals, and emotion engine. Below, we will explain the process flow of this system and each step in detail.

[0873] Acquiring and converting voice input

[0874] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text "The weather is nice today."

[0875] Text Analysis and Emotion Recognition

[0876] The converted text is sent from the device to the server. The server receives this text and first analyzes it using an emotion engine. The emotion engine analyzes the user's emotions from the tone of their voice and the content of their speech, and passes that emotion data to the generation AI. For example, if the user is speaking happily, the emotion engine will analyze the emotion as "joy."

[0877] Response Generation

[0878] The generative AI takes into account the text content and emotional data to generate an appropriate response: in this case, "What a lovely day. It would be nice to go for a walk on a day like this," in a bright, joyful tone of voice.

[0879] Audio Feedback

[0880] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless and emotionally relevant conversation.

[0881] Data storage and training

[0882] The server securely stores all conversation data, including the content of statements, responses, emotional data, and timestamps, which are then used in the next conversation. The generative AI and emotion engine use this data to learn the user's habits and emotional history and provide personalized responses in future conversations.

[0883] Specific examples

[0884] For example, if a user has previously spoken under stress, the emotion engine can use that history to generate responses that will reduce stress in the next conversation. If a user says, "I'm tired from work today," the system could provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0885] In this way, the present invention helps users maintain a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to significantly improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotion-sensitive communication.

[0886] The processing flow will be explained below.

[0887] Step 1:

[0888] The user launches an application on the device. When the application launches, the main interface is displayed and the device is ready for voice input.

[0889] Step 2:

[0890] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the voice through the microphone.

[0891] Step 3:

[0892] The device calls the speech recognition API and converts the user's speech into text in real time. Specifically, the speech "The weather is nice today" is converted into the text "The weather is nice today."

[0893] Step 4:

[0894] The terminal sends the converted text to the server, where it is queued for further processing.

[0895] Step 5:

[0896] The server passes the received text to the emotion engine for emotion analysis. The emotion engine analyzes the user's emotion (e.g., joy, sadness, anger, etc.) from the user's tone of voice and the content of the speech.

[0897] Step 6:

[0898] The server receives the emotion data generated by the emotion engine and passes the emotion data along with the text to the generation AI. The generation AI generates an appropriate response based on the user's statement and emotion. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0899] Step 7:

[0900] The server generates a response text and sends it to the device, which then passes the text to the speech synthesis API.

[0901] Step 8:

[0902] The device uses a speech synthesis API to convert the response text into speech, in this case the text "What a lovely day. It would be nice to go for a walk on a day like this" in a bright, happy tone.

[0903] Step 9:

[0904] The device plays the generated speech and provides feedback to the user, who hears the response, "What a lovely day. It would be nice to go for a walk on a day like this."

[0905] Step 10:

[0906] The server securely stores all conversation data with the user, including the content of the conversation, responses, emotional data, and timestamps.

[0907] Step 11:

[0908] The server uses the stored conversation data to train the generative AI and emotion engine so that in future conversations, it can provide personalized responses that take into account the user's habits and emotional history.

[0909] For example, if a user has previously spoken while feeling stressed, the system can generate a response that reduces stress in the next conversation based on that history. If the user says, "I'm tired from work today," the system can provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" In this way, users can enjoy a more personalized and emotionally relevant conversational experience.

[0910] These specific steps allow users to undergo natural and effective rehabilitation, and emotionally-sensitive responses can help users develop a more positive outlook.

[0911] Example 2

[0912] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0913] Conventional dialogue systems using speech recognition systems or generative AI models have difficulty in accurately grasping a user's emotions and reflecting them in responses, making it difficult to provide natural and personalized dialogue, especially for users with speech impediments. Furthermore, there are insufficient means to provide more personalized responses by utilizing a user's past conversation data or emotional history.

[0914] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotion and passing the emotion data to the generative AI model, a means for analyzing the user's emotion in real time and reflecting it in a response, and a means for learning the user's past conversation data and providing a personalized response. This enables natural and personalized dialogue that takes the user's emotion into consideration.

[0915] "User" refers to an individual who uses the system to provide voice input.

[0916] A "server" is a central computing device for performing processing, storing data, and communicating with other components.

[0917] A "terminal" is a device used by a user that has the functions of voice input / output and data transmission / reception.

[0918] The "means for acquiring voice input" is a method for collecting the user's voice through a device such as a microphone.

[0919] "Means for converting voice to text" refers to a method of converting acquired voice data into text information using a voice recognition API.

[0920] "Means of analyzing text using a generative AI model and generating an appropriate response" refers to a method that uses generative AI, an algorithm for providing an appropriate response based on text data.

[0921] The "means of converting the generated response into speech" refers to a method of converting text data into speech data using a speech synthesis API or the like.

[0922] "Means for providing audio feedback to the user" refers to a method of returning the converted audio data to the user via a speaker or the like.

[0923] "Means for saving conversation data and using it the next time" refers to a method for recording the content and emotional data of past conversations and using them in future interactions.

[0924] "Means for recognizing emotions and passing that emotional data to a generative AI model" refers to a method for analyzing emotions from user speech and providing that information to the generative AI.

[0925] "Means for analyzing emotions in real time and reflecting them in responses" refers to a method for analyzing emotions simultaneously with user utterances and immediately incorporating the results into responses.

[0926] This invention is a system that allows users to receive emotional feedback while undergoing speech rehabilitation at home. The system converts the user's voice input into text and generates an appropriate response based on that text and emotional data. It also converts the generated response into speech and provides feedback to the user. The system also incorporates an emotion recognition engine and has the function of analyzing the user's emotions. The specific processing flow of the system is described below.

[0927] Acquiring and converting voice input

[0928] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and calls a speech recognition API (e.g., a speech recognition API from a major cloud service provider) to convert the voice into text.

[0929] For example, if a user says, "The weather is nice today," the device captures the audio and uses a speech recognition API to convert it into text: "The weather is nice today."

[0930] Text analysis and emotion recognition

[0931] The converted text is sent from the device to a server, which then passes the received text to an emotion recognition engine (e.g., the API of a major emotion analysis service). The emotion recognition engine analyzes the user's emotions from the tone of voice and the content of the speech, and passes the emotional data to a generative AI model.

[0932] For example, if a user is speaking with a happy expression, the emotion recognition engine will analyze it as "joy." The server receives this analysis result and passes it to the generative AI model.

[0933] Response Generation

[0934] Generative AI models (e.g., advanced trained text generation models) generate appropriate responses based on text content and sentiment data.

[0935] In this case, the generated response would be "What a lovely day. It would be nice to go for a walk on a day like this," with a tone that reflects joy based on the emotion data.

[0936] Response transcription and feedback

[0937] The generated response text is sent from the server to the device, which then uses a speech synthesis API (e.g., a service that provides advanced speech synthesis technology) to convert the response text into speech, which is then played back through the device's speaker and provided as feedback to the user.

[0938] This process allows users to enjoy seamless and emotionally relevant interactions.

[0939] Data storage and training

[0940] The server securely stores all conversation data (statements, responses, emotional data, timestamps, etc.) The generative AI model and emotion recognition engine use the stored data to learn the user's habits and emotional history, providing personalized responses for future conversations.

[0941] Examples and prompts

[0942] For example, if a user has previously said, "I'm tired from work today," the emotion recognition engine can detect "stress," and in the next conversation, the generative AI model can generate a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[0943] Examples of prompts are:

[0944] User: I'm tired from work today.

[0945] Emotion: Stress

[0946] Prompt: The user says "I'm tired from work today" and is feeling stressed. Generate an appropriate response accordingly.

[0947] In this way, the present invention is a system that supports users in maintaining a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotionally sensitive communication.

[0948] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0949] Step 1: Getting voice input

[0950] When a user starts an application and speaks into the microphone, voice input begins. The device acquires the user's voice data through the microphone. The input is the user's voice, which is acquired by the device as voice data. For example, when a user says, "The weather is nice today," the device collects this voice data.

[0951] Step 2: Speech to text

[0952] The device sends the acquired voice data to a voice recognition API, which converts the voice into text. The input is voice data, and the output is the corresponding text data. For example, voice data such as "The weather is nice today" is converted into text data such as "The weather is nice today."

[0953] Step 3: Sending text data

[0954] The terminal sends the converted text data to the server. The input is text data, which is sent to the server as is. For example, the text "The weather is nice today" is sent.

[0955] Step 4: Emotion Recognition

[0956] The server passes the received text to an emotion recognition engine to analyze the emotion. The input is text data, and the output is emotion data. For example, the emotion recognition engine analyzes the utterance "The weather is nice today" and generates emotion data for "joy."

[0957] Step 5: Input to the generative AI

[0958] The server passes the analyzed text data and emotion data to the generative AI model. The input is text data and emotion data, which are passed to the generative AI model. For example, the text "The weather is nice today" and the emotion data "joy" are input to the generative AI model.

[0959] Step 6: Generate a response

[0960] The generative AI model generates an appropriate response based on text data and emotional data. The input to the generative AI model is text data and emotional data, and the output is a response text. For example, the generated response text might be, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0961] Step 7: Transcribing the response

[0962] The server sends the generated response text to the device, and the device sends the response text to the speech synthesis API to convert it into voice data. The input is the response text, and the output is voice data. For example, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into voice data.

[0963] Step 8: User Feedback

[0964] The device plays the converted voice data from the speaker and provides feedback to the user. The input is voice data, and the output is the voice played back to the user. For example, the user may receive feedback in a bright tone saying, "It's really nice weather. It would be nice to go for a walk on a day like this."

[0965] Step 9: Save your data

[0966] The server securely stores all conversation data. Inputs include utterances, responses, emotional data, and timestamps, and these are stored. For example, the utterance "The weather is nice today," the response "It really is nice weather," the emotional data "Joy," and the associated timestamps are stored.

[0967] Step 10: Data training

[0968] The generative AI model and emotion recognition engine use stored data to learn the user's habits and emotional history and reflect this in the next interaction. The input is stored interaction data, and learning from this improves the next response. For example, if a user previously said, "I'm tired from work today," the next time they make a similar statement, a more personalized response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" will be provided.

[0969] (Application example 2)

[0970] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0971] In modern society, it is not easy for users with speech disorders to undergo effective rehabilitation at home and maintain a positive attitude. Therefore, there is a need to develop systems that can recognize users' emotions and provide appropriate responses. It is also necessary to realize systems that are highly secure and can be used with peace of mind. It is particularly important to develop systems that can detect when a user feels anxiety or stress and respond appropriately.

[0972] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0973] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for recognizing the user's emotions, means for generating a response based on emotion data, and means for saving conversation data with the user and using it the next time the server is used. This enables users with speech disorders to undergo effective rehabilitation at home and communicate naturally in accordance with their emotions.

[0974] "Means for obtaining voice input" refers to a device or software that captures and records the user's speech.

[0975] "Means for converting voice to text" refers to technology that analyzes acquired voice data and converts it into a corresponding text format.

[0976] "Means of using generative AI to analyze text and generate appropriate responses" refers to a process of using artificial intelligence technology to analyze input text and generate appropriate responses.

[0977] "Means for converting the generated response to speech" refers to technology for converting the generated text response to speech.

[0978] "Audio feedback means" refers to a device or software that plays the generated audio to the user.

[0979] "Means for recognizing emotions" refers to technology that analyzes and understands emotions from a user's voice or text.

[0980] "Means for generating a response based on emotional data" refers to a technology that uses recognized emotional data to generate an appropriate response that is in line with the user's emotions.

[0981] "Means for saving conversation data and using it the next time" refers to a system that records conversations and emotional data with the user and saves it for reference the next time the user uses the system.

[0982] This invention relates to a system that receives voice input, converts it into text, generates appropriate responses using generative AI, and provides voice feedback to the user. Furthermore, by recognizing the user's emotions and generating responses based on them, it achieves more natural communication.

[0983] System configuration:

[0984] Hardware

[0985] Device: A device used by a user, such as a smartphone or computer, has a built-in microphone and speaker.

[0986] Server: A device that performs heavy processing, such as a cloud server.

[0987] software

[0988] Speech Recognition API: Installed on the device, it collects the user's voice and converts it into text.

[0989] Generative AI: Runs on a server and analyzes input text to generate appropriate responses, for example using natural language processing models.

[0990] Emotion engine: Software that analyzes text and voice tone to recognize user emotions.

[0991] Text-to-speech API: Software that converts generated text responses into speech.

[0992] Database: Data storage for saving conversation data and emotion data.

[0993] Data processing flow:

[0994] 1. User voice input and text conversion

[0995] When a user speaks into the microphone, the device's speech recognition API captures the speech and converts it into text. For example, if a user says, "I'm tired today," the speech is converted into text "I'm tired today."

[0996] 2. Text Analysis and Emotion Recognition

[0997] The text data is sent to a server, which then uses an emotion engine to analyze the user's emotions. For example, when a user says "I'm tired," the emotion engine recognizes the emotion "fatigue" from the tone of voice and choice of words.

[0998] 3. Response Generation

[0999] The server's generation AI generates an appropriate response based on the analyzed emotional data and text. In this case, the generation AI generates the text response, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[1000] 4. Audio Feedback

[1001] The generated text response is converted into speech using a speech synthesis API and played back through the device's speaker. The user is told, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[1002] 5. Data storage and learning

[1003] All conversation and emotion data is stored in a database on the server, allowing the system to provide personalized responses based on the user's past comments and emotional history in future conversations.

[1004] Specific examples

[1005] Example 1: If a user says, "I'm tired from work today," the system can provide a response such as, "Great work. Would you like some suggestions on how to relax?"

[1006] Example 2: If a user says, "I passed today," the system provides a joyful emotional response such as "Congratulations!"

[1007] Example prompt sentence:

[1008] "I'm tired from work today"

[1009] This invention can be used not only by users with speech impediments but also in everyday communication, and can provide more natural and emotional conversations.

[1010] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1011] Step 1:

[1012] Voice input begins when the user speaks into the microphone. The device uses the microphone to capture the user's voice. Voice data is input and stored in the device as a digital signal. The data captured here is the content of the conversation for rehabilitation purposes.

[1013] Step 2:

[1014] The device converts the acquired voice data into text. This is done by calling a voice recognition API. If the voice input is "I'm tired today," it will be output as text data saying "I'm tired today." This process converts the voice into text.

[1015] Step 3:

[1016] The converted text data is sent to the server. The server first passes this text data to the emotion engine for emotion analysis. For example, the text "I'm tired today" outputs the emotion data "fatigue." This emotion data is used in the next process of the generation AI.

[1017] Step 4:

[1018] The server calls the generation AI using the emotion data and text data received from the emotion engine to generate an appropriate response. The generation AI receives the emotion data "fatigue" and the text "I'm tired today" as input and outputs the text response "Thank you for your hard work. Would you like some suggestions for how to relax?". An appropriate response is generated at this step.

[1019] Step 5:

[1020] The generated text response is sent from the server to the device. The device passes this text data to a speech synthesis API and converts it into voice data. The text "Thank you for your hard work. Would you like some suggestions on how to relax?" is output as voice data and played through the device's speaker. This process provides feedback of the text response to the user as voice.

[1021] Step 6:

[1022] The server stores all conversation and emotion data in a database. The stored data includes voice input, converted text, generated responses, and emotion data. This stored data is used to learn the user's tendencies during future rehabilitation and conversations. This step improves the overall performance of the system.

[1023] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1024] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1025] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1026] [Fourth embodiment]

[1027] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1028] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1030] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1031] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1032] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1034] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1035] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1036] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1038] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1040] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. The system is composed of a means for acquiring the user's voice input, converting it into text, generating an appropriate response using a generative AI, and converting the response back into speech to provide feedback to the user. The system also stores conversation data with the user and makes it available for the next use, providing a personalized rehabilitation and communication experience.

[1041] The system program operates while the server and terminals exchange data with each other. Below, we will explain the flow of the system's processing and each step.

[1042] Acquiring and converting voice input

[1043] Voice input begins when a user launches an application and speaks into the microphone. The device captures the voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text, "The weather is nice today."

[1044] Text analysis and response generation

[1045] The converted text is sent from the device to the server. The server receives this text and analyzes it using a generation AI. The generation AI understands the context of the user's remarks and generates an appropriate response. For example, it generates a response such as, "It's really nice weather. It would be nice to go for a walk on a day like this."

[1046] Audio Feedback

[1047] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[1048] Data storage and training

[1049] The server securely stores all conversation data, including the user's utterances, responses, and conversation timestamps, and this data is used in the next conversation. The generative AI uses this data to learn the user's habits and past conversations, and provides personalized responses in the future.

[1050] Specific examples

[1051] For example, if a user launches the application at the same time every day, the generative AI can suggest topics related to that time of day. Also, if the user has previously talked about their favorite movies, the AI ​​can bring up new movies in the next conversation. This allows communication to be tailored to a specific user's interests and lifestyle, improving the effectiveness of rehabilitation.

[1052] In this way, the present invention helps users maintain a positive attitude while undergoing effective rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly and stress-free communication environment.

[1053] The processing flow will be explained below.

[1054] Step 1:

[1055] The user launches the application on the device. Upon launch, the main interface is displayed and the device is ready for voice input.

[1056] Step 2:

[1057] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the user's voice through the microphone.

[1058] Step 3:

[1059] The device calls the speech recognition API and converts the user's speech into text in real time. In this case, the speech "The weather is nice today" is recognized as the text "The weather is nice today."

[1060] Step 4:

[1061] The terminal sends the converted text to the server, where it is queued for analysis.

[1062] Step 5:

[1063] The server passes the received text to the generation AI, which analyzes the content of the text. The generation AI understands the context of the user's remarks and generates the most appropriate response. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[1064] Step 6:

[1065] The server generates a response text and sends it to the terminal, where it is prepared for speech output.

[1066] Step 7:

[1067] The device passes the received response text to the speech synthesis API, which converts the text into speech. In this case, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into speech.

[1068] Step 8:

[1069] The device plays the audio output from the speech synthesis API, and the user hears the response, "What a beautiful day. It would be nice to go for a walk on a day like this."

[1070] Step 9:

[1071] The server securely stores all conversation data with the user, including the content of comments, responses, timestamps, etc.

[1072] Step 10:

[1073] The server uses the stored data to train the generative AI to provide optimized responses for the next conversation based on the user's habits and preferences. This learning process results in a more personalized rehabilitation experience.

[1074] Through the above specific steps, the user can undergo rehabilitation through a natural conversation experience.

[1075] Example 1

[1076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1077] To effectively rehabilitate users with speech disorders at home and improve their communication skills, a system that generates personalized responses and provides a seamless conversation experience is needed. However, current technology does not adequately provide the means to analyze users' utterances in real time and generate appropriate responses. Furthermore, there are only a limited number of systems that utilize past conversation data to provide responses tailored to the user. This results in issues such as users not being able to fully benefit from rehabilitation and a decline in the quality of their communication.

[1078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1079] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generative AI model and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time, means for acquiring voice input in real time and performing preprocessing, means for converting voice into text using a voice recognition API, means for transmitting the text data to the server using a secure communication protocol, means for inputting a prompt sentence into the generative AI model for analysis and generating a response, means for transmitting the response data to the terminal using a secure communication protocol, means for converting the response text into voice using a voice synthesis API, and means for learning the user's habits and past conversation content based on the saved conversation data. This enables the user to undergo effective rehabilitation at home and improve the quality of communication by receiving personalized responses.

[1080] The "means for acquiring voice input" refers to a device or method for acquiring the voice uttered by the user as digital voice data through an input device such as a microphone.

[1081] "Means for converting captured speech to text" refers to a device or method that uses a speech recognition API or software to convert captured speech data into corresponding text data.

[1082] A "generative AI model" is an artificial intelligence that uses natural language processing techniques to analyze text and generate appropriate responses.

[1083] A "means for analyzing text and generating appropriate responses" is a device or method for analyzing input text data using a generative AI model and automatically generating appropriate responses.

[1084] A "means for converting generated responses into speech" is a device or method that uses a speech synthesis API or software to convert generated text responses into speech data.

[1085] The "means for providing audio feedback to the user" refers to a device or method for playing back the generated audio data to the user through an output device such as a speaker of the terminal.

[1086] The "means for storing conversation data with a user" refers to a device or method for recording and storing the user's utterances, responses, and related data in digital form.

[1087] A "next use method" is a device or method for referencing saved conversation data in a next session to generate a response that takes past interactions into account.

[1088] The "means for acquiring and pre-processing speech input in real time" refers to a device or method for instantly acquiring speech input from a user and performing pre-processing such as noise removal and data shaping.

[1089] A "means for converting speech to text using a speech recognition API" is a device or method for utilizing a speech recognition service to instantly convert captured speech data into corresponding text data.

[1090] "Means for transmitting text data to a server using a secure communication protocol" means a device or method for transmitting text data from a terminal to a server using an encrypted protocol (e.g., HTTPS) to ensure the security of the communication.

[1091] "Means for inputting a prompt sentence into a generative AI model, analyzing it, and generating a response" refers to a device or method for providing an appropriate input sentence (prompt sentence) to a generative AI model and generating a response based on the analysis results.

[1092] "Means for transmitting response data to a terminal using a secure communication protocol" refers to a device or method for transmitting generated response data from a server to a terminal using an encrypted protocol (e.g., HTTPS) to ensure the security of communications.

[1093] "Means for converting response text into speech using a speech synthesis API" refers to a device or method for converting generated response text into speech data using a speech synthesis service.

[1094] "Means for learning user habits and past conversation content based on saved conversation data" refers to a device or method for generating more appropriate responses by analyzing saved conversation history and learning the user's speech patterns and past conversation content.

[1095] This invention is a system that allows users with speech disorders to undergo rehabilitation at home at their own convenience. This system is composed of a series of means to acquire the user's voice input, convert it into text, analyze and generate a response using a generative AI model, and convert the response into speech and provide feedback to the user. Furthermore, conversation data with the user is saved and used the next time the system is used, providing a personalized rehabilitation and communication experience.

[1096] Hardware and software used

[1097] This system uses the following hardware and software:

[1098] Hardware: microphone, speaker, device (smartphone, tablet, PC, etc.)

[1099] Software: Speech recognition API (general name example: speech recognition service), generative AI model (general name example: natural language generation engine), speech synthesis API (general name example: speech synthesis service), secure communication protocol (e.g., HTTPS)

[1100] Acquiring and converting voice input

[1101] When a user launches an application on their device, voice input begins by speaking into the microphone. The device captures the voice through a built-in or external microphone and preprocesses the captured voice data in real time. Once preprocessed, the voice data is converted into text data using a speech recognition API.

[1102] Text analysis and response generation

[1103] The converted text data is sent from the device to the server using a secure communication protocol (HTTPS). The server inputs the received text data into the generative AI model for analysis. The generative AI model understands the context of the user's remarks and generates an appropriate response. For example, if the user says, "I'm tired today," the generative AI model will generate a response such as, "Please get plenty of rest. I think you'll feel better tomorrow."

[1104] Audio Feedback

[1105] The generated response text is sent from the server to the device using a secure communication protocol. The device inputs the received response text into a speech synthesis API and converts it into speech. This converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless conversation.

[1106] Data storage and training

[1107] The server securely stores all conversation data (user utterances, responses, timestamps, etc.). The stored data is used in the next conversation. The generative AI model uses the stored data to learn the user's habits and past conversations, and can provide personalized responses in future conversations. For example, by revisiting a movie the user previously mentioned, the model can create a continuous conversation and increase familiarity with the user.

[1108] Specific examples

[1109] When a user says, "Good morning, what shall we talk about today?", the voice is picked up through the microphone. The device's speech recognition API converts this voice into text, and the text data "Good morning, what shall we talk about today?" is sent to the server. The server's generative AI model analyzes it and generates a response: "Good morning! How about we talk about a new movie today?" This response is converted into speech by the speech synthesis API and played through the device's speaker.

[1110] Prompt Sentence Examples

[1111] Below are some example prompts to input to a generative AI model:

[1112] If the user says "I'm tired today," the prompt is:

[1113] Text format: "User said 'I am tired today'. Please generate an encouraging response."

[1114] In this way, the present invention provides users with effective rehabilitation at home and a seamless, personalized communication experience.

[1115] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1116] Step 1:

[1117] The user launches an application on the device and speaks into the microphone.

[1118] Input: User's voice data

[1119] How it works: The device captures audio in real time through the microphone.

[1120] Output: Audio data

[1121] Step 2:

[1122] Preprocessing is performed on the voice data acquired by the terminal.

[1123] Input: Acquired audio data

[1124] What it does: Performs preprocessing such as noise removal and volume normalization.

[1125] Output: Preprocessed audio data

[1126] Step 3:

[1127] The device uses a speech recognition API to convert the preprocessed voice data into text data.

[1128] Input: Preprocessed audio data

[1129] What it does: Calls a speech recognition API and converts the audio data into text data, for example, using the Google Cloud Speech-to-Text API.

[1130] Output: Text data

[1131] Step 4:

[1132] The terminal transmits the converted text data to the server using a secure communication protocol (HTTPS).

[1133] Input: Text data

[1134] How it works: Uses HTTPS to securely transmit text data.

[1135] Output: Text data received by the server

[1136] Step 5:

[1137] The text data received by the server is input into the generative AI model, where it is analyzed and an appropriate response is generated.

[1138] Input: Text data

[1139] How it works: A prompt is input to a generative AI model (e.g., a natural language generation engine), which then analyzes it and generates an appropriate response.

[1140] Output: Response text data

[1141] Step 6:

[1142] The server transmits the generated response text to the terminal using a secure communication protocol.

[1143] Input: Response text data

[1144] What it does: Uses HTTPS to securely transmit response text.

[1145] Output: Response text data received by the device

[1146] Step 7:

[1147] The device converts the received response text data into voice data using a speech synthesis API.

[1148] Input: Response text data

[1149] Behavior: Calls a speech synthesis API (e.g., a speech synthesis service) to convert text data into speech data.

[1150] Output: Audio data

[1151] Step 8:

[1152] The terminal plays the converted voice data through a speaker and provides feedback to the user.

[1153] Input: Audio data

[1154] What it does: Plays a sound through the speaker, providing feedback to the user.

[1155] Output: Audio feedback the user hears

[1156] Step 9:

[1157] The server securely stores all conversation data and uses it for the next conversation.

[1158] Input: Conversation data (user utterances, responses, timestamps)

[1159] What it does: Securely stores conversation data and keeps it in a format that can be learned and used by generative AI models.

[1160] Output: Saved conversation data

[1161] Step 10:

[1162] The server learns the user's habits and past conversation content based on the saved conversation data and personalizes the next response.

[1163] Input: Saved conversation data

[1164] How it works: Generative AI models analyze past conversation data, learn user habits and tendencies, and provide personalized responses the next time you talk to them.

[1165] Output: personalized response

[1166] Through these steps, users can undergo effective rehabilitation at home and experience seamless, personalized communication.

[1167] (Application example 1)

[1168] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1169] The goal of this project is to solve the problem that users with speech disorders have difficulty communicating smoothly with employees and customers in physical stores, which causes problems in their daily work. Furthermore, conventional rehabilitation systems were unable to flexibly respond to the individual circumstances of each user and were therefore unable to demonstrate sufficient effectiveness.

[1170] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1171] In this invention, the server includes means for acquiring user voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for saving conversation data with the user and using it the next time the server is used, and means for supporting real-time communication between employees and customers. This enables users with speech disabilities to communicate smoothly in physical stores and interact efficiently with employees and customers.

[1172] The "means for acquiring user's voice input" refers to a device or method for recognizing the voice spoken by the user and capturing it as a digital signal.

[1173] The term "means for converting captured speech to text" refers to a device or method for converting captured speech data into corresponding text data using speech recognition technology.

[1174] "Means for analyzing text and generating appropriate responses using generative AI" refers to devices and methods that use artificial intelligence technology to analyze input text and automatically generate appropriate responses to that text.

[1175] The "means for converting the generated response into speech" refers to a device or method including speech synthesis technology for converting text data into speech data and outputting it as speech.

[1176] A "means for providing audio feedback to a user" is a device or method for communicating the generated audio response to a user through a speaker or headset.

[1177] "Means for saving conversation data with the user and using it the next time" refers to a device or method for saving the dialogue history with the user in a database or the like, and referencing that data during future dialogues to provide personalized responses.

[1178] "Means for supporting real-time communication between employees and customers" refers to devices and methods for generating responses in real time and providing feedback to facilitate smooth dialogue between employees and customers in physical stores.

[1179] This invention is a system that enables users with speech impediments to communicate smoothly with employees and customers in brick-and-mortar stores. This system combines speech recognition, generative AI, and speech synthesis technologies, and is described in detail below.

[1180] Acquiring and converting voice input

[1181] First, when a user starts the application and speaks into the microphone, the device's microphone picks up the audio. The device uses the speech_recognition library to convert this audio data into text. For example, if a user says, "What products do you recommend?", the audio is converted into text data that reads, "What products do you recommend?"

[1182] Text analysis and response generation

[1183] The converted text data is sent from the device to the server, which uses the OpenAI library to analyze the text data and generate an appropriate response. The generation AI understands the context of the user's utterance and generates a response such as "The current recommended product is fresh fruit."

[1184] Audio Feedback

[1185] The generated text response is sent from the server to the device, which then uses the gTTS library to convert the response into speech, which is then played back through the device's speaker and fed back to the user, allowing for seamless communication.

[1186] Data storage and training

[1187] The server securely stores all conversation data, including user utterances, responses, and conversation timestamps, which are referenced the next time the conversation occurs. Using this data, the generative AI can provide personalized responses, making communication more user-friendly.

[1188] Specific examples

[1189] When a user speaks to the device, asking, "What is your recommended product?", the speech is converted into text and sent to the server. The server analyzes the speech and generates a response such as, "Our current recommended product is fresh fruit." This response is then converted into speech and fed back to the user. Using this system, users can communicate smoothly in physical stores.

[1190] Prompt Sentence Examples

[1191] "User Question: What products do you recommend?"

[1192] "Context: A customer and an employee are having a conversation in a store."

[1193] "Answer: Our current recommended product is fresh fruit."

[1194] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1195] Step 1:

[1196] When a user launches an application and speaks into the microphone, the device receives voice input. Specifically, the microphone device detects the user's speech and captures it as digital audio data. This audio data is then sent to the speech_recognition library and converted into text data using Google's speech recognition API. The input is the user's voice, and the output is the converted text data.

[1197] Step 2:

[1198] The acquired text data is sent from the device to the server. The server receives this text data and analyzes it using the OpenAI library. A generative AI model analyzes the text based on the prompt and generates an appropriate response. The input is the text data, and the output is the generated response text. For example, if the input is "What products do you recommend?", the output will be "Our current recommended products are fresh fruit."

[1199] Step 3:

[1200] The generated response text is sent from the server to the device. The device receives this text and converts it into voice data using the gTTS library. Specifically, the text is converted into a digital audio file using speech synthesis technology, and then converted into a format that can be played back through a speaker. The input is the generated response text, and the output is voice data.

[1201] Step 4:

[1202] The voice data is fed back to the user through the device's speaker. The user can listen to this voice and continue the conversation. Specifically, the device's speaker plays the generated voice data and transmits it to the user. The input is the voice data, and the output is the voice that the user hears.

[1203] Step 5:

[1204] The server securely stores all conversation data. The stored data includes the user's utterances, the generated responses, and the conversation timestamp. By referencing this data during the next interaction, the generative AI model can provide a more personalized response. The input is the conversation data with the user, and the output is the stored data.

[1205] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1206] This invention is a system for enabling users with speech disorders to effectively rehabilitate at home and maintain a positive attitude, and aims to achieve more human-like and natural communication by combining it with an emotion engine that recognizes the user's emotions. This system includes a means for acquiring the user's voice input, converting it into text, generating an appropriate response using generative AI, converting the response back into voice and providing feedback to the user, as well as an emotion engine that recognizes the user's emotions.

[1207] The system program operates by exchanging data between the server, terminals, and emotion engine. Below, we will explain the process flow of this system and each step in detail.

[1208] Acquiring and converting voice input

[1209] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and converts the voice into text by calling a speech recognition API. For example, if a user says, "The weather is nice today," the device converts this voice into the text "The weather is nice today."

[1210] Text Analysis and Emotion Recognition

[1211] The converted text is sent from the device to the server. The server receives this text and first analyzes it using an emotion engine. The emotion engine analyzes the user's emotions from the tone of their voice and the content of their speech, and passes that emotion data to the generation AI. For example, if the user is speaking happily, the emotion engine will analyze the emotion as "joy."

[1212] Response Generation

[1213] The generative AI takes into account the text content and emotional data to generate an appropriate response: in this case, "What a lovely day. It would be nice to go for a walk on a day like this," in a bright, joyful tone of voice.

[1214] Audio Feedback

[1215] The generated response text is sent from the server to the device. The device then uses a speech synthesis API to convert this text into speech. The converted speech is played back through the device's speaker and fed back to the user, allowing the user to enjoy a seamless and emotionally relevant conversation.

[1216] Data storage and training

[1217] The server securely stores all conversation data, including the content of statements, responses, emotional data, and timestamps, which are then used in the next conversation. The generative AI and emotion engine use this data to learn the user's habits and emotional history and provide personalized responses in future conversations.

[1218] Specific examples

[1219] For example, if a user has previously spoken under stress, the emotion engine can use that history to generate responses that will reduce stress in the next conversation. If a user says, "I'm tired from work today," the system could provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[1220] In this way, the present invention helps users maintain a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to significantly improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotion-sensitive communication.

[1221] The processing flow will be explained below.

[1222] Step 1:

[1223] The user launches an application on the device. When the application launches, the main interface is displayed and the device is ready for voice input.

[1224] Step 2:

[1225] The user speaks into the microphone, saying, "The weather is nice today." The device picks up the voice through the microphone.

[1226] Step 3:

[1227] The device calls the speech recognition API and converts the user's speech into text in real time. Specifically, the speech "The weather is nice today" is converted into the text "The weather is nice today."

[1228] Step 4:

[1229] The terminal sends the converted text to the server, where it is queued for further processing.

[1230] Step 5:

[1231] The server passes the received text to the emotion engine for emotion analysis. The emotion engine analyzes the user's emotion (e.g., joy, sadness, anger, etc.) from the user's tone of voice and the content of the speech.

[1232] Step 6:

[1233] The server receives the emotion data generated by the emotion engine and passes the emotion data along with the text to the generation AI. The generation AI generates an appropriate response based on the user's statement and emotion. In this case, the response generated is, "It's really nice weather. It would be nice to go for a walk on a day like this."

[1234] Step 7:

[1235] The server generates a response text and sends it to the device, which then passes the text to the speech synthesis API.

[1236] Step 8:

[1237] The device uses a speech synthesis API to convert the response text into speech, in this case the text "What a lovely day. It would be nice to go for a walk on a day like this" in a bright, happy tone.

[1238] Step 9:

[1239] The device plays the generated speech and provides feedback to the user, who hears the response, "What a lovely day. It would be nice to go for a walk on a day like this."

[1240] Step 10:

[1241] The server securely stores all conversation data with the user, including the content of the conversation, responses, emotional data, and timestamps.

[1242] Step 11:

[1243] The server uses the stored conversation data to train the generative AI and emotion engine so that in future conversations, it can provide personalized responses that take into account the user's habits and emotional history.

[1244] For example, if a user has previously spoken while feeling stressed, the system can generate a response that reduces stress in the next conversation based on that history. If the user says, "I'm tired from work today," the system can provide a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" In this way, users can enjoy a more personalized and emotionally relevant conversational experience.

[1245] These specific steps allow users to undergo natural and effective rehabilitation, and emotionally-sensitive responses can help users develop a more positive outlook.

[1246] Example 2

[1247] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1248] Conventional dialogue systems using speech recognition systems or generative AI models have difficulty in accurately grasping a user's emotions and reflecting them in responses, making it difficult to provide natural and personalized dialogue, especially for users with speech impediments. Furthermore, there are insufficient means to provide more personalized responses by utilizing a user's past conversation data or emotional history.

[1249] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recognizing the user's emotion and passing the emotion data to the generative AI model, a means for analyzing the user's emotion in real time and reflecting it in a response, and a means for learning the user's past conversation data and providing a personalized response. This enables natural and personalized dialogue that takes the user's emotion into consideration.

[1250] "User" refers to an individual who uses the system to provide voice input.

[1251] A "server" is a central computing device for performing processing, storing data, and communicating with other components.

[1252] A "terminal" is a device used by a user that has the functions of voice input / output and data transmission / reception.

[1253] The "means for acquiring voice input" is a method for collecting the user's voice through a device such as a microphone.

[1254] "Means for converting voice to text" refers to a method of converting acquired voice data into text information using a voice recognition API.

[1255] "Means of analyzing text using a generative AI model and generating an appropriate response" refers to a method that uses generative AI, an algorithm for providing an appropriate response based on text data.

[1256] The "means of converting the generated response into speech" refers to a method of converting text data into speech data using a speech synthesis API or the like.

[1257] "Means for providing audio feedback to the user" refers to a method of returning the converted audio data to the user via a speaker or the like.

[1258] "Means for saving conversation data and using it the next time" refers to a method for recording the content and emotional data of past conversations and using them in future interactions.

[1259] "Means for recognizing emotions and passing that emotional data to a generative AI model" refers to a method for analyzing emotions from user speech and providing that information to the generative AI.

[1260] "Means for analyzing emotions in real time and reflecting them in responses" refers to a method for analyzing emotions simultaneously with user utterances and immediately incorporating the results into responses.

[1261] This invention is a system that allows users to receive emotional feedback while undergoing speech rehabilitation at home. The system converts the user's voice input into text and generates an appropriate response based on that text and emotional data. The system also converts the generated response into speech and provides feedback to the user. The system also incorporates an emotion recognition engine and has the function of analyzing the user's emotions. The specific processing flow of the system is described below.

[1262] Acquiring and converting voice input

[1263] Voice input begins when a user launches an application and speaks into the microphone. The device captures the user's voice through the microphone and calls a speech recognition API (e.g., a speech recognition API from a major cloud service provider) to convert the voice into text.

[1264] For example, if a user says, "The weather is nice today," the device captures the audio and uses a speech recognition API to convert it into text: "The weather is nice today."

[1265] Text analysis and emotion recognition

[1266] The converted text is sent from the device to a server, which then passes the received text to an emotion recognition engine (e.g., the API of a major emotion analysis service). The emotion recognition engine analyzes the user's emotions from the tone of voice and the content of the speech, and passes the emotional data to a generative AI model.

[1267] For example, if a user is speaking with a happy expression, the emotion recognition engine will analyze it as "joy." The server receives this analysis result and passes it to the generative AI model.

[1268] Response Generation

[1269] Generative AI models (e.g., advanced trained text generation models) generate appropriate responses based on text content and sentiment data.

[1270] In this case, the generated response would be "What a lovely day. It would be nice to go for a walk on a day like this," with a tone that reflects joy based on the emotion data.

[1271] Response transcription and feedback

[1272] The generated response text is sent from the server to the device, which then uses a speech synthesis API (e.g., a service that provides advanced speech synthesis technology) to convert the response text into speech, which is then played back through the device's speaker and provided as feedback to the user.

[1273] This process allows users to enjoy seamless and emotionally relevant interactions.

[1274] Data storage and training

[1275] The server securely stores all conversation data (statements, responses, emotional data, timestamps, etc.) The generative AI model and emotion recognition engine use the stored data to learn the user's habits and emotional history, providing personalized responses for future conversations.

[1276] Examples and prompts

[1277] For example, if a user has previously said, "I'm tired from work today," the emotion recognition engine can detect "stress," and in the next conversation, the generative AI model can generate a response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[1278] Examples of prompts are:

[1279] User: I'm tired from work today.

[1280] Emotion: Stress

[1281] Prompt: The user says "I'm tired from work today" and is feeling stressed. Generate an appropriate response accordingly.

[1282] In this way, the present invention is a system that supports users in maintaining a positive attitude while undergoing emotional rehabilitation at home. The entire system aims to improve the quality of life for people with speech disorders by providing a user-friendly interface and realizing emotionally sensitive communication.

[1283] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1284] Step 1: Getting voice input

[1285] When a user starts an application and speaks into the microphone, voice input begins. The device acquires the user's voice data through the microphone. The input is the user's voice, which is acquired by the device as voice data. For example, when a user says, "The weather is nice today," the device collects this voice data.

[1286] Step 2: Speech to text

[1287] The device sends the acquired voice data to a voice recognition API, which converts the voice into text. The input is voice data, and the output is the corresponding text data. For example, voice data such as "The weather is nice today" is converted into text data such as "The weather is nice today."

[1288] Step 3: Sending text data

[1289] The terminal sends the converted text data to the server. The input is text data, which is sent to the server as is. For example, the text "The weather is nice today" is sent.

[1290] Step 4: Emotion Recognition

[1291] The server passes the received text to an emotion recognition engine to analyze the emotion. The input is text data, and the output is emotion data. For example, the emotion recognition engine analyzes the utterance "The weather is nice today" and generates emotion data for "joy."

[1292] Step 5: Input to the generative AI

[1293] The server passes the analyzed text data and emotion data to the generative AI model. The input is text data and emotion data, which are passed to the generative AI model. For example, the text "The weather is nice today" and the emotion data "joy" are input to the generative AI model.

[1294] Step 6: Generate a response

[1295] The generative AI model generates an appropriate response based on text data and emotional data. The input to the generative AI model is text data and emotional data, and the output is a response text. For example, the generated response text might be, "It's really nice weather. It would be nice to go for a walk on a day like this."

[1296] Step 7: Transcribing the response

[1297] The server sends the generated response text to the device, and the device sends the response text to the speech synthesis API to convert it into voice data. The input is the response text, and the output is voice data. For example, the text "What a lovely day. It would be nice to go for a walk on a day like this" is converted into voice data.

[1298] Step 8: User Feedback

[1299] The device plays the converted voice data from the speaker and provides feedback to the user. The input is voice data, and the output is the voice played back to the user. For example, the user may receive feedback in a bright tone saying, "It's really nice weather. It would be nice to go for a walk on a day like this."

[1300] Step 9: Save your data

[1301] The server securely stores all conversation data. Inputs include utterances, responses, emotional data, and timestamps, and these are stored. For example, the utterance "The weather is nice today," the response "It really is nice weather," the emotional data "Joy," and the associated timestamps are stored.

[1302] Step 10: Data training

[1303] The generative AI model and emotion recognition engine use stored data to learn the user's habits and emotional history and reflect this in the next interaction. The input is stored interaction data, and learning from this improves the next response. For example, if a user previously said, "I'm tired from work today," the next time they make a similar statement, a more personalized response such as, "Thank you for your hard work. Would you like some suggestions for how to relax?" will be provided.

[1304] (Application example 2)

[1305] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1306] In modern society, it is not easy for users with speech disorders to undergo effective rehabilitation at home and maintain a positive attitude. Therefore, there is a need to develop systems that can recognize users' emotions and provide appropriate responses. It is also necessary to realize systems that are highly secure and can be used with peace of mind. It is particularly important to develop systems that can detect when a user feels anxiety or stress and respond appropriately.

[1307] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1308] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice into text, means for analyzing the text using a generation AI and generating an appropriate response, means for converting the generated response into voice, means for providing voice feedback to the user, means for recognizing the user's emotions, means for generating a response based on emotion data, and means for saving conversation data with the user and using it the next time the server is used. This enables users with speech disorders to undergo effective rehabilitation at home and communicate naturally in accordance with their emotions.

[1309] "Means for obtaining voice input" refers to a device or software that captures and records the user's speech.

[1310] "Means for converting voice to text" refers to technology that analyzes acquired voice data and converts it into a corresponding text format.

[1311] "Means of using generative AI to analyze text and generate appropriate responses" refers to a process of using artificial intelligence technology to analyze input text and generate appropriate responses.

[1312] "Means for converting the generated response to speech" refers to technology for converting the generated text response to speech.

[1313] "Audio feedback means" refers to a device or software that plays the generated audio to the user.

[1314] "Means for recognizing emotions" refers to technology that analyzes and understands emotions from a user's voice or text.

[1315] "Means for generating a response based on emotional data" refers to a technology that uses recognized emotional data to generate an appropriate response that is in line with the user's emotions.

[1316] "Means for saving conversation data and using it the next time" refers to a system that records conversations and emotional data with the user and saves it for reference the next time the user uses the system.

[1317] This invention relates to a system that receives voice input, converts it into text, generates appropriate responses using generative AI, and provides voice feedback to the user. Furthermore, by recognizing the user's emotions and generating responses based on them, it achieves more natural communication.

[1318] System configuration:

[1319] Hardware

[1320] Device: A device used by a user, such as a smartphone or computer, has a built-in microphone and speaker.

[1321] Server: A device that performs heavy processing, such as a cloud server.

[1322] software

[1323] Speech Recognition API: Installed on the device, it collects the user's voice and converts it into text.

[1324] Generative AI: Runs on a server and analyzes input text to generate appropriate responses, for example using natural language processing models.

[1325] Emotion engine: Software that analyzes text and voice tone to recognize user emotions.

[1326] Text-to-speech API: Software that converts generated text responses into speech.

[1327] Database: Data storage for saving conversation data and emotion data.

[1328] Data processing flow:

[1329] 1. User voice input and text conversion

[1330] When a user speaks into the microphone, the device's speech recognition API captures the speech and converts it into text. For example, if a user says, "I'm tired today," the speech is converted into the text, "I'm tired today."

[1331] 2. Text Analysis and Emotion Recognition

[1332] The text data is sent to a server, which then uses an emotion engine to analyze the user's emotions. For example, when a user says "I'm tired," the emotion engine recognizes the emotion "fatigue" from the tone of voice and choice of words.

[1333] 3. Response Generation

[1334] The server's generation AI generates an appropriate response based on the analyzed emotional data and text. In this case, the generation AI generates the text response, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[1335] 4. Audio Feedback

[1336] The generated text response is converted into speech using a speech synthesis API and played back through the device's speaker. The user is told, "Thank you for your hard work. Would you like some suggestions for how to relax?"

[1337] 5. Data storage and learning

[1338] All conversation and emotion data is stored in a database on the server, allowing the system to provide personalized responses based on the user's past comments and emotional history in future conversations.

[1339] Specific examples

[1340] Example 1: If a user says, "I'm tired from work today," the system can provide a response such as, "Great work. Would you like some suggestions on how to relax?"

[1341] Example 2: If a user says, "I passed today," the system provides a joyful emotional response such as "Congratulations!"

[1342] Example prompt sentence:

[1343] "I'm tired from work today"

[1344] This invention can be used not only by users with speech impediments but also in everyday communication, and can provide more natural and emotional conversations.

[1345] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1346] Step 1:

[1347] Voice input begins when the user speaks into the microphone. The device uses the microphone to capture the user's voice. Voice data is input and stored in the device as a digital signal. The data captured here is the content of the conversation for rehabilitation purposes.

[1348] Step 2:

[1349] The device converts the acquired voice data into text. This is done by calling a voice recognition API. If the voice input is "I'm tired today," it will be output as text data saying "I'm tired today." This process converts the voice into text.

[1350] Step 3:

[1351] The converted text data is sent to the server. The server first passes this text data to the emotion engine for emotion analysis. For example, the text "I'm tired today" outputs the emotion data "fatigue." This emotion data is used in the next process of the generation AI.

[1352] Step 4:

[1353] The server calls the generation AI using the emotion data and text data received from the emotion engine to generate an appropriate response. The generation AI receives the emotion data "fatigue" and the text "I'm tired today" as input and outputs the text response "Thank you for your hard work. Would you like some suggestions for how to relax?". An appropriate response is generated at this step.

[1354] Step 5:

[1355] The generated text response is sent from the server to the device. The device passes this text data to a speech synthesis API and converts it into voice data. The text "Thank you for your hard work. Would you like some suggestions on how to relax?" is output as voice data and played through the device's speaker. This process provides feedback of the text response to the user as voice.

[1356] Step 6:

[1357] The server stores all conversation and emotion data in a database. The stored data includes voice input, converted text, generated responses, and emotion data. This stored data is used to learn the user's tendencies during future rehabilitation and conversations. This step improves the overall performance of the system.

[1358] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1359] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1360] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1361] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1362] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1363] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1364] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1365] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1366] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1367] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1368] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1369] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1370] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1371] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1372] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1373] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1374] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1375] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1376] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1377] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1378] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1379] The following is further disclosed regarding the above embodiment.

[1380] (Claim 1)

[1381] means for obtaining a user's voice input;

[1382] a means for converting the acquired speech into text;

[1383] a means for analyzing the text using generative AI to generate an appropriate response;

[1384] means for converting the generated response into speech;

[1385] means for providing audio feedback to the user;

[1386] A system that includes a means for saving conversation data with a user and using it the next time the system is used.

[1387] (Claim 2)

[1388] 10. The system of claim 1, wherein the generative AI includes means for learning from a user's past conversation data and providing personalized responses.

[1389] (Claim 3)

[1390] 2. The system of claim 1, wherein the generative AI includes means for analyzing user utterances in real time and generating appropriate responses based on the context of the conversation.

[1391] "Example 1"

[1392] (Claim 1)

[1393] means for obtaining a user's voice input;

[1394] a means for converting the acquired speech into text;

[1395] a means for analyzing the text using a generative AI model to generate an appropriate response; and

[1396] means for converting the generated response into speech;

[1397] means for providing audio feedback to the user;

[1398] A means for saving conversation data with the user and using it the next time the user uses the service;

[1399] a means for acquiring and pre-processing speech input in real time;

[1400] A means of converting speech to text using a speech recognition API;

[1401] means for transmitting text data to a server using a secure communication protocol;

[1402] A means for inputting a prompt sentence into a generative AI model, analyzing the sentence, and generating a response;

[1403] means for transmitting response data to the terminal using a secure communication protocol;

[1404] A means for converting the response text into speech using a speech synthesis API;

[1405] A system that includes a means for learning user habits and past conversation content based on saved conversation data.

[1406] (Claim 2)

[1407] 10. The system of claim 1, wherein the generative AI model includes means for learning from a user's past conversation data to provide personalized responses.

[1408] (Claim 3)

[1409] 10. The system of claim 1, wherein the generative AI model includes means for analyzing user utterances in real time and generating appropriate responses based on the context of the conversation.

[1410] "Application Example 1"

[1411] (Claim 1)

[1412] means for obtaining a user's voice input;

[1413] a means for converting the acquired speech into text;

[1414] a means for analyzing the text using generative AI to generate an appropriate response;

[1415] means for converting the generated response into speech;

[1416] means for providing audio feedback to the user;

[1417] A means for saving conversation data with the user and using it the next time the user uses the service;

[1418] A system that includes a means to support real-time communication between employees and customers.

[1419] (Claim 2)

[1420] The generating AI includes means for learning from the user's past conversation data and providing personalized responses.

[1421] 10. The system of claim 1.

[1422] (Claim 3)

[1423] The generating AI includes means for analyzing user utterances in real time and generating appropriate responses based on the context of the conversation.

[1424] 10. The system of claim 1.

[1425] "Example 2: Combining Emotion Engines"

[1426] (Claim 1)

[1427] means for obtaining a user's voice input;

[1428] a means for converting the acquired speech into text;

[1429] a means for analyzing the text using a generative AI model to generate an appropriate response; and

[1430] means for converting the generated response into speech;

[1431] means for providing audio feedback to the user;

[1432] A means for saving conversation data with the user and using it the next time the user uses the service;

[1433] A means of recognizing user emotions and passing that emotion data to a generative AI model;

[1434] A means of analyzing user sentiment in real time and reflecting it in responses

[1435] A system including:

[1436] (Claim 2)

[1437] 10. The system of claim 1, wherein the generative AI model includes means for learning from a user's past conversation data to provide personalized responses.

[1438] (Claim 3)

[1439] 10. The system of claim 1, wherein the generative AI model includes means for analyzing user utterances and sentiment in real time and generating appropriate responses based on the context of the conversation.

[1440] "Application example 2 when combining emotion engines"

[1441] (Claim 1)

[1442] means for obtaining a user's voice input;

[1443] a means for converting the acquired speech into text;

[1444] a means for analyzing the text using generative AI to generate an appropriate response;

[1445] means for converting the generated response into speech;

[1446] means for providing audio feedback to the user;

[1447] means for recognizing a user's emotion;

[1448] means for generating a response based on the emotion data;

[1449] A system that includes a means for saving conversation data with a user and using it the next time the system is used.

[1450] (Claim 2)

[1451] 10. The system of claim 1, wherein the generative AI includes means for learning from a user's past conversation data and providing personalized responses.

[1452] (Claim 3)

[1453] 2. The system of claim 1, wherein the generative AI includes means for analyzing user utterances in real time and generating appropriate responses based on the context of the conversation. [Explanation of symbols]

[1454] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for obtaining a user's voice input; a means for converting the acquired speech into text; a means for analyzing the text using generative AI to generate an appropriate response; means for converting the generated response into speech; means for providing audio feedback to the user; A system that includes a means for saving conversation data with a user and using it the next time the system is used.

2. 2. The system of claim 1, wherein the generative AI includes means for learning from a user's past conversation data and providing personalized responses.

3. The system of claim 1 , wherein the generative AI includes means for analyzing user utterances in real time and generating appropriate responses based on the context of the conversation.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A