System
A system leveraging generative AI and speech recognition technologies addresses the customization needs of elderly individuals with dementia, enhancing daily life support through personalized interactions and mobility assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-21
- Publication Date
- 2026-03-06
AI Technical Summary
Existing systems fail to effectively customize support for elderly individuals with dementia, particularly those with low IT literacy and mobility challenges, lacking comprehensive solutions for daily life assistance.
A system utilizing generative AI and speech recognition technology to create individually tailored models, enabling voice input, response generation, and interactive tutorials, along with transportation service booking, to enhance daily life support for the elderly.
Provides personalized support for conversations, improves IT literacy, and enhances mobility for elderly individuals with dementia, thereby improving their quality of life.
Smart Images

Figure 2026037221000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] To effectively support elderly people with dementia, it is important to customize the system to meet their individual needs, but existing systems lack the technological means to achieve this. Furthermore, the elderly face challenges such as low IT literacy and the inconvenience of mobility. A system that comprehensively solves these problems is needed, and this invention aims to solve these problems. [Means for solving the problem]
[0005] The present invention solves the above problems by the following means. A system is constructed that provides a means for saving information entered by a user in a database and generates individually customized generative AI models and speech recognition models based on the saved information. The system also provides a means for users to add voice input and includes a function for transmitting the voice data to a server. The server has a means for analyzing the received voice data and retraining the generative AI model and speech recognition model, and further includes a means for converting the voice input from the elderly person into text data and transmitting it to the server. The server provides a means for generating appropriate responses using the generative AI and transmitting it to the terminal, and the terminal has a means for converting the received response text into speech and playing it back to the elderly person. The system also includes a means for providing an interactive tutorial mode and a means for booking transportation services and adjusting transportation schedules.
[0006] A "user" is an individual who uses the system to input information about the senior citizen or add voice input.
[0007] "Information" refers to data entered by the user, such as the elderly person's name, background, hobbies, preferences, and dialect.
[0008] A "database" is an electronic storage system for storing user-entered information.
[0009] A "generative AI model" is an artificial intelligence model customized based on user input information, which generates appropriate responses to questions posed by elderly people.
[0010] The "voice recognition model" is an artificial intelligence model for converting elderly people's voices into text data.
[0011] "Voice input" refers to voice data that the elderly person or user utters to the system.
[0012] A "server" is a centralized computer system that manages and processes the entire system.
[0013] "Analysis" is the process by which the server processes the audio and text data it receives and understands its meaning.
[0014] "Retraining" is the process of retraining an existing generative AI model or speech recognition model to adapt to additional data or new conditions.
[0015] "Text data" is data in the form of sentences that is generated by a speech recognition model based on speech input.
[0016] A "response" is an answer created by the generative AI model in response to a question asked by an elderly person.
[0017] A "terminal" is a device such as a smartphone or tablet that is actually operated by a user or elderly person.
[0018] "Conversion to audio" is the process of converting text data back into audio format and playing it for the elderly.
[0019] "Interactive tutorial mode" is an interactive learning mode that makes it easier for seniors to learn how to operate IT.
[0020] "Transportation services" are means of transportation (e.g., vehicles or shuttle services) that can be reserved and used by seniors.
[0021] "Reservation" is the process by which seniors apply for transportation services in advance through the system.
[0022] A "travel schedule" is a schedule that plans the date, time, and route of travel for the elderly person. [Brief explanation of the drawings]
[0023] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] The present invention is a system for supporting dementia in the elderly, which combines generative AI and speech recognition technology to provide individually customized interactive support. Specific embodiments of this system are described in detail below.
[0045] System Overview
[0046] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[0047] Program processing
[0048] 1. User enters information
[0049] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0050] The device converts this information into JSON format and sends it to the server.
[0051] 2. The server saves the information and generates the model
[0052] The server stores the received information in a database.
[0053] The server customizes the generative AI model and speech recognition model based on the stored information.
[0054] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model with new audio datasets as needed.
[0055] 3. User adds voice input
[0056] Through the application, users record voice input and provide information about the senior's past events and favorite topics.
[0057] The device transmits the recorded audio data to the server.
[0058] 4. The server analyzes the audio data and updates the model.
[0059] The server analyzes the received voice data and retrains generative AI and speech recognition models optimized for specific dialects and individual speaking styles.
[0060] 5. Seniors start asking questions and initiating conversations
[0061] Seniors use the app to ask questions and initiate conversations.
[0062] The device collects the elderly person's voice in real time, converts it into text through a voice recognition engine, and sends it to a server.
[0063] 6. The server generates an appropriate response
[0064] The server analyzes the received text and uses generative AI to generate an appropriate response.
[0065] The generated response is sent to the terminal in text format.
[0066] 7. The device converts your reply into voice
[0067] The device converts the received response text into speech using a speech synthesis engine.
[0068] The device plays the generated audio to the elderly person.
[0069] Specific examples
[0070] For example, if an elderly person named Mr. Tanaka were to use this system, his son would enter information about Mr. Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, into the application. The son would then record and add information about Mr. Tanaka's past experiences and hobbies. This information and voice data is sent to a server, which then creates a generative AI model and a voice recognition model specifically for Mr. Tanaka.
[0071] When Tanaka asks "What's the weather like today?" through the application, the device converts the question into text and sends it to the server. The server uses generative AI to generate a response such as "It's sunny today, but the wind is expected to pick up a little in the afternoon," and sends it to the device. The device then converts this response into audio and plays it for Tanaka.
[0072] If Tanaka wants to travel to a specific location, he can reserve a transportation service through the application. The device sends the reservation information to the server, which then adjusts the travel schedule. As a result, Tanaka can travel safely and comfortably.
[0073] As described above, this invention aims to support elderly people with dementia by utilizing individually customized generative AI and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[0077] Step 2:
[0078] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[0079] Step 3:
[0080] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[0081] Step 4:
[0082] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[0083] Step 5:
[0084] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[0085] Step 6:
[0086] The server analyzes the received text and generates an appropriate response using a generation AI. The generated response text is then sent to the device.
[0087] Step 7:
[0088] The device converts the received response text into speech using a speech synthesis engine, and the device plays the generated speech back to the elderly person, providing an answer to their question.
[0089] Step 8:
[0090] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[0091] Step 9:
[0092] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[0093] Step 10:
[0094] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[0095] Example 1
[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0097] To support elderly people with dementia, there is a need for systems that provide interactive support tailored to individual needs. However, current technology does not have an established method for quickly generating and operating AI models and voice recognition models optimized for each elderly person, which means that elderly people are unable to receive appropriate support. In addition, there is a lack of comprehensive systems that also provide support for daily life, such as improving IT literacy and providing mobility support.
[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0099] In this invention, the server includes: a means for converting information entered by a user into a data format and transmitting it to the server; a means for storing the received information in a database; a means for generating individually customized generative AI models and speech recognition models based on the stored information; a means for the user to add voice input and transmit the voice data to the server; a means for the server to analyze the received voice data and retrain the models; a means for converting the voice to text and transmitting it to the server; a means for the server to analyze the text using the generative AI to generate an appropriate response and transmit the response text to the terminal; and a means for the terminal to convert the received response text into voice and play it back. This enables dementia support tailored to each elderly person. It also enables integrated support for elderly people's IT literacy and mobility, thereby improving the overall quality of their daily lives.
[0100] A "server" is a computer system that provides data and functionality to other computers and devices over a network.
[0101] A "terminal" is an electronic device that is directly operated by a user and that communicates with a server.
[0102] A "user" is a person who utilizes the system to enter information or add voice input.
[0103] A "data format" is a format for expressing information in a structured way, such as the JSON format.
[0104] A "database" is a system that stores and manages data, making it easy to search for and analyze information.
[0105] A "generative AI model" is a type of artificial intelligence that automatically responds and processes data, and is used for natural language processing, etc.
[0106] A "speech recognition model" is a model that analyzes voice data and converts it into text.
[0107] "Voice input" refers to voice data provided by a user speaking to the system.
[0108] "Text" means the written representation of voice input or other data.
[0109] A "speech synthesis engine" is software for converting text data into voice data.
[0110] The "interactive tutorial mode" is an interactive support function that helps users learn how to operate the system.
[0111] A "mobility service" is a service that provides assistance to a user moving to a specific location.
[0112] "Schedule adjustment" refers to the act of managing and optimizing the dates and times of transportation services booked by users.
[0113] This invention is a system aimed at supporting dementia in the elderly, combining generative AI and speech recognition technology to provide individually customized interactive support. This system is comprised of three main elements: a server, a terminal, and a user. Below, we will explain how to specifically implement this system.
[0114] server
[0115] The server is a high-performance computer system that manages and operates the database, generative AI model, and speech recognition model. The main software used is Tensorflow (registered trademark) or PyTorch, which uses Python, for learning and retraining the generative AI model and speech recognition model. MySQL (registered trademark) or PostgreSQL is used for the database.
[0116] Specific examples
[0117] When the user enters the elderly person's information, the server receives the JSON format data sent from the device and stores it in a database. Based on the received data, an individually customized generative AI model and voice recognition model are created.
[0118] Terminal
[0119] A terminal is a device that allows users to input information, record voice data, and communicate with a server. Terminals can be smartphones, tablets, or PCs. Audio recording and playback require a built-in or external microphone and speaker.
[0120] Specific examples
[0121] The user inputs information about Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, through the device, and this information is converted into JSON format and sent to the server. The user then speaks aloud about Tanaka's past events and hobbies, and the audio data is recorded by the device and sent to the server.
[0122] User
[0123] The user is the person who operates the system, inputting information and voice input for the elderly, learning IT operations through an interactive tutorial mode, and booking transportation services on behalf of the elderly.
[0124] Specific examples
[0125] For example, Tanaka's son uses an application to enter Tanaka's information and sends the voice data to the server. The server then customizes the generative AI model and voice recognition model based on the received information, creating a system specifically for Tanaka. When Tanaka asks, "What's the weather like today?", the server uses the generative AI to create an appropriate response and responds via voice via the device.
[0126] Prompt Sentence Examples
[0127] Information input prompt: "Please enter the senior's name, background, hobbies, favorite foods, and local dialect."
[0128] Voice prompt: "Record an older adult talking about past events and hobbies."
[0129] Conversation prompt: "What's the weather like today?"
[0130] Movement service prompt: "Enter details of your next move destination."
[0131] As described above, this system aims to support elderly people with dementia, and utilizes individually customized generative AI models and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[0132] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0133] Step 1:
[0134] User enters information
[0135] Specific operation:
[0136] Users launch a dedicated application and enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0137] input:
[0138] Information on the elderly person's name, background, hobbies, preferences, local dialect, etc.
[0139] Data processing / calculation:
[0140] The terminal converts the input information into JSON format.
[0141] output:
[0142] Elderly information data in JSON format.
[0143] Specific actions added:
[0144] When a user enters information into the input screen and clicks the "Submit" button, the information is converted into JSON format.
[0145] Step 2:
[0146] The server stores the information and generates the model
[0147] Specific operation:
[0148] The device sends the generated JSON data to the server.
[0149] The server stores the received data in a database.
[0150] input:
[0151] Elderly information data in JSON format.
[0152] Data processing / calculation:
[0153] Extract the necessary information from the JSON data and store it in the database.
[0154] Customize generative AI and speech recognition models using libraries such as Python and TensorFlow.
[0155] Adjust the model's learning parameters and possibly retrain it.
[0156] output:
[0157] Customized generative AI and speech recognition models.
[0158] Specific actions added:
[0159] Save the JSON data into a MySQL database and run a script to generate the model.
[0160] Step 3:
[0161] User adds voice input
[0162] Specific operation:
[0163] Users operate an application designed specifically for seniors and input information about past events and hobbies using voice.
[0164] input:
[0165] Voice input of elderly people's past events and hobbies.
[0166] Data processing / calculation:
[0167] The device records the audio and generates an audio file in WAV format or similar.
[0168] Send the audio file to the server.
[0169] output:
[0170] Audio files (WAV format, etc.).
[0171] Specific actions added:
[0172] Press the record button to record the audio, and then press the "send" button after recording is complete to send the audio file to the server.
[0173] Step 4:
[0174] The server analyzes the audio data and updates the model
[0175] Specific operation:
[0176] The server analyzes the received audio file and converts it into text using a speech-to-text engine.
[0177] The parameters of the generative AI model and speech recognition model are readjusted based on the converted text information.
[0178] input:
[0179] Audio file.
[0180] Data processing / calculation:
[0181] A speech-to-text engine is used to convert the audio into text and store it in a database.
[0182] Retrain the model based on the converted text and existing data.
[0183] output:
[0184] Updated generative AI and speech recognition models.
[0185] Specific actions added:
[0186] Input the audio file into the analysis script and run model retraining along with the text conversion results.
[0187] Step 5:
[0188] Seniors start asking questions and conversations
[0189] Specific operation:
[0190] The elderly person launches the application and speaks a prompt to start the conversation (e.g., "What's the weather like today?").
[0191] input:
[0192] Questions and conversations of the elderly.
[0193] Data processing / calculation:
[0194] The device collects audio in real time and records the audio data.
[0195] The speech is converted into text through a speech recognition engine and sent to the server.
[0196] output:
[0197] Conversation in text format.
[0198] Specific actions added:
[0199] The question is recorded, and after the recording is complete, it is automatically converted into text and sent to the server.
[0200] Step 6:
[0201] The server generates an appropriate response
[0202] Specific operation:
[0203] The server analyzes the received text and uses generative AI to generate an appropriate response.
[0204] The generated reply text is sent to the terminal.
[0205] input:
[0206] Textual conversation from the user.
[0207] Data processing / calculation:
[0208] Generative AI (e.g., GPT-3 (registered trademark)) is used to analyze the text and generate responses.
[0209] output:
[0210] The reply text.
[0211] Specific actions added:
[0212] It runs an algorithm to generate a response to the question and sends the generated text to the terminal.
[0213] Step 7:
[0214] The device converts your response into voice
[0215] Specific operation:
[0216] The device converts the received response text into speech using a speech synthesis engine.
[0217] The generated audio is played to the elderly person.
[0218] input:
[0219] The response text from the server.
[0220] Data processing / calculation:
[0221] Convert text to speech using a speech synthesis engine (e.g., Google® Text-to-Speech).
[0222] output:
[0223] Audio data.
[0224] Specific actions added:
[0225] The device starts the speech synthesis engine and plays the generated speech on the playback device.
[0226] (Application example 1)
[0227] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0228] Elderly people with dementia may face many challenges in their daily lives. These challenges include being unable to respond appropriately in emergencies, being unable to ensure daily safety, and having difficulty preventing accidents at home. Furthermore, dealing with these challenges places a burden on the individual and their family, which is problematic. The present invention aims to utilize generative AI and speech recognition technology to individually and effectively resolve these challenges.
[0229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0230] In this invention, the server includes means for saving information entered by a user in a database, means for generating individually customized generative AI models and voice recognition models based on the saved information, means for the user to add voice input and send the voice data to the server, means for the server to analyze the received voice data and retrain the generative AI model and voice recognition model, means for converting the elderly person's voice input into text data and sending it to the server, means for the server to generate an appropriate response using the generative AI and send it to the terminal, means for the terminal to convert the received response text into voice and play it back to the elderly person, means for notifying an emergency contact and providing voice instructions when the elderly person encounters an emergency, means for checking the safety of the elderly person at regular intervals and notifying if an abnormality occurs, and means for monitoring dangers within the home and issuing an alert if an abnormality occurs. This enables rapid response to emergencies, regular safety checks, and safety monitoring within the home.
[0231] "Means for saving information entered by the user in a database" refers to means for saving data obtained from the user, such as personal information, hobbies, and past events.
[0232] "Means for generating individually customized generative AI models and speech recognition models based on stored information" refers to means for creating generative AI and speech recognition models that are optimal for individual users by utilizing information stored in a database.
[0233] The "means for a user to add a voice input and transmit the voice data to a server" refers to a means for a user to record voice data and transmit the data to a server.
[0234] "Means for analyzing voice data received by the server and retraining the generative AI model and the voice recognition model" refers to means for analyzing voice data received by the server and retraining the generative AI and the voice recognition model.
[0235] The "means for converting the voice input of the elderly person into text data and transmitting it to the server" is a means for converting the voice uttered by the elderly person into text in real time and transmitting the text data to the server.
[0236] "Means for the server to use generation AI to generate an appropriate response and send it to the terminal" means a means for the server to use generation AI technology to generate an appropriate response to the received text data and send that response to the terminal.
[0237] The "means for converting the response text received by the terminal into speech and playing it back to the elderly" refers to a means for converting the response text received by the terminal from the server into speech using speech synthesis technology and playing it back to the elderly.
[0238] "Means for notifying emergency contacts and providing voice instructions when an elderly person encounters an emergency" refers to a means for automatically notifying pre-set emergency contacts and simultaneously providing appropriate voice guidance when an elderly person faces an emergency.
[0239] "Means for checking the safety of elderly people at regular intervals and notifying in the event of an abnormality" refers to a means for periodically checking the condition of elderly people and notifying emergency contacts if an abnormality is detected.
[0240] "Means for monitoring dangers within the home and issuing warnings in the event of an abnormality" refers to a means for monitoring smoke, gas leaks, etc. using sensors installed within the home and issuing warnings in the event of an abnormality.
[0241] The present invention is a system for supporting elderly people with dementia, which uses individually customized generative AI models and speech recognition models to provide support for elderly people in their daily lives. Specific embodiments of the system are described below.
[0242] System Configuration and Operation
[0243] Hardware and Software
[0244] Devices: Smartphones, smart glasses, head-mounted displays, etc.
[0245] Server: We use cloud-based servers to host the database, generative AI models, and speech recognition models.
[0246] Sensors: Smoke detectors and gas sensors are installed in the home, and data is sent to the device when an abnormality occurs.
[0247] Data processing and calculation
[0248] 1. Entering and saving information
[0249] Users enter personal information, hobbies, past events, etc. about the elderly person through a dedicated application, and this information is converted into JSON format and sent to the server, which then stores this data in a database.
[0250] 2. Generate and retrain the model
[0251] The server creates a generative AI model and a voice recognition model based on the stored information, and customizes them specifically for seniors. When additional voice input is received, the server analyzes the voice data and retrains the model.
[0252] 3. Speech Processing and Text Conversion
[0253] When the elderly person speaks, the device converts the speech into text in real time and sends it to the server, which then generates a response based on the received text.
[0254] 4. Response Generation and Audio Playback
[0255] The server uses generative AI to generate appropriate responses and sends them in text format to the device, which then converts the text into speech and plays it back to the elderly.
[0256] 5. Emergency response and safety monitoring
[0257] In the event of an emergency, the device will automatically notify emergency contacts and provide appropriate voice guidance. The device also monitors data from sensors and notifies users if an abnormality occurs.
[0258] Specific examples
[0259] For example, the present invention is useful in the following specific situations.
[0260] Example 1: Emergency response
[0261] If an elderly person falls and calls for help on their smartphone, the application will recognize the voice and quickly notify emergency contacts. At the same time, it will provide appropriate voice guidance to help the elderly person take immediate action.
[0262] Example prompt sentence:
[0263] "I have had a fall. Please notify your emergency contacts."
[0264] Example 2: Safety reminder
[0265] The application periodically asks, "How are you?" and waits for a response from the elderly person. If an abnormality is detected, it immediately notifies emergency contacts.
[0266] Example prompt sentence:
[0267] "How are you? If you don't respond, I'll notify my emergency contacts."
[0268] As described above, the present invention is a system that utilizes generative AI models and voice recognition technology to comprehensively support the daily lives of the elderly and provide a safe and secure environment.
[0269] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0270] Step 1:
[0271] The user enters information
[0272] Users enter personal information (such as name, address, and emergency contact information), lifestyle habits, hobbies, and past events through a dedicated application. The application converts the entered information into JSON format and sends it to the server via the network.
[0273] Input: Personal information and past event data entered by the user
[0274] Output: Data converted to JSON format
[0275] Step 2:
[0276] The server stores the information
[0277] The server parses the JSON format data received over the network and stores it in a database for later processing.
[0278] Input: JSON format user information data
[0279] Output: User information stored in the database
[0280] Step 3:
[0281] Generate a model based on the retained information
[0282] The server uses the information stored in the database to create individually customized generative AI and speech recognition models, which involves applying machine learning algorithms to the data.
[0283] Input: User information stored in the database
[0284] Output: Generated generative AI model and speech recognition model
[0285] Step 4:
[0286] User adds voice input
[0287] The user inputs voice data through the application, and the voice data is collected in real time by the device and sent to a server for analysis.
[0288] Input: User's voice data
[0289] Output: Audio data sent to the server
[0290] Step 5:
[0291] The server analyzes the audio data and retrains the model.
[0292] The server analyzes the received voice data and retrains the generative AI model and speech recognition model to improve their performance, using a speech recognition algorithm.
[0293] Input: Audio data
[0294] Output: Updated AI and speech recognition models after retraining
[0295] Step 6:
[0296] Converting elderly people's voice input into text
[0297] When an elderly person uses the application to input voice, the device converts the voice into text data in real time and sends it to the server, using a voice recognition engine.
[0298] Input: Elderly voice data
[0299] Output: Data converted to text
[0300] Step 7:
[0301] The server generates an appropriate response
[0302] The server uses AI to generate an appropriate response based on the received text data. This response is generated based on the user's individually customized information and sent to the device.
[0303] Input: Text data
[0304] Output: The generated response text
[0305] Step 8:
[0306] The device converts your response into speech
[0307] The device converts the response text received from the server into voice using a speech synthesis engine and plays it back to the elderly, who can then confirm the response by voice.
[0308] Input: Reply text
[0309] Output: The speech-transcribed reply
[0310] Step 9:
[0311] Emergency response
[0312] If an elderly person encounters an emergency, the device will instantly accept voice input, notify emergency contacts, and provide appropriate voice instructions from the server, using voice recognition and generation AI technology.
[0313] Input: Voice input in case of emergency
[0314] Output: Provided voice instructions and emergency notifications
[0315] Step 10:
[0316] Safety checks and home safety monitoring
[0317] The device periodically checks the elderly person's condition, and if an abnormality is detected, it immediately notifies the server and contacts the emergency contact. It also uses sensors in the home to monitor for smoke and gas leaks, and issues an alert if an abnormality is detected.
[0318] Input: Elderly status and home sensor data
[0319] Output: Safety check results and warning notifications
[0320] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0321] The present invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI, speech recognition technology, and an emotion engine. Specific embodiments of this system are described in detail below.
[0322] System Overview
[0323] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[0324] Furthermore, by combining it with an emotion engine, it becomes possible to recognize the user's emotions and generate responses based on those emotions. This emotion engine analyzes emotions from the user's voice and sends the results to a server, allowing the generative AI model to adjust the response content.
[0325] Program processing
[0326] 1. User enters information
[0327] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0328] The device converts this information into JSON format and sends it to the server by pressing the send button.
[0329] 2. The server saves the information and generates the model
[0330] The server stores the received information in a database.
[0331] The server customizes the generative AI model and speech recognition model based on the stored information.
[0332] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using new audio datasets as needed.
[0333] 3. User adds voice input
[0334] Users access the "Add Learning Data" section of the application and use the recording function to dictate events from the elderly person's past or favorite topics.
[0335] The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[0336] The device sends the encoded audio data to the server.
[0337] 4. The server analyzes the audio data and updates the model.
[0338] The server stores the received audio data in an appropriate format and analyzes it.
[0339] The server then retrains the generative AI model and speech recognition model based on this voice data, updating the models to be optimized for specific dialects and individual speaking styles.
[0340] 5. Seniors start asking questions and initiating conversations
[0341] Using the application, the elderly person presses the talk button to ask a question or start a conversation.
[0342] The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine, which then sends the text data to a server.
[0343] 6. Emotion engine recognizes user emotions
[0344] As the voice data is sent to the server, the emotion engine analyzes the emotion from the user's voice.
[0345] The emotion analysis results are sent to a server, which records the user's emotional state.
[0346] 7. The server generates an appropriate response based on the emotion.
[0347] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[0348] The generated reply text is sent to the terminal.
[0349] 8. The device converts your reply into voice
[0350] The response text received by the device is converted into speech using a speech synthesis engine.
[0351] The device plays the generated audio to the elderly person, providing answers to their questions and appropriate feedback.
[0352] Specific examples
[0353] For example, if Mr. Tanaka were to use the system, his son would input information about him, such as his name, background, hobbies, favorite foods, and local dialect, into the application, and provide additional voice data, which would then create a generative AI model and a voice recognition model specifically for him.
[0354] When Tanaka asks the application, "What's the weather like today?", the device converts this speech into text, while the emotion engine simultaneously analyzes Tanaka's emotions. The text data and emotion data are sent to the server, which uses this information to generate a response that takes emotion into account, such as, "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response is converted into audio and played to Tanaka in real time.
[0355] Also, if Tanaka is in a certain emotional state, for example, feeling anxious, the system will provide a reassuring response tailored to that emotion, such as, "The weather is nice today, and your favorite flowers might be in bloom."
[0356] In this way, the present invention aims to support dementia in the elderly, and by combining generative AI, voice recognition technology, and an emotion engine, it provides individually customized interactive support, realizing comprehensive support for the elderly and their families.
[0357] The processing flow will be explained below.
[0358] Step 1:
[0359] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[0360] Step 2:
[0361] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[0362] Step 3:
[0363] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[0364] Step 4:
[0365] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[0366] Step 5:
[0367] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[0368] Step 6:
[0369] While the voice data is being sent to the server, the device uses an emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions and sends the results to the server along with the text data.
[0370] Step 7:
[0371] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[0372] Step 8:
[0373] The server generates a response text and sends it to the device. The device converts the received response text into speech using a speech synthesis engine. The device plays the generated speech to the elderly person, providing answers to their questions and appropriate feedback.
[0374] Step 9:
[0375] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[0376] Step 10:
[0377] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[0378] Step 11:
[0379] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[0380] Example 2
[0381] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0382] There is a lack of interactive support optimized for each patient in dementia support for the elderly. In particular, it is difficult to provide support that meets the individual needs of the elderly by using individually customized generative AI models and speech recognition models. Another issue is the lack of technology that can recognize emotions from speech and generate appropriate responses based on them. Furthermore, there is a need for an interactive tutorial mode to improve the IT literacy of the elderly and for improved mobility service reservation functions.
[0383] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for storing information input by a user in a database, a means for generating individually customized generative AI models and voice recognition models based on the stored information, a means for converting the user's voice input into text data and transmitting it to the server, and a means for adjusting the content of responses generated based on the emotion analysis results. This makes it possible to provide optimal support according to the individual needs of elderly people. In addition, by analyzing the user's emotions using an emotion recognition engine and generating responses that take these emotions into consideration, more personalized interactive dialogue is realized.
[0384] "User" refers to an individual who uses the system to input information or voice input.
[0385] "Information" refers to data that users enter into the system, including the elderly person's name, background, hobbies, preferences, local dialect, etc.
[0386] A "database" is a storage device that stores input information and allows it to be retrieved when needed.
[0387] A "generative AI model" is an algorithmic model of artificial intelligence that is individually customized based on user input information.
[0388] A "speech recognition model" is an algorithmic model used to convert a user's speech data into text data.
[0389] "Audio data" refers to digital audio files generated by a user providing voice input.
[0390] A "server" is a computer system responsible for storing information and generating and training generative AI models and speech recognition models.
[0391] "Analysis" is the process in which the server analyzes information based on the voice data received and performs the necessary processing.
[0392] "Text data" is character string information converted from voice data by a voice recognition model.
[0393] An "emotion recognition engine" is software that analyzes emotions from the user's voice and sends the results to a server.
[0394] A "reply" is a response that the server generates in response to a user's question or request using a generative AI model.
[0395] The "interactive tutorial mode" is an interactive learning mode provided for seniors to learn how to use the system.
[0396] "Transportation service" is a function that allows elderly people to reserve available transportation and adjust their travel schedules.
[0397] "Model training" is the process by which a generative AI model or speech recognition model learns from stored data and improves its performance.
[0398] "Playback" is the process by which the device plays the generated audio to the elderly person.
[0399] This invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI models, speech recognition technology, and an emotion engine. Specific embodiments of the present invention are described below.
[0400] System Overview
[0401] Users enter information about the elderly person through a dedicated application, and the device converts this information into JSON format and sends it to the server. The server stores the received information in a database and individually customizes the generative AI model and speech recognition model. The generative AI model used is GPT-4 (registered trademark), and the speech recognition model is Google Speech-to-Text API. Furthermore, IBM Watson (registered trademark) Tone Analyzer is used for emotion recognition.
[0402] Users can improve the accuracy of the model by adding information about the senior's past experiences, hobbies, preferences, etc. This allows the generative AI model and speech recognition model to be individually tailored to the senior. The system also features an interactive tutorial mode and a mobility service booking function.
[0403] Furthermore, by combining it with an emotion recognition engine, it can recognize the user's emotions in real time and generate responses based on those emotions, enabling interactive dialogue that takes the user's emotional state into account.
[0404] Hardware and Software Details
[0405] server
[0406] Database: A relational database such as "MySQL" or "PostgreSQL" to store the information.
[0407] Generative AI models: Individually customized generative AI models (e.g., "GPT-4").
[0408] Speech recognition model: A speech recognition model specifically optimized for older adults (e.g., Google Speech-to-Text API).
[0409] Terminal
[0410] Conversion and transmission of input information: A function that converts information entered by the user into JSON format and sends it to the server. This is performed by software processing inside the device.
[0411] Encoding and sending audio data: The recorded audio data is encoded with appropriate sound quality parameters (sample rate, bit rate) and sent to the server.
[0412] Emotion Recognition Engine
[0413] IBM Watson Tone Analyzer: Analyzes emotions from the user's voice and sends the results to the server.
[0414] Response generation and speech synthesis
[0415] Generative AI model: Generates appropriate responses based on the user's text data and sentiment analysis results.
[0416] Speech synthesis engine: A speech synthesis engine such as "Amazon Polly" is used to convert response text into speech.
[0417] Specific examples
[0418] Specific examples are shown below.
[0419] For example, a user inputs information about an elderly person (Mr. Tanaka), such as his name, background, hobbies, favorite foods, and local dialect, into an application, which then sends this information to a server.
[0420] The server stores this information in a database and customizes the generative AI model and speech recognition model based on the input data. The user can then add voice input about Tanaka's past events and favorite topics, and send this voice data to the server.
[0421] The server analyzes the received voice data and retrains the generative AI model and speech recognition model. When the elderly person starts a question or conversation using the application, the device converts the voice into text and sends it to the server.
[0422] The emotion recognition engine analyzes emotions from the voice and sends the results to a server. The server uses a generative AI model based on the emotion analysis results and text data to generate an appropriate response and sends it to the device. The device then converts the response text into speech and plays it back to the elderly.
[0423] For example, if Tanaka asks, "How's the weather today?", the emotion analysis engine will determine that Tanaka is feeling a little anxious. The server will then generate a reassuring response: "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response will be converted into audio and played to Tanaka in real time.
[0424] In this way, the present invention aims to support dementia in the elderly, and by providing individually customized interactive support, it realizes comprehensive support for the elderly and their families.
[0425] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0426] Step 1: User Enters Information
[0427] What happens: The user opens a dedicated application and enters information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0428] Input: Information about the elderly person (such as name, background, hobbies, preferences, local dialect, etc.).
[0429] Data processing: The terminal converts this information into JSON format.
[0430] Output: Data converted to JSON format.
[0431] Specific operation: Enter data into the input form displayed on the terminal and press the send button. The data is converted into JSON format and the message "Submission successful" is displayed.
[0432] Step 2: The server saves the information and generates the model
[0433] Processing details: The server stores the received information in a database and customizes the generative AI model and speech recognition model individually.
[0434] Input: JSON formatted data.
[0435] Data processing: The received information is stored in a database and used to customize generative AI models (e.g., GPT-4) and speech recognition models (e.g., Google Speech-to-Text API).
[0436] Output: Customized generative AI model and speech recognition model.
[0437] What happens: The server parses the JSON data and stores the necessary information in a database. It then adjusts the parameters of the AI model based on the stored information and retrains the speech recognition model using the new audio dataset.
[0438] Step 3: User adds voice input
[0439] What happens: The user visits the "Add Learning Data" section of the app and uses the recording feature to dictate events from the senior's past or favorite topics.
[0440] Input: Speech data of elderly people.
[0441] Data processing: The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[0442] Output: The encoded audio data.
[0443] Specific operation: The user presses the record button and inputs voice information about the elderly person's past events or favorite topics. Once the recording is complete, the device encodes the voice data and sends it to the server.
[0444] Step 4: The server analyzes the audio data and updates the model
[0445] Processing details: The server saves the received audio data in an appropriate format and analyzes it.
[0446] Input: Encoded audio data.
[0447] Data processing: Analyze the voice data and retrain the generative AI model and speech recognition model.
[0448] Output: Updated generative AI model and speech recognition model.
[0449] What it does: The server converts the audio data into a different format, generates training data optimized for a specific dialect or speaking style, and updates the model based on the generated data.
[0450] Step 5: The senior initiates the question or conversation
[0451] What happens: Seniors use the application and press the talk button to ask questions or start a conversation.
[0452] Input: Speech data of elderly people.
[0453] Data processing: The device collects the elderly person's voice in real time and converts it into text using a voice recognition engine.
[0454] Output: Parsed text data.
[0455] Specific operation: When an elderly person presses the button to speak, the device collects voice data and converts it into text in real time. This text data is then sent to the server.
[0456] Step 6: The emotion engine recognizes the user's emotion
[0457] Processing details: As the voice data is sent to the server, the emotion engine analyzes the emotions from the user's voice.
[0458] Input: Audio data.
[0459] Data processing: The emotion analysis engine recognizes emotions from voice and converts the results into data.
[0460] Output: Sentiment analysis result data.
[0461] Specific operation: The voice data sent to the server is passed to the emotion engine, where emotion analysis is performed. The analysis results are stored in the database as the emotional state.
[0462] Step 7: The server generates an appropriate response based on the sentiment
[0463] Processing details: The server uses generative AI to generate an appropriate response based on the text received and the results of sentiment analysis.
[0464] Input: Text data and sentiment analysis result data.
[0465] Data processing: Generative AI models analyze text and emotional state to generate optimal responses.
[0466] Output: The generated response text.
[0467] Specific operation: The server uses the text data and the sentiment analysis results to query the generative AI model and generate an appropriate response, which is then queued for transmission to the device.
[0468] Step 8: Your device converts your response into audio
[0469] Processing details: The response text received by the device is converted into speech using a speech synthesis engine.
[0470] Input: Response text.
[0471] Data processing: Use a speech synthesis engine (e.g., Amazon Polly) to convert text data into speech.
[0472] Output: Synthesized speech data.
[0473] Specific operation: The device receives the response text and converts it into natural speech using a speech synthesis engine. The device then plays this speech back to the elderly person and provides appropriate feedback.
[0474] (Application example 2)
[0475] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0476] Current assistance systems for the elderly struggle to provide individually customized interactive support, and are unable to fully improve the convenience and sense of security of the elderly in their daily lives. Furthermore, seniors lack the means to receive appropriate real-time support when searching for products or answering questions in physical stores. To solve these problems and significantly improve the quality of life for the elderly, a new system is needed that combines generative AI models, speech recognition technology, and emotion engines to provide customized interactive assistance.
[0477] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for storing information input by a user in a database; means for generating individually customized generative AI models and speech recognition models based on the stored information; means for the user to add voice input and send the voice data to the server; means for the server to analyze the received voice data and retrain the generative AI model and speech recognition model; means for converting the elderly's voice input into text data and sending it to the server; means for the server to generate an appropriate response using the generative AI and send it to the terminal; means for the terminal to convert the received response text into speech and play it back to the elderly; means for analyzing emotions from the elderly's speech using an emotion engine and sending the result to the server; and means for providing real-time interactive support via a smart device worn by the elderly. This makes it possible to provide appropriate real-time support for elderly people when searching for products in physical stores, such as guidance and question-answering.
[0478] The "user" is a person who uses the system and is responsible for inputting information about the elderly person and adding voice input.
[0479] "Information" is a general term for data necessary for individual customization, such as the elderly person's name, background, hobbies, preferences, and local dialect.
[0480] "Database" refers to a storage device and system for storing user-entered information and elderly person's voice data.
[0481] A "generative AI model" is an artificial intelligence model that is generated based on information about the user or elderly person, and is used to provide appropriate responses and support.
[0482] The "voice recognition model" is an artificial intelligence model that analyzes the voices of elderly people and converts the content into text data.
[0483] The "server" is a computer system that receives data sent by users and elderly people, and stores, analyzes, trains, and generates responses.
[0484] "Voice data" refers to digitized data of the voices spoken by elderly people.
[0485] "Text data" is data of character information converted from voice data by a voice recognition model.
[0486] An "emotion engine" is software or a system for analyzing the emotions of elderly people from voice data.
[0487] A "smart device" is an electronic device worn by the elderly that provides real-time interactive support, such as smart glasses.
[0488] This invention is a system that combines generative AI models, speech recognition technology, and an emotion engine to provide customized interactive support to improve the experience of elderly people in physical stores. The system operates as follows: the user inputs information, and the server generates a generative AI model and a speech recognition model based on this information.
[0489] Hardware and Software Used
[0490] Hardware
[0491] Smart devices (e.g. smart glasses)
[0492] server
[0493] software
[0494] Speech recognition engine (Google Speech-to-Text API)
[0495] Generative AI model (OpenAI® GPT-3)
[0496] Emotion Engine (Affectiva SDK)
[0497] Speech synthesis engine (Google Text-to-Speech API)
[0498] Database (MongoDB)
[0499] What the program does
[0500] 1. User enters information
[0501] Users use a dedicated application to input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. This information is converted into JSON format and sent to the server, which then stores it in a database (MongoDB).
[0502] 2. Model generation
[0503] The server generates and customizes a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the received information. Voice data, including specific dialects and individual speaking patterns, is also used as training data for the speech recognition model.
[0504] 3. Real-time support in physical stores
[0505] When an elderly person wears a smart device and starts a conversation or asks a question, their voice is converted into text data using a voice recognition engine and sent to a server. At the same time, an emotion engine (Affectiva SDK) analyzes emotions from the voice data and sends the emotional information to the server.
[0506] 4. Reply Generation and Voice Response
[0507] The server uses a generative AI model (OpenAI GPT-3) to generate an appropriate response based on the received text data and emotion data. The generated response is then sent back to the smart device from the server and converted into speech using a speech synthesis engine (Google Text-to-Speech API). Appropriate guidance and question responses are provided to the elderly in real time.
[0508] Specific examples
[0509] For example, when an elderly person is shopping in a physical store, they can speak to their smart device and ask, "Where is the sugar?" This speech is converted into text data in real time and analyzed by an emotion engine. The server receives this and generates a response such as, "Sugar is at the far right of the food shelf. Your favorite snacks are also nearby." The response is converted into audio and provided to the elderly.
[0510] Prompt Sentence Examples
[0511] User: How's the weather today?
[0512] Assistant (considering emotions): It's sunny today, but it's expected to get a little windy this afternoon. Please be careful when you go out.
[0513] In this way, the present invention improves the quality of life of elderly people by providing them with real-time guidance and question answers that they need in physical stores.
[0514] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0515] Step 1:
[0516] The user enters information about the elderly person using a dedicated application. The information entered includes the elderly person's name, background, hobbies, preferences, local dialect, etc. This information is converted into JSON format and sent to the server. The server stores the received JSON data in a database (MongoDB).
[0517] Input: Basic information about the elderly person (name, background, hobbies, preferences, dialect)
[0518] Data processing: converting information into JSON format
[0519] Output: Information of elderly people stored in a database
[0520] Step 2:
[0521] The server customizes and generates a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the information stored in the database. Specifically, it builds training data for the model based on the stored information and adds new parameters to the existing model for learning.
[0522] Input: Elderly person information stored in a database
[0523] Data Computing: Customizing and retraining generative AI and speech recognition models
[0524] Output: Customized generative AI models and speech recognition models
[0525] Step 3:
[0526] The user uses the application's recording function to input voice data about the elderly person's past events or favorite topics. The voice data is encoded according to the sample rate and bit rate and sent to the server, which then stores the received voice data in an appropriate format.
[0527] Input: Speech data from elderly people
[0528] Data processing: Encoding of audio data
[0529] Output: Audio data stored on the server
[0530] Step 4:
[0531] The server analyzes the received audio data and retrains the speech recognition model, improving its accuracy based on the analyzed data and generating a model optimized for a particular dialect or individual speaking style.
[0532] Input: Audio data stored on the server
[0533] Data Computing: Analyzing speech data and retraining speech recognition models
[0534] Output: Optimized speech recognition model
[0535] Step 5:
[0536] Elderly people wear smart devices and start asking questions or having conversations. This speech is converted into text data in real time by a speech recognition engine (Google Speech-to-Text API) and sent to a server.
[0537] Input: Real-time voice data of elderly people
[0538] Data processing: Converting voice to text
[0539] Output: Send to server as text data
[0540] Step 6:
[0541] The server analyzes the elderly person's emotions through the received voice data using an emotion engine (Affectiva SDK). The emotion analysis results are also sent to the server.
[0542] Input: Elderly voice data
[0543] Data Computing: Emotion Analysis
[0544] Output: Emotion data of elderly people
[0545] Step 7:
[0546] The server generates an appropriate response based on the received text data and emotion data using a generative AI model (OpenAI GPT-3), and the generated response is then sent back to the smart device.
[0547] Input: Text data and elderly emotion data
[0548] Data Computing: Generating Responses with Generative AI Models
[0549] Output: Response data sent to the smart device
[0550] Step 8:
[0551] The response text received on the smart device is converted into speech using a speech synthesis engine (Google Text-to-Speech API), and a response is provided to the elderly person.
[0552] Input: Reply text sent to smart device
[0553] Data processing: Convert response text into speech
[0554] Output: Voice response provided to the senior
[0555] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0556] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0557] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0558] [Second embodiment]
[0559] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0560] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0561] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0562] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0563] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0564] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0565] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0566] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0567] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0568] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0569] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0570] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0571] The present invention is a system for supporting dementia in the elderly, which combines generative AI and speech recognition technology to provide individually customized interactive support. Specific embodiments of this system are described in detail below.
[0572] System Overview
[0573] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[0574] Program processing
[0575] 1. User enters information
[0576] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0577] The device converts this information into JSON format and sends it to the server.
[0578] 2. The server saves the information and generates the model
[0579] The server stores the received information in a database.
[0580] The server customizes the generative AI model and speech recognition model based on the stored information.
[0581] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model with new audio datasets as needed.
[0582] 3. User adds voice input
[0583] Through the application, users record voice input and provide information about the senior's past events and favorite topics.
[0584] The device transmits the recorded audio data to the server.
[0585] 4. The server analyzes the audio data and updates the model.
[0586] The server analyzes the received voice data and retrains generative AI and speech recognition models optimized for specific dialects and individual speaking styles.
[0587] 5. Seniors start asking questions and initiating conversations
[0588] Seniors use the app to ask questions and initiate conversations.
[0589] The device collects the elderly person's voice in real time, converts it into text through a voice recognition engine, and sends it to a server.
[0590] 6. The server generates an appropriate response
[0591] The server analyzes the received text and uses generative AI to generate an appropriate response.
[0592] The generated response is sent to the terminal in text format.
[0593] 7. The device converts your reply into voice
[0594] The device converts the received response text into speech using a speech synthesis engine.
[0595] The device plays the generated audio to the elderly person.
[0596] Specific examples
[0597] For example, if an elderly person named Mr. Tanaka were to use this system, his son would enter information about Mr. Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, into the application. The son would then record and add information about Mr. Tanaka's past experiences and hobbies. This information and voice data is sent to a server, which then creates a generative AI model and a voice recognition model specifically for Mr. Tanaka.
[0598] When Tanaka asks "What's the weather like today?" through the application, the device converts the question into text and sends it to the server. The server uses generative AI to generate a response such as "It's sunny today, but the wind is expected to pick up a little in the afternoon," and sends it to the device. The device then converts this response into audio and plays it for Tanaka.
[0599] If Tanaka wants to travel to a specific location, he can reserve a transportation service through the application. The device sends the reservation information to the server, which then adjusts the travel schedule. As a result, Tanaka can travel safely and comfortably.
[0600] As described above, this invention aims to support elderly people with dementia by utilizing individually customized generative AI and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[0601] The processing flow will be explained below.
[0602] Step 1:
[0603] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[0604] Step 2:
[0605] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[0606] Step 3:
[0607] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[0608] Step 4:
[0609] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[0610] Step 5:
[0611] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[0612] Step 6:
[0613] The server analyzes the received text and generates an appropriate response using a generation AI. The generated response text is then sent to the device.
[0614] Step 7:
[0615] The device converts the received response text into speech using a speech synthesis engine, and the device plays the generated speech back to the elderly person, providing an answer to their question.
[0616] Step 8:
[0617] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[0618] Step 9:
[0619] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[0620] Step 10:
[0621] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[0622] Example 1
[0623] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0624] To support elderly people with dementia, there is a need for systems that provide interactive support tailored to individual needs. However, current technology does not have an established method for quickly generating and operating AI models and voice recognition models optimized for each elderly person, which means that elderly people are unable to receive appropriate support. In addition, there is a lack of comprehensive systems that also provide support for daily life, such as improving IT literacy and providing mobility support.
[0625] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0626] In this invention, the server includes: a means for converting information entered by a user into a data format and transmitting it to the server; a means for storing the received information in a database; a means for generating individually customized generative AI models and speech recognition models based on the stored information; a means for the user to add voice input and transmit the voice data to the server; a means for the server to analyze the received voice data and retrain the models; a means for converting the voice to text and transmitting it to the server; a means for the server to analyze the text using the generative AI to generate an appropriate response and transmit the response text to the terminal; and a means for the terminal to convert the received response text into voice and play it back. This enables dementia support tailored to each elderly person. It also enables integrated support for elderly people's IT literacy and mobility, thereby improving the overall quality of their daily lives.
[0627] A "server" is a computer system that provides data and functionality to other computers and devices over a network.
[0628] A "terminal" is an electronic device that is directly operated by a user and that communicates with a server.
[0629] A "user" is a person who utilizes the system to enter information or add voice input.
[0630] A "data format" is a format for expressing information in a structured way, such as the JSON format.
[0631] A "database" is a system that stores and manages data, making it easy to search for and analyze information.
[0632] A "generative AI model" is a type of artificial intelligence that automatically responds and processes data, and is used for natural language processing, etc.
[0633] A "speech recognition model" is a model that analyzes voice data and converts it into text.
[0634] "Voice input" refers to voice data provided by a user speaking to the system.
[0635] "Text" means the written representation of voice input or other data.
[0636] A "speech synthesis engine" is software for converting text data into voice data.
[0637] The "interactive tutorial mode" is an interactive support function that helps users learn how to operate the system.
[0638] A "mobility service" is a service that provides assistance to a user moving to a specific location.
[0639] "Schedule adjustment" refers to the act of managing and optimizing the dates and times of transportation services booked by users.
[0640] This invention is a system aimed at supporting dementia in the elderly, combining generative AI and speech recognition technology to provide individually customized interactive support. This system is comprised of three main elements: a server, a terminal, and a user. Below, we will explain how to specifically implement this system.
[0641] server
[0642] The server is a high-performance computer system that manages and operates the database, generative AI model, and speech recognition model. The main software used is TensorFlow and PyTorch, which use Python, and is used to train and retrain the generative AI model and speech recognition model. MySQL, PostgreSQL, and other databases are used.
[0643] Specific examples
[0644] When the user enters the elderly person's information, the server receives the JSON format data sent from the device and stores it in a database. Based on the received data, an individually customized generative AI model and voice recognition model are created.
[0645] Terminal
[0646] A terminal is a device that allows users to input information, record voice data, and communicate with a server. Terminals can be smartphones, tablets, or PCs. Audio recording and playback require a built-in or external microphone and speaker.
[0647] Specific examples
[0648] The user inputs information about Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, through the device, and this information is converted into JSON format and sent to the server. The user then speaks aloud about Tanaka's past events and hobbies, and the audio data is recorded by the device and sent to the server.
[0649] User
[0650] The user is the person who operates the system, inputting information and voice input for the elderly, learning IT operations through an interactive tutorial mode, and booking transportation services on behalf of the elderly.
[0651] Specific examples
[0652] For example, Tanaka's son uses an application to enter Tanaka's information and sends the voice data to the server. The server then customizes the generative AI model and voice recognition model based on the received information, creating a system specifically for Tanaka. When Tanaka asks, "What's the weather like today?", the server uses the generative AI to create an appropriate response and responds via voice via the device.
[0653] Prompt Sentence Examples
[0654] Information input prompt: "Please enter the senior's name, background, hobbies, favorite foods, and local dialect."
[0655] Voice prompt: "Record an older adult talking about past events and hobbies."
[0656] Conversation prompt: "What's the weather like today?"
[0657] Movement service prompt: "Enter details of your next move destination."
[0658] As described above, this system aims to support elderly people with dementia, and utilizes individually customized generative AI models and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[0659] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0660] Step 1:
[0661] User enters information
[0662] Specific operation:
[0663] Users launch a dedicated application and enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0664] input:
[0665] Information on the elderly person's name, background, hobbies, preferences, local dialect, etc.
[0666] Data processing / calculation:
[0667] The terminal converts the input information into JSON format.
[0668] output:
[0669] Elderly information data in JSON format.
[0670] Specific actions added:
[0671] When a user enters information into the input screen and clicks the "Submit" button, the information is converted into JSON format.
[0672] Step 2:
[0673] The server stores the information and generates the model
[0674] Specific operation:
[0675] The device sends the generated JSON data to the server.
[0676] The server stores the received data in a database.
[0677] input:
[0678] Elderly information data in JSON format.
[0679] Data processing / calculation:
[0680] Extract the necessary information from the JSON data and store it in the database.
[0681] Customize generative AI and speech recognition models using libraries such as Python and TensorFlow.
[0682] Adjust the model's learning parameters and possibly retrain it.
[0683] output:
[0684] Customized generative AI and speech recognition models.
[0685] Specific actions added:
[0686] Save the JSON data into a MySQL database and run a script to generate the model.
[0687] Step 3:
[0688] User adds voice input
[0689] Specific operation:
[0690] Users operate an application designed specifically for seniors and input information about past events and hobbies using voice.
[0691] input:
[0692] Voice input of elderly people's past events and hobbies.
[0693] Data processing / calculation:
[0694] The device records the audio and generates an audio file in WAV format or similar.
[0695] Send the audio file to the server.
[0696] output:
[0697] Audio files (WAV format, etc.).
[0698] Specific actions added:
[0699] Press the record button to record the audio, and then press the "send" button after recording is complete to send the audio file to the server.
[0700] Step 4:
[0701] The server analyzes the audio data and updates the model
[0702] Specific operation:
[0703] The server analyzes the received audio file and converts it into text using a speech-to-text engine.
[0704] The parameters of the generative AI model and speech recognition model are readjusted based on the converted text information.
[0705] input:
[0706] Audio file.
[0707] Data processing / calculation:
[0708] A speech-to-text engine is used to convert the audio into text and store it in a database.
[0709] Retrain the model based on the converted text and existing data.
[0710] output:
[0711] Updated generative AI and speech recognition models.
[0712] Specific actions added:
[0713] Input the audio file into the analysis script and run model retraining along with the text conversion results.
[0714] Step 5:
[0715] Seniors start asking questions and conversations
[0716] Specific operation:
[0717] The elderly person launches the application and speaks a prompt to start the conversation (e.g., "What's the weather like today?").
[0718] input:
[0719] Questions and conversations of the elderly.
[0720] Data processing / calculation:
[0721] The device collects audio in real time and records the audio data.
[0722] The speech is converted into text through a speech recognition engine and sent to the server.
[0723] output:
[0724] Conversation in text format.
[0725] Specific actions added:
[0726] The question is recorded, and after the recording is complete, it is automatically converted into text and sent to the server.
[0727] Step 6:
[0728] The server generates an appropriate response
[0729] Specific operation:
[0730] The server analyzes the received text and uses generative AI to generate an appropriate response.
[0731] The generated reply text is sent to the terminal.
[0732] input:
[0733] Textual conversation from the user.
[0734] Data processing / calculation:
[0735] Generative AI (e.g., GPT-3) is used to analyze text and generate responses.
[0736] output:
[0737] The reply text.
[0738] Specific actions added:
[0739] It runs an algorithm to generate a response to the question and sends the generated text to the terminal.
[0740] Step 7:
[0741] The device converts your response into voice
[0742] Specific operation:
[0743] The device converts the received response text into speech using a speech synthesis engine.
[0744] The generated audio is played to the elderly person.
[0745] input:
[0746] The response text from the server.
[0747] Data processing / calculation:
[0748] Convert text to speech using a speech synthesis engine (e.g. Google Text-to-Speech).
[0749] output:
[0750] Audio data.
[0751] Specific actions added:
[0752] The device starts the speech synthesis engine and plays the generated speech on the playback device.
[0753] (Application example 1)
[0754] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0755] Elderly people with dementia may face many challenges in their daily lives. These challenges include being unable to respond appropriately in emergencies, being unable to ensure daily safety, and having difficulty preventing accidents at home. Furthermore, dealing with these challenges places a burden on the individual and their family, which is problematic. The present invention aims to utilize generative AI and speech recognition technology to individually and effectively resolve these challenges.
[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0757] In this invention, the server includes means for saving information entered by a user in a database, means for generating individually customized generative AI models and voice recognition models based on the saved information, means for the user to add voice input and send the voice data to the server, means for the server to analyze the received voice data and retrain the generative AI model and voice recognition model, means for converting the elderly person's voice input into text data and sending it to the server, means for the server to generate an appropriate response using the generative AI and send it to the terminal, means for the terminal to convert the received response text into voice and play it back to the elderly person, means for notifying an emergency contact and providing voice instructions when the elderly person encounters an emergency, means for checking the safety of the elderly person at regular intervals and notifying if an abnormality occurs, and means for monitoring dangers within the home and issuing an alert if an abnormality occurs. This enables rapid response to emergencies, regular safety checks, and safety monitoring within the home.
[0758] "Means for saving information entered by the user in a database" refers to means for saving data obtained from the user, such as personal information, hobbies, and past events.
[0759] "Means for generating individually customized generative AI models and speech recognition models based on stored information" refers to means for creating generative AI and speech recognition models that are optimal for individual users by utilizing information stored in a database.
[0760] The "means for a user to add a voice input and transmit the voice data to a server" refers to a means for a user to record voice data and transmit the data to a server.
[0761] "Means for analyzing voice data received by the server and retraining the generative AI model and the voice recognition model" refers to means for analyzing voice data received by the server and retraining the generative AI and the voice recognition model.
[0762] The "means for converting the voice input of the elderly person into text data and transmitting it to the server" is a means for converting the voice uttered by the elderly person into text in real time and transmitting the text data to the server.
[0763] "Means for the server to use generation AI to generate an appropriate response and send it to the terminal" means a means for the server to use generation AI technology to generate an appropriate response to the received text data and send that response to the terminal.
[0764] The "means for converting the response text received by the terminal into speech and playing it back to the elderly" refers to a means for converting the response text received by the terminal from the server into speech using speech synthesis technology and playing it back to the elderly.
[0765] "Means for notifying emergency contacts and providing voice instructions when an elderly person encounters an emergency" refers to a means for automatically notifying pre-set emergency contacts and simultaneously providing appropriate voice guidance when an elderly person faces an emergency.
[0766] "Means for checking the safety of elderly people at regular intervals and notifying in the event of an abnormality" refers to a means for periodically checking the condition of elderly people and notifying emergency contacts if an abnormality is detected.
[0767] "Means for monitoring dangers within the home and issuing warnings in the event of an abnormality" refers to a means for monitoring smoke, gas leaks, etc. using sensors installed within the home and issuing warnings in the event of an abnormality.
[0768] The present invention is a system for supporting elderly people with dementia, which uses individually customized generative AI models and speech recognition models to provide support for elderly people in their daily lives. Specific embodiments of the system are described below.
[0769] System Configuration and Operation
[0770] Hardware and Software
[0771] Devices: Smartphones, smart glasses, head-mounted displays, etc.
[0772] Server: We use cloud-based servers to host the database, generative AI models, and speech recognition models.
[0773] Sensors: Smoke detectors and gas sensors are installed in the home, and data is sent to the device when an abnormality occurs.
[0774] Data processing and calculation
[0775] 1. Entering and saving information
[0776] Users enter personal information, hobbies, past events, etc. about the elderly person through a dedicated application, and this information is converted into JSON format and sent to the server, which then stores this data in a database.
[0777] 2. Generate and retrain the model
[0778] The server creates a generative AI model and a voice recognition model based on the stored information, and customizes them specifically for seniors. When additional voice input is received, the server analyzes the voice data and retrains the model.
[0779] 3. Speech Processing and Text Conversion
[0780] When the elderly person speaks, the device converts the speech into text in real time and sends it to the server, which then generates a response based on the received text.
[0781] 4. Response Generation and Audio Playback
[0782] The server uses generative AI to generate appropriate responses and sends them in text format to the device, which then converts the text into speech and plays it back to the elderly.
[0783] 5. Emergency response and safety monitoring
[0784] In the event of an emergency, the device will automatically notify emergency contacts and provide appropriate voice guidance. The device also monitors data from sensors and notifies users if an abnormality occurs.
[0785] Specific examples
[0786] For example, the present invention is useful in the following specific situations.
[0787] Example 1: Emergency response
[0788] If an elderly person falls and calls for help on their smartphone, the application will recognize the voice and quickly notify emergency contacts. At the same time, it will provide appropriate voice guidance to help the elderly person take immediate action.
[0789] Example prompt sentence:
[0790] "I have had a fall. Please notify your emergency contacts."
[0791] Example 2: Safety reminder
[0792] The application periodically asks, "How are you?" and waits for a response from the elderly person. If an abnormality is detected, it immediately notifies emergency contacts.
[0793] Example prompt sentence:
[0794] "How are you? If you don't respond, I'll notify my emergency contacts."
[0795] As described above, the present invention is a system that utilizes generative AI models and voice recognition technology to comprehensively support the daily lives of the elderly and provide a safe and secure environment.
[0796] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0797] Step 1:
[0798] The user enters information
[0799] Users enter personal information (such as name, address, and emergency contact information), lifestyle habits, hobbies, and past events through a dedicated application. The application converts the entered information into JSON format and sends it to the server via the network.
[0800] Input: Personal information and past event data entered by the user
[0801] Output: Data converted to JSON format
[0802] Step 2:
[0803] The server stores the information
[0804] The server parses the JSON format data received over the network and stores it in a database for later processing.
[0805] Input: JSON format user information data
[0806] Output: User information stored in the database
[0807] Step 3:
[0808] Generate a model based on the retained information
[0809] The server uses the information stored in the database to create individually customized generative AI and speech recognition models, which involves applying machine learning algorithms to the data.
[0810] Input: User information stored in the database
[0811] Output: Generated generative AI model and speech recognition model
[0812] Step 4:
[0813] User adds voice input
[0814] The user inputs voice data through the application, and the voice data is collected in real time by the device and sent to a server for analysis.
[0815] Input: User's voice data
[0816] Output: Audio data sent to the server
[0817] Step 5:
[0818] The server analyzes the audio data and retrains the model.
[0819] The server analyzes the received voice data and retrains the generative AI model and speech recognition model to improve their performance, using a speech recognition algorithm.
[0820] Input: Audio data
[0821] Output: Updated AI and speech recognition models after retraining
[0822] Step 6:
[0823] Converting elderly people's voice input into text
[0824] When an elderly person uses the application to input voice, the device converts the voice into text data in real time and sends it to the server, using a voice recognition engine.
[0825] Input: Elderly voice data
[0826] Output: Data converted to text
[0827] Step 7:
[0828] The server generates an appropriate response
[0829] The server uses AI to generate an appropriate response based on the received text data. This response is generated based on the user's individually customized information and sent to the device.
[0830] Input: Text data
[0831] Output: The generated response text
[0832] Step 8:
[0833] The device converts your response into speech
[0834] The device converts the response text received from the server into voice using a speech synthesis engine and plays it back to the elderly, who can then confirm the response by voice.
[0835] Input: Reply text
[0836] Output: The speech-transcribed reply
[0837] Step 9:
[0838] Emergency response
[0839] If an elderly person encounters an emergency, the device will instantly accept voice input, notify emergency contacts, and provide appropriate voice instructions from the server, using voice recognition and generation AI technology.
[0840] Input: Voice input in case of emergency
[0841] Output: Provided voice instructions and emergency notifications
[0842] Step 10:
[0843] Safety checks and home safety monitoring
[0844] The device periodically checks the elderly person's condition, and if an abnormality is detected, it immediately notifies the server and contacts the emergency contact. It also uses sensors in the home to monitor for smoke and gas leaks, and issues an alert if an abnormality is detected.
[0845] Input: Elderly status and home sensor data
[0846] Output: Safety check results and warning notifications
[0847] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0848] The present invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI, speech recognition technology, and an emotion engine. Specific embodiments of this system are described in detail below.
[0849] System Overview
[0850] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[0851] Furthermore, by combining it with an emotion engine, it becomes possible to recognize the user's emotions and generate responses based on those emotions. This emotion engine analyzes emotions from the user's voice and sends the results to a server, allowing the generative AI model to adjust the response content.
[0852] Program processing
[0853] 1. User enters information
[0854] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0855] The device converts this information into JSON format and sends it to the server by pressing the send button.
[0856] 2. The server saves the information and generates the model
[0857] The server stores the received information in a database.
[0858] The server customizes the generative AI model and speech recognition model based on the stored information.
[0859] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using new audio datasets as needed.
[0860] 3. User adds voice input
[0861] Users access the "Add Learning Data" section of the application and use the recording function to dictate events from the elderly person's past or favorite topics.
[0862] The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[0863] The device sends the encoded audio data to the server.
[0864] 4. The server analyzes the audio data and updates the model.
[0865] The server stores the received audio data in an appropriate format and analyzes it.
[0866] The server then retrains the generative AI model and speech recognition model based on this voice data, updating the models to be optimized for specific dialects and individual speaking styles.
[0867] 5. Seniors start asking questions and initiating conversations
[0868] Using the application, the elderly person presses the talk button to ask a question or start a conversation.
[0869] The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine, which then sends the text data to a server.
[0870] 6. Emotion engine recognizes user emotions
[0871] As the voice data is sent to the server, the emotion engine analyzes the emotion from the user's voice.
[0872] The emotion analysis results are sent to a server, which records the user's emotional state.
[0873] 7. The server generates an appropriate response based on the emotion.
[0874] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[0875] The generated reply text is sent to the terminal.
[0876] 8. The device converts your reply into voice
[0877] The response text received by the device is converted into speech using a speech synthesis engine.
[0878] The device plays the generated audio to the elderly person, providing answers to their questions and appropriate feedback.
[0879] Specific examples
[0880] For example, if Mr. Tanaka were to use the system, his son would input information about him, such as his name, background, hobbies, favorite foods, and local dialect, into the application, and provide additional voice data, which would then create a generative AI model and a voice recognition model specifically for him.
[0881] When Tanaka asks the application, "What's the weather like today?", the device converts this speech into text, while the emotion engine simultaneously analyzes Tanaka's emotions. The text data and emotion data are sent to the server, which uses this information to generate a response that takes emotion into account, such as, "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response is converted into audio and played to Tanaka in real time.
[0882] Also, if Tanaka is in a certain emotional state, for example, feeling anxious, the system will provide a reassuring response tailored to that emotion, such as, "The weather is nice today, and your favorite flowers might be in bloom."
[0883] In this way, the present invention aims to support dementia in the elderly, and by combining generative AI, voice recognition technology, and an emotion engine, it provides individually customized interactive support, realizing comprehensive support for the elderly and their families.
[0884] The processing flow will be explained below.
[0885] Step 1:
[0886] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[0887] Step 2:
[0888] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[0889] Step 3:
[0890] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[0891] Step 4:
[0892] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[0893] Step 5:
[0894] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[0895] Step 6:
[0896] While the voice data is being sent to the server, the device uses an emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions and sends the results to the server along with the text data.
[0897] Step 7:
[0898] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[0899] Step 8:
[0900] The server generates a response text and sends it to the device. The device converts the received response text into speech using a speech synthesis engine. The device plays the generated speech to the elderly person, providing answers to their questions and appropriate feedback.
[0901] Step 9:
[0902] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[0903] Step 10:
[0904] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[0905] Step 11:
[0906] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[0907] Example 2
[0908] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0909] There is a lack of interactive support optimized for each patient in dementia support for the elderly. In particular, it is difficult to provide support that meets the individual needs of the elderly by using individually customized generative AI models and speech recognition models. Another issue is the lack of technology that can recognize emotions from speech and generate appropriate responses based on them. Furthermore, there is a need for an interactive tutorial mode to improve the IT literacy of the elderly and for improved mobility service reservation functions.
[0910] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for storing information input by a user in a database, a means for generating individually customized generative AI models and voice recognition models based on the stored information, a means for converting the user's voice input into text data and transmitting it to the server, and a means for adjusting the content of responses generated based on the emotion analysis results. This makes it possible to provide optimal support according to the individual needs of elderly people. In addition, by analyzing the user's emotions using an emotion recognition engine and generating responses that take these emotions into consideration, more personalized interactive dialogue is realized.
[0911] "User" refers to an individual who uses the system to input information or voice input.
[0912] "Information" refers to data that users enter into the system, including the elderly person's name, background, hobbies, preferences, local dialect, etc.
[0913] A "database" is a storage device that stores input information and allows it to be retrieved when needed.
[0914] A "generative AI model" is an algorithmic model of artificial intelligence that is individually customized based on user input information.
[0915] A "speech recognition model" is an algorithmic model used to convert a user's speech data into text data.
[0916] "Audio data" refers to digital audio files generated by a user providing voice input.
[0917] A "server" is a computer system responsible for storing information and generating and training generative AI models and speech recognition models.
[0918] "Analysis" is the process in which the server analyzes information based on the voice data received and performs the necessary processing.
[0919] "Text data" is character string information converted from voice data by a voice recognition model.
[0920] An "emotion recognition engine" is software that analyzes emotions from the user's voice and sends the results to a server.
[0921] A "reply" is a response that the server generates in response to a user's question or request using a generative AI model.
[0922] The "interactive tutorial mode" is an interactive learning mode provided for seniors to learn how to use the system.
[0923] "Transportation service" is a function that allows elderly people to reserve available transportation and adjust their travel schedules.
[0924] "Model training" is the process by which a generative AI model or speech recognition model learns from stored data and improves its performance.
[0925] "Playback" is the process by which the device plays the generated audio to the elderly person.
[0926] This invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI models, speech recognition technology, and an emotion engine. Specific embodiments of the present invention are described below.
[0927] System Overview
[0928] Users enter information about the elderly person through a dedicated application, and the device converts this information into JSON format and sends it to a server. The server stores the received information in a database and individually customizes the generative AI model and speech recognition model. The generative AI model used is GPT-4, and the speech recognition model is the Google Speech-to-Text API. The IBM Watson Tone Analyzer is used for emotion recognition.
[0929] Users can improve the accuracy of the model by adding information about the senior's past experiences, hobbies, preferences, etc. This allows the generative AI model and speech recognition model to be individually tailored to the senior. The system also features an interactive tutorial mode and a mobility service booking function.
[0930] Furthermore, by combining it with an emotion recognition engine, it can recognize the user's emotions in real time and generate responses based on those emotions, enabling interactive dialogue that takes the user's emotional state into account.
[0931] Hardware and Software Details
[0932] server
[0933] Database: A relational database such as "MySQL" or "PostgreSQL" to store the information.
[0934] Generative AI models: Individually customized generative AI models (e.g., "GPT-4").
[0935] Speech recognition model: A speech recognition model specifically optimized for older adults (e.g., Google Speech-to-Text API).
[0936] Terminal
[0937] Conversion and transmission of input information: A function that converts information entered by the user into JSON format and sends it to the server. This is performed by software processing inside the device.
[0938] Encoding and sending audio data: The recorded audio data is encoded with appropriate sound quality parameters (sample rate, bit rate) and sent to the server.
[0939] Emotion Recognition Engine
[0940] IBM Watson Tone Analyzer: Analyzes emotions from the user's voice and sends the results to the server.
[0941] Response generation and speech synthesis
[0942] Generative AI model: Generates appropriate responses based on the user's text data and sentiment analysis results.
[0943] Speech synthesis engine: A speech synthesis engine such as "Amazon Polly" is used to convert response text into speech.
[0944] Specific examples
[0945] Specific examples are shown below.
[0946] For example, a user inputs information about an elderly person (Mr. Tanaka), such as his name, background, hobbies, favorite foods, and local dialect, into an application, which then sends this information to a server.
[0947] The server stores this information in a database and customizes the generative AI model and speech recognition model based on the input data. The user can then add voice input about Tanaka's past events and favorite topics, and send this voice data to the server.
[0948] The server analyzes the received voice data and retrains the generative AI model and speech recognition model. When the elderly person starts a question or conversation using the application, the device converts the voice into text and sends it to the server.
[0949] The emotion recognition engine analyzes emotions from the voice and sends the results to a server. The server uses a generative AI model based on the emotion analysis results and text data to generate an appropriate response and sends it to the device. The device then converts the response text into speech and plays it back to the elderly.
[0950] For example, if Tanaka asks, "How's the weather today?", the emotion analysis engine will determine that Tanaka is feeling a little anxious. The server will then generate a reassuring response: "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response will be converted into audio and played to Tanaka in real time.
[0951] In this way, the present invention aims to support dementia in the elderly, and by providing individually customized interactive support, it realizes comprehensive support for the elderly and their families.
[0952] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0953] Step 1: User Enters Information
[0954] What happens: The user opens a dedicated application and enters information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[0955] Input: Information about the elderly person (such as name, background, hobbies, preferences, local dialect, etc.).
[0956] Data processing: The terminal converts this information into JSON format.
[0957] Output: Data converted to JSON format.
[0958] Specific operation: Enter data into the input form displayed on the terminal and press the send button. The data is converted into JSON format and the message "Submission successful" is displayed.
[0959] Step 2: The server saves the information and generates the model
[0960] Processing details: The server stores the received information in a database and customizes the generative AI model and speech recognition model individually.
[0961] Input: JSON formatted data.
[0962] Data processing: The received information is stored in a database and used to customize generative AI models (e.g., GPT-4) and speech recognition models (e.g., Google Speech-to-Text API).
[0963] Output: Customized generative AI model and speech recognition model.
[0964] What happens: The server parses the JSON data and stores the necessary information in a database. It then adjusts the parameters of the AI model based on the stored information and retrains the speech recognition model using the new audio dataset.
[0965] Step 3: User adds voice input
[0966] What happens: The user visits the "Add Learning Data" section of the app and uses the recording feature to dictate events from the senior's past or favorite topics.
[0967] Input: Speech data of elderly people.
[0968] Data processing: The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[0969] Output: The encoded audio data.
[0970] Specific operation: The user presses the record button and inputs voice information about the elderly person's past events or favorite topics. Once the recording is complete, the device encodes the voice data and sends it to the server.
[0971] Step 4: The server analyzes the audio data and updates the model
[0972] Processing details: The server saves the received audio data in an appropriate format and analyzes it.
[0973] Input: Encoded audio data.
[0974] Data processing: Analyze the voice data and retrain the generative AI model and speech recognition model.
[0975] Output: Updated generative AI model and speech recognition model.
[0976] What it does: The server converts the audio data into a different format, generates training data optimized for a specific dialect or speaking style, and updates the model based on the generated data.
[0977] Step 5: The senior initiates the question or conversation
[0978] What happens: Seniors use the application and press the talk button to ask questions or start a conversation.
[0979] Input: Speech data of elderly people.
[0980] Data processing: The device collects the elderly person's voice in real time and converts it into text using a voice recognition engine.
[0981] Output: Parsed text data.
[0982] Specific operation: When an elderly person presses the button to speak, the device collects voice data and converts it into text in real time. This text data is then sent to the server.
[0983] Step 6: The emotion engine recognizes the user's emotion
[0984] Processing details: As the voice data is sent to the server, the emotion engine analyzes the emotions from the user's voice.
[0985] Input: Audio data.
[0986] Data processing: The emotion analysis engine recognizes emotions from voice and converts the results into data.
[0987] Output: Sentiment analysis result data.
[0988] Specific operation: The voice data sent to the server is passed to the emotion engine, where emotion analysis is performed. The analysis results are stored in the database as the emotional state.
[0989] Step 7: The server generates an appropriate response based on the sentiment
[0990] Processing details: The server uses generative AI to generate an appropriate response based on the text received and the results of sentiment analysis.
[0991] Input: Text data and sentiment analysis result data.
[0992] Data processing: Generative AI models analyze text and emotional state to generate optimal responses.
[0993] Output: The generated response text.
[0994] Specific operation: The server uses the text data and the sentiment analysis results to query the generative AI model and generate an appropriate response, which is then queued for transmission to the device.
[0995] Step 8: Your device converts your response into audio
[0996] Processing details: The response text received by the device is converted into speech using a speech synthesis engine.
[0997] Input: Response text.
[0998] Data processing: Use a speech synthesis engine (e.g., Amazon Polly) to convert text data into speech.
[0999] Output: Synthesized speech data.
[1000] Specific operation: The device receives the response text and converts it into natural speech using a speech synthesis engine. The device then plays this speech back to the elderly person and provides appropriate feedback.
[1001] (Application example 2)
[1002] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1003] Current assistance systems for the elderly struggle to provide individually customized interactive support, and are unable to fully improve the convenience and sense of security of the elderly in their daily lives. Furthermore, seniors lack the means to receive appropriate real-time support when searching for products or answering questions in physical stores. To solve these problems and significantly improve the quality of life for the elderly, a new system is needed that combines generative AI models, speech recognition technology, and emotion engines to provide customized interactive assistance.
[1004] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for storing information input by a user in a database; means for generating individually customized generative AI models and speech recognition models based on the stored information; means for the user to add voice input and send the voice data to the server; means for the server to analyze the received voice data and retrain the generative AI model and speech recognition model; means for converting the elderly's voice input into text data and sending it to the server; means for the server to generate an appropriate response using the generative AI and send it to the terminal; means for the terminal to convert the received response text into speech and play it back to the elderly; means for analyzing emotions from the elderly's speech using an emotion engine and sending the result to the server; and means for providing real-time interactive support via a smart device worn by the elderly. This makes it possible to provide appropriate real-time support for elderly people when searching for products in physical stores, such as guidance and question-answering.
[1005] The "user" is a person who uses the system and is responsible for inputting information about the elderly person and adding voice input.
[1006] "Information" is a general term for data necessary for individual customization, such as the elderly person's name, background, hobbies, preferences, and local dialect.
[1007] "Database" refers to a storage device and system for storing user-entered information and elderly person's voice data.
[1008] A "generative AI model" is an artificial intelligence model that is generated based on information about the user or elderly person, and is used to provide appropriate responses and support.
[1009] The "voice recognition model" is an artificial intelligence model that analyzes the voices of elderly people and converts the content into text data.
[1010] The "server" is a computer system that receives data sent by users and elderly people, and stores, analyzes, trains, and generates responses.
[1011] "Voice data" refers to digitized data of the voices spoken by elderly people.
[1012] "Text data" is data of character information converted from voice data by a voice recognition model.
[1013] An "emotion engine" is software or a system for analyzing the emotions of elderly people from voice data.
[1014] A "smart device" is an electronic device worn by the elderly that provides real-time interactive support, such as smart glasses.
[1015] This invention is a system that combines generative AI models, speech recognition technology, and an emotion engine to provide customized interactive support to improve the experience of elderly people in physical stores. The system operates as follows: the user inputs information, and the server generates a generative AI model and a speech recognition model based on this information.
[1016] Hardware and Software Used
[1017] Hardware
[1018] Smart devices (e.g. smart glasses)
[1019] server
[1020] software
[1021] Speech recognition engine (Google Speech-to-Text API)
[1022] Generative AI model (OpenAI GPT-3)
[1023] Emotion Engine (Affectiva SDK)
[1024] Speech synthesis engine (Google Text-to-Speech API)
[1025] Database (MongoDB)
[1026] What the program does
[1027] 1. User enters information
[1028] Users use a dedicated application to input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. This information is converted into JSON format and sent to the server, which then stores it in a database (MongoDB).
[1029] 2. Model generation
[1030] The server generates and customizes a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the received information. Voice data, including specific dialects and individual speaking patterns, is also used as training data for the speech recognition model.
[1031] 3. Real-time support in physical stores
[1032] When an elderly person wears a smart device and starts a conversation or asks a question, their voice is converted into text data using a voice recognition engine and sent to a server. At the same time, an emotion engine (Affectiva SDK) analyzes emotions from the voice data and sends the emotional information to the server.
[1033] 4. Reply Generation and Voice Response
[1034] The server uses a generative AI model (OpenAI GPT-3) to generate an appropriate response based on the received text data and emotion data. The generated response is then sent back to the smart device from the server and converted into speech using a speech synthesis engine (Google Text-to-Speech API). Appropriate guidance and question responses are provided to the elderly in real time.
[1035] Specific examples
[1036] For example, when an elderly person is shopping in a physical store, they can speak to their smart device and ask, "Where is the sugar?" This speech is converted into text data in real time and analyzed by an emotion engine. The server receives this and generates a response such as, "Sugar is at the far right of the food shelf. Your favorite snacks are also nearby." The response is converted into audio and provided to the elderly.
[1037] Prompt Sentence Examples
[1038] User: How's the weather today?
[1039] Assistant (considering emotions): It's sunny today, but it's expected to get a little windy this afternoon. Please be careful when you go out.
[1040] In this way, the present invention improves the quality of life of elderly people by providing them with real-time guidance and question answers that they need in physical stores.
[1041] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1042] Step 1:
[1043] The user enters information about the elderly person using a dedicated application. The information entered includes the elderly person's name, background, hobbies, preferences, local dialect, etc. This information is converted into JSON format and sent to the server. The server stores the received JSON data in a database (MongoDB).
[1044] Input: Basic information about the elderly person (name, background, hobbies, preferences, dialect)
[1045] Data processing: converting information into JSON format
[1046] Output: Information of elderly people stored in a database
[1047] Step 2:
[1048] The server customizes and generates a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the information stored in the database. Specifically, it builds training data for the model based on the stored information and adds new parameters to the existing model for learning.
[1049] Input: Elderly person information stored in a database
[1050] Data Computing: Customizing and retraining generative AI and speech recognition models
[1051] Output: Customized generative AI models and speech recognition models
[1052] Step 3:
[1053] The user uses the application's recording function to input voice data about the elderly person's past events or favorite topics. The voice data is encoded according to the sample rate and bit rate and sent to the server, which then stores the received voice data in an appropriate format.
[1054] Input: Speech data from elderly people
[1055] Data processing: Encoding of audio data
[1056] Output: Audio data stored on the server
[1057] Step 4:
[1058] The server analyzes the received audio data and retrains the speech recognition model, improving its accuracy based on the analyzed data and generating a model optimized for a particular dialect or individual speaking style.
[1059] Input: Audio data stored on the server
[1060] Data Computing: Analyzing speech data and retraining speech recognition models
[1061] Output: Optimized speech recognition model
[1062] Step 5:
[1063] Elderly people wear smart devices and start asking questions or having conversations. This speech is converted into text data in real time by a speech recognition engine (Google Speech-to-Text API) and sent to a server.
[1064] Input: Real-time voice data of elderly people
[1065] Data processing: Converting voice to text
[1066] Output: Send to server as text data
[1067] Step 6:
[1068] The server analyzes the elderly person's emotions through the received voice data using an emotion engine (Affectiva SDK). The emotion analysis results are also sent to the server.
[1069] Input: Elderly voice data
[1070] Data Computing: Emotion Analysis
[1071] Output: Emotion data of elderly people
[1072] Step 7:
[1073] The server generates an appropriate response based on the received text data and emotion data using a generative AI model (OpenAI GPT-3), and the generated response is then sent back to the smart device.
[1074] Input: Text data and elderly emotion data
[1075] Data Computing: Generating Responses with Generative AI Models
[1076] Output: Response data sent to the smart device
[1077] Step 8:
[1078] The response text received on the smart device is converted into speech using a speech synthesis engine (Google Text-to-Speech API), and a response is provided to the elderly person.
[1079] Input: Reply text sent to smart device
[1080] Data processing: Convert response text into speech
[1081] Output: Voice response provided to the senior
[1082] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1083] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1084] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1085] [Third embodiment]
[1086] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1087] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1088] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1089] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1090] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1091] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1092] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1093] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1094] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1095] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1096] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1097] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1098] The present invention is a system for supporting dementia in the elderly, which combines generative AI and speech recognition technology to provide individually customized interactive support. Specific embodiments of this system are described in detail below.
[1099] System Overview
[1100] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[1101] Program processing
[1102] 1. User enters information
[1103] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1104] The device converts this information into JSON format and sends it to the server.
[1105] 2. The server saves the information and generates the model
[1106] The server stores the received information in a database.
[1107] The server customizes the generative AI model and speech recognition model based on the stored information.
[1108] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model with new audio datasets as needed.
[1109] 3. User adds voice input
[1110] Through the application, users record voice input and provide information about the senior's past events and favorite topics.
[1111] The device transmits the recorded audio data to the server.
[1112] 4. The server analyzes the audio data and updates the model.
[1113] The server analyzes the received voice data and retrains generative AI and speech recognition models optimized for specific dialects and individual speaking styles.
[1114] 5. Seniors start asking questions and initiating conversations
[1115] Seniors use the app to ask questions and initiate conversations.
[1116] The device collects the elderly person's voice in real time, converts it into text through a voice recognition engine, and sends it to a server.
[1117] 6. The server generates an appropriate response
[1118] The server analyzes the received text and uses generative AI to generate an appropriate response.
[1119] The generated response is sent to the terminal in text format.
[1120] 7. The device converts your reply into voice
[1121] The device converts the received response text into speech using a speech synthesis engine.
[1122] The device plays the generated audio to the elderly person.
[1123] Specific examples
[1124] For example, if an elderly person named Mr. Tanaka were to use this system, his son would enter information about Mr. Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, into the application. The son would then record and add information about Mr. Tanaka's past experiences and hobbies. This information and voice data is sent to a server, which then creates a generative AI model and a voice recognition model specifically for Mr. Tanaka.
[1125] When Tanaka asks "What's the weather like today?" through the application, the device converts the question into text and sends it to the server. The server uses generative AI to generate a response such as "It's sunny today, but the wind is expected to pick up a little in the afternoon," and sends it to the device. The device then converts this response into audio and plays it for Tanaka.
[1126] If Tanaka wants to travel to a specific location, he can reserve a transportation service through the application. The device sends the reservation information to the server, which then adjusts the travel schedule. As a result, Tanaka can travel safely and comfortably.
[1127] As described above, this invention aims to support elderly people with dementia by utilizing individually customized generative AI and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[1128] The processing flow will be explained below.
[1129] Step 1:
[1130] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[1131] Step 2:
[1132] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[1133] Step 3:
[1134] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[1135] Step 4:
[1136] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[1137] Step 5:
[1138] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[1139] Step 6:
[1140] The server analyzes the received text and generates an appropriate response using a generation AI. The generated response text is then sent to the device.
[1141] Step 7:
[1142] The device converts the received response text into speech using a speech synthesis engine, and the device plays the generated speech back to the elderly person, providing an answer to their question.
[1143] Step 8:
[1144] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[1145] Step 9:
[1146] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[1147] Step 10:
[1148] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[1149] Example 1
[1150] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1151] To support elderly people with dementia, there is a need for systems that provide interactive support tailored to individual needs. However, current technology does not have an established method for quickly generating and operating AI models and voice recognition models optimized for each elderly person, which means that elderly people are unable to receive appropriate support. In addition, there is a lack of comprehensive systems that also provide support for daily life, such as improving IT literacy and providing mobility support.
[1152] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1153] In this invention, the server includes: a means for converting information entered by a user into a data format and transmitting it to the server; a means for storing the received information in a database; a means for generating individually customized generative AI models and speech recognition models based on the stored information; a means for the user to add voice input and transmit the voice data to the server; a means for the server to analyze the received voice data and retrain the models; a means for converting the voice to text and transmitting it to the server; a means for the server to analyze the text using the generative AI to generate an appropriate response and transmit the response text to the terminal; and a means for the terminal to convert the received response text into voice and play it back. This enables dementia support tailored to each elderly person. It also enables integrated support for elderly people's IT literacy and mobility, thereby improving the overall quality of their daily lives.
[1154] A "server" is a computer system that provides data and functionality to other computers and devices over a network.
[1155] A "terminal" is an electronic device that is directly operated by a user and that communicates with a server.
[1156] A "user" is a person who utilizes the system to enter information or add voice input.
[1157] A "data format" is a format for expressing information in a structured way, such as the JSON format.
[1158] A "database" is a system that stores and manages data, making it easy to search for and analyze information.
[1159] A "generative AI model" is a type of artificial intelligence that automatically responds and processes data, and is used for natural language processing, etc.
[1160] A "speech recognition model" is a model that analyzes voice data and converts it into text.
[1161] "Voice input" refers to voice data provided by a user speaking to the system.
[1162] "Text" means the written representation of voice input or other data.
[1163] A "speech synthesis engine" is software for converting text data into voice data.
[1164] The "interactive tutorial mode" is an interactive support function that helps users learn how to operate the system.
[1165] A "mobility service" is a service that provides assistance to a user moving to a specific location.
[1166] "Schedule adjustment" refers to the act of managing and optimizing the dates and times of transportation services booked by users.
[1167] This invention is a system aimed at supporting dementia in the elderly, combining generative AI and speech recognition technology to provide individually customized interactive support. This system is comprised of three main elements: a server, a terminal, and a user. Below, we will explain how to specifically implement this system.
[1168] server
[1169] The server is a high-performance computer system that manages and operates the database, generative AI model, and speech recognition model. The main software used is TensorFlow and PyTorch, which use Python, and is used to train and retrain the generative AI model and speech recognition model. MySQL, PostgreSQL, and other databases are used.
[1170] Specific examples
[1171] When the user enters the elderly person's information, the server receives the JSON format data sent from the device and stores it in a database. Based on the received data, an individually customized generative AI model and voice recognition model are created.
[1172] Terminal
[1173] A terminal is a device that allows users to input information, record voice data, and communicate with a server. Terminals can be smartphones, tablets, or PCs. Audio recording and playback require a built-in or external microphone and speaker.
[1174] Specific examples
[1175] The user inputs information about Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, through the device, and this information is converted into JSON format and sent to the server. The user then speaks aloud about Tanaka's past events and hobbies, and the audio data is recorded by the device and sent to the server.
[1176] User
[1177] The user is the person who operates the system, inputting information and voice input for the elderly, learning IT operations through an interactive tutorial mode, and booking transportation services on behalf of the elderly.
[1178] Specific examples
[1179] For example, Tanaka's son uses an application to enter Tanaka's information and sends the voice data to the server. The server then customizes the generative AI model and voice recognition model based on the received information, creating a system specifically for Tanaka. When Tanaka asks, "What's the weather like today?", the server uses the generative AI to create an appropriate response and responds via voice via the device.
[1180] Prompt Sentence Examples
[1181] Information input prompt: "Please enter the senior's name, background, hobbies, favorite foods, and local dialect."
[1182] Voice prompt: "Record an older adult talking about past events and hobbies."
[1183] Conversation prompt: "What's the weather like today?"
[1184] Movement service prompt: "Enter details of your next move destination."
[1185] As described above, this system aims to support elderly people with dementia, and utilizes individually customized generative AI models and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[1186] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1187] Step 1:
[1188] User enters information
[1189] Specific operation:
[1190] Users launch a dedicated application and enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1191] input:
[1192] Information on the elderly person's name, background, hobbies, preferences, local dialect, etc.
[1193] Data processing / calculation:
[1194] The terminal converts the input information into JSON format.
[1195] output:
[1196] Elderly information data in JSON format.
[1197] Specific actions added:
[1198] When a user enters information into the input screen and clicks the "Submit" button, the information is converted into JSON format.
[1199] Step 2:
[1200] The server stores the information and generates the model
[1201] Specific operation:
[1202] The device sends the generated JSON data to the server.
[1203] The server stores the received data in a database.
[1204] input:
[1205] Elderly information data in JSON format.
[1206] Data processing / calculation:
[1207] Extract the necessary information from the JSON data and store it in the database.
[1208] Customize generative AI and speech recognition models using libraries such as Python and TensorFlow.
[1209] Adjust the model's learning parameters and possibly retrain it.
[1210] output:
[1211] Customized generative AI and speech recognition models.
[1212] Specific actions added:
[1213] Save the JSON data into a MySQL database and run a script to generate the model.
[1214] Step 3:
[1215] User adds voice input
[1216] Specific operation:
[1217] Users operate an application designed specifically for seniors and input information about past events and hobbies using voice.
[1218] input:
[1219] Voice input of elderly people's past events and hobbies.
[1220] Data processing / calculation:
[1221] The device records the audio and generates an audio file in WAV format or similar.
[1222] Send the audio file to the server.
[1223] output:
[1224] Audio files (WAV format, etc.).
[1225] Specific actions added:
[1226] Press the record button to record the audio, and then press the "send" button after recording is complete to send the audio file to the server.
[1227] Step 4:
[1228] The server analyzes the audio data and updates the model
[1229] Specific operation:
[1230] The server analyzes the received audio file and converts it into text using a speech-to-text engine.
[1231] The parameters of the generative AI model and speech recognition model are readjusted based on the converted text information.
[1232] input:
[1233] Audio file.
[1234] Data processing / calculation:
[1235] A speech-to-text engine is used to convert the audio into text and store it in a database.
[1236] Retrain the model based on the converted text and existing data.
[1237] output:
[1238] Updated generative AI and speech recognition models.
[1239] Specific actions added:
[1240] Input the audio file into the analysis script and run model retraining along with the text conversion results.
[1241] Step 5:
[1242] Seniors start asking questions and conversations
[1243] Specific operation:
[1244] The elderly person launches the application and speaks a prompt to start the conversation (e.g., "What's the weather like today?").
[1245] input:
[1246] Questions and conversations of the elderly.
[1247] Data processing / calculation:
[1248] The device collects audio in real time and records the audio data.
[1249] The speech is converted into text through a speech recognition engine and sent to the server.
[1250] output:
[1251] Conversation in text format.
[1252] Specific actions added:
[1253] The question is recorded, and after the recording is complete, it is automatically converted into text and sent to the server.
[1254] Step 6:
[1255] The server generates an appropriate response
[1256] Specific operation:
[1257] The server analyzes the received text and uses generative AI to generate an appropriate response.
[1258] The generated reply text is sent to the terminal.
[1259] input:
[1260] Textual conversation from the user.
[1261] Data processing / calculation:
[1262] Generative AI (e.g., GPT-3) is used to analyze text and generate responses.
[1263] output:
[1264] The reply text.
[1265] Specific actions added:
[1266] It runs an algorithm to generate a response to the question and sends the generated text to the terminal.
[1267] Step 7:
[1268] The device converts your response into voice
[1269] Specific operation:
[1270] The device converts the received response text into speech using a speech synthesis engine.
[1271] The generated audio is played to the elderly person.
[1272] input:
[1273] The response text from the server.
[1274] Data processing / calculation:
[1275] Convert text to speech using a speech synthesis engine (e.g. Google Text-to-Speech).
[1276] output:
[1277] Audio data.
[1278] Specific actions added:
[1279] The device starts the speech synthesis engine and plays the generated speech on the playback device.
[1280] (Application example 1)
[1281] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1282] Elderly people with dementia may face many challenges in their daily lives. These challenges include being unable to respond appropriately in emergencies, being unable to ensure daily safety, and having difficulty preventing accidents at home. Furthermore, dealing with these challenges places a burden on the individual and their family, which is problematic. The present invention aims to utilize generative AI and speech recognition technology to individually and effectively resolve these challenges.
[1283] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1284] In this invention, the server includes means for saving information entered by a user in a database, means for generating individually customized generative AI models and voice recognition models based on the saved information, means for the user to add voice input and send the voice data to the server, means for the server to analyze the received voice data and retrain the generative AI model and voice recognition model, means for converting the elderly person's voice input into text data and sending it to the server, means for the server to generate an appropriate response using the generative AI and send it to the terminal, means for the terminal to convert the received response text into voice and play it back to the elderly person, means for notifying an emergency contact and providing voice instructions when the elderly person encounters an emergency, means for checking the safety of the elderly person at regular intervals and notifying if an abnormality occurs, and means for monitoring dangers within the home and issuing an alert if an abnormality occurs. This enables rapid response to emergencies, regular safety checks, and safety monitoring within the home.
[1285] "Means for saving information entered by the user in a database" refers to means for saving data obtained from the user, such as personal information, hobbies, and past events.
[1286] "Means for generating individually customized generative AI models and speech recognition models based on stored information" refers to means for creating generative AI and speech recognition models that are optimal for individual users by utilizing information stored in a database.
[1287] The "means for a user to add a voice input and transmit the voice data to a server" refers to a means for a user to record voice data and transmit the data to a server.
[1288] "Means for analyzing voice data received by the server and retraining the generative AI model and the voice recognition model" refers to means for analyzing voice data received by the server and retraining the generative AI and the voice recognition model.
[1289] The "means for converting the voice input of the elderly person into text data and transmitting it to the server" is a means for converting the voice uttered by the elderly person into text in real time and transmitting the text data to the server.
[1290] "Means for the server to use generation AI to generate an appropriate response and send it to the terminal" means a means for the server to use generation AI technology to generate an appropriate response to the received text data and send that response to the terminal.
[1291] The "means for converting the response text received by the terminal into speech and playing it back to the elderly" refers to a means for converting the response text received by the terminal from the server into speech using speech synthesis technology and playing it back to the elderly.
[1292] "Means for notifying emergency contacts and providing voice instructions when an elderly person encounters an emergency" refers to a means for automatically notifying pre-set emergency contacts and simultaneously providing appropriate voice guidance when an elderly person faces an emergency.
[1293] "Means for checking the safety of elderly people at regular intervals and notifying in the event of an abnormality" refers to a means for periodically checking the condition of elderly people and notifying emergency contacts if an abnormality is detected.
[1294] "Means for monitoring dangers within the home and issuing warnings in the event of an abnormality" refers to a means for monitoring smoke, gas leaks, etc. using sensors installed within the home and issuing warnings in the event of an abnormality.
[1295] The present invention is a system for supporting elderly people with dementia, which uses individually customized generative AI models and speech recognition models to provide support for elderly people in their daily lives. Specific embodiments of the system are described below.
[1296] System Configuration and Operation
[1297] Hardware and Software
[1298] Devices: Smartphones, smart glasses, head-mounted displays, etc.
[1299] Server: We use cloud-based servers to host the database, generative AI models, and speech recognition models.
[1300] Sensors: Smoke detectors and gas sensors are installed in the home, and data is sent to the device when an abnormality occurs.
[1301] Data processing and calculation
[1302] 1. Entering and saving information
[1303] Users enter personal information, hobbies, past events, etc. about the elderly person through a dedicated application, and this information is converted into JSON format and sent to the server, which then stores this data in a database.
[1304] 2. Generate and retrain the model
[1305] The server creates a generative AI model and a voice recognition model based on the stored information, and customizes them specifically for seniors. When additional voice input is received, the server analyzes the voice data and retrains the model.
[1306] 3. Speech Processing and Text Conversion
[1307] When the elderly person speaks, the device converts the speech into text in real time and sends it to the server, which then generates a response based on the received text.
[1308] 4. Response Generation and Audio Playback
[1309] The server uses generative AI to generate appropriate responses and sends them in text format to the device, which then converts the text into speech and plays it back to the elderly.
[1310] 5. Emergency response and safety monitoring
[1311] In the event of an emergency, the device will automatically notify emergency contacts and provide appropriate voice guidance. The device also monitors data from sensors and notifies users if an abnormality occurs.
[1312] Specific examples
[1313] For example, the present invention is useful in the following specific situations.
[1314] Example 1: Emergency response
[1315] If an elderly person falls and calls for help on their smartphone, the application will recognize the voice and quickly notify emergency contacts. At the same time, it will provide appropriate voice guidance to help the elderly person take immediate action.
[1316] Example prompt sentence:
[1317] "I have had a fall. Please notify your emergency contacts."
[1318] Example 2: Safety reminder
[1319] The application periodically asks, "How are you?" and waits for a response from the elderly person. If an abnormality is detected, it immediately notifies emergency contacts.
[1320] Example prompt sentence:
[1321] "How are you? If you don't respond, I'll notify my emergency contacts."
[1322] As described above, the present invention is a system that utilizes generative AI models and voice recognition technology to comprehensively support the daily lives of the elderly and provide a safe and secure environment.
[1323] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1324] Step 1:
[1325] The user enters information
[1326] Users enter personal information (such as name, address, and emergency contact information), lifestyle habits, hobbies, and past events through a dedicated application. The application converts the entered information into JSON format and sends it to the server via the network.
[1327] Input: Personal information and past event data entered by the user
[1328] Output: Data converted to JSON format
[1329] Step 2:
[1330] The server stores the information
[1331] The server parses the JSON format data received over the network and stores it in a database for later processing.
[1332] Input: JSON format user information data
[1333] Output: User information stored in the database
[1334] Step 3:
[1335] Generate a model based on the retained information
[1336] The server uses the information stored in the database to create individually customized generative AI and speech recognition models, which involves applying machine learning algorithms to the data.
[1337] Input: User information stored in the database
[1338] Output: Generated generative AI model and speech recognition model
[1339] Step 4:
[1340] User adds voice input
[1341] The user inputs voice data through the application, and the voice data is collected in real time by the device and sent to a server for analysis.
[1342] Input: User's voice data
[1343] Output: Audio data sent to the server
[1344] Step 5:
[1345] The server analyzes the audio data and retrains the model.
[1346] The server analyzes the received voice data and retrains the generative AI model and speech recognition model to improve their performance, using a speech recognition algorithm.
[1347] Input: Audio data
[1348] Output: Updated AI and speech recognition models after retraining
[1349] Step 6:
[1350] Converting elderly people's voice input into text
[1351] When an elderly person uses the application to input voice, the device converts the voice into text data in real time and sends it to the server, using a voice recognition engine.
[1352] Input: Elderly voice data
[1353] Output: Data converted to text
[1354] Step 7:
[1355] The server generates an appropriate response
[1356] The server uses AI to generate an appropriate response based on the received text data. This response is generated based on the user's individually customized information and sent to the device.
[1357] Input: Text data
[1358] Output: The generated response text
[1359] Step 8:
[1360] The device converts your response into speech
[1361] The device converts the response text received from the server into voice using a speech synthesis engine and plays it back to the elderly, who can then confirm the response by voice.
[1362] Input: Reply text
[1363] Output: The speech-transcribed reply
[1364] Step 9:
[1365] Emergency response
[1366] If an elderly person encounters an emergency, the device will instantly accept voice input, notify emergency contacts, and provide appropriate voice instructions from the server, using voice recognition and generation AI technology.
[1367] Input: Voice input in case of emergency
[1368] Output: Provided voice instructions and emergency notifications
[1369] Step 10:
[1370] Safety checks and home safety monitoring
[1371] The device periodically checks the elderly person's condition, and if an abnormality is detected, it immediately notifies the server and contacts the emergency contact. It also uses sensors in the home to monitor for smoke and gas leaks, and issues an alert if an abnormality is detected.
[1372] Input: Elderly status and home sensor data
[1373] Output: Safety check results and warning notifications
[1374] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1375] The present invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI, speech recognition technology, and an emotion engine. Specific embodiments of this system are described in detail below.
[1376] System Overview
[1377] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[1378] Furthermore, by combining it with an emotion engine, it becomes possible to recognize the user's emotions and generate responses based on those emotions. This emotion engine analyzes emotions from the user's voice and sends the results to a server, allowing the generative AI model to adjust the response content.
[1379] Program processing
[1380] 1. User enters information
[1381] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1382] The device converts this information into JSON format and sends it to the server by pressing the send button.
[1383] 2. The server saves the information and generates the model
[1384] The server stores the received information in a database.
[1385] The server customizes the generative AI model and speech recognition model based on the stored information.
[1386] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using new audio datasets as needed.
[1387] 3. User adds voice input
[1388] Users access the "Add Learning Data" section of the application and use the recording function to dictate events from the elderly person's past or favorite topics.
[1389] The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[1390] The device sends the encoded audio data to the server.
[1391] 4. The server analyzes the audio data and updates the model.
[1392] The server stores the received audio data in an appropriate format and analyzes it.
[1393] The server then retrains the generative AI model and speech recognition model based on this voice data, updating the models to be optimized for specific dialects and individual speaking styles.
[1394] 5. Seniors start asking questions and initiating conversations
[1395] Using the application, the elderly person presses the talk button to ask a question or start a conversation.
[1396] The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine, which then sends the text data to a server.
[1397] 6. Emotion engine recognizes user emotions
[1398] As the voice data is sent to the server, the emotion engine analyzes the emotion from the user's voice.
[1399] The emotion analysis results are sent to a server, which records the user's emotional state.
[1400] 7. The server generates an appropriate response based on the emotion.
[1401] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[1402] The generated reply text is sent to the terminal.
[1403] 8. The device converts your reply into voice
[1404] The response text received by the device is converted into speech using a speech synthesis engine.
[1405] The device plays the generated audio to the elderly person, providing answers to their questions and appropriate feedback.
[1406] Specific examples
[1407] For example, if Mr. Tanaka were to use the system, his son would input information about him, such as his name, background, hobbies, favorite foods, and local dialect, into the application, and provide additional voice data, which would then create a generative AI model and a voice recognition model specifically for him.
[1408] When Tanaka asks the application, "What's the weather like today?", the device converts this speech into text, while the emotion engine simultaneously analyzes Tanaka's emotions. The text data and emotion data are sent to the server, which uses this information to generate a response that takes emotion into account, such as, "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response is converted into audio and played to Tanaka in real time.
[1409] Also, if Tanaka is in a certain emotional state, for example, feeling anxious, the system will provide a reassuring response tailored to that emotion, such as, "The weather is nice today, and your favorite flowers might be in bloom."
[1410] In this way, the present invention aims to support dementia in the elderly, and by combining generative AI, voice recognition technology, and an emotion engine, it provides individually customized interactive support, realizing comprehensive support for the elderly and their families.
[1411] The processing flow will be explained below.
[1412] Step 1:
[1413] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[1414] Step 2:
[1415] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[1416] Step 3:
[1417] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[1418] Step 4:
[1419] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[1420] Step 5:
[1421] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[1422] Step 6:
[1423] While the voice data is being sent to the server, the device uses an emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions and sends the results to the server along with the text data.
[1424] Step 7:
[1425] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[1426] Step 8:
[1427] The server generates a response text and sends it to the device. The device converts the received response text into speech using a speech synthesis engine. The device plays the generated speech to the elderly person, providing answers to their questions and appropriate feedback.
[1428] Step 9:
[1429] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[1430] Step 10:
[1431] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[1432] Step 11:
[1433] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[1434] Example 2
[1435] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1436] There is a lack of interactive support optimized for each patient in dementia support for the elderly. In particular, it is difficult to provide support that meets the individual needs of the elderly by using individually customized generative AI models and speech recognition models. Another issue is the lack of technology that can recognize emotions from speech and generate appropriate responses based on them. Furthermore, there is a need for an interactive tutorial mode to improve the IT literacy of the elderly and for improved mobility service reservation functions.
[1437] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for storing information input by a user in a database, a means for generating individually customized generative AI models and voice recognition models based on the stored information, a means for converting the user's voice input into text data and transmitting it to the server, and a means for adjusting the content of responses generated based on the emotion analysis results. This makes it possible to provide optimal support according to the individual needs of elderly people. In addition, by analyzing the user's emotions using an emotion recognition engine and generating responses that take these emotions into consideration, more personalized interactive dialogue is realized.
[1438] "User" refers to an individual who uses the system to input information or voice input.
[1439] "Information" refers to data that users enter into the system, including the elderly person's name, background, hobbies, preferences, local dialect, etc.
[1440] A "database" is a storage device that stores input information and allows it to be retrieved when needed.
[1441] A "generative AI model" is an algorithmic model of artificial intelligence that is individually customized based on user input information.
[1442] A "speech recognition model" is an algorithmic model used to convert a user's speech data into text data.
[1443] "Audio data" refers to digital audio files generated by a user providing voice input.
[1444] A "server" is a computer system responsible for storing information and generating and training generative AI models and speech recognition models.
[1445] "Analysis" is the process in which the server analyzes information based on the voice data received and performs the necessary processing.
[1446] "Text data" is character string information converted from voice data by a voice recognition model.
[1447] An "emotion recognition engine" is software that analyzes emotions from the user's voice and sends the results to a server.
[1448] A "reply" is a response that the server generates in response to a user's question or request using a generative AI model.
[1449] The "interactive tutorial mode" is an interactive learning mode provided for seniors to learn how to use the system.
[1450] "Transportation service" is a function that allows elderly people to reserve available transportation and adjust their travel schedules.
[1451] "Model training" is the process by which a generative AI model or speech recognition model learns from stored data and improves its performance.
[1452] "Playback" is the process by which the device plays the generated audio to the elderly person.
[1453] This invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI models, speech recognition technology, and an emotion engine. Specific embodiments of the present invention are described below.
[1454] System Overview
[1455] Users enter information about the elderly person through a dedicated application, and the device converts this information into JSON format and sends it to a server. The server stores the received information in a database and individually customizes the generative AI model and speech recognition model. The generative AI model used is GPT-4, and the speech recognition model is the Google Speech-to-Text API. The IBM Watson Tone Analyzer is used for emotion recognition.
[1456] Users can improve the accuracy of the model by adding information about the senior's past experiences, hobbies, preferences, etc. This allows the generative AI model and speech recognition model to be individually tailored to the senior. The system also features an interactive tutorial mode and a mobility service booking function.
[1457] Furthermore, by combining it with an emotion recognition engine, it can recognize the user's emotions in real time and generate responses based on those emotions, enabling interactive dialogue that takes the user's emotional state into account.
[1458] Hardware and Software Details
[1459] server
[1460] Database: A relational database such as "MySQL" or "PostgreSQL" to store the information.
[1461] Generative AI models: Individually customized generative AI models (e.g., "GPT-4").
[1462] Speech recognition model: A speech recognition model specifically optimized for older adults (e.g., Google Speech-to-Text API).
[1463] Terminal
[1464] Conversion and transmission of input information: A function that converts information entered by the user into JSON format and sends it to the server. This is performed by software processing inside the device.
[1465] Encoding and sending audio data: The recorded audio data is encoded with appropriate sound quality parameters (sample rate, bit rate) and sent to the server.
[1466] Emotion Recognition Engine
[1467] IBM Watson Tone Analyzer: Analyzes emotions from the user's voice and sends the results to the server.
[1468] Response generation and speech synthesis
[1469] Generative AI model: Generates appropriate responses based on the user's text data and sentiment analysis results.
[1470] Speech synthesis engine: A speech synthesis engine such as "Amazon Polly" is used to convert response text into speech.
[1471] Specific examples
[1472] Specific examples are shown below.
[1473] For example, a user inputs information about an elderly person (Mr. Tanaka), such as his name, background, hobbies, favorite foods, and local dialect, into an application, which then sends this information to a server.
[1474] The server stores this information in a database and customizes the generative AI model and speech recognition model based on the input data. The user can then add voice input about Tanaka's past events and favorite topics, and send this voice data to the server.
[1475] The server analyzes the received voice data and retrains the generative AI model and speech recognition model. When the elderly person starts a question or conversation using the application, the device converts the voice into text and sends it to the server.
[1476] The emotion recognition engine analyzes emotions from the voice and sends the results to a server. The server uses a generative AI model based on the emotion analysis results and text data to generate an appropriate response and sends it to the device. The device then converts the response text into speech and plays it back to the elderly.
[1477] For example, if Tanaka asks, "How's the weather today?", the emotion analysis engine will determine that Tanaka is feeling a little anxious. The server will then generate a reassuring response: "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response will be converted into audio and played to Tanaka in real time.
[1478] In this way, the present invention aims to support dementia in the elderly, and by providing individually customized interactive support, it realizes comprehensive support for the elderly and their families.
[1479] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1480] Step 1: User Enters Information
[1481] What happens: The user opens a dedicated application and enters information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1482] Input: Information about the elderly person (such as name, background, hobbies, preferences, local dialect, etc.).
[1483] Data processing: The terminal converts this information into JSON format.
[1484] Output: Data converted to JSON format.
[1485] Specific operation: Enter data into the input form displayed on the terminal and press the send button. The data is converted into JSON format and the message "Submission successful" is displayed.
[1486] Step 2: The server saves the information and generates the model
[1487] Processing details: The server stores the received information in a database and customizes the generative AI model and speech recognition model individually.
[1488] Input: JSON formatted data.
[1489] Data processing: The received information is stored in a database and used to customize generative AI models (e.g., GPT-4) and speech recognition models (e.g., Google Speech-to-Text API).
[1490] Output: Customized generative AI model and speech recognition model.
[1491] What happens: The server parses the JSON data and stores the necessary information in a database. It then adjusts the parameters of the AI model based on the stored information and retrains the speech recognition model using the new audio dataset.
[1492] Step 3: User adds voice input
[1493] What happens: The user visits the "Add Learning Data" section of the app and uses the recording feature to dictate events from the senior's past or favorite topics.
[1494] Input: Speech data of elderly people.
[1495] Data processing: The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[1496] Output: The encoded audio data.
[1497] Specific operation: The user presses the record button and inputs voice information about the elderly person's past events or favorite topics. Once the recording is complete, the device encodes the voice data and sends it to the server.
[1498] Step 4: The server analyzes the audio data and updates the model
[1499] Processing details: The server saves the received audio data in an appropriate format and analyzes it.
[1500] Input: Encoded audio data.
[1501] Data processing: Analyze the voice data and retrain the generative AI model and speech recognition model.
[1502] Output: Updated generative AI model and speech recognition model.
[1503] What it does: The server converts the audio data into a different format, generates training data optimized for a specific dialect or speaking style, and updates the model based on the generated data.
[1504] Step 5: The senior initiates the question or conversation
[1505] What happens: Seniors use the application and press the talk button to ask questions or start a conversation.
[1506] Input: Speech data of elderly people.
[1507] Data processing: The device collects the elderly person's voice in real time and converts it into text using a voice recognition engine.
[1508] Output: Parsed text data.
[1509] Specific operation: When an elderly person presses the button to speak, the device collects voice data and converts it into text in real time. This text data is then sent to the server.
[1510] Step 6: The emotion engine recognizes the user's emotion
[1511] Processing details: As the voice data is sent to the server, the emotion engine analyzes the emotions from the user's voice.
[1512] Input: Audio data.
[1513] Data processing: The emotion analysis engine recognizes emotions from voice and converts the results into data.
[1514] Output: Sentiment analysis result data.
[1515] Specific operation: The voice data sent to the server is passed to the emotion engine, where emotion analysis is performed. The analysis results are stored in the database as the emotional state.
[1516] Step 7: The server generates an appropriate response based on the sentiment
[1517] Processing details: The server uses generative AI to generate an appropriate response based on the text received and the results of sentiment analysis.
[1518] Input: Text data and sentiment analysis result data.
[1519] Data processing: Generative AI models analyze text and emotional state to generate optimal responses.
[1520] Output: The generated response text.
[1521] Specific operation: The server uses the text data and the sentiment analysis results to query the generative AI model and generate an appropriate response, which is then queued for transmission to the device.
[1522] Step 8: Your device converts your response into audio
[1523] Processing details: The response text received by the device is converted into speech using a speech synthesis engine.
[1524] Input: Response text.
[1525] Data processing: Use a speech synthesis engine (e.g., Amazon Polly) to convert text data into speech.
[1526] Output: Synthesized speech data.
[1527] Specific operation: The device receives the response text and converts it into natural speech using a speech synthesis engine. The device then plays this speech back to the elderly person and provides appropriate feedback.
[1528] (Application example 2)
[1529] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1530] Current assistance systems for the elderly struggle to provide individually customized interactive support, and are unable to fully improve the convenience and sense of security of the elderly in their daily lives. Furthermore, seniors lack the means to receive appropriate real-time support when searching for products or answering questions in physical stores. To solve these problems and significantly improve the quality of life for the elderly, a new system is needed that combines generative AI models, speech recognition technology, and emotion engines to provide customized interactive assistance.
[1531] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for storing information input by a user in a database; means for generating individually customized generative AI models and speech recognition models based on the stored information; means for the user to add voice input and send the voice data to the server; means for the server to analyze the received voice data and retrain the generative AI model and speech recognition model; means for converting the elderly's voice input into text data and sending it to the server; means for the server to generate an appropriate response using the generative AI and send it to the terminal; means for the terminal to convert the received response text into speech and play it back to the elderly; means for analyzing emotions from the elderly's speech using an emotion engine and sending the result to the server; and means for providing real-time interactive support via a smart device worn by the elderly. This makes it possible to provide appropriate real-time support for elderly people when searching for products in physical stores, such as guidance and question-answering.
[1532] The "user" is a person who uses the system and is responsible for inputting information about the elderly person and adding voice input.
[1533] "Information" is a general term for data necessary for individual customization, such as the elderly person's name, background, hobbies, preferences, and local dialect.
[1534] "Database" refers to a storage device and system for storing user-entered information and elderly person's voice data.
[1535] A "generative AI model" is an artificial intelligence model that is generated based on information about the user or elderly person, and is used to provide appropriate responses and support.
[1536] The "voice recognition model" is an artificial intelligence model that analyzes the voices of elderly people and converts the content into text data.
[1537] The "server" is a computer system that receives data sent by users and elderly people, and stores, analyzes, trains, and generates responses.
[1538] "Voice data" refers to digitized data of the voices spoken by elderly people.
[1539] "Text data" is data of character information converted from voice data by a voice recognition model.
[1540] An "emotion engine" is software or a system for analyzing the emotions of elderly people from voice data.
[1541] A "smart device" is an electronic device worn by the elderly that provides real-time interactive support, such as smart glasses.
[1542] This invention is a system that combines generative AI models, speech recognition technology, and an emotion engine to provide customized interactive support to improve the experience of elderly people in physical stores. The system operates as follows: the user inputs information, and the server generates a generative AI model and a speech recognition model based on this information.
[1543] Hardware and Software Used
[1544] Hardware
[1545] Smart devices (e.g. smart glasses)
[1546] server
[1547] software
[1548] Speech recognition engine (Google Speech-to-Text API)
[1549] Generative AI model (OpenAI GPT-3)
[1550] Emotion Engine (Affectiva SDK)
[1551] Speech synthesis engine (Google Text-to-Speech API)
[1552] Database (MongoDB)
[1553] What the program does
[1554] 1. User enters information
[1555] Users use a dedicated application to input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. This information is converted into JSON format and sent to the server, which then stores it in a database (MongoDB).
[1556] 2. Model generation
[1557] The server generates and customizes a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the received information. Voice data, including specific dialects and individual speaking patterns, is also used as training data for the speech recognition model.
[1558] 3. Real-time support in physical stores
[1559] When an elderly person wears a smart device and starts a conversation or asks a question, their voice is converted into text data using a voice recognition engine and sent to a server. At the same time, an emotion engine (Affectiva SDK) analyzes emotions from the voice data and sends the emotional information to the server.
[1560] 4. Reply Generation and Voice Response
[1561] The server uses a generative AI model (OpenAI GPT-3) to generate an appropriate response based on the received text data and emotion data. The generated response is then sent back to the smart device from the server and converted into speech using a speech synthesis engine (Google Text-to-Speech API). Appropriate guidance and question responses are provided to the elderly in real time.
[1562] Specific examples
[1563] For example, when an elderly person is shopping in a physical store, they can speak to their smart device and ask, "Where is the sugar?" This speech is converted into text data in real time and analyzed by an emotion engine. The server receives this and generates a response such as, "Sugar is at the far right of the food shelf. Your favorite snacks are also nearby." The response is converted into audio and provided to the elderly.
[1564] Prompt Sentence Examples
[1565] User: How's the weather today?
[1566] Assistant (considering emotions): It's sunny today, but it's expected to get a little windy this afternoon. Please be careful when you go out.
[1567] In this way, the present invention improves the quality of life of elderly people by providing them with real-time guidance and question answers that they need in physical stores.
[1568] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1569] Step 1:
[1570] The user enters information about the elderly person using a dedicated application. The information entered includes the elderly person's name, background, hobbies, preferences, local dialect, etc. This information is converted into JSON format and sent to the server. The server stores the received JSON data in a database (MongoDB).
[1571] Input: Basic information about the elderly person (name, background, hobbies, preferences, dialect)
[1572] Data processing: converting information into JSON format
[1573] Output: Information of elderly people stored in a database
[1574] Step 2:
[1575] The server customizes and generates a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the information stored in the database. Specifically, it builds training data for the model based on the stored information and adds new parameters to the existing model for learning.
[1576] Input: Elderly person information stored in a database
[1577] Data Computing: Customizing and retraining generative AI and speech recognition models
[1578] Output: Customized generative AI models and speech recognition models
[1579] Step 3:
[1580] The user uses the application's recording function to input voice data about the elderly person's past events or favorite topics. The voice data is encoded according to the sample rate and bit rate and sent to the server, which then stores the received voice data in an appropriate format.
[1581] Input: Speech data from elderly people
[1582] Data processing: Encoding of audio data
[1583] Output: Audio data stored on the server
[1584] Step 4:
[1585] The server analyzes the received audio data and retrains the speech recognition model, improving its accuracy based on the analyzed data and generating a model optimized for a particular dialect or individual speaking style.
[1586] Input: Audio data stored on the server
[1587] Data Computing: Analyzing speech data and retraining speech recognition models
[1588] Output: Optimized speech recognition model
[1589] Step 5:
[1590] Elderly people wear smart devices and start asking questions or having conversations. This speech is converted into text data in real time by a speech recognition engine (Google Speech-to-Text API) and sent to a server.
[1591] Input: Real-time voice data of elderly people
[1592] Data processing: Converting voice to text
[1593] Output: Send to server as text data
[1594] Step 6:
[1595] The server analyzes the elderly person's emotions through the received voice data using an emotion engine (Affectiva SDK). The emotion analysis results are also sent to the server.
[1596] Input: Elderly voice data
[1597] Data Computing: Emotion Analysis
[1598] Output: Emotion data of elderly people
[1599] Step 7:
[1600] The server generates an appropriate response based on the received text data and emotion data using a generative AI model (OpenAI GPT-3), and the generated response is then sent back to the smart device.
[1601] Input: Text data and elderly emotion data
[1602] Data Computing: Generating Responses with Generative AI Models
[1603] Output: Response data sent to the smart device
[1604] Step 8:
[1605] The response text received on the smart device is converted into speech using a speech synthesis engine (Google Text-to-Speech API), and a response is provided to the elderly person.
[1606] Input: Reply text sent to smart device
[1607] Data processing: Convert response text into speech
[1608] Output: Voice response provided to the senior
[1609] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1610] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1611] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1612] [Fourth embodiment]
[1613] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1614] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1615] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1616] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1617] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1618] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1619] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1620] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1621] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1622] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1623] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1624] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1625] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1626] The present invention is a system for supporting dementia in the elderly, which combines generative AI and speech recognition technology to provide individually customized interactive support. Specific embodiments of this system are described in detail below.
[1627] System Overview
[1628] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[1629] Program processing
[1630] 1. User enters information
[1631] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1632] The device converts this information into JSON format and sends it to the server.
[1633] 2. The server saves the information and generates the model
[1634] The server stores the received information in a database.
[1635] The server customizes the generative AI model and speech recognition model based on the stored information.
[1636] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model with new audio datasets as needed.
[1637] 3. User adds voice input
[1638] Through the application, users record voice input and provide information about the senior's past events and favorite topics.
[1639] The device transmits the recorded audio data to the server.
[1640] 4. The server analyzes the audio data and updates the model.
[1641] The server analyzes the received voice data and retrains generative AI and speech recognition models optimized for specific dialects and individual speaking styles.
[1642] 5. Seniors start asking questions and initiating conversations
[1643] Seniors use the app to ask questions and initiate conversations.
[1644] The device collects the elderly person's voice in real time, converts it into text through a voice recognition engine, and sends it to a server.
[1645] 6. The server generates an appropriate response
[1646] The server analyzes the received text and uses generative AI to generate an appropriate response.
[1647] The generated response is sent to the terminal in text format.
[1648] 7. The device converts your reply into voice
[1649] The device converts the received response text into speech using a speech synthesis engine.
[1650] The device plays the generated audio to the elderly person.
[1651] Specific examples
[1652] For example, if an elderly person named Mr. Tanaka were to use this system, his son would enter information about Mr. Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, into the application. The son would then record and add information about Mr. Tanaka's past experiences and hobbies. This information and voice data is sent to a server, which then creates a generative AI model and a voice recognition model specifically for Mr. Tanaka.
[1653] When Tanaka asks "What's the weather like today?" through the application, the device converts the question into text and sends it to the server. The server uses generative AI to generate a response such as "It's sunny today, but the wind is expected to pick up a little in the afternoon," and sends it to the device. The device then converts this response into audio and plays it for Tanaka.
[1654] If Tanaka wants to travel to a specific location, he can reserve a transportation service through the application. The device sends the reservation information to the server, which then adjusts the travel schedule. As a result, Tanaka can travel safely and comfortably.
[1655] As described above, this invention aims to support elderly people with dementia by utilizing individually customized generative AI and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[1656] The processing flow will be explained below.
[1657] Step 1:
[1658] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[1659] Step 2:
[1660] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[1661] Step 3:
[1662] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[1663] Step 4:
[1664] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[1665] Step 5:
[1666] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[1667] Step 6:
[1668] The server analyzes the received text and generates an appropriate response using a generation AI. The generated response text is then sent to the device.
[1669] Step 7:
[1670] The device converts the received response text into speech using a speech synthesis engine, and the device plays the generated speech back to the elderly person, providing an answer to their question.
[1671] Step 8:
[1672] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[1673] Step 9:
[1674] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[1675] Step 10:
[1676] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[1677] Example 1
[1678] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1679] To support elderly people with dementia, there is a need for systems that provide interactive support tailored to individual needs. However, current technology does not have an established method for quickly generating and operating AI models and voice recognition models optimized for each elderly person, which means that elderly people are unable to receive appropriate support. In addition, there is a lack of comprehensive systems that also provide support for daily life, such as improving IT literacy and providing mobility support.
[1680] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1681] In this invention, the server includes: a means for converting information entered by a user into a data format and transmitting it to the server; a means for storing the received information in a database; a means for generating individually customized generative AI models and speech recognition models based on the stored information; a means for the user to add voice input and transmit the voice data to the server; a means for the server to analyze the received voice data and retrain the models; a means for converting the voice to text and transmitting it to the server; a means for the server to analyze the text using the generative AI to generate an appropriate response and transmit the response text to the terminal; and a means for the terminal to convert the received response text into voice and play it back. This enables dementia support tailored to each elderly person. It also enables integrated support for elderly people's IT literacy and mobility, thereby improving the overall quality of their daily lives.
[1682] A "server" is a computer system that provides data and functionality to other computers and devices over a network.
[1683] A "terminal" is an electronic device that is directly operated by a user and that communicates with a server.
[1684] A "user" is a person who utilizes the system to enter information or add voice input.
[1685] A "data format" is a format for expressing information in a structured way, such as the JSON format.
[1686] A "database" is a system that stores and manages data, making it easy to search for and analyze information.
[1687] A "generative AI model" is a type of artificial intelligence that automatically responds and processes data, and is used for natural language processing, etc.
[1688] A "speech recognition model" is a model that analyzes voice data and converts it into text.
[1689] "Voice input" refers to voice data provided by a user speaking to the system.
[1690] "Text" means the written representation of voice input or other data.
[1691] A "speech synthesis engine" is software for converting text data into voice data.
[1692] The "interactive tutorial mode" is an interactive support function that helps users learn how to operate the system.
[1693] A "mobility service" is a service that provides assistance to a user moving to a specific location.
[1694] "Schedule adjustment" refers to the act of managing and optimizing the dates and times of transportation services booked by users.
[1695] This invention is a system aimed at supporting dementia in the elderly, combining generative AI and speech recognition technology to provide individually customized interactive support. This system is comprised of three main elements: a server, a terminal, and a user. Below, we will explain how to specifically implement this system.
[1696] server
[1697] The server is a high-performance computer system that manages and operates the database, generative AI model, and speech recognition model. The main software used is TensorFlow and PyTorch, which use Python, and is used to train and retrain the generative AI model and speech recognition model. MySQL, PostgreSQL, and other databases are used.
[1698] Specific examples
[1699] When the user enters the elderly person's information, the server receives the JSON format data sent from the device and stores it in a database. Based on the received data, an individually customized generative AI model and voice recognition model are created.
[1700] Terminal
[1701] A terminal is a device that allows users to input information, record voice data, and communicate with a server. Terminals can be smartphones, tablets, or PCs. Audio recording and playback require a built-in or external microphone and speaker.
[1702] Specific examples
[1703] The user inputs information about Tanaka, such as his name, background, hobbies, favorite foods, and local dialect, through the device, and this information is converted into JSON format and sent to the server. The user then speaks aloud about Tanaka's past events and hobbies, and the audio data is recorded by the device and sent to the server.
[1704] User
[1705] The user is the person who operates the system, inputting information and voice input for the elderly, learning IT operations through an interactive tutorial mode, and booking transportation services on behalf of the elderly.
[1706] Specific examples
[1707] For example, Tanaka's son uses an application to enter Tanaka's information and sends the voice data to the server. The server then customizes the generative AI model and voice recognition model based on the received information, creating a system specifically for Tanaka. When Tanaka asks, "What's the weather like today?", the server uses the generative AI to create an appropriate response and responds via voice via the device.
[1708] Prompt Sentence Examples
[1709] Information input prompt: "Please enter the senior's name, background, hobbies, favorite foods, and local dialect."
[1710] Voice prompt: "Record an older adult talking about past events and hobbies."
[1711] Conversation prompt: "What's the weather like today?"
[1712] Movement service prompt: "Enter details of your next move destination."
[1713] As described above, this system aims to support elderly people with dementia, and utilizes individually customized generative AI models and voice recognition technology to provide support for conversations and questions. It is a comprehensive support system that also includes support for improving IT literacy and mobility.
[1714] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1715] Step 1:
[1716] User enters information
[1717] Specific operation:
[1718] Users launch a dedicated application and enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1719] input:
[1720] Information on the elderly person's name, background, hobbies, preferences, local dialect, etc.
[1721] Data processing / calculation:
[1722] The terminal converts the input information into JSON format.
[1723] output:
[1724] Elderly information data in JSON format.
[1725] Specific actions added:
[1726] When a user enters information into the input screen and clicks the "Submit" button, the information is converted into JSON format.
[1727] Step 2:
[1728] The server stores the information and generates the model
[1729] Specific operation:
[1730] The device sends the generated JSON data to the server.
[1731] The server stores the received data in a database.
[1732] input:
[1733] Elderly information data in JSON format.
[1734] Data processing / calculation:
[1735] Extract the necessary information from the JSON data and store it in the database.
[1736] Customize generative AI and speech recognition models using libraries such as Python and TensorFlow.
[1737] Adjust the model's learning parameters and possibly retrain it.
[1738] output:
[1739] Customized generative AI and speech recognition models.
[1740] Specific actions added:
[1741] Save the JSON data into a MySQL database and run a script to generate the model.
[1742] Step 3:
[1743] User adds voice input
[1744] Specific operation:
[1745] Users operate an application designed specifically for seniors and input information about past events and hobbies using voice.
[1746] input:
[1747] Voice input of elderly people's past events and hobbies.
[1748] Data processing / calculation:
[1749] The device records the audio and generates an audio file in WAV format or similar.
[1750] Send the audio file to the server.
[1751] output:
[1752] Audio files (WAV format, etc.).
[1753] Specific actions added:
[1754] Press the record button to record the audio, and then press the "send" button after recording is complete to send the audio file to the server.
[1755] Step 4:
[1756] The server analyzes the audio data and updates the model
[1757] Specific operation:
[1758] The server analyzes the received audio file and converts it into text using a speech-to-text engine.
[1759] The parameters of the generative AI model and speech recognition model are readjusted based on the converted text information.
[1760] input:
[1761] Audio file.
[1762] Data processing / calculation:
[1763] A speech-to-text engine is used to convert the audio into text and store it in a database.
[1764] Retrain the model based on the converted text and existing data.
[1765] output:
[1766] Updated generative AI and speech recognition models.
[1767] Specific actions added:
[1768] Input the audio file into the analysis script and run model retraining along with the text conversion results.
[1769] Step 5:
[1770] Seniors start asking questions and conversations
[1771] Specific operation:
[1772] The elderly person launches the application and speaks a prompt to start the conversation (e.g., "What's the weather like today?").
[1773] input:
[1774] Questions and conversations of the elderly.
[1775] Data processing / calculation:
[1776] The device collects audio in real time and records the audio data.
[1777] The speech is converted into text through a speech recognition engine and sent to the server.
[1778] output:
[1779] Conversation in text format.
[1780] Specific actions added:
[1781] The question is recorded, and after the recording is complete, it is automatically converted into text and sent to the server.
[1782] Step 6:
[1783] The server generates an appropriate response
[1784] Specific operation:
[1785] The server analyzes the received text and uses generative AI to generate an appropriate response.
[1786] The generated reply text is sent to the terminal.
[1787] input:
[1788] Textual conversation from the user.
[1789] Data processing / calculation:
[1790] Generative AI (e.g., GPT-3) is used to analyze text and generate responses.
[1791] output:
[1792] The reply text.
[1793] Specific actions added:
[1794] It runs an algorithm to generate a response to the question and sends the generated text to the terminal.
[1795] Step 7:
[1796] The device converts your response into voice
[1797] Specific operation:
[1798] The device converts the received response text into speech using a speech synthesis engine.
[1799] The generated audio is played to the elderly person.
[1800] input:
[1801] The response text from the server.
[1802] Data processing / calculation:
[1803] Convert text to speech using a speech synthesis engine (e.g. Google Text-to-Speech).
[1804] output:
[1805] Audio data.
[1806] Specific actions added:
[1807] The device starts the speech synthesis engine and plays the generated speech on the playback device.
[1808] (Application example 1)
[1809] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1810] Elderly people with dementia may face many challenges in their daily lives. These challenges include being unable to respond appropriately in emergencies, being unable to ensure daily safety, and having difficulty preventing accidents at home. Furthermore, dealing with these challenges places a burden on the individual and their family, which is problematic. The present invention aims to utilize generative AI and speech recognition technology to individually and effectively resolve these challenges.
[1811] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1812] In this invention, the server includes means for saving information entered by a user in a database, means for generating individually customized generative AI models and voice recognition models based on the saved information, means for the user to add voice input and send the voice data to the server, means for the server to analyze the received voice data and retrain the generative AI model and voice recognition model, means for converting the elderly person's voice input into text data and sending it to the server, means for the server to generate an appropriate response using the generative AI and send it to the terminal, means for the terminal to convert the received response text into voice and play it back to the elderly person, means for notifying an emergency contact and providing voice instructions when the elderly person encounters an emergency, means for checking the safety of the elderly person at regular intervals and notifying if an abnormality occurs, and means for monitoring dangers within the home and issuing an alert if an abnormality occurs. This enables rapid response to emergencies, regular safety checks, and safety monitoring within the home.
[1813] "Means for saving information entered by the user in a database" refers to means for saving data obtained from the user, such as personal information, hobbies, and past events.
[1814] "Means for generating individually customized generative AI models and speech recognition models based on stored information" refers to means for creating generative AI and speech recognition models that are optimal for individual users by utilizing information stored in a database.
[1815] The "means for a user to add a voice input and transmit the voice data to a server" refers to a means for a user to record voice data and transmit the data to a server.
[1816] "Means for analyzing voice data received by the server and retraining the generative AI model and the voice recognition model" refers to means for analyzing voice data received by the server and retraining the generative AI and the voice recognition model.
[1817] The "means for converting the voice input of the elderly person into text data and transmitting it to the server" is a means for converting the voice uttered by the elderly person into text in real time and transmitting the text data to the server.
[1818] "Means for the server to use generation AI to generate an appropriate response and send it to the terminal" means a means for the server to use generation AI technology to generate an appropriate response to the received text data and send that response to the terminal.
[1819] The "means for converting the response text received by the terminal into speech and playing it back to the elderly" refers to a means for converting the response text received by the terminal from the server into speech using speech synthesis technology and playing it back to the elderly.
[1820] "Means for notifying emergency contacts and providing voice instructions when an elderly person encounters an emergency" refers to a means for automatically notifying pre-set emergency contacts and simultaneously providing appropriate voice guidance when an elderly person faces an emergency.
[1821] "Means for checking the safety of elderly people at regular intervals and notifying in the event of an abnormality" refers to a means for periodically checking the condition of elderly people and notifying emergency contacts if an abnormality is detected.
[1822] "Means for monitoring dangers within the home and issuing warnings in the event of an abnormality" refers to a means for monitoring smoke, gas leaks, etc. using sensors installed within the home and issuing warnings in the event of an abnormality.
[1823] The present invention is a system for supporting elderly people with dementia, which uses individually customized generative AI models and speech recognition models to provide support for elderly people in their daily lives. Specific embodiments of the system are described below.
[1824] System Configuration and Operation
[1825] Hardware and Software
[1826] Devices: Smartphones, smart glasses, head-mounted displays, etc.
[1827] Server: We use cloud-based servers to host the database, generative AI models, and speech recognition models.
[1828] Sensors: Smoke detectors and gas sensors are installed in the home, and data is sent to the device when an abnormality occurs.
[1829] Data processing and calculation
[1830] 1. Entering and saving information
[1831] Users enter personal information, hobbies, past events, etc. about the elderly person through a dedicated application, and this information is converted into JSON format and sent to the server, which then stores this data in a database.
[1832] 2. Generate and retrain the model
[1833] The server creates a generative AI model and a voice recognition model based on the stored information, and customizes them specifically for seniors. When additional voice input is received, the server analyzes the voice data and retrains the model.
[1834] 3. Speech Processing and Text Conversion
[1835] When the elderly person speaks, the device converts the speech into text in real time and sends it to the server, which then generates a response based on the received text.
[1836] 4. Response Generation and Audio Playback
[1837] The server uses generative AI to generate appropriate responses and sends them in text format to the device, which then converts the text into speech and plays it back to the elderly.
[1838] 5. Emergency response and safety monitoring
[1839] In the event of an emergency, the device will automatically notify emergency contacts and provide appropriate voice guidance. The device also monitors data from sensors and notifies users if an abnormality occurs.
[1840] Specific examples
[1841] For example, the present invention is useful in the following specific situations.
[1842] Example 1: Emergency response
[1843] If an elderly person falls and calls for help on their smartphone, the application will recognize the voice and quickly notify emergency contacts. At the same time, it will provide appropriate voice guidance to help the elderly person take immediate action.
[1844] Example prompt sentence:
[1845] "I have had a fall. Please notify your emergency contacts."
[1846] Example 2: Safety reminder
[1847] The application periodically asks, "How are you?" and waits for a response from the elderly person. If an abnormality is detected, it immediately notifies emergency contacts.
[1848] Example prompt sentence:
[1849] "How are you? If you don't respond, I'll notify my emergency contacts."
[1850] As described above, the present invention is a system that utilizes generative AI models and voice recognition technology to comprehensively support the daily lives of the elderly and provide a safe and secure environment.
[1851] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1852] Step 1:
[1853] The user enters information
[1854] Users enter personal information (such as name, address, and emergency contact information), lifestyle habits, hobbies, and past events through a dedicated application. The application converts the entered information into JSON format and sends it to the server via the network.
[1855] Input: Personal information and past event data entered by the user
[1856] Output: Data converted to JSON format
[1857] Step 2:
[1858] The server stores the information
[1859] The server parses the JSON format data received over the network and stores it in a database for later processing.
[1860] Input: JSON format user information data
[1861] Output: User information stored in the database
[1862] Step 3:
[1863] Generate a model based on the retained information
[1864] The server uses the information stored in the database to create individually customized generative AI and speech recognition models, which involves applying machine learning algorithms to the data.
[1865] Input: User information stored in the database
[1866] Output: Generated generative AI model and speech recognition model
[1867] Step 4:
[1868] User adds voice input
[1869] The user inputs voice data through the application, and the voice data is collected in real time by the device and sent to a server for analysis.
[1870] Input: User's voice data
[1871] Output: Audio data sent to the server
[1872] Step 5:
[1873] The server analyzes the audio data and retrains the model.
[1874] The server analyzes the received voice data and retrains the generative AI model and speech recognition model to improve their performance, using a speech recognition algorithm.
[1875] Input: Audio data
[1876] Output: Updated AI and speech recognition models after retraining
[1877] Step 6:
[1878] Converting elderly people's voice input into text
[1879] When an elderly person uses the application to input voice, the device converts the voice into text data in real time and sends it to the server, using a voice recognition engine.
[1880] Input: Elderly voice data
[1881] Output: Data converted to text
[1882] Step 7:
[1883] The server generates an appropriate response
[1884] The server uses AI to generate an appropriate response based on the received text data. This response is generated based on the user's individually customized information and sent to the device.
[1885] Input: Text data
[1886] Output: The generated response text
[1887] Step 8:
[1888] The device converts your response into speech
[1889] The device converts the response text received from the server into voice using a speech synthesis engine and plays it back to the elderly, who can then confirm the response by voice.
[1890] Input: Reply text
[1891] Output: The speech-transcribed reply
[1892] Step 9:
[1893] Emergency response
[1894] If an elderly person encounters an emergency, the device will instantly accept voice input, notify emergency contacts, and provide appropriate voice instructions from the server, using voice recognition and generation AI technology.
[1895] Input: Voice input in case of emergency
[1896] Output: Provided voice instructions and emergency notifications
[1897] Step 10:
[1898] Safety checks and home safety monitoring
[1899] The device periodically checks the elderly person's condition, and if an abnormality is detected, it immediately notifies the server and contacts the emergency contact. It also uses sensors in the home to monitor for smoke and gas leaks, and issues an alert if an abnormality is detected.
[1900] Input: Elderly status and home sensor data
[1901] Output: Safety check results and warning notifications
[1902] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1903] The present invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI, speech recognition technology, and an emotion engine. Specific embodiments of this system are described in detail below.
[1904] System Overview
[1905] Users input information about the senior through the app, and the system stores this information in a database to individually customize the generative AI model and voice recognition model. Users can improve the accuracy of the model by adding information about the senior's past events, hobbies, preferences, etc. The app also features an interactive tutorial mode and a transportation service booking function, helping seniors develop IT literacy and making their daily lives more convenient.
[1906] Furthermore, by combining it with an emotion engine, it becomes possible to recognize the user's emotions and generate responses based on those emotions. This emotion engine analyzes emotions from the user's voice and sends the results to a server, allowing the generative AI model to adjust the response content.
[1907] Program processing
[1908] 1. User enters information
[1909] Using a dedicated application, users enter information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[1910] The device converts this information into JSON format and sends it to the server by pressing the send button.
[1911] 2. The server saves the information and generates the model
[1912] The server stores the received information in a database.
[1913] The server customizes the generative AI model and speech recognition model based on the stored information.
[1914] The server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using new audio datasets as needed.
[1915] 3. User adds voice input
[1916] Users access the "Add Learning Data" section of the application and use the recording function to dictate events from the elderly person's past or favorite topics.
[1917] The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[1918] The device sends the encoded audio data to the server.
[1919] 4. The server analyzes the audio data and updates the model.
[1920] The server stores the received audio data in an appropriate format and analyzes it.
[1921] The server then retrains the generative AI model and speech recognition model based on this voice data, updating the models to be optimized for specific dialects and individual speaking styles.
[1922] 5. Seniors start asking questions and initiating conversations
[1923] Using the application, the elderly person presses the talk button to ask a question or start a conversation.
[1924] The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine, which then sends the text data to a server.
[1925] 6. Emotion engine recognizes user emotions
[1926] As the voice data is sent to the server, the emotion engine analyzes the emotion from the user's voice.
[1927] The emotion analysis results are sent to a server, which records the user's emotional state.
[1928] 7. The server generates an appropriate response based on the emotion.
[1929] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[1930] The generated reply text is sent to the terminal.
[1931] 8. The device converts your reply into voice
[1932] The response text received by the device is converted into speech using a speech synthesis engine.
[1933] The device plays the generated audio to the elderly person, providing answers to their questions and appropriate feedback.
[1934] Specific examples
[1935] For example, if Mr. Tanaka were to use the system, his son would input information about him, such as his name, background, hobbies, favorite foods, and local dialect, into the application, and provide additional voice data, which would then create a generative AI model and a voice recognition model specifically for him.
[1936] When Tanaka asks the application, "What's the weather like today?", the device converts this speech into text, while the emotion engine simultaneously analyzes Tanaka's emotions. The text data and emotion data are sent to the server, which uses this information to generate a response that takes emotion into account, such as, "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response is converted into audio and played to Tanaka in real time.
[1937] Also, if Tanaka is in a certain emotional state, for example, feeling anxious, the system will provide a reassuring response tailored to that emotion, such as, "The weather is nice today, and your favorite flowers might be in bloom."
[1938] In this way, the present invention aims to support dementia in the elderly, and by combining generative AI, voice recognition technology, and an emotion engine, it provides individually customized interactive support, realizing comprehensive support for the elderly and their families.
[1939] The processing flow will be explained below.
[1940] Step 1:
[1941] Using a dedicated application, users input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. The device converts this information into JSON format and sends it to the server by pressing the send button.
[1942] Step 2:
[1943] The server stores the received information in a database, and then customizes the generative AI model and speech recognition model based on the stored information. Specifically, the server adjusts the learning parameters of the generative AI model and retrains the speech recognition model using the new audio dataset.
[1944] Step 3:
[1945] Users access the "Add Learning Data" section of the application and use the recording function to input voice data about the elderly person's past events or favorite topics. The device then encodes the recorded voice data according to sound quality parameters such as sample rate and bit rate.
[1946] Step 4:
[1947] The device sends the encoded voice data to the server, which stores it in the appropriate format and analyzes it. The server then uses this voice data to retrain the generative AI model and speech recognition model, updating the models to be optimized for specific dialects and individual speaking styles.
[1948] Step 5:
[1949] The elderly person uses the application to press the talk button to ask a question or start a conversation. The device collects the elderly person's voice in real time and converts it into text using a speech recognition engine. This text data is then sent to the server.
[1950] Step 6:
[1951] While the voice data is being sent to the server, the device uses an emotion engine to analyze the user's emotions. The emotion engine recognizes the emotions and sends the results to the server along with the text data.
[1952] Step 7:
[1953] The server analyzes the received text and the results of sentiment analysis, and uses generative AI to generate an appropriate response that takes into account the user's emotional state.
[1954] Step 8:
[1955] The server generates a response text and sends it to the device. The device converts the received response text into speech using a speech synthesis engine. The device plays the generated speech to the elderly person, providing answers to their questions and appropriate feedback.
[1956] Step 9:
[1957] Seniors attend IT classes to learn basic IT operations and how to use applications. After the class, the device automatically switches to interactive tutorial mode, allowing seniors to review what they learned.
[1958] Step 10:
[1959] The elderly person accesses the transportation service reservation screen from within the application and enters their destination and desired travel time. The device converts the reservation information into JSON format and sends it to the server.
[1960] Step 11:
[1961] The server checks the received reservation information and adjusts the appropriate transportation service schedule. The server then sends the transportation schedule to the terminal, which then displays the reservation confirmation information and notifies the elderly, allowing them to use the transportation service.
[1962] Example 2
[1963] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1964] There is a lack of interactive support optimized for each patient in dementia support for the elderly. In particular, it is difficult to provide support that meets the individual needs of the elderly by using individually customized generative AI models and speech recognition models. Another issue is the lack of technology that can recognize emotions from speech and generate appropriate responses based on them. Furthermore, there is a need for an interactive tutorial mode to improve the IT literacy of the elderly and for improved mobility service reservation functions.
[1965] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for storing information input by a user in a database, a means for generating individually customized generative AI models and voice recognition models based on the stored information, a means for converting the user's voice input into text data and transmitting it to the server, and a means for adjusting the content of responses generated based on the emotion analysis results. This makes it possible to provide optimal support according to the individual needs of elderly people. In addition, by analyzing the user's emotions using an emotion recognition engine and generating responses that take these emotions into consideration, more personalized interactive dialogue is realized.
[1966] "User" refers to an individual who uses the system to input information or voice input.
[1967] "Information" refers to data that users enter into the system, including the elderly person's name, background, hobbies, preferences, local dialect, etc.
[1968] A "database" is a storage device that stores input information and allows it to be retrieved when needed.
[1969] A "generative AI model" is an algorithmic model of artificial intelligence that is individually customized based on user input information.
[1970] A "speech recognition model" is an algorithmic model used to convert a user's speech data into text data.
[1971] "Audio data" refers to digital audio files generated by a user providing voice input.
[1972] A "server" is a computer system responsible for storing information and generating and training generative AI models and speech recognition models.
[1973] "Analysis" is the process in which the server analyzes information based on the voice data received and performs the necessary processing.
[1974] "Text data" is character string information converted from voice data by a voice recognition model.
[1975] An "emotion recognition engine" is software that analyzes emotions from the user's voice and sends the results to a server.
[1976] A "reply" is a response that the server generates in response to a user's question or request using a generative AI model.
[1977] The "interactive tutorial mode" is an interactive learning mode provided for seniors to learn how to use the system.
[1978] "Transportation service" is a function that allows elderly people to reserve available transportation and adjust their travel schedules.
[1979] "Model training" is the process by which a generative AI model or speech recognition model learns from stored data and improves its performance.
[1980] "Playback" is the process by which the device plays the generated audio to the elderly person.
[1981] This invention is a system for supporting dementia in the elderly, which provides individually customized interactive support by combining generative AI models, speech recognition technology, and an emotion engine. Specific embodiments of the present invention are described below.
[1982] System Overview
[1983] Users enter information about the elderly person through a dedicated application, and the device converts this information into JSON format and sends it to a server. The server stores the received information in a database and individually customizes the generative AI model and speech recognition model. The generative AI model used is GPT-4, and the speech recognition model is the Google Speech-to-Text API. The IBM Watson Tone Analyzer is used for emotion recognition.
[1984] Users can improve the accuracy of the model by adding information about the senior's past experiences, hobbies, preferences, etc. This allows the generative AI model and speech recognition model to be individually tailored to the senior. The system also features an interactive tutorial mode and a mobility service booking function.
[1985] Furthermore, by combining it with an emotion recognition engine, it can recognize the user's emotions in real time and generate responses based on those emotions, enabling interactive dialogue that takes the user's emotional state into account.
[1986] Hardware and Software Details
[1987] server
[1988] Database: A relational database such as "MySQL" or "PostgreSQL" to store the information.
[1989] Generative AI models: Individually customized generative AI models (e.g., "GPT-4").
[1990] Speech recognition model: A speech recognition model specifically optimized for older adults (e.g., Google Speech-to-Text API).
[1991] Terminal
[1992] Conversion and transmission of input information: A function that converts information entered by the user into JSON format and sends it to the server. This is performed by software processing inside the device.
[1993] Encoding and sending audio data: The recorded audio data is encoded with appropriate sound quality parameters (sample rate, bit rate) and sent to the server.
[1994] Emotion Recognition Engine
[1995] IBM Watson Tone Analyzer: Analyzes emotions from the user's voice and sends the results to the server.
[1996] Response generation and speech synthesis
[1997] Generative AI model: Generates appropriate responses based on the user's text data and sentiment analysis results.
[1998] Speech synthesis engine: A speech synthesis engine such as "Amazon Polly" is used to convert response text into speech.
[1999] Specific examples
[2000] Specific examples are shown below.
[2001] For example, a user inputs information about an elderly person (Mr. Tanaka), such as his name, background, hobbies, favorite foods, and local dialect, into an application, which then sends this information to a server.
[2002] The server stores this information in a database and customizes the generative AI model and speech recognition model based on the input data. The user can then add voice input about Tanaka's past events and favorite topics, and send this voice data to the server.
[2003] The server analyzes the received voice data and retrains the generative AI model and speech recognition model. When the elderly person starts a question or conversation using the application, the device converts the voice into text and sends it to the server.
[2004] The emotion recognition engine analyzes emotions from the voice and sends the results to a server. The server uses a generative AI model based on the emotion analysis results and text data to generate an appropriate response and sends it to the device. The device then converts the response text into speech and plays it back to the elderly.
[2005] For example, if Tanaka asks, "How's the weather today?", the emotion analysis engine will determine that Tanaka is feeling a little anxious. The server will then generate a reassuring response: "It's sunny today, but the wind is expected to pick up a little in the afternoon." This response will be converted into audio and played to Tanaka in real time.
[2006] In this way, the present invention aims to support dementia in the elderly, and by providing individually customized interactive support, it realizes comprehensive support for the elderly and their families.
[2007] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2008] Step 1: User Enters Information
[2009] What happens: The user opens a dedicated application and enters information about the elderly person, such as their name, background, hobbies, preferences, and local dialect.
[2010] Input: Information about the elderly person (such as name, background, hobbies, preferences, local dialect, etc.).
[2011] Data processing: The terminal converts this information into JSON format.
[2012] Output: Data converted to JSON format.
[2013] Specific operation: Enter data into the input form displayed on the terminal and press the send button. The data is converted into JSON format and the message "Submission successful" is displayed.
[2014] Step 2: The server saves the information and generates the model
[2015] Processing details: The server stores the received information in a database and customizes the generative AI model and speech recognition model individually.
[2016] Input: JSON formatted data.
[2017] Data processing: The received information is stored in a database and used to customize generative AI models (e.g., GPT-4) and speech recognition models (e.g., Google Speech-to-Text API).
[2018] Output: Customized generative AI model and speech recognition model.
[2019] What happens: The server parses the JSON data and stores the necessary information in a database. It then adjusts the parameters of the AI model based on the stored information and retrains the speech recognition model using the new audio dataset.
[2020] Step 3: User adds voice input
[2021] What happens: The user visits the "Add Learning Data" section of the app and uses the recording feature to dictate events from the senior's past or favorite topics.
[2022] Input: Speech data of elderly people.
[2023] Data processing: The device encodes the recorded audio data according to sound quality parameters such as sample rate and bit rate.
[2024] Output: The encoded audio data.
[2025] Specific operation: The user presses the record button and inputs voice information about the elderly person's past events or favorite topics. Once the recording is complete, the device encodes the voice data and sends it to the server.
[2026] Step 4: The server analyzes the audio data and updates the model
[2027] Processing details: The server saves the received audio data in an appropriate format and analyzes it.
[2028] Input: Encoded audio data.
[2029] Data processing: Analyze the voice data and retrain the generative AI model and speech recognition model.
[2030] Output: Updated generative AI model and speech recognition model.
[2031] What it does: The server converts the audio data into a different format, generates training data optimized for a specific dialect or speaking style, and updates the model based on the generated data.
[2032] Step 5: The senior initiates the question or conversation
[2033] What happens: Seniors use the application and press the talk button to ask questions or start a conversation.
[2034] Input: Speech data of elderly people.
[2035] Data processing: The device collects the elderly person's voice in real time and converts it into text using a voice recognition engine.
[2036] Output: Parsed text data.
[2037] Specific operation: When an elderly person presses the button to speak, the device collects voice data and converts it into text in real time. This text data is then sent to the server.
[2038] Step 6: The emotion engine recognizes the user's emotion
[2039] Processing details: As the voice data is sent to the server, the emotion engine analyzes the emotions from the user's voice.
[2040] Input: Audio data.
[2041] Data processing: The emotion analysis engine recognizes emotions from voice and converts the results into data.
[2042] Output: Sentiment analysis result data.
[2043] Specific operation: The voice data sent to the server is passed to the emotion engine, where emotion analysis is performed. The analysis results are stored in the database as the emotional state.
[2044] Step 7: The server generates an appropriate response based on the sentiment
[2045] Processing details: The server uses generative AI to generate an appropriate response based on the text received and the results of sentiment analysis.
[2046] Input: Text data and sentiment analysis result data.
[2047] Data processing: Generative AI models analyze text and emotional state to generate optimal responses.
[2048] Output: The generated response text.
[2049] Specific operation: The server uses the text data and the sentiment analysis results to query the generative AI model and generate an appropriate response, which is then queued for transmission to the device.
[2050] Step 8: Your device converts your response into audio
[2051] Processing details: The response text received by the device is converted into speech using a speech synthesis engine.
[2052] Input: Response text.
[2053] Data processing: Use a speech synthesis engine (e.g., Amazon Polly) to convert text data into speech.
[2054] Output: Synthesized speech data.
[2055] Specific operation: The device receives the response text and converts it into natural speech using a speech synthesis engine. The device then plays this speech back to the elderly person and provides appropriate feedback.
[2056] (Application example 2)
[2057] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2058] Current assistance systems for the elderly struggle to provide individually customized interactive support, and are unable to fully improve the convenience and sense of security of the elderly in their daily lives. Furthermore, seniors lack the means to receive appropriate real-time support when searching for products or answering questions in physical stores. To solve these problems and significantly improve the quality of life for the elderly, a new system is needed that combines generative AI models, speech recognition technology, and emotion engines to provide customized interactive assistance.
[2059] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for storing information input by a user in a database; means for generating individually customized generative AI models and speech recognition models based on the stored information; means for the user to add voice input and send the voice data to the server; means for the server to analyze the received voice data and retrain the generative AI model and speech recognition model; means for converting the elderly's voice input into text data and sending it to the server; means for the server to generate an appropriate response using the generative AI and send it to the terminal; means for the terminal to convert the received response text into speech and play it back to the elderly; means for analyzing emotions from the elderly's speech using an emotion engine and sending the result to the server; and means for providing real-time interactive support via a smart device worn by the elderly. This makes it possible to provide appropriate real-time support for elderly people when searching for products in physical stores, such as guidance and question-answering.
[2060] The "user" is a person who uses the system and is responsible for inputting information about the elderly person and adding voice input.
[2061] "Information" is a general term for data necessary for individual customization, such as the elderly person's name, background, hobbies, preferences, and local dialect.
[2062] "Database" refers to a storage device and system for storing user-entered information and elderly person's voice data.
[2063] A "generative AI model" is an artificial intelligence model that is generated based on information about the user or elderly person, and is used to provide appropriate responses and support.
[2064] The "voice recognition model" is an artificial intelligence model that analyzes the voices of elderly people and converts the content into text data.
[2065] The "server" is a computer system that receives data sent by users and elderly people, and stores, analyzes, trains, and generates responses.
[2066] "Voice data" refers to digitized data of the voices spoken by elderly people.
[2067] "Text data" is data of character information converted from voice data by a voice recognition model.
[2068] An "emotion engine" is software or a system for analyzing the emotions of elderly people from voice data.
[2069] A "smart device" is an electronic device worn by the elderly that provides real-time interactive support, such as smart glasses.
[2070] This invention is a system that combines generative AI models, speech recognition technology, and an emotion engine to provide customized interactive support to improve the experience of elderly people in physical stores. The system operates as follows: the user inputs information, and the server generates a generative AI model and a speech recognition model based on this information.
[2071] Hardware and Software Used
[2072] Hardware
[2073] Smart devices (e.g. smart glasses)
[2074] server
[2075] software
[2076] Speech recognition engine (Google Speech-to-Text API)
[2077] Generative AI model (OpenAI GPT-3)
[2078] Emotion Engine (Affectiva SDK)
[2079] Speech synthesis engine (Google Text-to-Speech API)
[2080] Database (MongoDB)
[2081] What the program does
[2082] 1. User enters information
[2083] Users use a dedicated application to input information about the elderly person, such as their name, background, hobbies, preferences, and local dialect. This information is converted into JSON format and sent to the server, which then stores it in a database (MongoDB).
[2084] 2. Model generation
[2085] The server generates and customizes a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the received information. Voice data, including specific dialects and individual speaking patterns, is also used as training data for the speech recognition model.
[2086] 3. Real-time support in physical stores
[2087] When an elderly person wears a smart device and starts a conversation or asks a question, their voice is converted into text data using a voice recognition engine and sent to a server. At the same time, an emotion engine (Affectiva SDK) analyzes emotions from the voice data and sends the emotional information to the server.
[2088] 4. Reply Generation and Voice Response
[2089] The server uses a generative AI model (OpenAI GPT-3) to generate an appropriate response based on the received text data and emotion data. The generated response is then sent back to the smart device from the server and converted into speech using a speech synthesis engine (Google Text-to-Speech API). Appropriate guidance and question responses are provided to the elderly in real time.
[2090] Specific examples
[2091] For example, when an elderly person is shopping in a physical store, they can speak to their smart device and ask, "Where is the sugar?" This speech is converted into text data in real time and analyzed by an emotion engine. The server receives this and generates a response such as, "Sugar is at the far right of the food shelf. Your favorite snacks are also nearby." The response is converted into audio and provided to the elderly.
[2092] Prompt Sentence Examples
[2093] User: How's the weather today?
[2094] Assistant (considering emotions): It's sunny today, but it's expected to get a little windy this afternoon. Please be careful when you go out.
[2095] In this way, the present invention improves the quality of life of elderly people by providing them with real-time guidance and question answers that they need in physical stores.
[2096] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2097] Step 1:
[2098] The user enters information about the elderly person using a dedicated application. The information entered includes the elderly person's name, background, hobbies, preferences, local dialect, etc. This information is converted into JSON format and sent to the server. The server stores the received JSON data in a database (MongoDB).
[2099] Input: Basic information about the elderly person (name, background, hobbies, preferences, dialect)
[2100] Data processing: converting information into JSON format
[2101] Output: Information of elderly people stored in a database
[2102] Step 2:
[2103] The server customizes and generates a generative AI model (OpenAI GPT-3) and a speech recognition model (Google Speech-to-Text API) based on the information stored in the database. Specifically, it builds training data for the model based on the stored information and adds new parameters to the existing model for learning.
[2104] Input: Elderly person information stored in a database
[2105] Data Computing: Customizing and retraining generative AI and speech recognition models
[2106] Output: Customized generative AI models and speech recognition models
[2107] Step 3:
[2108] The user uses the application's recording function to input voice data about the elderly person's past events or favorite topics. The voice data is encoded according to the sample rate and bit rate and sent to the server, which then stores the received voice data in an appropriate format.
[2109] Input: Speech data from elderly people
[2110] Data processing: Encoding of audio data
[2111] Output: Audio data stored on the server
[2112] Step 4:
[2113] The server analyzes the received audio data and retrains the speech recognition model, improving its accuracy based on the analyzed data and generating a model optimized for a particular dialect or individual speaking style.
[2114] Input: Audio data stored on the server
[2115] Data Computing: Analyzing speech data and retraining speech recognition models
[2116] Output: Optimized speech recognition model
[2117] Step 5:
[2118] Elderly people wear smart devices and start asking questions or having conversations. This speech is converted into text data in real time by a speech recognition engine (Google Speech-to-Text API) and sent to a server.
[2119] Input: Real-time voice data of elderly people
[2120] Data processing: Converting voice to text
[2121] Output: Send to server as text data
[2122] Step 6:
[2123] The server analyzes the elderly person's emotions through the received voice data using an emotion engine (Affectiva SDK). The emotion analysis results are also sent to the server.
[2124] Input: Elderly voice data
[2125] Data Computing: Emotion Analysis
[2126] Output: Emotion data of elderly people
[2127] Step 7:
[2128] The server generates an appropriate response based on the received text data and emotion data using a generative AI model (OpenAI GPT-3), and the generated response is then sent back to the smart device.
[2129] Input: Text data and elderly emotion data
[2130] Data Computing: Generating Responses with Generative AI Models
[2131] Output: Response data sent to the smart device
[2132] Step 8:
[2133] The response text received on the smart device is converted into speech using a speech synthesis engine (Google Text-to-Speech API), and a response is provided to the elderly person.
[2134] Input: Reply text sent to smart device
[2135] Data processing: Convert response text into speech
[2136] Output: Voice response provided to the senior
[2137] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2138] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2139] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2140] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2141] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles...
Claims
1. means for storing the information entered by the user in a database; A means for generating individually customized generative AI models and speech recognition models based on the stored information; means for a user to add voice input and transmit the voice data to a server; A means for analyzing the received voice data by the server and retraining the generative AI model and the voice recognition model; A means for converting the elderly person's voice input into text data and sending it to a server; A means for the server to generate an appropriate response using a generation AI and transmit it to the terminal; A means for converting the received response text into audio and playing it back to the elderly person; A system including:
2. 2. The system according to claim 1, further comprising means for providing an interactive tutorial mode for elderly people to learn IT operations.
3. The system of claim 1 further comprising means for reserving transportation services available to the elderly and adjusting transportation schedules.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A