System
The generative AI Elderly Companion system addresses the challenges of loneliness and emergency response for the elderly by managing personal information, converting voice input to text, providing reminders, and facilitating emergency contacts, enhancing daily life safety and fulfillment.
Patent Information
- Application Number
- JP2024115247
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Elderly individuals face challenges with loneliness, difficulty managing their health, and rapid response in emergencies due to the lack of reminder functions and dialogue interactions in daily life, leading to issues such as forgetting medication or missing doctor's appointments, and inadequate support during emergencies.
A system that manages and stores user basic information and individual needs, converts voice input into text data, displays or plays back response data, provides reminders, and sends emergency contact requests to a server for notification, utilizing a generative AI Elderly Companion system with features like speech recognition, reminder functions, and emergency contact capabilities.
The system supports elderly individuals by enabling intuitive operation, personalized information provision, and emergency contact functions, reducing loneliness and ensuring a safe and fulfilling daily life.
Smart Images

Figure 2026014250000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Elderly people face significant challenges, including loneliness, difficulty managing their health, and rapid response in emergencies. In particular, the lack of reminder functions and dialogue interactions in daily life can lead to forgetting to take medication or miss a doctor's appointment. Another serious problem is the lack of prompt and appropriate support in the event of an emergency. [Means for solving the problem]
[0005] The present invention solves the above problems by providing a system including: means for managing and saving a user's basic information and individual needs; means for converting a user's voice input into text data and sending it to a server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information in the server and displaying a notification to the user at a specified time; and means for sending an emergency contact request to the server, which then sends an emergency notification to a pre-set emergency contact.
[0006] "Basic user information" refers to personal identification information and information related to the life of the elderly person using the system, such as their name, age, address, health status, and emergency contact information.
[0007] "Individual needs" refers to each user's preferences, lifestyle, and health care requirements and desires.
[0008] "Means for management and storage" refers to a mechanism that provides the functionality to store basic information and individual needs of users in a database and enable them to refer to and update the information as needed.
[0009] "Voice input" refers to a method in which a user inputs information or instructions to a system by speaking aloud.
[0010] "Convert to text data" refers to the process of analyzing speech input and converting it into corresponding text data.
[0011] "Means of sending to the server" refers to the method or technology for sending the converted text data to the server via a network.
[0012] "Response data" refers to the response information returned in response to a request from a server.
[0013] "Display or playback means" refers to a display device or audio output device for visually or audibly notifying the user of the response data.
[0014] "Reminder information" refers to information that should be notified to the user at a specific time or date, such as when to take medicine or when to make a doctor's appointment.
[0015] "Means for displaying a notification to the user" refers to a technology that notifies the user of reminder information at a predetermined time by a pop-up message or voice message.
[0016] An "emergency contact request" refers to a signal or message that a user sends to request help when they are faced with an emergency situation.
[0017] "Emergency contacts" refers to contact information for family members or caregivers who should be contacted in an emergency.
[0018] "Means for sending emergency notifications" refers to the technology or method for quickly contacting registered emergency contacts when an emergency contact request occurs. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The "Generative AI Elderly Companion" system of this invention has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back the response data from the server to the user via their device. This system can also support the lives of the elderly through reminder and emergency contact functions.
[0041] Manage and store basic information and needs
[0042] server:
[0043] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[0044] Device:
[0045] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[0046] Converts voice input into text and sends it to the server
[0047] Device:
[0048] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[0049] server:
[0050] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[0051] Receiving and displaying response data
[0052] Device:
[0053] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[0054] Reminder function
[0055] server:
[0056] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[0057] Device:
[0058] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[0059] Emergency contact function
[0060] Device:
[0061] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[0062] server:
[0063] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[0064] Specific use cases
[0065] Usage example 1:
[0066] The user asks, "When is my next hospital appointment?" The device recognizes the speech and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "My next hospital appointment is tomorrow at 10:00." The device plays this back as audio.
[0067] Usage example 2:
[0068] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately."
[0069] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and enable them to live a safe and fulfilling daily life.
[0070] The processing flow will be explained below.
[0071] Step 1:
[0072] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[0073] Step 2:
[0074] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[0075] Step 3:
[0076] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[0077] Step 4:
[0078] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[0079] Step 5:
[0080] The device receives the voice input, converts it into text data using a speech recognition engine, and sends the converted text data to the server.
[0081] Step 6:
[0082] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[0083] Step 7:
[0084] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[0085] Step 8:
[0086] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[0087] Step 9:
[0088] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker.
[0089] Step 10:
[0090] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[0091] Step 11:
[0092] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[0093] Step 12:
[0094] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[0095] Step 13:
[0096] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[0097] Step 14:
[0098] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[0099] Step 15:
[0100] The user presses the emergency button. The emergency button on the device is operated.
[0101] Step 16:
[0102] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[0103] Step 17:
[0104] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[0105] The above are the specific processing steps for carrying out the invention.
[0106] Example 1
[0107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0108] To help seniors live safe and fulfilling daily lives without feeling lonely, systems with personalized information provision, reminder functions, and emergency contact functions are needed. However, existing systems are often difficult for seniors to use, and information provision is fragmented and unintegrated. Furthermore, the accuracy of voice input and dialogue interaction is low, and external information acquisition is often difficult. This makes it difficult for seniors to receive appropriate support.
[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0110] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for the server to obtain information from an external information service, and means for converting voice into text data using a voice recognition engine. This allows elderly people to intuitively operate the device, smoothly utilize personalized information provision, reminder functions, and emergency contact functions, and also makes it easy to obtain external information, enabling them to live a safe and fulfilling life.
[0111] "User's basic information and individual needs" refers collectively to individually personalized data such as the user's name, age, address, health status, emergency contact information, medication times, and schedule for regular health checks.
[0112] "Means for converting voice input into text data" is a general term for devices and software that use voice recognition technology to convert a user's voice into text data and send that text data to a server.
[0113] "Response data from the server" refers to information and messages analyzed and generated by the server, including data obtained from external information services.
[0114] "Reminder information" refers to notification content that a user should receive at a specific time or date, such as when to take medicine or a schedule for regular health checks.
[0115] An "emergency contact request" is a request made by a user to quickly seek assistance in an emergency, and is an emergency signal sent from the terminal to the server.
[0116] "Means of obtaining information from external information services" is a general term for the technologies and methods that allow a server to retrieve necessary information from public APIs and databases on the Internet.
[0117] A "speech recognition engine" is a piece of software or hardware that understands a user's voice input and converts it into text data.
[0118] "Dialogue interaction" refers to two-way communication between a user and a system based on voice input and text data.
[0119] This invention relates to a generative AI "Elderly Companion" system that supports the lives of the elderly, and has a series of functions to manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. This system also supports the lives of the elderly through reminder functions and emergency contact functions.
[0120] In this embodiment of the present invention, the system is configured by a server, terminals, and users communicating with each other. The server is hosted on the cloud and manages user information through a database. Specifically, a MySQL database is used, and data security is ensured by SSL encryption.
[0121] Manage and store basic information and needs
[0122] Device:
[0123] Users enter basic information such as their name, age, address, health condition, and emergency contact details, as well as their individual needs such as medication times and schedules for regular health checks, through the device's interface. Either voice input or text input is possible, using a smartphone or a dedicated device.
[0124] server:
[0125] The server receives the basic information and needs information sent from the device and stores it in a MySQL database. The server uses this data to provide personalized support to the user.
[0126] Converts voice input into text and sends it to the server
[0127] Device:
[0128] When a user asks a question or gives an instruction by voice, such as "When is my next doctor's appointment?", the device converts the voice into text data using Google's speech recognition API. The converted text data is then sent to the server. The user can use a headset or built-in microphone to do this.
[0129] server:
[0130] The server analyzes the received text data and obtains the necessary information (for example, reservation information or weather information). To obtain the weather information, it uses a public API such as the OpenWeatherMap API. The server then sends the analysis results to the terminal as response data.
[0131] Receiving and displaying response data
[0132] Device:
[0133] The device that receives the response data from the server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text data into speech and tells the user, "Your next hospital appointment is tomorrow at 10:00." If the device has a display, it can also display the information on the screen.
[0134] Reminder function
[0135] server:
[0136] Reminder information is stored on the server based on a schedule specified by the user, and the server sends the reminder information to the device at the specified time.
[0137] Device:
[0138] The device will receive a reminder notification at the specified time, informing the user via voice or pop-up notification, "It's 8 o'clock. Time to take your medicine."
[0139] Emergency contact function
[0140] Device:
[0141] When a user faces an emergency, they press the emergency button on their device, which instantly sends an emergency signal to the server.
[0142] server:
[0143] When the server receives an emergency signal, it sends a message to pre-registered emergency contacts (family members, caregivers) via SMS, phone, or email saying, "The user has pressed the emergency button. Please check immediately."
[0144] Specific examples of actions and prompts
[0145] Usage example 1:
[0146] The user asks, "What's the weather like today?" The device converts the speech to text and sends it to the server. The server retrieves weather information using a weather information service, and then sends the information to the device: "Today's weather is sunny. The temperature is 24 degrees."
[0147] Usage example 2:
[0148] The user asks, "When is my next hospital appointment?" The device converts the speech to text and sends it to the server. The server retrieves the appointment information from its database and responds, "My next hospital appointment is tomorrow at 10:00." The device plays this back aloud.
[0149] This invention enables elderly people to easily access personalized information, reminder functions, and emergency contact functions through an interface that they can operate intuitively. This system reduces the sense of loneliness felt by elderly people and supports safe and fulfilling lives.
[0150] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0151] Step 1:
[0152] The user enters basic information and individual needs into the device, such as name, age, address, health condition, emergency contact information, medication times, and schedule for regular health checks, by voice or text. This information is temporarily stored on the device and then sent to the server.
[0153] Input: Name, age, address, health status, emergency contact, needs
[0154] Output: Basic information and needs sent to the server
[0155] Step 2:
[0156] The device receives voice input and converts it into text data. For example, if a user says, "What's the weather like today?", the device converts the voice into text using a speech recognition engine (Google Speech-to-Text API). The text data is then sent to the server.
[0157] Input: Voice input (e.g. "What's the weather like today?")
[0158] Output: Text data (e.g. "What's the weather like today?")
[0159] Step 3:
[0160] The server analyzes the received text data and obtains the corresponding information. For example, if the text data "What's the weather like today?" is received, the server will call an external weather information API (e.g., OpenWeatherMap API) to obtain the weather information and obtain the results.
[0161] Input: Text data (e.g., "What's the weather like today?")
[0162] Output: Weather information (e.g. sunny, temperature 24 degrees)
[0163] Step 4:
[0164] The server sends the acquired information to the terminal as text data. The server generates response data based on the analysis results and sends it to the terminal.
[0165] Input: Weather information (e.g. sunny, temperature 24 degrees)
[0166] Output: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[0167] Step 5:
[0168] The device will play back the response data received from the server as audio. The device will convert the text data into audio using a speech synthesis engine (Google Text-to-Speech API) and convey it to the user. If the device has a display, it can also display it.
[0169] Input: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[0170] Output: Voice playback (e.g. "Today's weather is sunny. The temperature is 24 degrees.") and display
[0171] Step 6:
[0172] The device receives a reminder notification at the specified time and notifies the user. The server sends pre-set reminder information to the device at the specified time. The device notifies the user with a voice or pop-up notification such as, "It's 8 o'clock. It's time to take your medicine."
[0173] Input: Reminder information (e.g., every day at 8:00)
[0174] Output: Audio and popup notification (e.g. "It's 8 o'clock. Time to take your medicine.")
[0175] Step 7:
[0176] When a user presses the emergency button, the device sends an emergency signal to the server. When the device sends the emergency signal, the server notifies the emergency contact. The emergency contact is sent a message via SMS, phone call, or email saying, "The user has pressed the emergency button. Please check immediately."
[0177] Input: Press emergency button
[0178] Output: Emergency notification message (e.g. "The user pressed the emergency button. Please check immediately.")
[0179] The above is the specific flow of processing steps for the "Generative AI Elderly Companion" system.
[0180] (Application example 1)
[0181] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0182] There is a need for support systems that can reduce the difficulty elderly people have in finding products in physical stores and enable them to shop safely and efficiently. Conventional support systems have difficulty providing the information elderly people need quickly, and do not have sufficient reminder or emergency contact functions.
[0183] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0184] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for providing product location information to support the elderly in a physical store, and means for notifying the user of reminders set in the physical store, thereby reducing the difficulty for the elderly when searching for products in a physical store and enabling them to shop safely and efficiently.
[0185] "Basic information" refers to individual information such as the user's name, age, address, health status, and emergency contact information.
[0186] "Individual needs" are information about a user's specific requirements or preferences, such as when to take medication or when to schedule regular health checks.
[0187] "Voice input" refers to data input to the system by a user speaking.
[0188] "Text data" is character information generated by analyzing voice input.
[0189] A "server" is a computer system that manages and processes data sent by users.
[0190] "Response data" is reply data that the server generates and sends in response to a user request.
[0191] "Reminder information" is information for notifying the user at a specified time.
[0192] "Emergency contacts" are information about people or organizations that the user wants to contact in the event of an emergency.
[0193] An "emergency notification" is a message sent from the server when an emergency occurs.
[0194] A "brick and mortar store" is a sales or service establishment located in a physical location.
[0195] "Product location information" refers to information about the location where a specific product is located within a physical store.
[0196] A "generative AI model" is an artificial intelligence algorithm model that generates text data in response to user requests.
[0197] MODE FOR CARRYING OUT THE INVENTION
[0198] The following describes an embodiment of the present invention.
[0199] System Program Overview
[0200] The "Generative AI Elderly Companion" system, which is the subject of this invention, has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. Among the functions of this system, we will explain the important functions that support the elderly in brick-and-mortar stores in particular.
[0201] Hardware and software used
[0202] Smartphone (iOS, Android): A device that allows users to input and receive information.
[0203] In-store robots (e.g., Pepper): Provide an interface for users to obtain and receive information within the store.
[0204] Server: Manages user information, performs voice recognition, and trains and updates generative AI models.
[0205] Speech recognition engine (Google Cloud Speech-to-Text): Converts voice input into text data.
[0206] Generative AI model (GPT-4 by OpenAI): Analyzes and generates text data based on user requests.
[0207] Database (MySQL): Stores user basic information, individual needs, reminder information, emergency contacts, etc.
[0208] Weather API (OpenWeatherMap): Used to retrieve the weather information desired by the user.
[0209] System procedures and operations
[0210] Manage and store basic information and needs
[0211] The server stores basic information and individual needs entered by the user through their device (smartphone or robot) in a MySQL database. Users can enter information by voice or text.
[0212] Converts voice input into text and sends it to the server
[0213] The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the user's voice input into text data and send it to the server.
[0214] Receiving and displaying response data
[0215] The server analyzes the received text data using a generative AI model (GPT-4) and generates appropriate response data. For example, when a user asks, "Where is this product?", the server obtains the product's location information and sends it to the device, saying, "Product A is in aisle 3." The device then provides this to the user via voice or a screen display.
[0216] Reminder function
[0217] The reminder information set by the user is stored on the server, and when the specified time comes, the server sends the reminder information to the device and generates a notification, which then notifies the device with a voice or a pop-up message saying, "It's time to take your medicine."
[0218] Emergency contact function
[0219] In the event of an emergency, the user presses the emergency button on the device, and the device sends an emergency signal to the server, which then sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[0220] Examples of specific examples and prompts
[0221] Specific examples
[0222] For example, when a user asks a smartphone or robot, "Where is this product?", the speech recognition engine (Google Cloud Speech-to-Text) converts the speech into text and sends it to the server. The server then analyzes it using a generative AI model (GPT-4), obtains the product location information, and generates a response such as "Product A is in aisle 3," which the smartphone or robot then plays back.
[0223] Prompt Sentence Examples
[0224] User: "Where is this item?"
[0225] System: "Converts speech to text and sends it to the server..."
[0226] Generative AI model: "Analyzing product information for user search..."
[0227] Response: "Item A is in aisle 3"
[0228] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to enjoy shopping in brick-and-mortar stores more comfortably and safely.
[0229] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0230] Step 1: Manage and save your basic information and needs
[0231] Users use their smartphones or in-store robots to input their basic information (name, age, address, health condition, emergency contact information) and individual needs (medicine dosing times, regular health check schedules) by voice or text. This input data is converted into text by a speech recognition engine (Google Cloud Speech-to-Text) on the device. The converted text data is sent to the server and stored in a MySQL database.
[0232] Input: User voice or text input
[0233] Data processing: Converting voice to text
[0234] Output: Send and save text data to the server
[0235] Step 2: Convert speech to text and send to server
[0236] The user asks a question by voice, such as "Where is this item?" The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the voice into text data and sends the text data to the server. The server receives this text data.
[0237] Input: User voice input
[0238] Data processing: Converting voice to text
[0239] Output: Send text data to the server
[0240] Step 3: Parsing text data and generating a response
[0241] The server uses a generative AI model (GPT-4) to analyze the received text data and generate an appropriate response. For example, it retrieves product location information from a database based on the user's question and generates a response such as "Product A is in aisle 3."
[0242] Input: Text data
[0243] Data computation: Analysis using generative AI models
[0244] Output: Generated response data
[0245] Step 4: Send and display response data
[0246] The server sends the generated response data to the terminal. The terminal receives this response data and uses a speech synthesis engine to play it back as a voice such as "Product A is in aisle 3," and also displays it on the screen.
[0247] Input: Response data
[0248] Data processing: voice synthesis and screen display
[0249] Output: Speech and text display
[0250] Step 5: Execute the reminder function
[0251] The server sends the reminder information set by the user to the device at the specified time. The device receives the reminder notification and notifies the user by voice, saying "It's time to take your medicine." A pop-up notification is also displayed.
[0252] Input: Reminder information and designated time
[0253] Data Processing: Notification Generation
[0254] Output: Sound and popup notification
[0255] Step 6: Implementing emergency contact functions
[0256] When a user presses the emergency button, the device sends an emergency signal to the server, which receives the signal and sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[0257] Input: Emergency signal
[0258] Data Processing: Emergency notification generation and transmission
[0259] Output:Notify emergency contacts
[0260] The above processing steps enable users to quickly obtain the information they need in a physical store, allowing them to enjoy shopping safely and comfortably.
[0261] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0262] The "Generative AI Elderly Companion" system, which is the subject of this invention, manages and stores the user's basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, making conversational interactions more personalized. It also supports the daily lives of the elderly through reminder and emergency contact functions.
[0263] Manage and store basic information and needs
[0264] server:
[0265] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[0266] Device:
[0267] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[0268] Converts voice input into text and sends it to the server
[0269] Device:
[0270] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[0271] server:
[0272] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[0273] Receiving and displaying response data
[0274] Device:
[0275] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[0276] Reminder function
[0277] server:
[0278] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[0279] Device:
[0280] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[0281] Emergency contact function
[0282] Device:
[0283] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[0284] server:
[0285] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[0286] Emotion recognition function
[0287] Device:
[0288] The device sends the user's voice input to the emotion engine, which analyzes the tone, speed, volume, etc. of the voice to determine the user's emotional state.
[0289] server:
[0290] The server analyzes the emotional information received from the emotion engine and generates an appropriate response based on it. For example, if a user says in a sad voice, "I'm not feeling well today," the server will generate an empathetic response such as, "What's wrong? Tell me your story."
[0291] Device:
[0292] Once the response data based on emotion recognition is sent from the server, the device displays or plays it back to the user, making interactions with the user more natural and personalized.
[0293] Specific use cases
[0294] Usage example 1:
[0295] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[0296] Usage example 2:
[0297] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[0298] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and lead a safe and fulfilling daily life. By incorporating an emotion engine, the system will be able to respond more sensitively to the user's feelings, providing a higher level of satisfaction.
[0299] The processing flow will be explained below.
[0300] Step 1:
[0301] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[0302] Step 2:
[0303] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[0304] Step 3:
[0305] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[0306] Step 4:
[0307] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[0308] Step 5:
[0309] The device receives voice input, converts the voice into text data using a speech recognition engine, and sends the converted text data to the server.
[0310] Step 6:
[0311] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[0312] Step 7:
[0313] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[0314] Step 8:
[0315] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[0316] Step 9:
[0317] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker. The device plays back "Today's weather is sunny. The temperature is 24 degrees."
[0318] Step 10:
[0319] The emotion engine analyzes the tone, rate, and volume of the user's voice to identify the user's emotional state: if the user speaks in an anxious voice, it will be identified as in an "anxious" state.
[0320] Step 11:
[0321] The server analyzes the emotional information received from the emotion engine and generates an appropriate response. For example, if a user says in a sad voice, "I'm not feeling well today," the server generates a response that shows empathy, such as, "What's wrong? Tell me your story."
[0322] Step 12:
[0323] The response data containing the generated emotional response is sent from the server to the device, which then plays it back to the user, saying, "What's wrong? Tell me what you think."
[0324] Step 13:
[0325] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[0326] Step 14:
[0327] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[0328] Step 15:
[0329] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[0330] Step 16:
[0331] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[0332] Step 17:
[0333] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[0334] Step 18:
[0335] The user presses the emergency button. The emergency button on the device is operated.
[0336] Step 19:
[0337] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[0338] Step 20:
[0339] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[0340] The above are the specific processing steps for carrying out the invention.
[0341] Example 2
[0342] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0343] Elderly people require various forms of support in their daily lives, but current technology makes it difficult to provide personalized assistance that fully reflects the individual needs and emotional state of each elderly person. Furthermore, systems for responding quickly and appropriately in emergencies are inadequate. Therefore, there is a need to create an environment where elderly people can live safely and without feeling lonely.
[0344] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0345] In this invention, the server includes: means for managing and saving basic information and individual needs of a user; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying or playing a notification to the user at a specified time; means for sending an emergency contact request to the server and for the server to send an emergency notification to a pre-set emergency contact; means for analyzing the tone, speed and volume of the user's voice to identify the user's emotional state; and means for generating a personalized response based on the emotional state and providing it to the user.
[0346] This will enable personalized assistance to be provided in response to the individual needs and emotional state of the elderly, as well as a rapid and appropriate response in emergencies.
[0347] "Basic user information" refers to basic information about a user, such as name, age, address, health status, and emergency contact information.
[0348] "Individual needs" refers to information about the support and services that a user individually needs, such as when to take medication or schedule regular health checks.
[0349] "Voice input" refers to voice data input when a user speaks to the system through a microphone.
[0350] "Text data" refers to voice input converted into character string data using a voice recognition engine or the like.
[0351] "Server" refers to a computer system that receives data from users and stores, analyzes, and processes it.
[0352] "Response data" refers to information that the server generates based on a user request and returns to the user.
[0353] "Reminder information" refers to information for notifying the user at a specific time or timing.
[0354] "Emergency contact request" refers to an emergency signal sent through the system when a user faces an emergency.
[0355] "Emergency notification" refers to an emergency message sent from the server to a pre-defined emergency contact.
[0356] "Emotional state" refers to the user's emotions estimated based on an analysis of the tone, rate, volume, etc. of the voice.
[0357] "Personalized response" refers to an individualized response that is generated based on the user's basic information, individual needs, and emotional state.
[0358] The present invention, the "Generative AI Elderly Companion" system, provides personalized support that reflects the individual needs and emotional state of elderly people, reducing feelings of loneliness in daily life and creating a safe living environment.
[0359] Hardware and software used
[0360] The server is a high-performance computer system, such as an AWS EC2 instance or Microsoft Azure VM. MySQL or PostgreSQL is used as the database management system. Google Cloud Speech-to-Text or Amazon Transcribe is used as the speech recognition engine, and IBM Watson Tone Analyzer or Microsoft Azure Cognitive Services is used as the emotion recognition engine.
[0361] Manage and store basic information and needs
[0362] The user uses the device to enter basic information (name, age, address, health status, emergency contact information). The device provides voice or text input, verifies the entered information, and sends it to the server. The server stores the received basic information in a database and provides personalized support based on this information.
[0363] Converts voice input into text and sends it to the server
[0364] When a user asks "What's the weather like today?", the device uses a voice recognition engine to convert the voice into text data and sends it to the server. The server then analyzes the received text data and obtains weather information using a weather API provided by a third party.
[0365] Receiving and displaying response data
[0366] The response data (such as weather information) sent from the server is received by the terminal, which then provides it to the user through voice or a display device. For example, it may say, "Today's weather is sunny. The temperature is 24 degrees."
[0367] Reminder function
[0368] When a user instructs the device to "set a reminder to take my medicine at 8 o'clock every morning," the device converts this instruction into text data and sends it to the server. The server stores the reminder information in a database and sends a reminder notification to the device at the specified time (8 o'clock every morning). The device then notifies the user by voice or a pop-up notification, saying, "It's 8 o'clock. It's time to take your medicine."
[0369] Emergency contact function
[0370] In an emergency, the user presses the emergency button on the device. The device then sends an emergency signal to the server. The server then contacts pre-registered emergency contacts (family members or caregivers) via SMS, phone, email, etc. It then sends a message saying, "The user has pressed the emergency button. Please check immediately."
[0371] Emotion recognition function
[0372] If a user says "I'm feeling bad today" in a sad voice, the device sends this voice data to an emotion recognition engine, which analyzes the tone, speed, and volume of the voice to identify the emotional state. The server uses the emotional information to generate a personalized response that shows empathy, such as "What's wrong? Tell me your story," and the device plays this aloud to the user.
[0373] Specific use cases
[0374] Usage example 1:
[0375] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[0376] Usage example 2:
[0377] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[0378] The above is an embodiment of the present invention. This allows elderly people to live a safe and fulfilling daily life and reduce their sense of loneliness. Furthermore, the introduction of an emotion engine enables the system to respond in a way that is sensitive to the user's feelings, providing a higher level of satisfaction.
[0379] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0380] Steps for entering and saving basic information
[0381] Step 1:
[0382] The user uses the device to input basic information (name, age, address, health status, emergency contact information). The input method can be selected as voice input or text input. For example, if the user inputs "My name is Ichiro Tanaka, and I'm 75 years old" by voice, the device will capture the voice data.
[0383] input:
[0384] Basic information entered by the user (name, age, address, health status, emergency contact information)
[0385] output:
[0386] Audio or text data
[0387] Step 2:
[0388] When the device receives voice input, it uses a speech recognition engine (such as Google Cloud Speech-to-Text) to convert the voice into text data. For example, voice data such as "My name is Tanaka Ichiro and I'm 75 years old" is converted into text data such as "Name: Tanaka Ichiro, Age: 75."
[0389] input:
[0390] Audio data
[0391] output:
[0392] Text data
[0393] Step 3:
[0394] The device sends the converted text data to the server. Specifically, data transmission is performed using an API request.
[0395] input:
[0396] Text data
[0397] output:
[0398] API request (text data)
[0399] Step 4:
[0400] The server stores the received text data in a database, for example, using MySQL or PostgreSQL as data records.
[0401] input:
[0402] Text data
[0403] output:
[0404] Data Records
[0405] Processing steps for converting voice input to text and sending it to the server
[0406] Step 1:
[0407] The user asks a question by voice, "What's the weather like today?" This voice data is acquired by the terminal.
[0408] input:
[0409] User voice input
[0410] output:
[0411] Audio data
[0412] Step 2:
[0413] The device uses a voice recognition engine to convert the acquired voice data into text data. Specifically, the voice data "What's the weather like today?" is converted into text data "What's the weather like today?"
[0414] input:
[0415] Audio data
[0416] output:
[0417] Text data
[0418] Step 3:
[0419] The device sends the converted text data to the server via an API request.
[0420] input:
[0421] Text data
[0422] output:
[0423] API request (text data)
[0424] Step 4:
[0425] The server analyzes the received text data and retrieves the corresponding information. For example, to retrieve weather information, the server sends a request to a weather API provided by a third party and receives response data such as "Today's weather is sunny and the temperature is 24 degrees."
[0426] input:
[0427] Text data
[0428] output:
[0429] Response data (weather information)
[0430] Processing steps for receiving and displaying response data
[0431] Step 1:
[0432] The server sends the acquired response data (such as weather information) to the terminal.
[0433] input:
[0434] Response data (weather information)
[0435] output:
[0436] API response (response data)
[0437] Step 2:
[0438] The device analyzes the received response data and provides it to the user through a voice output device, for example, by playing back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[0439] input:
[0440] API response (response data)
[0441] output:
[0442] Audio data
[0443] Reminder function processing steps
[0444] Step 1:
[0445] The user instructs, "Set a reminder to take my medicine at 8 o'clock every morning." The device receives this voice data.
[0446] input:
[0447] User voice instructions
[0448] output:
[0449] Audio data
[0450] Step 2:
[0451] The device uses a speech recognition engine to convert voice data into text data. For example, voice data such as "Set a reminder to take my medicine at 8 o'clock every morning" is converted into text data such as "Take my medicine at 8 o'clock every morning."
[0452] input:
[0453] Audio data
[0454] output:
[0455] Text data
[0456] Step 3:
[0457] The device sends the converted text data to the server via an API request.
[0458] input:
[0459] Text data
[0460] output:
[0461] API request (text data)
[0462] Step 4:
[0463] The server stores the received reminder information in a database and configures it to generate notifications at a specified time, for example, every morning at 8:00.
[0464] input:
[0465] Text data
[0466] output:
[0467] Database records and timer settings
[0468] Step 5:
[0469] When the specified time arrives, the server sends a notification to the device, such as "It's time to take your medicine."
[0470] input:
[0471] Timer Event
[0472] output:
[0473] API response (notification data)
[0474] Step 6:
[0475] The device will then provide the received notification data to the user via voice or pop-up notification, for example, "It's 8 o'clock. Time to take your medicine."
[0476] input:
[0477] API response (notification data)
[0478] output:
[0479] Audio data or popup notification
[0480] Emergency Contact Function Processing Steps
[0481] Step 1:
[0482] The user presses the emergency button, and the device detects this emergency signal.
[0483] input:
[0484] Emergency button input
[0485] output:
[0486] Emergency Signal Data
[0487] Step 2:
[0488] The device sends an emergency signal to the server, which is done via an API request.
[0489] input:
[0490] Emergency Signal Data
[0491] output:
[0492] API Request (Emergency Signal Data)
[0493] Step 3:
[0494] The server analyzes the received emergency signal and sends an emergency notification to pre-defined emergency contacts, for example, by SMS, phone call, or email, with a message saying, "The user has pressed the emergency button. Please check immediately."
[0495] input:
[0496] API Request (Emergency Signal Data)
[0497] output:
[0498] Emergency notification data (SMS, phone, email)
[0499] Emotion Recognition Processing Steps
[0500] Step 1:
[0501] The user inputs "I feel bad today" in a sad voice. This voice data is acquired by the terminal.
[0502] input:
[0503] User voice input
[0504] output:
[0505] Audio data
[0506] Step 2:
[0507] The device sends the captured voice data to an emotion recognition engine, which analyzes the tone, rate, and volume of the voice to identify the emotional state.
[0508] input:
[0509] Audio data
[0510] output:
[0511] Emotional state data
[0512] Step 3:
[0513] The server generates a personalized response based on the emotional state data received from the emotion recognition engine, for example, an empathetic response such as "What's wrong? Tell me your story."
[0514] input:
[0515] Emotional state data
[0516] output:
[0517] Response data
[0518] Step 4:
[0519] The server transmits the generated response data to the terminal.
[0520] input:
[0521] Response data
[0522] output:
[0523] API response (response data)
[0524] Step 5:
[0525] The device will then play back the received response data in voice, for example, "What's wrong? Tell me what happened."
[0526] input:
[0527] API response (response data)
[0528] output:
[0529] Audio data
[0530] (Application example 2)
[0531] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0532] Elderly people have greater difficulty navigating physical stores, searching for products, and responding to emergencies. They also often feel anxious about asking store staff questions and navigating the store. Current systems struggle to adequately address these issues, preventing seniors from enjoying shopping with peace of mind.
[0533] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for managing and saving a user's basic information and individual needs; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying a notification to the user at a specified time; means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact; means for navigating the store using the user's current location and store map data; and means for analyzing the user's voice input and generating a personalized response based on the user's emotional state. This enables seniors to move confidently through physical stores, easily find products and services, and quickly respond to anxieties or emergencies.
[0534] Below are definitions of important terms included in the patent claims, rewritten to suit the application example.
[0535] "Means for managing and storing a user's basic information and individual needs" refers to a means for storing basic information such as the user's name, age, address, health condition, and emergency contact information, as well as individual needs such as medication times and regular health check schedules, in a database, and for managing and updating this information as needed.
[0536] "Means for converting user's voice input into text data and sending it to the server" refers to means for converting information input by the user by voice into text data using voice recognition technology and sending that text data to the server.
[0537] "Means for receiving response data from the server and displaying or playing it back to the user" refers to means for receiving information sent from the server (e.g., weather information or store directions) and providing it to the user visually or audibly.
[0538] The "means for saving specified reminder information on a server and displaying a notification to the user at a specified time" refers to a means for saving reminder information set by a user on a server and sending a notification to the user at a specified time based on that reminder information.
[0539] "Means for sending an emergency contact request to a server, and for the server to send an emergency notification to pre-registered emergency contacts" refers to a means for a user to send an emergency contact request to a server by pressing an emergency button, etc., and for the server to send an emergency notification by phone or SMS to pre-registered emergency contacts based on that request.
[0540] "Means for navigating within a store using the user's current location and store map data" refers to a means for guiding the user to their desired location by using the location information of the user's smartphone or device and combining it with map data within the store.
[0541] "Means for analyzing a user's voice input and generating a personalized response based on the user's emotional state" refers to means for analyzing the tone, rate, and content of a user's voice to identify the user's emotional state, and generating and providing an appropriate response or message accordingly.
[0542] System Program
[0543] The purpose of this invention, the "Generative AI Elderly Companion Shopper" system, is to help seniors shop comfortably in brick-and-mortar stores. This system manages and stores users' basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. It also supports the daily lives of seniors through reminder functions, emergency contact functions, and emotion recognition functions.
[0544] Hardware and Software Use
[0545] Hardware:
[0546] Smartphone (with camera, microphone, and speaker)
[0547] software:
[0548] Google Cloud Speech-to-Text API (voice recognition)
[0549] Google Cloud Natural Language API (Text Analysis)
[0550] Twilio API (emergency contact function)
[0551] Emotion Engine (e.g. Affectiva's SDK)
[0552] Firebase (database)
[0553] Data processing and calculation
[0554] Managing and storing your basic information and individual needs:
[0555] The server stores the basic information entered by the user (name, age, address, health status, emergency contacts) and individual needs (medication times, schedule for regular health checks) in a Firebase database, allowing the system to provide personalized support to the user.
[0556] Convert speech to text and send to server:
[0557] When a user asks, "Where is the restroom?", the device uses the Google Cloud Speech-to-Text API to convert speech to text data and send it to the server, which then uses the Google Cloud Natural Language API to parse the text and generate an appropriate response.
[0558] Receive and display response data:
[0559] The response data sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may provide voice guidance such as, "To find the restroom, go straight down the corridor on the right."
[0560] Reminder function:
[0561] The server stores the reminder information set by the user (e.g., take medicine at 8 o'clock every day) in Firebase. When the specified time arrives, the device sends a voice or pop-up notification saying, "It's 8 o'clock. Time to take your medicine."
[0562] Emergency contact features:
[0563] If a user is about to fall inside the store, they can press the emergency button on their device, which sends an emergency signal to the server. The server then uses the Twilio API to send a notification to registered emergency contacts (family members or caregivers). The emergency notification is sent via SMS or phone.
[0564] Emotion recognition function:
[0565] If a user says, "I'm not feeling well today," the device uses its emotion engine to analyze the speech and identify the user's emotional state. The server then generates an appropriate response based on this, saying, "What's wrong? Tell me."
[0566] Examples of specific examples and prompts
[0567] As a specific example, consider the case where an elderly person named Mr. A gets lost in a store.
[0568] Person A asks the app, "Where is the restroom?" inside the store. Voice recognition is activated, and the speech is converted into text and sent to the server.
[0569] The server identifies the location of the restroom from the store's map data and generates a navigation plan along with Mr. A's current location.
[0570] The app will provide voice guidance saying, "To get to the restroom, go straight down the corridor on the right."
[0571] If the emotion engine detects Mr. A's anxiety, the app will ask, "If you can't find it, I'll call a store employee."
[0572] Example prompt sentence:
[0573] The user asks the smartphone microphone, "Where is the restroom?" The voice input is converted into text and sent to the server. The system should refer to the user's current location and a map of the store, calculate the shortest route, and provide voice guidance. If the user asks in an anxious voice, the system should generate an additional response asking, "If you can't find it, would you like me to call a store employee?"
[0574] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0575] Step 1:
[0576] A user asks into the microphone of their smartphone, "Where is the restroom?" A voice input occurs, and the device collects this voice input data.
[0577] Input: User voice input
[0578] Output: Audio data
[0579] Step 2:
[0580] The device uses the Google Cloud Speech-to-Text API to convert the voice data into text data, which is then sent to the server.
[0581] Input: Audio data
[0582] Output: Text data
[0583] Step 3:
[0584] The server analyzes the received text data using the Google Cloud Natural Language API, and as a result, it understands that the user is looking for a restroom.
[0585] Input: Text data
[0586] Output: Analysis result (intent)
[0587] Step 4:
[0588] The server identifies the user's location by referencing the user's current location and the store's map data, and calculates the shortest route based on this.
[0589] Input: User's current location data, store map data
[0590] Output: Shortest route data
[0591] Step 5:
[0592] The server uses the calculated shortest route data to generate a navigation plan to help the user reach their destination easily. The generated navigation plan is created in voice and text format.
[0593] Input: Shortest route data
[0594] Output: Navigation plan (voice and text)
[0595] Step 6:
[0596] The server sends the generated navigation plan to the device, which receives it and provides voice and text guidance to the user.
[0597] Input: Navigation plan
[0598] Output: User instructions
[0599] Step 7:
[0600] At the same time, the device analyzes the user's voice input with an emotion engine to identify their emotional state. If the device detects that the user is asking a question in an anxious voice, the server generates an additional response (e.g., "If you can't find it, should I call a store clerk?").
[0601] Input: User's voice data
[0602] Output: Emotional state, additional responses
[0603] Step 8:
[0604] As additional responses are generated, the server sends them to the terminal, which displays or plays them audibly to the user.
[0605] Input: Additional response data
[0606] Output: Additional instructions for the user
[0607] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0608] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0609] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0610] [Second embodiment]
[0611] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0612] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0613] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0614] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0615] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0616] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0617] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0618] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0619] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0620] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0621] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0622] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0623] The "Generative AI Elderly Companion" system of this invention has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back the response data from the server to the user via their device. This system can also support the lives of the elderly through reminder and emergency contact functions.
[0624] Manage and store basic information and needs
[0625] server:
[0626] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[0627] Device:
[0628] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[0629] Converts voice input into text and sends it to the server
[0630] Device:
[0631] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[0632] server:
[0633] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[0634] Receiving and displaying response data
[0635] Device:
[0636] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[0637] Reminder function
[0638] server:
[0639] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[0640] Device:
[0641] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[0642] Emergency contact function
[0643] Device:
[0644] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[0645] server:
[0646] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[0647] Specific use cases
[0648] Usage example 1:
[0649] The user asks, "When is my next hospital appointment?" The device recognizes the speech and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "My next hospital appointment is tomorrow at 10:00." The device plays this back as audio.
[0650] Usage example 2:
[0651] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately."
[0652] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and enable them to live a safe and fulfilling daily life.
[0653] The processing flow will be explained below.
[0654] Step 1:
[0655] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[0656] Step 2:
[0657] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[0658] Step 3:
[0659] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[0660] Step 4:
[0661] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[0662] Step 5:
[0663] The device receives the voice input, converts it into text data using a speech recognition engine, and sends the converted text data to the server.
[0664] Step 6:
[0665] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[0666] Step 7:
[0667] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[0668] Step 8:
[0669] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[0670] Step 9:
[0671] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker.
[0672] Step 10:
[0673] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[0674] Step 11:
[0675] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[0676] Step 12:
[0677] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[0678] Step 13:
[0679] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[0680] Step 14:
[0681] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[0682] Step 15:
[0683] The user presses the emergency button. The emergency button on the device is operated.
[0684] Step 16:
[0685] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[0686] Step 17:
[0687] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[0688] The above are the specific processing steps for carrying out the invention.
[0689] Example 1
[0690] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0691] To help seniors live safe and fulfilling daily lives without feeling lonely, systems with personalized information provision, reminder functions, and emergency contact functions are needed. However, existing systems are often difficult for seniors to use, and information provision is fragmented and unintegrated. Furthermore, the accuracy of voice input and dialogue interaction is low, and external information acquisition is often difficult. This makes it difficult for seniors to receive appropriate support.
[0692] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0693] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for the server to obtain information from an external information service, and means for converting voice into text data using a voice recognition engine. This allows elderly people to intuitively operate the device, smoothly utilize personalized information provision, reminder functions, and emergency contact functions, and also makes it easy to obtain external information, enabling them to live a safe and fulfilling life.
[0694] "User's basic information and individual needs" refers collectively to individually personalized data such as the user's name, age, address, health status, emergency contact information, medication times, and schedule for regular health checks.
[0695] "Means for converting voice input into text data" is a general term for devices and software that use voice recognition technology to convert a user's voice into text data and send that text data to a server.
[0696] "Response data from the server" refers to information and messages analyzed and generated by the server, including data obtained from external information services.
[0697] "Reminder information" refers to notification content that a user should receive at a specific time or date, such as when to take medicine or a schedule for regular health checks.
[0698] An "emergency contact request" is a request made by a user to quickly seek assistance in an emergency, and is an emergency signal sent from the terminal to the server.
[0699] "Means of obtaining information from external information services" is a general term for the technologies and methods that allow a server to retrieve necessary information from public APIs and databases on the Internet.
[0700] A "speech recognition engine" is a piece of software or hardware that understands a user's voice input and converts it into text data.
[0701] "Dialogue interaction" refers to two-way communication between a user and a system based on voice input and text data.
[0702] This invention relates to a generative AI "Elderly Companion" system that supports the lives of the elderly, and has a series of functions to manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. This system also supports the lives of the elderly through reminder functions and emergency contact functions.
[0703] In this embodiment of the present invention, the system is configured by a server, terminals, and users communicating with each other. The server is hosted on the cloud and manages user information through a database. Specifically, a MySQL database is used, and data security is ensured by SSL encryption.
[0704] Manage and store basic information and needs
[0705] Device:
[0706] Users enter basic information such as their name, age, address, health condition, and emergency contact details, as well as their individual needs such as medication times and schedules for regular health checks, through the device's interface. Either voice input or text input is possible, using a smartphone or a dedicated device.
[0707] server:
[0708] The server receives the basic information and needs information sent from the device and stores it in a MySQL database. The server uses this data to provide personalized support to the user.
[0709] Converts voice input into text and sends it to the server
[0710] Device:
[0711] When a user asks a question or gives an instruction by voice, such as "When is my next doctor's appointment?", the device converts the voice into text data using Google's speech recognition API. The converted text data is then sent to the server. The user can use a headset or built-in microphone to do this.
[0712] server:
[0713] The server analyzes the received text data and obtains the necessary information (for example, reservation information or weather information). To obtain the weather information, it uses a public API such as the OpenWeatherMap API. The server then sends the analysis results to the terminal as response data.
[0714] Receiving and displaying response data
[0715] Device:
[0716] The device that receives the response data from the server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text data into speech and tells the user, "Your next hospital appointment is tomorrow at 10:00." If the device has a display, it can also display the information on the screen.
[0717] Reminder function
[0718] server:
[0719] Reminder information is stored on the server based on a schedule specified by the user, and the server sends the reminder information to the device at the specified time.
[0720] Device:
[0721] The device will receive a reminder notification at the specified time, informing the user via voice or pop-up notification, "It's 8 o'clock. Time to take your medicine."
[0722] Emergency contact function
[0723] Device:
[0724] When a user faces an emergency, they press the emergency button on their device, which instantly sends an emergency signal to the server.
[0725] server:
[0726] When the server receives an emergency signal, it sends a message to pre-registered emergency contacts (family members, caregivers) via SMS, phone, or email saying, "The user has pressed the emergency button. Please check immediately."
[0727] Specific examples of actions and prompts
[0728] Usage example 1:
[0729] The user asks, "What's the weather like today?" The device converts the speech to text and sends it to the server. The server retrieves weather information using a weather information service, and then sends the information to the device: "Today's weather is sunny. The temperature is 24 degrees."
[0730] Usage example 2:
[0731] The user asks, "When is my next hospital appointment?" The device converts the speech to text and sends it to the server. The server retrieves the appointment information from its database and responds, "My next hospital appointment is tomorrow at 10:00." The device plays this back aloud.
[0732] This invention enables elderly people to easily access personalized information, reminder functions, and emergency contact functions through an interface that they can operate intuitively. This system reduces the sense of loneliness felt by elderly people and supports safe and fulfilling lives.
[0733] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0734] Step 1:
[0735] The user enters basic information and individual needs into the device, such as name, age, address, health condition, emergency contact information, medication times, and schedule for regular health checks, by voice or text. This information is temporarily stored on the device and then sent to the server.
[0736] Input: Name, age, address, health status, emergency contact, needs
[0737] Output: Basic information and needs sent to the server
[0738] Step 2:
[0739] The device receives voice input and converts it into text data. For example, if a user says, "What's the weather like today?", the device converts the voice into text using a speech recognition engine (Google Speech-to-Text API). The text data is then sent to the server.
[0740] Input: Voice input (e.g. "What's the weather like today?")
[0741] Output: Text data (e.g. "What's the weather like today?")
[0742] Step 3:
[0743] The server analyzes the received text data and obtains the corresponding information. For example, if the text data "What's the weather like today?" is received, the server will call an external weather information API (e.g., OpenWeatherMap API) to obtain the weather information and obtain the results.
[0744] Input: Text data (e.g., "What's the weather like today?")
[0745] Output: Weather information (e.g. sunny, temperature 24 degrees)
[0746] Step 4:
[0747] The server sends the acquired information to the terminal as text data. The server generates response data based on the analysis results and sends it to the terminal.
[0748] Input: Weather information (e.g. sunny, temperature 24 degrees)
[0749] Output: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[0750] Step 5:
[0751] The device will play back the response data received from the server as audio. The device will convert the text data into audio using a speech synthesis engine (Google Text-to-Speech API) and convey it to the user. If the device has a display, it can also display it.
[0752] Input: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[0753] Output: Voice playback (e.g. "Today's weather is sunny. The temperature is 24 degrees.") and display
[0754] Step 6:
[0755] The device receives a reminder notification at the specified time and notifies the user. The server sends pre-set reminder information to the device at the specified time. The device notifies the user with a voice or pop-up notification such as, "It's 8 o'clock. It's time to take your medicine."
[0756] Input: Reminder information (e.g., every day at 8:00)
[0757] Output: Audio and popup notification (e.g. "It's 8 o'clock. Time to take your medicine.")
[0758] Step 7:
[0759] When a user presses the emergency button, the device sends an emergency signal to the server. When the device sends the emergency signal, the server notifies the emergency contact. The emergency contact is sent a message via SMS, phone call, or email saying, "The user has pressed the emergency button. Please check immediately."
[0760] Input: Press emergency button
[0761] Output: Emergency notification message (e.g. "The user pressed the emergency button. Please check immediately.")
[0762] The above is the specific flow of processing steps for the "Generative AI Elderly Companion" system.
[0763] (Application example 1)
[0764] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0765] There is a need for support systems that can reduce the difficulty elderly people have in finding products in physical stores and enable them to shop safely and efficiently. Conventional support systems have difficulty providing the information elderly people need quickly, and do not have sufficient reminder or emergency contact functions.
[0766] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0767] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for providing product location information to support the elderly in a physical store, and means for notifying the user of reminders set in the physical store, thereby reducing the difficulty for the elderly when searching for products in a physical store and enabling them to shop safely and efficiently.
[0768] "Basic information" refers to individual information such as the user's name, age, address, health status, and emergency contact information.
[0769] "Individual needs" are information about a user's specific requirements or preferences, such as when to take medication or when to schedule regular health checks.
[0770] "Voice input" refers to data input to the system by a user speaking.
[0771] "Text data" is character information generated by analyzing voice input.
[0772] A "server" is a computer system that manages and processes data sent by users.
[0773] "Response data" is reply data that the server generates and sends in response to a user request.
[0774] "Reminder information" is information for notifying the user at a specified time.
[0775] "Emergency contacts" are information about people or organizations that the user wants to contact in the event of an emergency.
[0776] An "emergency notification" is a message sent from the server when an emergency occurs.
[0777] A "brick and mortar store" is a sales or service establishment located in a physical location.
[0778] "Product location information" refers to information about the location where a specific product is located within a physical store.
[0779] A "generative AI model" is an artificial intelligence algorithm model that generates text data in response to user requests.
[0780] MODE FOR CARRYING OUT THE INVENTION
[0781] The following describes an embodiment of the present invention.
[0782] System Program Overview
[0783] The "Generative AI Elderly Companion" system, which is the subject of this invention, has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. Among the functions of this system, we will explain the important functions that support the elderly in brick-and-mortar stores in particular.
[0784] Hardware and software used
[0785] Smartphone (iOS, Android): A device that allows users to input and receive information.
[0786] In-store robots (e.g., Pepper): Provide an interface for users to obtain and receive information within the store.
[0787] Server: Manages user information, performs voice recognition, and trains and updates generative AI models.
[0788] Speech recognition engine (Google Cloud Speech-to-Text): Converts voice input into text data.
[0789] Generative AI model (GPT-4 by OpenAI): Analyzes and generates text data based on user requests.
[0790] Database (MySQL): Stores user basic information, individual needs, reminder information, emergency contacts, etc.
[0791] Weather API (OpenWeatherMap): Used to retrieve the weather information desired by the user.
[0792] System procedures and operations
[0793] Manage and store basic information and needs
[0794] The server stores basic information and individual needs entered by the user through their device (smartphone or robot) in a MySQL database. Users can enter information by voice or text.
[0795] Converts voice input into text and sends it to the server
[0796] The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the user's voice input into text data and send it to the server.
[0797] Receiving and displaying response data
[0798] The server analyzes the received text data using a generative AI model (GPT-4) and generates appropriate response data. For example, when a user asks, "Where is this product?", the server obtains the product's location information and sends it to the device, saying, "Product A is in aisle 3." The device then provides this to the user via voice or a screen display.
[0799] Reminder function
[0800] The reminder information set by the user is stored on the server, and when the specified time comes, the server sends the reminder information to the device and generates a notification, which then notifies the device with a voice or a pop-up message saying, "It's time to take your medicine."
[0801] Emergency contact function
[0802] In the event of an emergency, the user presses the emergency button on the device, and the device sends an emergency signal to the server, which then sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[0803] Examples of specific examples and prompts
[0804] Specific examples
[0805] For example, when a user asks a smartphone or robot, "Where is this product?", the speech recognition engine (Google Cloud Speech-to-Text) converts the speech into text and sends it to the server. The server then analyzes it using a generative AI model (GPT-4), obtains the product location information, and generates a response such as "Product A is in aisle 3," which the smartphone or robot then plays back.
[0806] Prompt Sentence Examples
[0807] User: "Where is this item?"
[0808] System: "Converts speech to text and sends it to the server..."
[0809] Generative AI model: "Analyzing product information for user search..."
[0810] Response: "Item A is in aisle 3"
[0811] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to enjoy shopping in brick-and-mortar stores more comfortably and safely.
[0812] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0813] Step 1: Manage and save your basic information and needs
[0814] Users use their smartphones or in-store robots to input their basic information (name, age, address, health condition, emergency contact information) and individual needs (medicine dosing times, regular health check schedules) by voice or text. This input data is converted into text by a speech recognition engine (Google Cloud Speech-to-Text) on the device. The converted text data is sent to the server and stored in a MySQL database.
[0815] Input: User voice or text input
[0816] Data processing: Converting voice to text
[0817] Output: Send and save text data to the server
[0818] Step 2: Convert speech to text and send to server
[0819] The user asks a question by voice, such as "Where is this item?" The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the voice into text data and sends the text data to the server. The server receives this text data.
[0820] Input: User voice input
[0821] Data processing: Converting voice to text
[0822] Output: Send text data to the server
[0823] Step 3: Parsing text data and generating a response
[0824] The server uses a generative AI model (GPT-4) to analyze the received text data and generate an appropriate response. For example, it retrieves product location information from a database based on the user's question and generates a response such as "Product A is in aisle 3."
[0825] Input: Text data
[0826] Data computation: Analysis using generative AI models
[0827] Output: Generated response data
[0828] Step 4: Send and display response data
[0829] The server sends the generated response data to the terminal. The terminal receives this response data and uses a speech synthesis engine to play it back as a voice such as "Product A is in aisle 3," and also displays it on the screen.
[0830] Input: Response data
[0831] Data processing: voice synthesis and screen display
[0832] Output: Speech and text display
[0833] Step 5: Execute the reminder function
[0834] The server sends the reminder information set by the user to the device at the specified time. The device receives the reminder notification and notifies the user by voice, saying "It's time to take your medicine." A pop-up notification is also displayed.
[0835] Input: Reminder information and designated time
[0836] Data Processing: Notification Generation
[0837] Output: Sound and popup notification
[0838] Step 6: Implementing emergency contact functions
[0839] When a user presses the emergency button, the device sends an emergency signal to the server, which receives the signal and sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[0840] Input: Emergency signal
[0841] Data Processing: Emergency notification generation and transmission
[0842] Output:Notify emergency contacts
[0843] The above processing steps enable users to quickly obtain the information they need in a physical store, allowing them to enjoy shopping safely and comfortably.
[0844] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0845] The "Generative AI Elderly Companion" system, which is the subject of this invention, manages and stores the user's basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, making conversational interactions more personalized. It also supports the daily lives of the elderly through reminder and emergency contact functions.
[0846] Manage and store basic information and needs
[0847] server:
[0848] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[0849] Device:
[0850] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[0851] Converts voice input into text and sends it to the server
[0852] Device:
[0853] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[0854] server:
[0855] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[0856] Receiving and displaying response data
[0857] Device:
[0858] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[0859] Reminder function
[0860] server:
[0861] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[0862] Device:
[0863] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[0864] Emergency contact function
[0865] Device:
[0866] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[0867] server:
[0868] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[0869] Emotion recognition function
[0870] Device:
[0871] The device sends the user's voice input to the emotion engine, which analyzes the tone, speed, volume, etc. of the voice to determine the user's emotional state.
[0872] server:
[0873] The server analyzes the emotional information received from the emotion engine and generates an appropriate response based on it. For example, if a user says in a sad voice, "I'm not feeling well today," the server will generate an empathetic response such as, "What's wrong? Tell me your story."
[0874] Device:
[0875] Once the response data based on emotion recognition is sent from the server, the device displays or plays it back to the user, making interactions with the user more natural and personalized.
[0876] Specific use cases
[0877] Usage example 1:
[0878] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[0879] Usage example 2:
[0880] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[0881] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and lead a safe and fulfilling daily life. By incorporating an emotion engine, the system will be able to respond more sensitively to the user's feelings, providing a higher level of satisfaction.
[0882] The processing flow will be explained below.
[0883] Step 1:
[0884] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[0885] Step 2:
[0886] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[0887] Step 3:
[0888] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[0889] Step 4:
[0890] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[0891] Step 5:
[0892] The device receives voice input, converts the voice into text data using a speech recognition engine, and sends the converted text data to the server.
[0893] Step 6:
[0894] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[0895] Step 7:
[0896] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[0897] Step 8:
[0898] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[0899] Step 9:
[0900] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker. The device plays back "Today's weather is sunny. The temperature is 24 degrees."
[0901] Step 10:
[0902] The emotion engine analyzes the tone, rate, and volume of the user's voice to identify the user's emotional state: if the user speaks in an anxious voice, it will be identified as in an "anxious" state.
[0903] Step 11:
[0904] The server analyzes the emotional information received from the emotion engine and generates an appropriate response. For example, if a user says in a sad voice, "I'm not feeling well today," the server generates a response that shows empathy, such as, "What's wrong? Tell me your story."
[0905] Step 12:
[0906] The response data containing the generated emotional response is sent from the server to the device, which then plays it back to the user, saying, "What's wrong? Tell me what you think."
[0907] Step 13:
[0908] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[0909] Step 14:
[0910] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[0911] Step 15:
[0912] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[0913] Step 16:
[0914] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[0915] Step 17:
[0916] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[0917] Step 18:
[0918] The user presses the emergency button. The emergency button on the device is operated.
[0919] Step 19:
[0920] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[0921] Step 20:
[0922] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[0923] The above are the specific processing steps for carrying out the invention.
[0924] Example 2
[0925] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0926] Elderly people require various forms of support in their daily lives, but current technology makes it difficult to provide personalized assistance that fully reflects the individual needs and emotional state of each elderly person. Furthermore, systems for responding quickly and appropriately in emergencies are inadequate. Therefore, there is a need to create an environment where elderly people can live safely and without feeling lonely.
[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0928] In this invention, the server includes: means for managing and saving basic information and individual needs of a user; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying or playing a notification to the user at a specified time; means for sending an emergency contact request to the server and for the server to send an emergency notification to a pre-set emergency contact; means for analyzing the tone, speed and volume of the user's voice to identify the user's emotional state; and means for generating a personalized response based on the emotional state and providing it to the user.
[0929] This will enable personalized assistance to be provided in response to the individual needs and emotional state of the elderly, as well as a rapid and appropriate response in emergencies.
[0930] "Basic user information" refers to basic information about a user, such as name, age, address, health status, and emergency contact information.
[0931] "Individual needs" refers to information about the support and services that a user individually needs, such as when to take medication or schedule regular health checks.
[0932] "Voice input" refers to voice data input when a user speaks to the system through a microphone.
[0933] "Text data" refers to voice input converted into character string data using a voice recognition engine or the like.
[0934] "Server" refers to a computer system that receives data from users and stores, analyzes, and processes it.
[0935] "Response data" refers to information that the server generates based on a user request and returns to the user.
[0936] "Reminder information" refers to information for notifying the user at a specific time or timing.
[0937] "Emergency contact request" refers to an emergency signal sent through the system when a user faces an emergency.
[0938] "Emergency notification" refers to an emergency message sent from the server to a pre-defined emergency contact.
[0939] "Emotional state" refers to the user's emotions estimated based on an analysis of the tone, rate, volume, etc. of the voice.
[0940] "Personalized response" refers to an individualized response that is generated based on the user's basic information, individual needs, and emotional state.
[0941] The present invention, the "Generative AI Elderly Companion" system, provides personalized support that reflects the individual needs and emotional state of elderly people, reducing feelings of loneliness in daily life and creating a safe living environment.
[0942] Hardware and software used
[0943] The server is a high-performance computer system, such as an AWS EC2 instance or Microsoft Azure VM. MySQL or PostgreSQL is used as the database management system. Google Cloud Speech-to-Text or Amazon Transcribe is used as the speech recognition engine, and IBM Watson Tone Analyzer or Microsoft Azure Cognitive Services is used as the emotion recognition engine.
[0944] Manage and store basic information and needs
[0945] The user uses the device to enter basic information (name, age, address, health status, emergency contact information). The device provides voice or text input, verifies the entered information, and sends it to the server. The server stores the received basic information in a database and provides personalized support based on this information.
[0946] Converts voice input into text and sends it to the server
[0947] When a user asks "What's the weather like today?", the device uses a voice recognition engine to convert the voice into text data and sends it to the server. The server then analyzes the received text data and obtains weather information using a weather API provided by a third party.
[0948] Receiving and displaying response data
[0949] The response data (such as weather information) sent from the server is received by the terminal, which then provides it to the user through voice or a display device. For example, it may say, "Today's weather is sunny. The temperature is 24 degrees."
[0950] Reminder function
[0951] When a user instructs the device to "set a reminder to take my medicine at 8 o'clock every morning," the device converts this instruction into text data and sends it to the server. The server stores the reminder information in a database and sends a reminder notification to the device at the specified time (8 o'clock every morning). The device then notifies the user by voice or a pop-up notification, saying, "It's 8 o'clock. It's time to take your medicine."
[0952] Emergency contact function
[0953] In an emergency, the user presses the emergency button on the device. The device then sends an emergency signal to the server. The server then contacts pre-registered emergency contacts (family members or caregivers) via SMS, phone, email, etc. It then sends a message saying, "The user has pressed the emergency button. Please check immediately."
[0954] Emotion recognition function
[0955] If a user says "I'm feeling bad today" in a sad voice, the device sends this voice data to an emotion recognition engine, which analyzes the tone, speed, and volume of the voice to identify the emotional state. The server uses the emotional information to generate a personalized response that shows empathy, such as "What's wrong? Tell me your story," and the device plays this aloud to the user.
[0956] Specific use cases
[0957] Usage example 1:
[0958] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[0959] Usage example 2:
[0960] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[0961] The above is an embodiment of the present invention. This allows elderly people to live a safe and fulfilling daily life and reduce their sense of loneliness. Furthermore, the introduction of an emotion engine enables the system to respond in a way that is sensitive to the user's feelings, providing a higher level of satisfaction.
[0962] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0963] Steps for entering and saving basic information
[0964] Step 1:
[0965] The user uses the device to input basic information (name, age, address, health status, emergency contact information). The input method can be selected as voice input or text input. For example, if the user inputs "My name is Ichiro Tanaka, and I'm 75 years old" by voice, the device will capture the voice data.
[0966] input:
[0967] Basic information entered by the user (name, age, address, health status, emergency contact information)
[0968] output:
[0969] Audio or text data
[0970] Step 2:
[0971] When the device receives voice input, it uses a speech recognition engine (such as Google Cloud Speech-to-Text) to convert the voice into text data. For example, voice data such as "My name is Tanaka Ichiro and I'm 75 years old" is converted into text data such as "Name: Tanaka Ichiro, Age: 75."
[0972] input:
[0973] Audio data
[0974] output:
[0975] Text data
[0976] Step 3:
[0977] The device sends the converted text data to the server. Specifically, data transmission is performed using an API request.
[0978] input:
[0979] Text data
[0980] output:
[0981] API request (text data)
[0982] Step 4:
[0983] The server stores the received text data in a database, for example, using MySQL or PostgreSQL as data records.
[0984] input:
[0985] Text data
[0986] output:
[0987] Data Records
[0988] Processing steps for converting voice input to text and sending it to the server
[0989] Step 1:
[0990] The user asks a question by voice, "What's the weather like today?" This voice data is acquired by the terminal.
[0991] input:
[0992] User voice input
[0993] output:
[0994] Audio data
[0995] Step 2:
[0996] The device uses a voice recognition engine to convert the acquired voice data into text data. Specifically, the voice data "What's the weather like today?" is converted into text data "What's the weather like today?"
[0997] input:
[0998] Audio data
[0999] output:
[1000] Text data
[1001] Step 3:
[1002] The device sends the converted text data to the server via an API request.
[1003] input:
[1004] Text data
[1005] output:
[1006] API request (text data)
[1007] Step 4:
[1008] The server analyzes the received text data and retrieves the corresponding information. For example, to retrieve weather information, the server sends a request to a weather API provided by a third party and receives response data such as "Today's weather is sunny and the temperature is 24 degrees."
[1009] input:
[1010] Text data
[1011] output:
[1012] Response data (weather information)
[1013] Processing steps for receiving and displaying response data
[1014] Step 1:
[1015] The server sends the acquired response data (such as weather information) to the terminal.
[1016] input:
[1017] Response data (weather information)
[1018] output:
[1019] API response (response data)
[1020] Step 2:
[1021] The device analyzes the received response data and provides it to the user through a voice output device, for example, by playing back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[1022] input:
[1023] API response (response data)
[1024] output:
[1025] Audio data
[1026] Reminder function processing steps
[1027] Step 1:
[1028] The user instructs, "Set a reminder to take my medicine at 8 o'clock every morning." The device receives this voice data.
[1029] input:
[1030] User voice instructions
[1031] output:
[1032] Audio data
[1033] Step 2:
[1034] The device uses a speech recognition engine to convert voice data into text data. For example, voice data such as "Set a reminder to take my medicine at 8 o'clock every morning" is converted into text data such as "Take my medicine at 8 o'clock every morning."
[1035] input:
[1036] Audio data
[1037] output:
[1038] Text data
[1039] Step 3:
[1040] The device sends the converted text data to the server via an API request.
[1041] input:
[1042] Text data
[1043] output:
[1044] API request (text data)
[1045] Step 4:
[1046] The server stores the received reminder information in a database and configures it to generate notifications at a specified time, for example, every morning at 8:00.
[1047] input:
[1048] Text data
[1049] output:
[1050] Database records and timer settings
[1051] Step 5:
[1052] When the specified time arrives, the server sends a notification to the device, such as "It's time to take your medicine."
[1053] input:
[1054] Timer Event
[1055] output:
[1056] API response (notification data)
[1057] Step 6:
[1058] The device will then provide the received notification data to the user via voice or pop-up notification, for example, "It's 8 o'clock. Time to take your medicine."
[1059] input:
[1060] API response (notification data)
[1061] output:
[1062] Audio data or popup notification
[1063] Emergency Contact Function Processing Steps
[1064] Step 1:
[1065] The user presses the emergency button, and the device detects this emergency signal.
[1066] input:
[1067] Emergency button input
[1068] output:
[1069] Emergency Signal Data
[1070] Step 2:
[1071] The device sends an emergency signal to the server, which is done via an API request.
[1072] input:
[1073] Emergency Signal Data
[1074] output:
[1075] API Request (Emergency Signal Data)
[1076] Step 3:
[1077] The server analyzes the received emergency signal and sends an emergency notification to pre-defined emergency contacts, for example, by SMS, phone call, or email, with a message saying, "The user has pressed the emergency button. Please check immediately."
[1078] input:
[1079] API Request (Emergency Signal Data)
[1080] output:
[1081] Emergency notification data (SMS, phone, email)
[1082] Emotion Recognition Processing Steps
[1083] Step 1:
[1084] The user inputs "I feel bad today" in a sad voice. This voice data is acquired by the terminal.
[1085] input:
[1086] User voice input
[1087] output:
[1088] Audio data
[1089] Step 2:
[1090] The device sends the captured voice data to an emotion recognition engine, which analyzes the tone, rate, and volume of the voice to identify the emotional state.
[1091] input:
[1092] Audio data
[1093] output:
[1094] Emotional state data
[1095] Step 3:
[1096] The server generates a personalized response based on the emotional state data received from the emotion recognition engine, for example, an empathetic response such as "What's wrong? Tell me your story."
[1097] input:
[1098] Emotional state data
[1099] output:
[1100] Response data
[1101] Step 4:
[1102] The server transmits the generated response data to the terminal.
[1103] input:
[1104] Response data
[1105] output:
[1106] API response (response data)
[1107] Step 5:
[1108] The device will then play back the received response data in voice, for example, "What's wrong? Tell me what happened."
[1109] input:
[1110] API response (response data)
[1111] output:
[1112] Audio data
[1113] (Application example 2)
[1114] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1115] Elderly people have greater difficulty navigating physical stores, searching for products, and responding to emergencies. They also often feel anxious about asking store staff questions and navigating the store. Current systems struggle to adequately address these issues, preventing seniors from enjoying shopping with peace of mind.
[1116] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for managing and saving a user's basic information and individual needs; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying a notification to the user at a specified time; means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact; means for navigating the store using the user's current location and store map data; and means for analyzing the user's voice input and generating a personalized response based on the user's emotional state. This enables seniors to move confidently through physical stores, easily find products and services, and quickly respond to anxieties or emergencies.
[1117] Below are definitions of important terms included in the patent claims, rewritten to suit the application example.
[1118] "Means for managing and storing a user's basic information and individual needs" refers to a means for storing basic information such as the user's name, age, address, health condition, and emergency contact information, as well as individual needs such as medication times and regular health check schedules, in a database, and for managing and updating this information as needed.
[1119] "Means for converting user's voice input into text data and sending it to the server" refers to means for converting information input by the user by voice into text data using voice recognition technology and sending that text data to the server.
[1120] "Means for receiving response data from the server and displaying or playing it back to the user" refers to means for receiving information sent from the server (e.g., weather information or store directions) and providing it to the user visually or audibly.
[1121] The "means for saving specified reminder information on a server and displaying a notification to the user at a specified time" refers to a means for saving reminder information set by a user on a server and sending a notification to the user at a specified time based on that reminder information.
[1122] "Means for sending an emergency contact request to a server, and for the server to send an emergency notification to pre-registered emergency contacts" refers to a means for a user to send an emergency contact request to a server by pressing an emergency button, etc., and for the server to send an emergency notification by phone or SMS to pre-registered emergency contacts based on that request.
[1123] "Means for navigating within a store using the user's current location and store map data" refers to a means for guiding the user to their desired location by using the location information of the user's smartphone or device and combining it with map data within the store.
[1124] "Means for analyzing a user's voice input and generating a personalized response based on the user's emotional state" refers to means for analyzing the tone, rate, and content of a user's voice to identify the user's emotional state, and generating and providing an appropriate response or message accordingly.
[1125] System Program
[1126] The purpose of this invention, the "Generative AI Elderly Companion Shopper" system, is to help seniors shop comfortably in brick-and-mortar stores. This system manages and stores users' basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. It also supports the daily lives of seniors through reminder functions, emergency contact functions, and emotion recognition functions.
[1127] Hardware and Software Use
[1128] Hardware:
[1129] Smartphone (with camera, microphone, and speaker)
[1130] software:
[1131] Google Cloud Speech-to-Text API (voice recognition)
[1132] Google Cloud Natural Language API (Text Analysis)
[1133] Twilio API (emergency contact function)
[1134] Emotion Engine (e.g. Affectiva's SDK)
[1135] Firebase (database)
[1136] Data processing and calculation
[1137] Managing and storing your basic information and individual needs:
[1138] The server stores the basic information entered by the user (name, age, address, health status, emergency contacts) and individual needs (medication times, schedule for regular health checks) in a Firebase database, allowing the system to provide personalized support to the user.
[1139] Convert speech to text and send to server:
[1140] When a user asks, "Where is the restroom?", the device uses the Google Cloud Speech-to-Text API to convert speech to text data and send it to the server, which then uses the Google Cloud Natural Language API to parse the text and generate an appropriate response.
[1141] Receive and display response data:
[1142] The response data sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may provide voice guidance such as, "To find the restroom, go straight down the corridor on the right."
[1143] Reminder function:
[1144] The server stores the reminder information set by the user (e.g., take medicine at 8 o'clock every day) in Firebase. When the specified time arrives, the device sends a voice or pop-up notification saying, "It's 8 o'clock. Time to take your medicine."
[1145] Emergency contact features:
[1146] If a user is about to fall inside the store, they can press the emergency button on their device, which sends an emergency signal to the server. The server then uses the Twilio API to send a notification to registered emergency contacts (family members or caregivers). The emergency notification is sent via SMS or phone.
[1147] Emotion recognition function:
[1148] If a user says, "I'm not feeling well today," the device uses its emotion engine to analyze the speech and identify the user's emotional state. The server then generates an appropriate response based on this, saying, "What's wrong? Tell me."
[1149] Examples of specific examples and prompts
[1150] As a specific example, consider the case where an elderly person named Mr. A gets lost in a store.
[1151] Person A asks the app, "Where is the restroom?" inside the store. Voice recognition is activated, and the speech is converted into text and sent to the server.
[1152] The server identifies the location of the restroom from the store's map data and generates a navigation plan along with Mr. A's current location.
[1153] The app will provide voice guidance saying, "To get to the restroom, go straight down the corridor on the right."
[1154] If the emotion engine detects Mr. A's anxiety, the app will ask, "If you can't find it, I'll call a store employee."
[1155] Example prompt sentence:
[1156] The user asks the smartphone microphone, "Where is the restroom?" The voice input is converted into text and sent to the server. The system should refer to the user's current location and a map of the store, calculate the shortest route, and provide voice guidance. If the user asks in an anxious voice, the system should generate an additional response asking, "If you can't find it, would you like me to call a store employee?"
[1157] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1158] Step 1:
[1159] A user asks into the microphone of their smartphone, "Where is the restroom?" A voice input occurs, and the device collects this voice input data.
[1160] Input: User voice input
[1161] Output: Audio data
[1162] Step 2:
[1163] The device uses the Google Cloud Speech-to-Text API to convert the voice data into text data, which is then sent to the server.
[1164] Input: Audio data
[1165] Output: Text data
[1166] Step 3:
[1167] The server analyzes the received text data using the Google Cloud Natural Language API, and as a result, it understands that the user is looking for a restroom.
[1168] Input: Text data
[1169] Output: Analysis result (intent)
[1170] Step 4:
[1171] The server identifies the user's location by referencing the user's current location and the store's map data, and calculates the shortest route based on this.
[1172] Input: User's current location data, store map data
[1173] Output: Shortest route data
[1174] Step 5:
[1175] The server uses the calculated shortest route data to generate a navigation plan to help the user reach their destination easily. The generated navigation plan is created in voice and text format.
[1176] Input: Shortest route data
[1177] Output: Navigation plan (voice and text)
[1178] Step 6:
[1179] The server sends the generated navigation plan to the device, which receives it and provides voice and text guidance to the user.
[1180] Input: Navigation plan
[1181] Output: User instructions
[1182] Step 7:
[1183] At the same time, the device analyzes the user's voice input with an emotion engine to identify their emotional state. If the device detects that the user is asking a question in an anxious voice, the server generates an additional response (e.g., "If you can't find it, should I call a store clerk?").
[1184] Input: User's voice data
[1185] Output: Emotional state, additional responses
[1186] Step 8:
[1187] As additional responses are generated, the server sends them to the terminal, which displays or plays them audibly to the user.
[1188] Input: Additional response data
[1189] Output: Additional instructions for the user
[1190] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1191] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1192] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1193] [Third embodiment]
[1194] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1195] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1196] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1197] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1198] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1199] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1200] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1201] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1202] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1203] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1204] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1205] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1206] The "Generative AI Elderly Companion" system of this invention has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back the response data from the server to the user via their device. This system can also support the lives of the elderly through reminder and emergency contact functions.
[1207] Manage and store basic information and needs
[1208] server:
[1209] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[1210] Device:
[1211] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[1212] Converts voice input into text and sends it to the server
[1213] Device:
[1214] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[1215] server:
[1216] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[1217] Receiving and displaying response data
[1218] Device:
[1219] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[1220] Reminder function
[1221] server:
[1222] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[1223] Device:
[1224] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[1225] Emergency contact function
[1226] Device:
[1227] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[1228] server:
[1229] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[1230] Specific use cases
[1231] Usage example 1:
[1232] The user asks, "When is my next hospital appointment?" The device recognizes the speech and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "My next hospital appointment is tomorrow at 10:00." The device plays this back as audio.
[1233] Usage example 2:
[1234] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately."
[1235] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and enable them to live a safe and fulfilling daily life.
[1236] The processing flow will be explained below.
[1237] Step 1:
[1238] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[1239] Step 2:
[1240] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[1241] Step 3:
[1242] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[1243] Step 4:
[1244] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[1245] Step 5:
[1246] The device receives the voice input, converts it into text data using a speech recognition engine, and sends the converted text data to the server.
[1247] Step 6:
[1248] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[1249] Step 7:
[1250] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[1251] Step 8:
[1252] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[1253] Step 9:
[1254] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker.
[1255] Step 10:
[1256] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[1257] Step 11:
[1258] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[1259] Step 12:
[1260] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[1261] Step 13:
[1262] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[1263] Step 14:
[1264] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[1265] Step 15:
[1266] The user presses the emergency button. The emergency button on the device is operated.
[1267] Step 16:
[1268] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[1269] Step 17:
[1270] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[1271] The above are the specific processing steps for carrying out the invention.
[1272] Example 1
[1273] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1274] To help seniors live safe and fulfilling daily lives without feeling lonely, systems with personalized information provision, reminder functions, and emergency contact functions are needed. However, existing systems are often difficult for seniors to use, and information provision is fragmented and unintegrated. Furthermore, the accuracy of voice input and dialogue interaction is low, and external information acquisition is often difficult. This makes it difficult for seniors to receive appropriate support.
[1275] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1276] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for the server to obtain information from an external information service, and means for converting voice into text data using a voice recognition engine. This allows elderly people to intuitively operate the device, smoothly utilize personalized information provision, reminder functions, and emergency contact functions, and also makes it easy to obtain external information, enabling them to live a safe and fulfilling life.
[1277] "User's basic information and individual needs" refers collectively to individually personalized data such as the user's name, age, address, health status, emergency contact information, medication times, and schedule for regular health checks.
[1278] "Means for converting voice input into text data" is a general term for devices and software that use voice recognition technology to convert a user's voice into text data and send that text data to a server.
[1279] "Response data from the server" refers to information and messages analyzed and generated by the server, including data obtained from external information services.
[1280] "Reminder information" refers to notification content that a user should receive at a specific time or date, such as when to take medicine or a schedule for regular health checks.
[1281] An "emergency contact request" is a request made by a user to quickly seek assistance in an emergency, and is an emergency signal sent from the terminal to the server.
[1282] "Means of obtaining information from external information services" is a general term for the technologies and methods that allow a server to retrieve necessary information from public APIs and databases on the Internet.
[1283] A "speech recognition engine" is a piece of software or hardware that understands a user's voice input and converts it into text data.
[1284] "Dialogue interaction" refers to two-way communication between a user and a system based on voice input and text data.
[1285] This invention relates to a generative AI "Elderly Companion" system that supports the lives of the elderly, and has a series of functions to manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. This system also supports the lives of the elderly through reminder functions and emergency contact functions.
[1286] In this embodiment of the present invention, the system is configured by a server, terminals, and users communicating with each other. The server is hosted on the cloud and manages user information through a database. Specifically, a MySQL database is used, and data security is ensured by SSL encryption.
[1287] Manage and store basic information and needs
[1288] Device:
[1289] Users enter basic information such as their name, age, address, health condition, and emergency contact details, as well as their individual needs such as medication times and schedules for regular health checks, through the device's interface. Either voice input or text input is possible, using a smartphone or a dedicated device.
[1290] server:
[1291] The server receives the basic information and needs information sent from the device and stores it in a MySQL database. The server uses this data to provide personalized support to the user.
[1292] Converts voice input into text and sends it to the server
[1293] Device:
[1294] When a user asks a question or gives an instruction by voice, such as "When is my next doctor's appointment?", the device converts the voice into text data using Google's speech recognition API. The converted text data is then sent to the server. The user can use a headset or built-in microphone to do this.
[1295] server:
[1296] The server analyzes the received text data and obtains the necessary information (for example, reservation information or weather information). To obtain the weather information, it uses a public API such as the OpenWeatherMap API. The server then sends the analysis results to the terminal as response data.
[1297] Receiving and displaying response data
[1298] Device:
[1299] The device that receives the response data from the server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text data into speech and tells the user, "Your next hospital appointment is tomorrow at 10:00." If the device has a display, it can also display the information on the screen.
[1300] Reminder function
[1301] server:
[1302] Reminder information is stored on the server based on a schedule specified by the user, and the server sends the reminder information to the device at the specified time.
[1303] Device:
[1304] The device will receive a reminder notification at the specified time, informing the user via voice or pop-up notification, "It's 8 o'clock. Time to take your medicine."
[1305] Emergency contact function
[1306] Device:
[1307] When a user faces an emergency, they press the emergency button on their device, which instantly sends an emergency signal to the server.
[1308] server:
[1309] When the server receives an emergency signal, it sends a message to pre-registered emergency contacts (family members, caregivers) via SMS, phone, or email saying, "The user has pressed the emergency button. Please check immediately."
[1310] Specific examples of actions and prompts
[1311] Usage example 1:
[1312] The user asks, "What's the weather like today?" The device converts the speech to text and sends it to the server. The server retrieves weather information using a weather information service, and then sends the information to the device: "Today's weather is sunny. The temperature is 24 degrees."
[1313] Usage example 2:
[1314] The user asks, "When is my next hospital appointment?" The device converts the speech to text and sends it to the server. The server retrieves the appointment information from its database and responds, "My next hospital appointment is tomorrow at 10:00." The device plays this back aloud.
[1315] This invention enables elderly people to easily access personalized information, reminder functions, and emergency contact functions through an interface that they can operate intuitively. This system reduces the sense of loneliness felt by elderly people and supports safe and fulfilling lives.
[1316] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1317] Step 1:
[1318] The user enters basic information and individual needs into the device, such as name, age, address, health condition, emergency contact information, medication times, and schedule for regular health checks, by voice or text. This information is temporarily stored on the device and then sent to the server.
[1319] Input: Name, age, address, health status, emergency contact, needs
[1320] Output: Basic information and needs sent to the server
[1321] Step 2:
[1322] The device receives voice input and converts it into text data. For example, if a user says, "What's the weather like today?", the device converts the voice into text using a speech recognition engine (Google Speech-to-Text API). The text data is then sent to the server.
[1323] Input: Voice input (e.g. "What's the weather like today?")
[1324] Output: Text data (e.g. "What's the weather like today?")
[1325] Step 3:
[1326] The server analyzes the received text data and obtains the corresponding information. For example, if the text data "What's the weather like today?" is received, the server will call an external weather information API (e.g., OpenWeatherMap API) to obtain the weather information and obtain the results.
[1327] Input: Text data (e.g., "What's the weather like today?")
[1328] Output: Weather information (e.g. sunny, temperature 24 degrees)
[1329] Step 4:
[1330] The server sends the acquired information to the terminal as text data. The server generates response data based on the analysis results and sends it to the terminal.
[1331] Input: Weather information (e.g. sunny, temperature 24 degrees)
[1332] Output: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[1333] Step 5:
[1334] The device will play back the response data received from the server as audio. The device will convert the text data into audio using a speech synthesis engine (Google Text-to-Speech API) and convey it to the user. If the device has a display, it can also display it.
[1335] Input: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[1336] Output: Voice playback (e.g. "Today's weather is sunny. The temperature is 24 degrees.") and display
[1337] Step 6:
[1338] The device receives a reminder notification at the specified time and notifies the user. The server sends pre-set reminder information to the device at the specified time. The device notifies the user with a voice or pop-up notification such as, "It's 8 o'clock. It's time to take your medicine."
[1339] Input: Reminder information (e.g., every day at 8:00)
[1340] Output: Audio and popup notification (e.g. "It's 8 o'clock. Time to take your medicine.")
[1341] Step 7:
[1342] When a user presses the emergency button, the device sends an emergency signal to the server. When the device sends the emergency signal, the server notifies the emergency contact. The emergency contact is sent a message via SMS, phone call, or email saying, "The user has pressed the emergency button. Please check immediately."
[1343] Input: Press emergency button
[1344] Output: Emergency notification message (e.g. "The user pressed the emergency button. Please check immediately.")
[1345] The above is the specific flow of processing steps for the "Generative AI Elderly Companion" system.
[1346] (Application example 1)
[1347] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1348] There is a need for support systems that can reduce the difficulty elderly people have in finding products in physical stores and enable them to shop safely and efficiently. Conventional support systems have difficulty providing the information elderly people need quickly, and do not have sufficient reminder or emergency contact functions.
[1349] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1350] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for providing product location information to support the elderly in a physical store, and means for notifying the user of reminders set in the physical store, thereby reducing the difficulty for the elderly when searching for products in a physical store and enabling them to shop safely and efficiently.
[1351] "Basic information" refers to individual information such as the user's name, age, address, health status, and emergency contact information.
[1352] "Individual needs" are information about a user's specific requirements or preferences, such as when to take medication or when to schedule regular health checks.
[1353] "Voice input" refers to data input to the system by a user speaking.
[1354] "Text data" is character information generated by analyzing voice input.
[1355] A "server" is a computer system that manages and processes data sent by users.
[1356] "Response data" is reply data that the server generates and sends in response to a user request.
[1357] "Reminder information" is information for notifying the user at a specified time.
[1358] "Emergency contacts" are information about people or organizations that the user wants to contact in the event of an emergency.
[1359] An "emergency notification" is a message sent from the server when an emergency occurs.
[1360] A "brick and mortar store" is a sales or service establishment located in a physical location.
[1361] "Product location information" refers to information about the location where a specific product is located within a physical store.
[1362] A "generative AI model" is an artificial intelligence algorithm model that generates text data in response to user requests.
[1363] MODE FOR CARRYING OUT THE INVENTION
[1364] The following describes an embodiment of the present invention.
[1365] System Program Overview
[1366] The "Generative AI Elderly Companion" system, which is the subject of this invention, has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. Among the functions of this system, we will explain the important functions that support the elderly in brick-and-mortar stores in particular.
[1367] Hardware and software used
[1368] Smartphone (iOS, Android): A device that allows users to input and receive information.
[1369] In-store robots (e.g., Pepper): Provide an interface for users to obtain and receive information within the store.
[1370] Server: Manages user information, performs voice recognition, and trains and updates generative AI models.
[1371] Speech recognition engine (Google Cloud Speech-to-Text): Converts voice input into text data.
[1372] Generative AI model (GPT-4 by OpenAI): Analyzes and generates text data based on user requests.
[1373] Database (MySQL): Stores user basic information, individual needs, reminder information, emergency contacts, etc.
[1374] Weather API (OpenWeatherMap): Used to retrieve the weather information desired by the user.
[1375] System procedures and operations
[1376] Manage and store basic information and needs
[1377] The server stores basic information and individual needs entered by the user through their device (smartphone or robot) in a MySQL database. Users can enter information by voice or text.
[1378] Converts voice input into text and sends it to the server
[1379] The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the user's voice input into text data and send it to the server.
[1380] Receiving and displaying response data
[1381] The server analyzes the received text data using a generative AI model (GPT-4) and generates appropriate response data. For example, when a user asks, "Where is this product?", the server obtains the product's location information and sends it to the device, saying, "Product A is in aisle 3." The device then provides this to the user via voice or a screen display.
[1382] Reminder function
[1383] The reminder information set by the user is stored on the server, and when the specified time comes, the server sends the reminder information to the device and generates a notification, which then notifies the device with a voice or a pop-up message saying, "It's time to take your medicine."
[1384] Emergency contact function
[1385] In the event of an emergency, the user presses the emergency button on the device, and the device sends an emergency signal to the server, which then sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[1386] Examples of specific examples and prompts
[1387] Specific examples
[1388] For example, when a user asks a smartphone or robot, "Where is this product?", the speech recognition engine (Google Cloud Speech-to-Text) converts the speech into text and sends it to the server. The server then analyzes it using a generative AI model (GPT-4), obtains the product location information, and generates a response such as "Product A is in aisle 3," which the smartphone or robot then plays back.
[1389] Prompt Sentence Examples
[1390] User: "Where is this item?"
[1391] System: "Converts speech to text and sends it to the server..."
[1392] Generative AI model: "Analyzing product information for user search..."
[1393] Response: "Item A is in aisle 3"
[1394] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to enjoy shopping in brick-and-mortar stores more comfortably and safely.
[1395] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1396] Step 1: Manage and save your basic information and needs
[1397] Users use their smartphones or in-store robots to input their basic information (name, age, address, health condition, emergency contact information) and individual needs (medicine dosing times, regular health check schedules) by voice or text. This input data is converted into text by a speech recognition engine (Google Cloud Speech-to-Text) on the device. The converted text data is sent to the server and stored in a MySQL database.
[1398] Input: User voice or text input
[1399] Data processing: Converting voice to text
[1400] Output: Send and save text data to the server
[1401] Step 2: Convert speech to text and send to server
[1402] The user asks a question by voice, such as "Where is this item?" The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the voice into text data and sends the text data to the server. The server receives this text data.
[1403] Input: User voice input
[1404] Data processing: Converting voice to text
[1405] Output: Send text data to the server
[1406] Step 3: Parsing text data and generating a response
[1407] The server uses a generative AI model (GPT-4) to analyze the received text data and generate an appropriate response. For example, it retrieves product location information from a database based on the user's question and generates a response such as "Product A is in aisle 3."
[1408] Input: Text data
[1409] Data computation: Analysis using generative AI models
[1410] Output: Generated response data
[1411] Step 4: Send and display response data
[1412] The server sends the generated response data to the terminal. The terminal receives this response data and uses a speech synthesis engine to play it back as a voice such as "Product A is in aisle 3," and also displays it on the screen.
[1413] Input: Response data
[1414] Data processing: voice synthesis and screen display
[1415] Output: Speech and text display
[1416] Step 5: Execute the reminder function
[1417] The server sends the reminder information set by the user to the device at the specified time. The device receives the reminder notification and notifies the user by voice, saying "It's time to take your medicine." A pop-up notification is also displayed.
[1418] Input: Reminder information and designated time
[1419] Data Processing: Notification Generation
[1420] Output: Sound and popup notification
[1421] Step 6: Implementing emergency contact functions
[1422] When a user presses the emergency button, the device sends an emergency signal to the server, which receives the signal and sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[1423] Input: Emergency signal
[1424] Data Processing: Emergency notification generation and transmission
[1425] Output:Notify emergency contacts
[1426] The above processing steps enable users to quickly obtain the information they need in a physical store, allowing them to enjoy shopping safely and comfortably.
[1427] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1428] The "Generative AI Elderly Companion" system, which is the subject of this invention, manages and stores the user's basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, making conversational interactions more personalized. It also supports the daily lives of the elderly through reminder and emergency contact functions.
[1429] Manage and store basic information and needs
[1430] server:
[1431] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[1432] Device:
[1433] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[1434] Converts voice input into text and sends it to the server
[1435] Device:
[1436] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[1437] server:
[1438] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[1439] Receiving and displaying response data
[1440] Device:
[1441] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[1442] Reminder function
[1443] server:
[1444] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[1445] Device:
[1446] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[1447] Emergency contact function
[1448] Device:
[1449] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[1450] server:
[1451] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[1452] Emotion recognition function
[1453] Device:
[1454] The device sends the user's voice input to the emotion engine, which analyzes the tone, speed, volume, etc. of the voice to determine the user's emotional state.
[1455] server:
[1456] The server analyzes the emotional information received from the emotion engine and generates an appropriate response based on it. For example, if a user says in a sad voice, "I'm not feeling well today," the server will generate an empathetic response such as, "What's wrong? Tell me your story."
[1457] Device:
[1458] Once the response data based on emotion recognition is sent from the server, the device displays or plays it back to the user, making interactions with the user more natural and personalized.
[1459] Specific use cases
[1460] Usage example 1:
[1461] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[1462] Usage example 2:
[1463] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[1464] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and lead a safe and fulfilling daily life. By incorporating an emotion engine, the system will be able to respond more sensitively to the user's feelings, providing a higher level of satisfaction.
[1465] The processing flow will be explained below.
[1466] Step 1:
[1467] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[1468] Step 2:
[1469] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[1470] Step 3:
[1471] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[1472] Step 4:
[1473] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[1474] Step 5:
[1475] The device receives voice input, converts the voice into text data using a speech recognition engine, and sends the converted text data to the server.
[1476] Step 6:
[1477] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[1478] Step 7:
[1479] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[1480] Step 8:
[1481] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[1482] Step 9:
[1483] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker. The device plays back "Today's weather is sunny. The temperature is 24 degrees."
[1484] Step 10:
[1485] The emotion engine analyzes the tone, rate, and volume of the user's voice to identify the user's emotional state: if the user speaks in an anxious voice, it will be identified as in an "anxious" state.
[1486] Step 11:
[1487] The server analyzes the emotional information received from the emotion engine and generates an appropriate response. For example, if a user says in a sad voice, "I'm not feeling well today," the server generates a response that shows empathy, such as, "What's wrong? Tell me your story."
[1488] Step 12:
[1489] The response data containing the generated emotional response is sent from the server to the device, which then plays it back to the user, saying, "What's wrong? Tell me what you think."
[1490] Step 13:
[1491] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[1492] Step 14:
[1493] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[1494] Step 15:
[1495] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[1496] Step 16:
[1497] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[1498] Step 17:
[1499] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[1500] Step 18:
[1501] The user presses the emergency button. The emergency button on the device is operated.
[1502] Step 19:
[1503] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[1504] Step 20:
[1505] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[1506] The above are the specific processing steps for carrying out the invention.
[1507] Example 2
[1508] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1509] Elderly people require various forms of support in their daily lives, but current technology makes it difficult to provide personalized assistance that fully reflects the individual needs and emotional state of each elderly person. Furthermore, systems for responding quickly and appropriately in emergencies are inadequate. Therefore, there is a need to create an environment where elderly people can live safely and without feeling lonely.
[1510] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1511] In this invention, the server includes: means for managing and saving basic information and individual needs of a user; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying or playing a notification to the user at a specified time; means for sending an emergency contact request to the server and for the server to send an emergency notification to a pre-set emergency contact; means for analyzing the tone, speed and volume of the user's voice to identify the user's emotional state; and means for generating a personalized response based on the emotional state and providing it to the user.
[1512] This will enable personalized assistance to be provided in response to the individual needs and emotional state of the elderly, as well as a rapid and appropriate response in emergencies.
[1513] "Basic user information" refers to basic information about a user, such as name, age, address, health status, and emergency contact information.
[1514] "Individual needs" refers to information about the support and services that a user individually needs, such as when to take medication or schedule regular health checks.
[1515] "Voice input" refers to voice data input when a user speaks to the system through a microphone.
[1516] "Text data" refers to voice input converted into character string data using a voice recognition engine or the like.
[1517] "Server" refers to a computer system that receives data from users and stores, analyzes, and processes it.
[1518] "Response data" refers to information that the server generates based on a user request and returns to the user.
[1519] "Reminder information" refers to information for notifying the user at a specific time or timing.
[1520] "Emergency contact request" refers to an emergency signal sent through the system when a user faces an emergency.
[1521] "Emergency notification" refers to an emergency message sent from the server to a pre-defined emergency contact.
[1522] "Emotional state" refers to the user's emotions estimated based on an analysis of the tone, rate, volume, etc. of the voice.
[1523] "Personalized response" refers to an individualized response that is generated based on the user's basic information, individual needs, and emotional state.
[1524] The present invention, the "Generative AI Elderly Companion" system, provides personalized support that reflects the individual needs and emotional state of elderly people, reducing feelings of loneliness in daily life and creating a safe living environment.
[1525] Hardware and software used
[1526] The server is a high-performance computer system, such as an AWS EC2 instance or Microsoft Azure VM. MySQL or PostgreSQL is used as the database management system. Google Cloud Speech-to-Text or Amazon Transcribe is used as the speech recognition engine, and IBM Watson Tone Analyzer or Microsoft Azure Cognitive Services is used as the emotion recognition engine.
[1527] Manage and store basic information and needs
[1528] The user uses the device to enter basic information (name, age, address, health status, emergency contact information). The device provides voice or text input, verifies the entered information, and sends it to the server. The server stores the received basic information in a database and provides personalized support based on this information.
[1529] Converts voice input into text and sends it to the server
[1530] When a user asks "What's the weather like today?", the device uses a voice recognition engine to convert the voice into text data and sends it to the server. The server then analyzes the received text data and obtains weather information using a weather API provided by a third party.
[1531] Receiving and displaying response data
[1532] The response data (such as weather information) sent from the server is received by the terminal, which then provides it to the user through voice or a display device. For example, it may say, "Today's weather is sunny. The temperature is 24 degrees."
[1533] Reminder function
[1534] When a user instructs the device to "set a reminder to take my medicine at 8 o'clock every morning," the device converts this instruction into text data and sends it to the server. The server stores the reminder information in a database and sends a reminder notification to the device at the specified time (8 o'clock every morning). The device then notifies the user by voice or a pop-up notification, saying, "It's 8 o'clock. It's time to take your medicine."
[1535] Emergency contact function
[1536] In an emergency, the user presses the emergency button on the device. The device then sends an emergency signal to the server. The server then contacts pre-registered emergency contacts (family members or caregivers) via SMS, phone, email, etc. It then sends a message saying, "The user has pressed the emergency button. Please check immediately."
[1537] Emotion recognition function
[1538] If a user says "I'm feeling bad today" in a sad voice, the device sends this voice data to an emotion recognition engine, which analyzes the tone, speed, and volume of the voice to identify the emotional state. The server uses the emotional information to generate a personalized response that shows empathy, such as "What's wrong? Tell me your story," and the device plays this aloud to the user.
[1539] Specific use cases
[1540] Usage example 1:
[1541] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[1542] Usage example 2:
[1543] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[1544] The above is an embodiment of the present invention. This allows elderly people to live a safe and fulfilling daily life and reduce their sense of loneliness. Furthermore, the introduction of an emotion engine enables the system to respond in a way that is sensitive to the user's feelings, providing a higher level of satisfaction.
[1545] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1546] Steps for entering and saving basic information
[1547] Step 1:
[1548] The user uses the device to input basic information (name, age, address, health status, emergency contact information). The input method can be selected as voice input or text input. For example, if the user inputs "My name is Ichiro Tanaka, and I'm 75 years old" by voice, the device will capture the voice data.
[1549] input:
[1550] Basic information entered by the user (name, age, address, health status, emergency contact information)
[1551] output:
[1552] Audio or text data
[1553] Step 2:
[1554] When the device receives voice input, it uses a speech recognition engine (such as Google Cloud Speech-to-Text) to convert the voice into text data. For example, voice data such as "My name is Tanaka Ichiro and I'm 75 years old" is converted into text data such as "Name: Tanaka Ichiro, Age: 75."
[1555] input:
[1556] Audio data
[1557] output:
[1558] Text data
[1559] Step 3:
[1560] The device sends the converted text data to the server. Specifically, data transmission is performed using an API request.
[1561] input:
[1562] Text data
[1563] output:
[1564] API request (text data)
[1565] Step 4:
[1566] The server stores the received text data in a database, for example, using MySQL or PostgreSQL as data records.
[1567] input:
[1568] Text data
[1569] output:
[1570] Data Records
[1571] Processing steps for converting voice input to text and sending it to the server
[1572] Step 1:
[1573] The user asks a question by voice, "What's the weather like today?" This voice data is acquired by the terminal.
[1574] input:
[1575] User voice input
[1576] output:
[1577] Audio data
[1578] Step 2:
[1579] The device uses a voice recognition engine to convert the acquired voice data into text data. Specifically, the voice data "What's the weather like today?" is converted into text data "What's the weather like today?"
[1580] input:
[1581] Audio data
[1582] output:
[1583] Text data
[1584] Step 3:
[1585] The device sends the converted text data to the server via an API request.
[1586] input:
[1587] Text data
[1588] output:
[1589] API request (text data)
[1590] Step 4:
[1591] The server analyzes the received text data and retrieves the corresponding information. For example, to retrieve weather information, the server sends a request to a weather API provided by a third party and receives response data such as "Today's weather is sunny and the temperature is 24 degrees."
[1592] input:
[1593] Text data
[1594] output:
[1595] Response data (weather information)
[1596] Processing steps for receiving and displaying response data
[1597] Step 1:
[1598] The server sends the acquired response data (such as weather information) to the terminal.
[1599] input:
[1600] Response data (weather information)
[1601] output:
[1602] API response (response data)
[1603] Step 2:
[1604] The device analyzes the received response data and provides it to the user through a voice output device, for example, by playing back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[1605] input:
[1606] API response (response data)
[1607] output:
[1608] Audio data
[1609] Reminder function processing steps
[1610] Step 1:
[1611] The user instructs, "Set a reminder to take my medicine at 8 o'clock every morning." The device receives this voice data.
[1612] input:
[1613] User voice instructions
[1614] output:
[1615] Audio data
[1616] Step 2:
[1617] The device uses a speech recognition engine to convert voice data into text data. For example, voice data such as "Set a reminder to take my medicine at 8 o'clock every morning" is converted into text data such as "Take my medicine at 8 o'clock every morning."
[1618] input:
[1619] Audio data
[1620] output:
[1621] Text data
[1622] Step 3:
[1623] The device sends the converted text data to the server via an API request.
[1624] input:
[1625] Text data
[1626] output:
[1627] API request (text data)
[1628] Step 4:
[1629] The server stores the received reminder information in a database and configures it to generate notifications at a specified time, for example, every morning at 8:00.
[1630] input:
[1631] Text data
[1632] output:
[1633] Database records and timer settings
[1634] Step 5:
[1635] When the specified time arrives, the server sends a notification to the device, such as "It's time to take your medicine."
[1636] input:
[1637] Timer Event
[1638] output:
[1639] API response (notification data)
[1640] Step 6:
[1641] The device will then provide the received notification data to the user via voice or pop-up notification, for example, "It's 8 o'clock. Time to take your medicine."
[1642] input:
[1643] API response (notification data)
[1644] output:
[1645] Audio data or popup notification
[1646] Emergency Contact Function Processing Steps
[1647] Step 1:
[1648] The user presses the emergency button, and the device detects this emergency signal.
[1649] input:
[1650] Emergency button input
[1651] output:
[1652] Emergency Signal Data
[1653] Step 2:
[1654] The device sends an emergency signal to the server, which is done via an API request.
[1655] input:
[1656] Emergency Signal Data
[1657] output:
[1658] API Request (Emergency Signal Data)
[1659] Step 3:
[1660] The server analyzes the received emergency signal and sends an emergency notification to pre-defined emergency contacts, for example, by SMS, phone call, or email, with a message saying, "The user has pressed the emergency button. Please check immediately."
[1661] input:
[1662] API Request (Emergency Signal Data)
[1663] output:
[1664] Emergency notification data (SMS, phone, email)
[1665] Emotion Recognition Processing Steps
[1666] Step 1:
[1667] The user inputs "I feel bad today" in a sad voice. This voice data is acquired by the terminal.
[1668] input:
[1669] User voice input
[1670] output:
[1671] Audio data
[1672] Step 2:
[1673] The device sends the captured voice data to an emotion recognition engine, which analyzes the tone, rate, and volume of the voice to identify the emotional state.
[1674] input:
[1675] Audio data
[1676] output:
[1677] Emotional state data
[1678] Step 3:
[1679] The server generates a personalized response based on the emotional state data received from the emotion recognition engine, for example, an empathetic response such as "What's wrong? Tell me your story."
[1680] input:
[1681] Emotional state data
[1682] output:
[1683] Response data
[1684] Step 4:
[1685] The server transmits the generated response data to the terminal.
[1686] input:
[1687] Response data
[1688] output:
[1689] API response (response data)
[1690] Step 5:
[1691] The device will then play back the received response data in voice, for example, "What's wrong? Tell me what happened."
[1692] input:
[1693] API response (response data)
[1694] output:
[1695] Audio data
[1696] (Application example 2)
[1697] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1698] Elderly people have greater difficulty navigating physical stores, searching for products, and responding to emergencies. They also often feel anxious about asking store staff questions and navigating the store. Current systems struggle to adequately address these issues, preventing seniors from enjoying shopping with peace of mind.
[1699] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for managing and saving a user's basic information and individual needs; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying a notification to the user at a specified time; means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact; means for navigating the store using the user's current location and store map data; and means for analyzing the user's voice input and generating a personalized response based on the user's emotional state. This enables seniors to move confidently through physical stores, easily find products and services, and quickly respond to anxieties or emergencies.
[1700] Below are definitions of important terms included in the patent claims, rewritten to suit the application example.
[1701] "Means for managing and storing a user's basic information and individual needs" refers to a means for storing basic information such as the user's name, age, address, health condition, and emergency contact information, as well as individual needs such as medication times and regular health check schedules, in a database, and for managing and updating this information as needed.
[1702] "Means for converting user's voice input into text data and sending it to the server" refers to means for converting information input by the user by voice into text data using voice recognition technology and sending that text data to the server.
[1703] "Means for receiving response data from the server and displaying or playing it back to the user" refers to means for receiving information sent from the server (e.g., weather information or store directions) and providing it to the user visually or audibly.
[1704] The "means for saving specified reminder information on a server and displaying a notification to the user at a specified time" refers to a means for saving reminder information set by a user on a server and sending a notification to the user at a specified time based on that reminder information.
[1705] "Means for sending an emergency contact request to a server, and for the server to send an emergency notification to pre-registered emergency contacts" refers to a means for a user to send an emergency contact request to a server by pressing an emergency button, etc., and for the server to send an emergency notification by phone or SMS to pre-registered emergency contacts based on that request.
[1706] "Means for navigating within a store using the user's current location and store map data" refers to a means for guiding the user to their desired location by using the location information of the user's smartphone or device and combining it with map data within the store.
[1707] "Means for analyzing a user's voice input and generating a personalized response based on the user's emotional state" refers to means for analyzing the tone, rate, and content of a user's voice to identify the user's emotional state, and generating and providing an appropriate response or message accordingly.
[1708] System Program
[1709] The purpose of this invention, the "Generative AI Elderly Companion Shopper" system, is to help seniors shop comfortably in brick-and-mortar stores. This system manages and stores users' basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. It also supports the daily lives of seniors through reminder functions, emergency contact functions, and emotion recognition functions.
[1710] Hardware and Software Use
[1711] Hardware:
[1712] Smartphone (with camera, microphone, and speaker)
[1713] software:
[1714] Google Cloud Speech-to-Text API (voice recognition)
[1715] Google Cloud Natural Language API (Text Analysis)
[1716] Twilio API (emergency contact function)
[1717] Emotion Engine (e.g. Affectiva's SDK)
[1718] Firebase (database)
[1719] Data processing and calculation
[1720] Managing and storing your basic information and individual needs:
[1721] The server stores the basic information entered by the user (name, age, address, health status, emergency contacts) and individual needs (medication times, schedule for regular health checks) in a Firebase database, allowing the system to provide personalized support to the user.
[1722] Convert speech to text and send to server:
[1723] When a user asks, "Where is the restroom?", the device uses the Google Cloud Speech-to-Text API to convert speech to text data and send it to the server, which then uses the Google Cloud Natural Language API to parse the text and generate an appropriate response.
[1724] Receive and display response data:
[1725] The response data sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may provide voice guidance such as, "To find the restroom, go straight down the corridor on the right."
[1726] Reminder function:
[1727] The server stores the reminder information set by the user (e.g., take medicine at 8 o'clock every day) in Firebase. When the specified time arrives, the device sends a voice or pop-up notification saying, "It's 8 o'clock. Time to take your medicine."
[1728] Emergency contact features:
[1729] If a user is about to fall inside the store, they can press the emergency button on their device, which sends an emergency signal to the server. The server then uses the Twilio API to send a notification to registered emergency contacts (family members or caregivers). The emergency notification is sent via SMS or phone.
[1730] Emotion recognition function:
[1731] If a user says, "I'm not feeling well today," the device uses its emotion engine to analyze the speech and identify the user's emotional state. The server then generates an appropriate response based on this, saying, "What's wrong? Tell me."
[1732] Examples of specific examples and prompts
[1733] As a specific example, consider the case where an elderly person named Mr. A gets lost in a store.
[1734] Person A asks the app, "Where is the restroom?" inside the store. Voice recognition is activated, and the speech is converted into text and sent to the server.
[1735] The server identifies the location of the restroom from the store's map data and generates a navigation plan along with Mr. A's current location.
[1736] The app will provide voice guidance saying, "To get to the restroom, go straight down the corridor on the right."
[1737] If the emotion engine detects Mr. A's anxiety, the app will ask, "If you can't find it, I'll call a store employee."
[1738] Example prompt sentence:
[1739] The user asks the smartphone microphone, "Where is the restroom?" The voice input is converted into text and sent to the server. The system should refer to the user's current location and a map of the store, calculate the shortest route, and provide voice guidance. If the user asks in an anxious voice, the system should generate an additional response asking, "If you can't find it, would you like me to call a store employee?"
[1740] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1741] Step 1:
[1742] A user asks into the microphone of their smartphone, "Where is the restroom?" A voice input occurs, and the device collects this voice input data.
[1743] Input: User voice input
[1744] Output: Audio data
[1745] Step 2:
[1746] The device uses the Google Cloud Speech-to-Text API to convert the voice data into text data, which is then sent to the server.
[1747] Input: Audio data
[1748] Output: Text data
[1749] Step 3:
[1750] The server analyzes the received text data using the Google Cloud Natural Language API, and as a result, it understands that the user is looking for a restroom.
[1751] Input: Text data
[1752] Output: Analysis result (intent)
[1753] Step 4:
[1754] The server identifies the user's location by referencing the user's current location and the store's map data, and calculates the shortest route based on this.
[1755] Input: User's current location data, store map data
[1756] Output: Shortest route data
[1757] Step 5:
[1758] The server uses the calculated shortest route data to generate a navigation plan to help the user reach their destination easily. The generated navigation plan is created in voice and text format.
[1759] Input: Shortest route data
[1760] Output: Navigation plan (voice and text)
[1761] Step 6:
[1762] The server sends the generated navigation plan to the device, which receives it and provides voice and text guidance to the user.
[1763] Input: Navigation plan
[1764] Output: User instructions
[1765] Step 7:
[1766] At the same time, the device analyzes the user's voice input with an emotion engine to identify their emotional state. If the device detects that the user is asking a question in an anxious voice, the server generates an additional response (e.g., "If you can't find it, should I call a store clerk?").
[1767] Input: User's voice data
[1768] Output: Emotional state, additional responses
[1769] Step 8:
[1770] As additional responses are generated, the server sends them to the terminal, which displays or plays them audibly to the user.
[1771] Input: Additional response data
[1772] Output: Additional instructions for the user
[1773] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1774] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1775] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1776] [Fourth embodiment]
[1777] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1778] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1779] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1780] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1781] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1782] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1783] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1784] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1785] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1786] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1787] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1788] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1789] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1790] The "Generative AI Elderly Companion" system of this invention has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back the response data from the server to the user via their device. This system can also support the lives of the elderly through reminder and emergency contact functions.
[1791] Manage and store basic information and needs
[1792] server:
[1793] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[1794] Device:
[1795] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[1796] Converts voice input into text and sends it to the server
[1797] Device:
[1798] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[1799] server:
[1800] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[1801] Receiving and displaying response data
[1802] Device:
[1803] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[1804] Reminder function
[1805] server:
[1806] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[1807] Device:
[1808] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[1809] Emergency contact function
[1810] Device:
[1811] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[1812] server:
[1813] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[1814] Specific use cases
[1815] Usage example 1:
[1816] The user asks, "When is my next hospital appointment?" The device recognizes the speech and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "My next hospital appointment is tomorrow at 10:00." The device plays this back as audio.
[1817] Usage example 2:
[1818] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately."
[1819] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and enable them to live a safe and fulfilling daily life.
[1820] The processing flow will be explained below.
[1821] Step 1:
[1822] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[1823] Step 2:
[1824] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[1825] Step 3:
[1826] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[1827] Step 4:
[1828] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[1829] Step 5:
[1830] The device receives the voice input, converts it into text data using a speech recognition engine, and sends the converted text data to the server.
[1831] Step 6:
[1832] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[1833] Step 7:
[1834] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[1835] Step 8:
[1836] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[1837] Step 9:
[1838] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker.
[1839] Step 10:
[1840] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[1841] Step 11:
[1842] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[1843] Step 12:
[1844] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[1845] Step 13:
[1846] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[1847] Step 14:
[1848] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[1849] Step 15:
[1850] The user presses the emergency button. The emergency button on the device is operated.
[1851] Step 16:
[1852] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[1853] Step 17:
[1854] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[1855] The above are the specific processing steps for carrying out the invention.
[1856] Example 1
[1857] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1858] To help seniors live safe and fulfilling daily lives without feeling lonely, systems with personalized information provision, reminder functions, and emergency contact functions are needed. However, existing systems are often difficult for seniors to use, and information provision is fragmented and unintegrated. Furthermore, the accuracy of voice input and dialogue interaction is low, and external information acquisition is often difficult. This makes it difficult for seniors to receive appropriate support.
[1859] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1860] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for the server to obtain information from an external information service, and means for converting voice into text data using a voice recognition engine. This allows elderly people to intuitively operate the device, smoothly utilize personalized information provision, reminder functions, and emergency contact functions, and also makes it easy to obtain external information, enabling them to live a safe and fulfilling life.
[1861] "User's basic information and individual needs" refers collectively to individually personalized data such as the user's name, age, address, health status, emergency contact information, medication times, and schedule for regular health checks.
[1862] "Means for converting voice input into text data" is a general term for devices and software that use voice recognition technology to convert a user's voice into text data and send that text data to a server.
[1863] "Response data from the server" refers to information and messages analyzed and generated by the server, including data obtained from external information services.
[1864] "Reminder information" refers to notification content that a user should receive at a specific time or date, such as when to take medicine or a schedule for regular health checks.
[1865] An "emergency contact request" is a request made by a user to quickly seek assistance in an emergency, and is an emergency signal sent from the terminal to the server.
[1866] "Means of obtaining information from external information services" is a general term for the technologies and methods that allow a server to retrieve necessary information from public APIs and databases on the Internet.
[1867] A "speech recognition engine" is a piece of software or hardware that understands a user's voice input and converts it into text data.
[1868] "Dialogue interaction" refers to two-way communication between a user and a system based on voice input and text data.
[1869] This invention relates to a generative AI "Elderly Companion" system that supports the lives of the elderly, and has a series of functions to manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. This system also supports the lives of the elderly through reminder functions and emergency contact functions.
[1870] In this embodiment of the present invention, the system is configured by a server, terminals, and users communicating with each other. The server is hosted on the cloud and manages user information through a database. Specifically, a MySQL database is used, and data security is ensured by SSL encryption.
[1871] Manage and store basic information and needs
[1872] Device:
[1873] Users enter basic information such as their name, age, address, health condition, and emergency contact details, as well as their individual needs such as medication times and schedules for regular health checks, through the device's interface. Either voice input or text input is possible, using a smartphone or a dedicated device.
[1874] server:
[1875] The server receives the basic information and needs information sent from the device and stores it in a MySQL database. The server uses this data to provide personalized support to the user.
[1876] Converts voice input into text and sends it to the server
[1877] Device:
[1878] When a user asks a question or gives an instruction by voice, such as "When is my next doctor's appointment?", the device converts the voice into text data using Google's speech recognition API. The converted text data is then sent to the server. The user can use a headset or built-in microphone to do this.
[1879] server:
[1880] The server analyzes the received text data and obtains the necessary information (for example, reservation information or weather information). To obtain the weather information, it uses a public API such as the OpenWeatherMap API. The server then sends the analysis results to the terminal as response data.
[1881] Receiving and displaying response data
[1882] Device:
[1883] The device that receives the response data from the server uses a speech synthesis engine (e.g., Google Text-to-Speech API) to convert the text data into speech and tells the user, "Your next hospital appointment is tomorrow at 10:00." If the device has a display, it can also display the information on the screen.
[1884] Reminder function
[1885] server:
[1886] Reminder information is stored on the server based on a schedule specified by the user, and the server sends the reminder information to the device at the specified time.
[1887] Device:
[1888] The device will receive a reminder notification at the specified time, informing the user via voice or pop-up notification, "It's 8 o'clock. Time to take your medicine."
[1889] Emergency contact function
[1890] Device:
[1891] When a user faces an emergency, they press the emergency button on their device, which instantly sends an emergency signal to the server.
[1892] server:
[1893] When the server receives an emergency signal, it sends a message to pre-registered emergency contacts (family members, caregivers) via SMS, phone, or email saying, "The user has pressed the emergency button. Please check immediately."
[1894] Specific examples of actions and prompts
[1895] Usage example 1:
[1896] The user asks, "What's the weather like today?" The device converts the speech to text and sends it to the server. The server retrieves weather information using a weather information service, and then sends the information to the device: "Today's weather is sunny. The temperature is 24 degrees."
[1897] Usage example 2:
[1898] The user asks, "When is my next hospital appointment?" The device converts the speech to text and sends it to the server. The server retrieves the appointment information from its database and responds, "My next hospital appointment is tomorrow at 10:00." The device plays this back aloud.
[1899] This invention enables elderly people to easily access personalized information, reminder functions, and emergency contact functions through an interface that they can operate intuitively. This system reduces the sense of loneliness felt by elderly people and supports safe and fulfilling lives.
[1900] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1901] Step 1:
[1902] The user enters basic information and individual needs into the device, such as name, age, address, health condition, emergency contact information, medication times, and schedule for regular health checks, by voice or text. This information is temporarily stored on the device and then sent to the server.
[1903] Input: Name, age, address, health status, emergency contact, needs
[1904] Output: Basic information and needs sent to the server
[1905] Step 2:
[1906] The device receives voice input and converts it into text data. For example, if a user says, "What's the weather like today?", the device converts the voice into text using a speech recognition engine (Google Speech-to-Text API). The text data is then sent to the server.
[1907] Input: Voice input (e.g. "What's the weather like today?")
[1908] Output: Text data (e.g. "What's the weather like today?")
[1909] Step 3:
[1910] The server analyzes the received text data and obtains the corresponding information. For example, if the text data "What's the weather like today?" is received, the server will call an external weather information API (e.g., OpenWeatherMap API) to obtain the weather information and obtain the results.
[1911] Input: Text data (e.g., "What's the weather like today?")
[1912] Output: Weather information (e.g. sunny, temperature 24 degrees)
[1913] Step 4:
[1914] The server sends the acquired information to the terminal as text data. The server generates response data based on the analysis results and sends it to the terminal.
[1915] Input: Weather information (e.g. sunny, temperature 24 degrees)
[1916] Output: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[1917] Step 5:
[1918] The device will play back the response data received from the server as audio. The device will convert the text data into audio using a speech synthesis engine (Google Text-to-Speech API) and convey it to the user. If the device has a display, it can also display it.
[1919] Input: Response data (e.g. "Today's weather is sunny. The temperature is 24 degrees.")
[1920] Output: Voice playback (e.g. "Today's weather is sunny. The temperature is 24 degrees.") and display
[1921] Step 6:
[1922] The device receives a reminder notification at the specified time and notifies the user. The server sends pre-set reminder information to the device at the specified time. The device notifies the user with a voice or pop-up notification such as, "It's 8 o'clock. It's time to take your medicine."
[1923] Input: Reminder information (e.g., every day at 8:00)
[1924] Output: Audio and popup notification (e.g. "It's 8 o'clock. Time to take your medicine.")
[1925] Step 7:
[1926] When a user presses the emergency button, the device sends an emergency signal to the server. When the device sends the emergency signal, the server notifies the emergency contact. The emergency contact is sent a message via SMS, phone call, or email saying, "The user has pressed the emergency button. Please check immediately."
[1927] Input: Press emergency button
[1928] Output: Emergency notification message (e.g. "The user pressed the emergency button. Please check immediately.")
[1929] The above is the specific flow of processing steps for the "Generative AI Elderly Companion" system.
[1930] (Application example 1)
[1931] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1932] There is a need for support systems that can reduce the difficulty elderly people have in finding products in physical stores and enable them to shop safely and efficiently. Conventional support systems have difficulty providing the information elderly people need quickly, and do not have sufficient reminder or emergency contact functions.
[1933] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1934] In this invention, the server includes means for managing and saving the user's basic information and individual needs, means for converting the user's voice input into text data and sending it to the server, means for receiving response data from the server and displaying or playing it back to the user, means for saving designated reminder information to the server and displaying a notification to the user at a designated time, means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact, means for providing product location information to support the elderly in a physical store, and means for notifying the user of reminders set in the physical store, thereby reducing the difficulty for the elderly when searching for products in a physical store and enabling them to shop safely and efficiently.
[1935] "Basic information" refers to individual information such as the user's name, age, address, health status, and emergency contact information.
[1936] "Individual needs" are information about a user's specific requirements or preferences, such as when to take medication or when to schedule regular health checks.
[1937] "Voice input" refers to data input to the system by a user speaking.
[1938] "Text data" is character information generated by analyzing voice input.
[1939] A "server" is a computer system that manages and processes data sent by users.
[1940] "Response data" is reply data that the server generates and sends in response to a user request.
[1941] "Reminder information" is information for notifying the user at a specified time.
[1942] "Emergency contacts" are information about people or organizations that the user wants to contact in the event of an emergency.
[1943] An "emergency notification" is a message sent from the server when an emergency occurs.
[1944] A "brick and mortar store" is a sales or service establishment located in a physical location.
[1945] "Product location information" refers to information about the location where a specific product is located within a physical store.
[1946] A "generative AI model" is an artificial intelligence algorithm model that generates text data in response to user requests.
[1947] MODE FOR CARRYING OUT THE INVENTION
[1948] The following describes an embodiment of the present invention.
[1949] System Program Overview
[1950] The "Generative AI Elderly Companion" system, which is the subject of this invention, has a series of functions that manage and store the user's basic information and individual needs, convert voice input into text data and send it to a server, and display or play back response data from the server to the user via a terminal. Among the functions of this system, we will explain the important functions that support the elderly in brick-and-mortar stores in particular.
[1951] Hardware and software used
[1952] Smartphone (iOS, Android): A device that allows users to input and receive information.
[1953] In-store robots (e.g., Pepper): Provide an interface for users to obtain and receive information within the store.
[1954] Server: Manages user information, performs voice recognition, and trains and updates generative AI models.
[1955] Speech recognition engine (Google Cloud Speech-to-Text): Converts voice input into text data.
[1956] Generative AI model (GPT-4 by OpenAI): Analyzes and generates text data based on user requests.
[1957] Database (MySQL): Stores user basic information, individual needs, reminder information, emergency contacts, etc.
[1958] Weather API (OpenWeatherMap): Used to retrieve the weather information desired by the user.
[1959] System procedures and operations
[1960] Manage and store basic information and needs
[1961] The server stores basic information and individual needs entered by the user through their device (smartphone or robot) in a MySQL database. Users can enter information by voice or text.
[1962] Converts voice input into text and sends it to the server
[1963] The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the user's voice input into text data and send it to the server.
[1964] Receiving and displaying response data
[1965] The server analyzes the received text data using a generative AI model (GPT-4) and generates appropriate response data. For example, when a user asks, "Where is this product?", the server obtains the product's location information and sends it to the device, saying, "Product A is in aisle 3." The device then provides this to the user via voice or a screen display.
[1966] Reminder function
[1967] The reminder information set by the user is stored on the server, and when the specified time comes, the server sends the reminder information to the device and generates a notification, which then notifies the device with a voice or a pop-up message saying, "It's time to take your medicine."
[1968] Emergency contact function
[1969] In the event of an emergency, the user presses the emergency button on the device, and the device sends an emergency signal to the server, which then sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[1970] Examples of specific examples and prompts
[1971] Specific examples
[1972] For example, when a user asks a smartphone or robot, "Where is this product?", the speech recognition engine (Google Cloud Speech-to-Text) converts the speech into text and sends it to the server. The server then analyzes it using a generative AI model (GPT-4), obtains the product location information, and generates a response such as "Product A is in aisle 3," which the smartphone or robot then plays back.
[1973] Prompt Sentence Examples
[1974] User: "Where is this item?"
[1975] System: "Converts speech to text and sends it to the server..."
[1976] Generative AI model: "Analyzing product information for user search..."
[1977] Response: "Item A is in aisle 3"
[1978] The above is a specific embodiment for carrying out the present invention. This system allows elderly people to enjoy shopping in brick-and-mortar stores more comfortably and safely.
[1979] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1980] Step 1: Manage and save your basic information and needs
[1981] Users use their smartphones or in-store robots to input their basic information (name, age, address, health condition, emergency contact information) and individual needs (medicine dosing times, regular health check schedules) by voice or text. This input data is converted into text by a speech recognition engine (Google Cloud Speech-to-Text) on the device. The converted text data is sent to the server and stored in a MySQL database.
[1982] Input: User voice or text input
[1983] Data processing: Converting voice to text
[1984] Output: Send and save text data to the server
[1985] Step 2: Convert speech to text and send to server
[1986] The user asks a question by voice, such as "Where is this item?" The device uses a speech recognition engine (Google Cloud Speech-to-Text) to convert the voice into text data and sends the text data to the server. The server receives this text data.
[1987] Input: User voice input
[1988] Data processing: Converting voice to text
[1989] Output: Send text data to the server
[1990] Step 3: Parsing text data and generating a response
[1991] The server uses a generative AI model (GPT-4) to analyze the received text data and generate an appropriate response. For example, it retrieves product location information from a database based on the user's question and generates a response such as "Product A is in aisle 3."
[1992] Input: Text data
[1993] Data computation: Analysis using generative AI models
[1994] Output: Generated response data
[1995] Step 4: Send and display response data
[1996] The server sends the generated response data to the terminal. The terminal receives this response data and uses a speech synthesis engine to play it back as a voice such as "Product A is in aisle 3," and also displays it on the screen.
[1997] Input: Response data
[1998] Data processing: voice synthesis and screen display
[1999] Output: Speech and text display
[2000] Step 5: Execute the reminder function
[2001] The server sends the reminder information set by the user to the device at the specified time. The device receives the reminder notification and notifies the user by voice, saying "It's time to take your medicine." A pop-up notification is also displayed.
[2002] Input: Reminder information and designated time
[2003] Data Processing: Notification Generation
[2004] Output: Sound and popup notification
[2005] Step 6: Implementing emergency contact functions
[2006] When a user presses the emergency button, the device sends an emergency signal to the server, which receives the signal and sends an emergency notification via SMS, phone call, or email to registered emergency contacts.
[2007] Input: Emergency signal
[2008] Data Processing: Emergency notification generation and transmission
[2009] Output:Notify emergency contacts
[2010] The above processing steps enable users to quickly obtain the information they need in a physical store, allowing them to enjoy shopping safely and comfortably.
[2011] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2012] The "Generative AI Elderly Companion" system, which is the subject of this invention, manages and stores the user's basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, making conversational interactions more personalized. It also supports the daily lives of the elderly through reminder and emergency contact functions.
[2013] Manage and store basic information and needs
[2014] server:
[2015] The basic information entered by the user (name, age, address, health condition, emergency contact information) and individual needs (medication times, schedule for regular health checks, etc.) are stored in a database on the server side, and the server uses this data to provide personalized support.
[2016] Device:
[2017] The terminal provides an interface for users to input information. Voice input or text input is possible, and the input information is sent to a server where it is stored and managed.
[2018] Converts voice input into text and sends it to the server
[2019] Device:
[2020] When a user asks "What's the weather like today?", the device uses a speech recognition engine to convert the voice into text data and send it to the server.
[2021] server:
[2022] The server analyzes the received text data and retrieves the corresponding information (weather information in this case). Weather information is typically retrieved using a third-party weather API.
[2023] Receiving and displaying response data
[2024] Device:
[2025] The response data (such as weather information) sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may play back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[2026] Reminder function
[2027] server:
[2028] Reminder information set by the user (e.g., take medicine at 8 o'clock every day) is stored on the server. When the specified time arrives, the server sends the reminder information to the device and generates a notification.
[2029] Device:
[2030] At the specified time, the device will notify the user via voice or pop-up message saying, "It's 8 o'clock. Time to take your medicine."
[2031] Emergency contact function
[2032] Device:
[2033] In the event of an emergency, the user presses the emergency button on the device, which then sends an emergency signal to the server.
[2034] server:
[2035] When the server receives an emergency signal, it immediately contacts registered emergency contacts (family members or caregivers) and sends emergency notifications via SMS, phone, email, etc.
[2036] Emotion recognition function
[2037] Device:
[2038] The device sends the user's voice input to the emotion engine, which analyzes the tone, speed, volume, etc. of the voice to determine the user's emotional state.
[2039] server:
[2040] The server analyzes the emotional information received from the emotion engine and generates an appropriate response based on it. For example, if a user says in a sad voice, "I'm not feeling well today," the server will generate an empathetic response such as, "What's wrong? Tell me your story."
[2041] Device:
[2042] Once the response data based on emotion recognition is sent from the server, the device displays or plays it back to the user, making interactions with the user more natural and personalized.
[2043] Specific use cases
[2044] Usage example 1:
[2045] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[2046] Usage example 2:
[2047] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[2048] The above is a specific embodiment for carrying out the present invention. This system will help elderly people to reduce their sense of loneliness and lead a safe and fulfilling daily life. By incorporating an emotion engine, the system will be able to respond more sensitively to the user's feelings, providing a higher level of satisfaction.
[2049] The processing flow will be explained below.
[2050] Step 1:
[2051] The user enters basic information (name, age, address, health status, emergency contact information) through the terminal. The user enters the information on the input screen and presses the send button.
[2052] Step 2:
[2053] The terminal sends the input information to the server, which generates a data packet in text format and sends it to the server via the network.
[2054] Step 3:
[2055] The server stores the received user information in a database, verifies the accuracy of the information, and stores it in the appropriate database fields.
[2056] Step 4:
[2057] The user attempts to obtain information through voice input, for example, "What's the weather like today?"
[2058] Step 5:
[2059] The device receives voice input, converts the voice into text data using a speech recognition engine, and sends the converted text data to the server.
[2060] Step 6:
[2061] The server analyzes the received text data, retrieves corresponding information from appropriate sources (e.g., weather API), and sends queries based on the analysis to external services.
[2062] Step 7:
[2063] The server generates response data based on the information it has obtained, building a response in text format such as "Today's weather is sunny. The temperature is 24 degrees."
[2064] Step 8:
[2065] The server sends the generated response data to the terminal, encapsulating the response data in a packet in text format and sending it to the terminal.
[2066] Step 9:
[2067] The device displays the received response data to the user or plays it aloud. The text data is converted by a speech synthesis engine and transmitted to the user through the speaker. The device plays back "Today's weather is sunny. The temperature is 24 degrees."
[2068] Step 10:
[2069] The emotion engine analyzes the tone, rate, and volume of the user's voice to identify the user's emotional state: if the user speaks in an anxious voice, it will be identified as in an "anxious" state.
[2070] Step 11:
[2071] The server analyzes the emotional information received from the emotion engine and generates an appropriate response. For example, if a user says in a sad voice, "I'm not feeling well today," the server generates a response that shows empathy, such as, "What's wrong? Tell me your story."
[2072] Step 12:
[2073] The response data containing the generated emotional response is sent from the server to the device, which then plays it back to the user, saying, "What's wrong? Tell me what you think."
[2074] Step 13:
[2075] The user enters reminder information, for example, "Take my medicine at 8 o'clock every day."
[2076] Step 14:
[2077] The terminal transmits the reminder information to the server, which then packets the reminder information in text format and transmits it to the server.
[2078] Step 15:
[2079] The server stores the received reminder information in a database and sets up a schedule to generate reminder notifications at specified times.
[2080] Step 16:
[2081] When the specified reminder time arrives, the server generates a reminder notification and sends it to the device. The reminder content is sent in text format to the device.
[2082] Step 17:
[2083] The device will display or play the received reminder notification to the user. For example, "It's 8 o'clock. Time to take your medicine."
[2084] Step 18:
[2085] The user presses the emergency button. The emergency button on the device is operated.
[2086] Step 19:
[2087] The terminal transmits an emergency signal to the server, generates an emergency data packet, and transmits it to the server via the network.
[2088] Step 20:
[2089] When the server receives an emergency signal, it sends an emergency notification to the registered emergency contacts via SMS, phone call, or email informing them that "a user has pressed the emergency button."
[2090] The above are the specific processing steps for carrying out the invention.
[2091] Example 2
[2092] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2093] Elderly people require various forms of support in their daily lives, but current technology makes it difficult to provide personalized assistance that fully reflects the individual needs and emotional state of each elderly person. Furthermore, systems for responding quickly and appropriately in emergencies are inadequate. Therefore, there is a need to create an environment where elderly people can live safely and without feeling lonely.
[2094] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2095] In this invention, the server includes: means for managing and saving basic information and individual needs of a user; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying or playing a notification to the user at a specified time; means for sending an emergency contact request to the server and for the server to send an emergency notification to a pre-set emergency contact; means for analyzing the tone, speed and volume of the user's voice to identify the user's emotional state; and means for generating a personalized response based on the emotional state and providing it to the user.
[2096] This will enable personalized assistance to be provided in response to the individual needs and emotional state of the elderly, as well as a rapid and appropriate response in emergencies.
[2097] "Basic user information" refers to basic information about a user, such as name, age, address, health status, and emergency contact information.
[2098] "Individual needs" refers to information about the support and services that a user individually needs, such as when to take medication or schedule regular health checks.
[2099] "Voice input" refers to voice data input when a user speaks to the system through a microphone.
[2100] "Text data" refers to voice input converted into character string data using a voice recognition engine or the like.
[2101] "Server" refers to a computer system that receives data from users and stores, analyzes, and processes it.
[2102] "Response data" refers to information that the server generates based on a user request and returns to the user.
[2103] "Reminder information" refers to information for notifying the user at a specific time or timing.
[2104] "Emergency contact request" refers to an emergency signal sent through the system when a user faces an emergency.
[2105] "Emergency notification" refers to an emergency message sent from the server to a pre-defined emergency contact.
[2106] "Emotional state" refers to the user's emotions estimated based on an analysis of the tone, rate, volume, etc. of the voice.
[2107] "Personalized response" refers to an individualized response that is generated based on the user's basic information, individual needs, and emotional state.
[2108] The present invention, the "Generative AI Elderly Companion" system, provides personalized support that reflects the individual needs and emotional state of elderly people, reducing feelings of loneliness in daily life and creating a safe living environment.
[2109] Hardware and software used
[2110] The server is a high-performance computer system, such as an AWS EC2 instance or Microsoft Azure VM. MySQL or PostgreSQL is used as the database management system. Google Cloud Speech-to-Text or Amazon Transcribe is used as the speech recognition engine, and IBM Watson Tone Analyzer or Microsoft Azure Cognitive Services is used as the emotion recognition engine.
[2111] Manage and store basic information and needs
[2112] The user uses the device to enter basic information (name, age, address, health status, emergency contact information). The device provides voice or text input, verifies the entered information, and sends it to the server. The server stores the received basic information in a database and provides personalized support based on this information.
[2113] Converts voice input into text and sends it to the server
[2114] When a user asks "What's the weather like today?", the device uses a voice recognition engine to convert the voice into text data and sends it to the server. The server then analyzes the received text data and obtains weather information using a weather API provided by a third party.
[2115] Receiving and displaying response data
[2116] The response data (such as weather information) sent from the server is received by the terminal, which then provides it to the user through voice or a display device. For example, it may say, "Today's weather is sunny. The temperature is 24 degrees."
[2117] Reminder function
[2118] When a user instructs the device to "set a reminder to take my medicine at 8 o'clock every morning," the device converts this instruction into text data and sends it to the server. The server stores the reminder information in a database and sends a reminder notification to the device at the specified time (8 o'clock every morning). The device then notifies the user by voice or a pop-up notification, saying, "It's 8 o'clock. It's time to take your medicine."
[2119] Emergency contact function
[2120] In an emergency, the user presses the emergency button on the device. The device then sends an emergency signal to the server. The server then contacts pre-registered emergency contacts (family members or caregivers) via SMS, phone, email, etc. It then sends a message saying, "The user has pressed the emergency button. Please check immediately."
[2121] Emotion recognition function
[2122] If a user says "I'm feeling bad today" in a sad voice, the device sends this voice data to an emotion recognition engine, which analyzes the tone, speed, and volume of the voice to identify the emotional state. The server uses the emotional information to generate a personalized response that shows empathy, such as "What's wrong? Tell me your story," and the device plays this aloud to the user.
[2123] Specific use cases
[2124] Usage example 1:
[2125] The user asks, "When is my next hospital appointment?" The device recognizes the voice and sends the text data to the server. The server retrieves the appointment information from the database and sends a response saying, "Your next hospital appointment is tomorrow at 10:00." The device plays this back aloud. If the user sounds anxious, the emotion engine detects this and the server generates an additional response saying, "Is there something you're worried about?"
[2126] Usage example 2:
[2127] If a user is about to fall at home, they press the emergency button on their device. The device immediately sends an emergency signal to the server. The server then makes an emergency call to pre-registered family members and sends them a message saying, "The user has pressed the emergency button. Please check immediately." If the emotion engine detects a panicked state in the user's voice, the server notifies the family members as well.
[2128] The above is an embodiment of the present invention. This allows elderly people to live a safe and fulfilling daily life and reduce their sense of loneliness. Furthermore, the introduction of an emotion engine enables the system to respond in a way that is sensitive to the user's feelings, providing a higher level of satisfaction.
[2129] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2130] Steps for entering and saving basic information
[2131] Step 1:
[2132] The user uses the device to input basic information (name, age, address, health status, emergency contact information). The input method can be selected as voice input or text input. For example, if the user inputs "My name is Ichiro Tanaka, and I'm 75 years old" by voice, the device will capture the voice data.
[2133] input:
[2134] Basic information entered by the user (name, age, address, health status, emergency contact information)
[2135] output:
[2136] Audio or text data
[2137] Step 2:
[2138] When the device receives voice input, it uses a speech recognition engine (such as Google Cloud Speech-to-Text) to convert the voice into text data. For example, voice data such as "My name is Tanaka Ichiro and I'm 75 years old" is converted into text data such as "Name: Tanaka Ichiro, Age: 75."
[2139] input:
[2140] Audio data
[2141] output:
[2142] Text data
[2143] Step 3:
[2144] The device sends the converted text data to the server. Specifically, data transmission is performed using an API request.
[2145] input:
[2146] Text data
[2147] output:
[2148] API request (text data)
[2149] Step 4:
[2150] The server stores the received text data in a database, for example, using MySQL or PostgreSQL as data records.
[2151] input:
[2152] Text data
[2153] output:
[2154] Data Records
[2155] Processing steps for converting voice input to text and sending it to the server
[2156] Step 1:
[2157] The user asks a question by voice, "What's the weather like today?" This voice data is acquired by the terminal.
[2158] input:
[2159] User voice input
[2160] output:
[2161] Audio data
[2162] Step 2:
[2163] The device uses a voice recognition engine to convert the acquired voice data into text data. Specifically, the voice data "What's the weather like today?" is converted into text data "What's the weather like today?"
[2164] input:
[2165] Audio data
[2166] output:
[2167] Text data
[2168] Step 3:
[2169] The device sends the converted text data to the server via an API request.
[2170] input:
[2171] Text data
[2172] output:
[2173] API request (text data)
[2174] Step 4:
[2175] The server analyzes the received text data and retrieves the corresponding information. For example, to retrieve weather information, the server sends a request to a weather API provided by a third party and receives response data such as "Today's weather is sunny and the temperature is 24 degrees."
[2176] input:
[2177] Text data
[2178] output:
[2179] Response data (weather information)
[2180] Processing steps for receiving and displaying response data
[2181] Step 1:
[2182] The server sends the acquired response data (such as weather information) to the terminal.
[2183] input:
[2184] Response data (weather information)
[2185] output:
[2186] API response (response data)
[2187] Step 2:
[2188] The device analyzes the received response data and provides it to the user through a voice output device, for example, by playing back a voice message saying, "Today's weather is sunny. The temperature is 24 degrees."
[2189] input:
[2190] API response (response data)
[2191] output:
[2192] Audio data
[2193] Reminder function processing steps
[2194] Step 1:
[2195] The user instructs, "Set a reminder to take my medicine at 8 o'clock every morning." The device receives this voice data.
[2196] input:
[2197] User voice instructions
[2198] output:
[2199] Audio data
[2200] Step 2:
[2201] The device uses a speech recognition engine to convert voice data into text data. For example, voice data such as "Set a reminder to take my medicine at 8 o'clock every morning" is converted into text data such as "Take my medicine at 8 o'clock every morning."
[2202] input:
[2203] Audio data
[2204] output:
[2205] Text data
[2206] Step 3:
[2207] The device sends the converted text data to the server via an API request.
[2208] input:
[2209] Text data
[2210] output:
[2211] API request (text data)
[2212] Step 4:
[2213] The server stores the received reminder information in a database and configures it to generate notifications at a specified time, for example, every morning at 8:00.
[2214] input:
[2215] Text data
[2216] output:
[2217] Database records and timer settings
[2218] Step 5:
[2219] When the specified time arrives, the server sends a notification to the device, such as "It's time to take your medicine."
[2220] input:
[2221] Timer Event
[2222] output:
[2223] API response (notification data)
[2224] Step 6:
[2225] The device will then provide the received notification data to the user via voice or pop-up notification, for example, "It's 8 o'clock. Time to take your medicine."
[2226] input:
[2227] API response (notification data)
[2228] output:
[2229] Audio data or popup notification
[2230] Emergency Contact Function Processing Steps
[2231] Step 1:
[2232] The user presses the emergency button, and the device detects this emergency signal.
[2233] input:
[2234] Emergency button input
[2235] output:
[2236] Emergency Signal Data
[2237] Step 2:
[2238] The device sends an emergency signal to the server, which is done via an API request.
[2239] input:
[2240] Emergency Signal Data
[2241] output:
[2242] API Request (Emergency Signal Data)
[2243] Step 3:
[2244] The server analyzes the received emergency signal and sends an emergency notification to pre-defined emergency contacts, for example, by SMS, phone call, or email, with a message saying, "The user has pressed the emergency button. Please check immediately."
[2245] input:
[2246] API Request (Emergency Signal Data)
[2247] output:
[2248] Emergency notification data (SMS, phone, email)
[2249] Emotion Recognition Processing Steps
[2250] Step 1:
[2251] The user inputs "I feel bad today" in a sad voice. This voice data is acquired by the terminal.
[2252] input:
[2253] User voice input
[2254] output:
[2255] Audio data
[2256] Step 2:
[2257] The device sends the captured voice data to an emotion recognition engine, which analyzes the tone, rate, and volume of the voice to identify the emotional state.
[2258] input:
[2259] Audio data
[2260] output:
[2261] Emotional state data
[2262] Step 3:
[2263] The server generates a personalized response based on the emotional state data received from the emotion recognition engine, for example, an empathetic response such as "What's wrong? Tell me your story."
[2264] input:
[2265] Emotional state data
[2266] output:
[2267] Response data
[2268] Step 4:
[2269] The server transmits the generated response data to the terminal.
[2270] input:
[2271] Response data
[2272] output:
[2273] API response (response data)
[2274] Step 5:
[2275] The device will then play back the received response data in voice, for example, "What's wrong? Tell me what happened."
[2276] input:
[2277] API response (response data)
[2278] output:
[2279] Audio data
[2280] (Application example 2)
[2281] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2282] Elderly people have greater difficulty navigating physical stores, searching for products, and responding to emergencies. They also often feel anxious about asking store staff questions and navigating the store. Current systems struggle to adequately address these issues, preventing seniors from enjoying shopping with peace of mind.
[2283] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for managing and saving a user's basic information and individual needs; means for converting the user's voice input into text data and sending it to the server; means for receiving response data from the server and displaying or playing it back to the user; means for saving specified reminder information to the server and displaying a notification to the user at a specified time; means for sending an emergency contact request to the server and having the server send an emergency notification to a pre-set emergency contact; means for navigating the store using the user's current location and store map data; and means for analyzing the user's voice input and generating a personalized response based on the user's emotional state. This enables seniors to move confidently through physical stores, easily find products and services, and quickly respond to anxieties or emergencies.
[2284] Below are definitions of important terms included in the patent claims, rewritten to suit the application example.
[2285] "Means for managing and storing a user's basic information and individual needs" refers to a means for storing basic information such as the user's name, age, address, health condition, and emergency contact information, as well as individual needs such as medication times and regular health check schedules, in a database, and for managing and updating this information as needed.
[2286] "Means for converting user's voice input into text data and sending it to the server" refers to means for converting information input by the user by voice into text data using voice recognition technology and sending that text data to the server.
[2287] "Means for receiving response data from the server and displaying or playing it back to the user" refers to means for receiving information sent from the server (e.g., weather information or store directions) and providing it to the user visually or audibly.
[2288] The "means for saving specified reminder information on a server and displaying a notification to the user at a specified time" refers to a means for saving reminder information set by a user on a server and sending a notification to the user at a specified time based on that reminder information.
[2289] "Means for sending an emergency contact request to a server, and for the server to send an emergency notification to pre-registered emergency contacts" refers to a means for a user to send an emergency contact request to a server by pressing an emergency button, etc., and for the server to send an emergency notification by phone or SMS to pre-registered emergency contacts based on that request.
[2290] "Means for navigating within a store using the user's current location and store map data" refers to a means for guiding the user to their desired location by using the location information of the user's smartphone or device and combining it with map data within the store.
[2291] "Means for analyzing a user's voice input and generating a personalized response based on the user's emotional state" refers to means for analyzing the tone, rate, and content of a user's voice to identify the user's emotional state, and generating and providing an appropriate response or message accordingly.
[2292] System Program
[2293] The purpose of this invention, the "Generative AI Elderly Companion Shopper" system, is to help seniors shop comfortably in brick-and-mortar stores. This system manages and stores users' basic information and individual needs, converts voice input into text data and sends it to a server, and displays or plays back response data from the server to the user via their device. It also supports the daily lives of seniors through reminder functions, emergency contact functions, and emotion recognition functions.
[2294] Hardware and Software Use
[2295] Hardware:
[2296] Smartphone (with camera, microphone, and speaker)
[2297] software:
[2298] Google Cloud Speech-to-Text API (voice recognition)
[2299] Google Cloud Natural Language API (Text Analysis)
[2300] Twilio API (emergency contact function)
[2301] Emotion Engine (e.g. Affectiva's SDK)
[2302] Firebase (database)
[2303] Data processing and calculation
[2304] Managing and storing your basic information and individual needs:
[2305] The server stores the basic information entered by the user (name, age, address, health status, emergency contacts) and individual needs (medication times, schedule for regular health checks) in a Firebase database, allowing the system to provide personalized support to the user.
[2306] Convert speech to text and send to server:
[2307] When a user asks, "Where is the restroom?", the device uses the Google Cloud Speech-to-Text API to convert speech to text data and send it to the server, which then uses the Google Cloud Natural Language API to parse the text and generate an appropriate response.
[2308] Receive and display response data:
[2309] The response data sent from the server is received by the terminal and provided to the user through voice or a display device. For example, the terminal may provide voice guidance such as, "To find the restroom, go straight down the corridor on the right."
[2310] Reminder function:
[2311] The server stores the reminder information set by the user (e.g., take medicine at 8 o'clock every day) in Firebase. When the specified time arrives, the device sends a voice or pop-up notification saying, "It's 8 o'clock. Time to take your medicine."
[2312] Emergency contact features:
[2313] If a user is about to fall inside the store, they can press the emergency button on their device, which sends an emergency signal to the server. The server then uses the Twilio API to send a notification to registered emergency contacts (family members or caregivers). The emergency notification is sent via SMS or phone.
[2314] Emotion recognition function:
[2315] If a user says, "I'm not feeling well today," the device uses its emotion engine to analyze the speech and identify the user's emotional state. The server then generates an appropriate response based on this, saying, "What's wrong? Tell me."
[2316] Examples of specific examples and prompts
[2317] As a specific example, consider the case where an elderly person named Mr. A gets lost in a store.
[2318] Person A asks the app, "Where is the restroom?" inside the store. Voice recognition is activated, and the speech is converted into text and sent to the server.
[2319] The server identifies the location of the restroom from the store's map data and generates a navigation plan along with Mr. A's current location.
[2320] The app will provide voice guidance saying, "To get to the restroom, go straight down the corridor on the right."
[2321] If the emotion engine detects Mr. A's anxiety, the app will ask, "If you can't find it, I'll call a store employee."
[2322] Example prompt sentence:
[2323] The user asks the smartphone microphone, "Where is the restroom?" The voice input is converted into text and sent to the server. The system should refer to the user's current location and a map of the store, calculate the shortest route, and provide voice guidance. If the user asks in an anxious voice, the system should generate an additional response asking, "If you can't find it, would you like me to call a store employee?"
[2324] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2325] Step 1:
[2326] A user asks into the microphone of their smartphone, "Where is the restroom?" A voice input occurs, and the device collects this voice input data.
[2327] Input: User voice input
[2328] Output: Audio data
[2329] Step 2:
[2330] The device uses the Google Cloud Speech-to-Text API to convert the voice data into text data, which is then sent to the server.
[2331] Input: Audio data
[2332] Output: Text data
[2333] Step 3:
[2334] The server analyzes the received text data using the Google Cloud Natural Language API, and as a result, it understands that the user is looking for a restroom.
[2335] Input: Text data
[2336] Output: Analysis result (intent)
[2337] Step 4:
[2338] The server identifies the user's location by referencing the user's current location and the store's map data, and calculates the shortest route based on this.
[2339] Input: User's current location data, store map data
[2340] Output: Shortest route data
[2341] Step 5:
[2342] The server uses the calculated shortest route data to generate a navigation plan to help the user reach their destination easily. The generated navigation plan is created in voice and text format.
[2343] Input: Shortest route data
[2344] Output: Navigation plan (voice and text)
[2345] Step 6:
[2346] The server sends the generated navigation plan to the device, which receives it and provides voice and text guidance to the user.
[2347] Input: Navigation plan
[2348] Output: User instructions
[2349] Step 7:
[2350] At the same time, the device analyzes the user's voice input with an emotion engine to identify their emotional state. If the device detects that the user is asking a question in an anxious voice, the server generates an additional response (e.g., "If you can't find it, should I call a store clerk?").
[2351] Input: User's voice data
[2352] Output: Emotional state, additional responses
[2353] Step 8:
[2354] As additional responses are generated, the server sends them to the terminal, which displays or plays them audibly to the user.
[2355] Input: Additional response data
[2356] Output: Additional instructions for the user
[2357] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2358] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2359] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2360] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2361] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2362] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2363] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2364] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach...
Claims
1. A means of managing and storing basic information about users and their individual needs; means for converting the user's voice input into text data and transmitting the text data to a server; means for receiving response data from the server and displaying or reproducing the data to a user; means for storing the designated reminder information in the server and displaying a notification to the user at a designated time; The system includes means for sending an emergency contact request to the server, and for the server to send an emergency notification to a pre-defined emergency contact.
2. The system of claim 1 , wherein the server further comprises means for periodically training and updating the AI model.
3. The system of claim 1 , further comprising means for providing a dialogue interaction based on a voice input of the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A