system
A voice-activated system helps elderly individuals manage daily schedules and reduce loneliness by converting voice inputs to text, analyzing and registering appointments, and providing interactive reminders and assistance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-04-08
AI Technical Summary
Elderly individuals face challenges in managing daily schedules and appointments, often leading to forgotten commitments and feelings of loneliness, with limited support for independent living.
A system that receives voice input, converts it to text, analyzes schedule information, registers it in a database, generates reminder notifications, and provides interactive assistance using generative AI to manage daily tasks and reduce loneliness.
Enables elderly individuals to manage their schedules effectively, reduces the risk of missed appointments, and provides conversational support to enhance their sense of independence and confidence.
Smart Images

Figure 2026060611000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The elderly, especially the elderly living alone, may have difficulties in many parts of their lives, such as daily schedule management and forgetting appointments. In addition, there is a lack of support for maintaining an independent life, and loneliness is a particularly big problem for the elderly who have lost their partners. The present invention proposes a system that provides schedule management by voice input, reminder notifications, and interactive assistance functions so that the elderly can live an independent life with peace of mind.
Means for Solving the Problems
[0005] The present invention provides a system that includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, and means for generating reminder notifications based on the schedule information and sending them to the user's terminal. The system further includes means for comprehensively analyzing the voice input and providing additional information to the user in an interactive format, and means for automatically adjusting the reminder notifications based on specific times or conditions. This allows elderly people to easily manage their daily schedules through voice and ensure they don't miss easily forgotten appointments thanks to reminder notifications. Furthermore, the use of interactive functions can reduce feelings of loneliness and enable them to live with confidence.
[0006] "Users" refer to elderly people and their caregivers who use this system.
[0007] "Voice input" refers to words or instructions spoken by the user through a microphone.
[0008] "Means of converting to text" refers to the process of analyzing voice input and converting it into textual information.
[0009] "Schedule information" refers to the detailed information of appointments and reminders set by the user.
[0010] A "database" refers to a storage device used to store and manage schedule information and user data.
[0011] A "reminder notification" refers to a message or alert sent to inform a user of their schedule.
[0012] "Terminal" refers to electronic devices used to run this system, such as smartphones and tablets.
[0013] "Conversational assistance features" refer to functions that allow users to obtain information or give instructions through natural conversations with the system. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. The system's program processing is described in detail below in natural language.
[0036] Server-side processing
[0037] User authentication and data storage
[0038] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[0039] Model operation of generative AI
[0040] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[0041] Schedule notification
[0042] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[0043] Terminal-side processing
[0044] Voice input function
[0045] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[0046] Display reminders
[0047] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[0048] Conversational interface
[0049] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[0050] User-side actions
[0051] Inputting voice commands
[0052] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[0053] Notification confirmation
[0054] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[0055] Conversational communication
[0056] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[0057] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. The simple operation via voice input and the provision of appropriate information by the generated AI create an efficient and secure living environment for the user.
[0058] The following describes the processing flow.
[0059] User authentication and data storage
[0060] Step 1:
[0061] The user launches the app on their smartphone.
[0062] The user taps the smartphone icon to open the app.
[0063] Step 2:
[0064] The device displays the login screen to the user.
[0065] The device displays a screen prompting the user to enter their user ID and password.
[0066] Step 3:
[0067] The user enters their user ID and password and taps the "Login" button.
[0068] The user enters the correct authentication information.
[0069] Step 4:
[0070] The device sends authentication information to the server.
[0071] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[0072] Step 5:
[0073] The server checks the authentication information received and searches the database for matching user information.
[0074] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[0075] Step 6:
[0076] The server sends the authentication result to the terminal.
[0077] The server returns the authentication success / failure result as a response.
[0078] Step 7:
[0079] The device displays a screen corresponding to the authentication result.
[0080] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[0081] Voice input and schedule registration
[0082] Step 1:
[0083] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[0084] The user taps the microphone button and enters their schedule by voice.
[0085] Step 2:
[0086] The device sends voice data to the server.
[0087] The device acquires voice data and sends it to the server.
[0088] Step 3:
[0089] The server receives the audio data and converts it into text using a generative AI model.
[0090] The server converts the audio data into text and then analyzes its content.
[0091] Step 4:
[0092] The server analyzes the text and extracts information related to the schedule.
[0093] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[0094] Step 5:
[0095] The server sends the registration results to the terminal.
[0096] The server notifies the terminal that the appointment registration was successful.
[0097] Step 6:
[0098] The device notifies the user of the registration result.
[0099] The device displays the message "The appointment has been successfully registered" to the user.
[0100] Reminder notifications and confirmations
[0101] Step 1:
[0102] The server sets the timing of reminders based on a schedule.
[0103] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[0104] Step 2:
[0105] When the server reaches the reminder time, it sends a reminder notification to the device.
[0106] The server will issue a reminder notification at the time set by the server.
[0107] Step 3:
[0108] The device receives a reminder notification.
[0109] The device receives the push notification and displays it on the screen.
[0110] Step 4:
[0111] The user checks the notification and takes action if necessary.
[0112] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[0113] Conversational interface
[0114] Step 1:
[0115] Users ask the app questions such as, "What's the weather like today?"
[0116] The user presses the microphone button to input their question by voice.
[0117] Step 2:
[0118] The device sends voice data to the server.
[0119] The device acquires voice data and sends it to the server.
[0120] Step 3:
[0121] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[0122] The server analyzes the audio data, retrieves weather information, and generates a response.
[0123] Step 4:
[0124] The server sends the generated response to the terminal.
[0125] The server sends a response to the terminal saying, "It's sunny today."
[0126] Step 5:
[0127] The device plays back the response it received as audio.
[0128] The device plays a voice message to the user saying, "It's sunny today."
[0129] The above is a detailed explanation of the program's processing flow, broken down into individual steps.
[0130] (Example 1)
[0131] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0132] In systems designed to support elderly individuals in living independent and safe daily lives, it is essential to efficiently analyze voice input, appropriately manage user schedules and provide reminder notifications, and enhance user convenience and peace of mind by offering interactive support. Furthermore, utilizing generative AI models to extract appropriate tasks from voice input and enabling natural dialogue with the user is also crucial.
[0133] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0134] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for operating a generative AI model and analyzing voice data to match it with specific tasks and information, and means for generating dialogue with the user based on the analysis results of the generative AI model. This makes it possible for elderly people to easily operate the system through voice input and to receive appropriate reminder notifications and interactive feedback.
[0135] "Means for receiving user voice input" refers to devices or systems for receiving information entered by a user via voice.
[0136] "Means for converting voice input to text" refers to methods and technologies for analyzing received voice data and converting it into text format.
[0137] "Means for analyzing user schedule information and registering it in a database" refers to a system or process that interprets information obtained as text and saves it in a database as the user's schedule or appointments.
[0138] "A means of generating and sending reminder notifications to the user's device" refers to a method for creating notifications based on registered schedule information and sending them to the user's device.
[0139] "A means of operating a generative AI model to analyze voice data and match it to specific tasks or information" refers to a technology that uses generative AI to process voice data and identify appropriate tasks or information.
[0140] "A means of generating user dialogue based on the analysis results of a generative AI model" refers to a system that generates and provides information for natural dialogue with the user based on the results of analysis by the generative AI.
[0141] This invention is an interactive lifestyle support system that helps elderly people live independently and safely, and provides its functions using a device such as a smartphone. This system accepts voice input from the user, analyzes it, and provides appropriate reminders and notifications, as well as conversational support utilizing a generative AI model.
[0142] Server-side processing
[0143] User Authentication
[0144] When a user logs into the app, the server authenticates them by comparing the entered user ID and password with the database. It also saves the data of successfully authenticated users to the database.
[0145] Specific operation: The user enters their user ID and password on their smartphone, and the server verifies this against the database to perform authentication.
[0146] Data storage
[0147] The server stores the schedule information and settings of authenticated users in a database. For example, if there are new appointments or changes to existing ones, this information is added to the database.
[0148] Specific operation: When a user voice-inputs "I want to change my dentist appointment to 3 PM tomorrow," the server saves this information to the database.
[0149] Operation of Generative AI Models
[0150] The server operates a generative AI model and analyzes the user's voice data. It extracts specific tasks and information from the voice input and registers them in a database.
[0151] Specific operation: When a user voice-inputs "I will take my medicine at 8 AM tomorrow," the generating AI model recognizes this as an appointment and saves it to the database.
[0152] Sending schedule notifications
[0153] The server generates reminder notifications based on the registered schedule information and sends them to the user's smartphone.
[0154] Specific action: A reminder notification to "take your medicine" is sent to the user's smartphone at 8:00 AM.
[0155] Terminal-side processing
[0156] Acquiring voice input
[0157] The device captures the user's voice and sends it to the server to be converted into text.
[0158] Specific operation: When the user uses voice input to say "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[0159] Display reminders
[0160] The device displays reminder notifications received from the server.
[0161] Specific action: A notification appears on the user's smartphone stating, "I'm going shopping at 2 PM."
[0162] Providing a conversational interface
[0163] The device plays the response from the generating AI as audio, allowing the user to interact with the app.
[0164] Specific operation: When the user voice-inputs "What's the weather like today?", the device plays the response "It's sunny today" from the server's AI generation system.
[0165] User-side processing
[0166] Inputting voice commands
[0167] Users use their smartphones to input their daily schedules and questions by voice.
[0168] Specific action: The user gives a voice command saying, "I will go to the doctor at 10 AM tomorrow."
[0169] Notification confirmation
[0170] Users check reminders and notifications displayed on their smartphones and take action based on them.
[0171] Specific actions: Check notifications such as "It's 12 o'clock now. It's lunchtime," and prepare the meal.
[0172] Conversational communication
[0173] Users can check and adjust their schedules by interacting with the app.
[0174] Specific action: When the user asks "What were your plans for today?", the app responds "I was planning to go shopping at 2pm today."
[0175] Examples of prompts for generative AI models
[0176] Example of a prompt
[0177] User voice input: "What's the weather like today?"
[0178] Generated AI model prompt: "The user wants to know today's weather. Please provide weather information such as sunny, cloudy, or rainy."
[0179] This system, with its interconnected components, provides support for elderly individuals to live independently through daily schedule management and dialogue. Its simple voice input operation and the provision of appropriate information by AI create an efficient and reassuring living environment for users.
[0180] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0181] Step 1:
[0182] Acquisition of user voice input
[0183] The user launches the app on their smartphone and enters voice commands. For example, they might voice-inform an appointment such as, "I'm going to the doctor tomorrow at 10 AM."
[0184] The terminal receives this audio and performs initial processing. The terminal receives the input as an audio signal and prepares it for conversion to text.
[0185] Input: Voice command (Example: "I will go to the doctor at 10 AM tomorrow")
[0186] Output: Audio data
[0187] Step 2:
[0188] Sending audio data to the server
[0189] The terminal uses a communication protocol to send the acquired voice data to the server. This transfers the user's voice information to the server.
[0190] Input: Audio data
[0191] Output: Sending audio data to the server
[0192] Specific operation: The device uses an internet connection to send voice data to the server. For example, it uses Wi-Fi or mobile data communication.
[0193] Step 3:
[0194] Voice input to text conversion
[0195] The server analyzes the received audio data and converts it into text. Speech recognition technology is used to convert the audio signal into linguistic information. Speech recognition APIs and software are used for this process.
[0196] Input: Audio data
[0197] Output: Text data (Example: "I'm going to the doctor at 10 AM tomorrow")
[0198] Specific operation: The server calls a speech recognition API to convert the audio data into text data. This API is often a cloud-based service.
[0199] Step 4:
[0200] Text data analysis and database registration
[0201] The server analyzes text data and extracts user schedule information. This information is then registered in a database. Natural language processing technology is used for the analysis, and the schedule information is added to the database as a new record.
[0202] Input: Text data
[0203] Output: Schedule information registered in the database
[0204] Specific operation: The server uses a natural language processing algorithm to extract date and time information from the text. Then, it stores the extracted data in a relational database.
[0205] Step 5:
[0206] Voice analysis of generative AI models
[0207] The server uses a generative AI model to analyze text data and match it with relevant tasks and information. The generative AI model understands the context from voice input and determines the appropriate action.
[0208] Input: Text data
[0209] Output: Matched task information or answers
[0210] Specific operation: The server runs the generated AI model, and the text "I will go to the doctor at 10 AM tomorrow" is matched as a scheduled task.
[0211] Step 6:
[0212] Generating and sending scheduled notifications
[0213] The server generates reminder notifications based on registered schedule information and sends these notifications to the user's device. Appropriate reminders are generated based on the specified time or conditions.
[0214] Input: Schedule information in the database
[0215] Output: Reminder notification to user's device
[0216] Specific operation: The server creates a reminder notification based on the schedule information and sends it to the user's smartphone using a push notification service. For example, a notification might be set saying, "I have a doctor's appointment tomorrow at 10 AM."
[0217] Step 7:
[0218] Display reminders
[0219] The device displays reminder notifications received from the server. The user checks the notification and takes action.
[0220] Input: Reminder notification from server
[0221] Output: Notification displayed on the smartphone screen
[0222] Specific action: The smartphone receives a push notification and displays a message on the screen saying, "Go to the doctor tomorrow at 10 AM." The user can tap it to view details or set an alarm.
[0223] Step 8:
[0224] Providing a conversational interface
[0225] The device plays back the responses from the generated AI as audio, allowing the user to communicate with the app in a conversational format.
[0226] Input: Generated AI response data from the server
[0227] Output: Generated AI response played back as audio.
[0228] Specific operation: When a user asks the device "What's the weather like today?", the device sends this to the server, and the AI-generated response "It's sunny today" is played aloud. This allows the user to have a natural conversation with the system.
[0229] As described above, through each processing step, the present invention provides a system that supports elderly people in living independently and safely, and realizes a convenient and secure living environment through voice input.
[0230] (Application Example 1)
[0231] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0232] To improve work efficiency within factories, there is a need for means to provide workers with timely and appropriate information and task reminders. Furthermore, an interactive user interface that is easy for elderly or inexperienced workers to use is required.
[0233] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0234] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for analyzing voice input from factory workers and matching it to specific tasks, and means for a robot to send push notifications to workers based on the work schedule. This makes it possible for workers to perform their work efficiently and obtain necessary information in a timely manner.
[0235] "User voice input" refers to voice commands that a user gives to the system through a voice input device such as a microphone.
[0236] "Means of converting to text" refers to software or hardware that has the function of converting voice-input data into text information.
[0237] "User schedule information" refers to data that records the user's planned activities and task schedules.
[0238] "Means of registering in the database" refers to the function of storing the analyzed schedule information in the database.
[0239] A "reminder notification" is a notification sent to a user based on a pre-set schedule.
[0240] "User's device" refers to electronic devices used by the user, such as smartphones, tablets, and personal computers.
[0241] A "factory worker" is an employee who performs physical or mechanical tasks within a factory.
[0242] "Means for analyzing voice input" refers to software or hardware that has the function of appropriately understanding and classifying received voice data.
[0243] "Means for matching to specific tasks" refers to a function that links information obtained through voice analysis to related tasks and work.
[0244] "Means of sending push notifications" refers to a function that sends information to a device in real time and displays a notification to the user.
[0245] Program generation
[0246] To implement this application, a program is first needed that accepts user voice input and converts it into text. This program uses a speech recognition API (e.g., Google® Cloud Speech-to-Text). After the voice data is converted to text, that text data is sent to a generating AI model (e.g., OpenAI® GPT-4®) to be matched with specific tasks or information.
[0247] Hardware and software usage
[0248] Hardware:
[0249] Microphone: Used to capture the user's voice input.
[0250] Speaker: Used to play the generated audio response.
[0251] Display (built into the robot): Used to display text information and reminders.
[0252] software:
[0253] Speech recognition APIs (e.g., Google Cloud Speech-to-Text): Used to convert speech data into text.
[0254] Generative AI models (e.g., OpenAI GPT-4): Used to analyze text data and match it to specific tasks or information.
[0255] Database management systems (e.g., MySQL®): Used to store and manage schedule information.
[0256] Program processing
[0257] 1. Voice input:
[0258] The server captures the user's voice using the microphone and temporarily stores it locally.
[0259] For example, if a factory worker says, "Tell me the inventory of parts," the voice data is sent to the server and temporarily stored.
[0260] 2. Converting and transmitting audio data:
[0261] The speech recognition API is used to convert the audio data into text, which is then sent to a server where the generation AI model is running.
[0262] Example: The audio "Tell me the stock of parts" is converted to the text "Tell me the stock of parts".
[0263] 3. Processing of Generative AI Models:
[0264] The generating AI analyzes the received text data and extracts related tasks and information.
[0265] Example: Based on the instruction "Tell me the inventory status of the parts," the system retrieves the necessary information from the inventory database.
[0266] 4. Provision of information and notification:
[0267] It generates necessary information and reminders and sends them from the server to the robot.
[0268] For example, information such as "We have very little stock left of part A" is generated and communicated to the worker as a notification via the robot's display or speaker.
[0269] 5. Output:
[0270] The robot provides information to workers through voice and displays.
[0271] For example, if you ask "What's the next step?", you might get the answer "The next step is to install part D."
[0272] Examples of specific cases and prompt statements
[0273] As a concrete example, if a factory worker asks a robot, "When is my break time?", the robot will respond, "Today's break times are 10 AM and 3 PM." An example of the prompt in this case would be as follows:
[0274] "Please generate today's break times based on the voice instructions from the factory workers."
[0275] This enables workers to work efficiently and obtain the necessary information in a timely manner.
[0276] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0277] Step 1:
[0278] Receive the user's voice input
[0279] Input: The user instructs the microphone verbally "Tell me the inventory of parts".
[0280] Processing content: The terminal recognizes this voice input and temporarily stores it locally.
[0281] Output: A voice data file is generated and passed to the next processing step.
[0282] Step 2:
[0283] Convert voice data to text
[0284] Input: The voice data file generated in Step 1.
[0285] Processing content: Use a voice recognition API (e.g., Google Cloud Speech-to-Text) to convert voice data into character information.
[0286] Output: Text data "Tell me the inventory of parts" is generated.
[0287] Step 3:
[0288] Send the text data to the generative AI model
[0289] Input: The text data generated in Step 2.
[0290] Processing details: Text data is sent to a server where an AI model (e.g., OpenAI GPT-4) is running.
[0291] Output: Text data is received by the generating AI model.
[0292] Step 4:
[0293] The generative AI model analyzes the text data.
[0294] Input: Text data received in Step 3.
[0295] Processing details: The generating AI model analyzes text data to identify tasks and information. In this case, it extracts the necessary information from the inventory database.
[0296] Output: Information such as "We have very little stock left of part A."
[0297] Step 5:
[0298] Create information and generate notifications
[0299] Input: Information obtained in Step 4.
[0300] Processing details: The server generates a notification message based on this information and sends it to the terminal.
[0301] Output: A notification message is generated stating, "We have very little stock left of part A."
[0302] Step 6:
[0303] Display and play audible notification messages.
[0304] Input: The notification message generated in Step 5.
[0305] Processing details: The device displays this notification message on its screen and plays it aloud through its speaker.
[0306] Output: The user is visually and audibly notified that "there is only a little inventory of part A left".
[0307] Through the above processing steps, the worker can efficiently receive information and act based on it.
[0308] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0309] The present invention is an interactive life support system that supports the elderly to live independently and safely, and provides its functions using a terminal such as a smartphone. Furthermore, by combining an emotion engine that recognizes the user's emotion, more personalized support is made possible. Hereinafter, the processing of the system program will be specifically described in natural language.
[0310] Server-side processing
[0311] User authentication and data storage
[0312] When the user logs in to the application, the server performs authentication. If the authentication is successful, the user's data is saved on the server. For example, the user opens the application on a smartphone and logs in. The server checks the user ID and password and performs authentication. Then, if there are new schedules or changes, that information is saved in the database.
[0313] Operation of the generation AI model
[0314] The server operates the generation AI model and analyzes the input of voice data. When the user performs voice input, the server recognizes the input and matches it to a specific task or information. For example, if there is a voice input of "take medicine at 8 o'clock tomorrow morning", the generation AI model recognizes this as a schedule and registers it in the database.
[0315] Operation of the Emotion Engine
[0316] The server operates an emotion engine that analyzes the user's emotional state from voice data. This allows it to adjust schedule information and reminder notifications based on the user's emotional state. For example, if it detects that the user is feeling down, it can adjust the timing and content of reminder notifications.
[0317] Schedule notification
[0318] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[0319] Terminal-side processing
[0320] Voice input function
[0321] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[0322] Display reminders
[0323] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[0324] Conversational interface
[0325] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[0326] User-side actions
[0327] Inputting voice commands
[0328] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[0329] Notification confirmation
[0330] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[0331] Conversational communication
[0332] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[0333] Recognition and response to emotions
[0334] When a user asks a question or gives a command by voice, the device sends the audio to the server, where the emotion engine analyzes the user's emotional state. For example, if the user says, "I'm not feeling very well," the server uses the emotion engine to analyze the user's emotional state and adjusts the content and timing of reminder notifications or generates comforting messages based on the results.
[0335] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. By combining it with an emotion engine, personalized support based on the user's emotional state becomes possible, further enhancing their sense of security and satisfaction.
[0336] The following describes the processing flow.
[0337] User authentication and data storage
[0338] Step 1:
[0339] The user launches the app on their smartphone.
[0340] The user taps the smartphone icon to open the app.
[0341] Step 2:
[0342] The device displays the login screen to the user.
[0343] The device displays a screen prompting the user to enter their user ID and password.
[0344] Step 3:
[0345] The user enters their user ID and password and taps the "Login" button.
[0346] The user enters the correct authentication information.
[0347] Step 4:
[0348] The device sends authentication information to the server.
[0349] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[0350] Step 5:
[0351] The server checks the authentication information received and searches the database for matching user information.
[0352] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[0353] Step 6:
[0354] The server sends the authentication result to the terminal.
[0355] The server returns the authentication success / failure result as a response.
[0356] Step 7:
[0357] The device displays a screen corresponding to the authentication result.
[0358] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[0359] Voice input and schedule registration
[0360] Step 1:
[0361] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[0362] The user taps the microphone button and enters their schedule by voice.
[0363] Step 2:
[0364] The device sends voice data to the server.
[0365] The device acquires voice data and sends it to the server.
[0366] Step 3:
[0367] The server receives the audio data and converts it into text using a generative AI model.
[0368] The server converts the audio data into text and then analyzes its content.
[0369] Step 4:
[0370] The server analyzes the text and extracts information related to the schedule.
[0371] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[0372] Step 5:
[0373] The server sends the registration results to the terminal.
[0374] The server notifies the terminal that the appointment registration was successful.
[0375] Step 6:
[0376] The device notifies the user of the registration result.
[0377] The device displays the message "The appointment has been successfully registered" to the user.
[0378] Reminder notifications and confirmations
[0379] Step 1:
[0380] The server sets the timing of reminders based on a schedule.
[0381] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[0382] Step 2:
[0383] When the server reaches the reminder time, it sends a reminder notification to the device.
[0384] The server will issue a reminder notification at the time set by the server.
[0385] Step 3:
[0386] The device receives a reminder notification.
[0387] The device receives the push notification and displays it on the screen.
[0388] Step 4:
[0389] The user checks the notification and takes action if necessary.
[0390] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[0391] Conversational interface
[0392] Step 1:
[0393] Users ask the app questions such as, "What's the weather like today?"
[0394] The user presses the microphone button to input their question by voice.
[0395] Step 2:
[0396] The device sends voice data to the server.
[0397] The device acquires voice data and sends it to the server.
[0398] Step 3:
[0399] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[0400] The server analyzes the audio data, retrieves weather information, and generates a response.
[0401] Step 4:
[0402] The server sends the generated response to the terminal.
[0403] The server sends a response to the terminal saying, "It's sunny today."
[0404] Step 5:
[0405] The device plays back the response it received as audio.
[0406] The device plays a voice message to the user saying, "It's sunny today."
[0407] Operation of the Emotion Engine
[0408] Step 1:
[0409] When a user voices a question or gives instructions, that voice data is sent to the server.
[0410] For example, the user might say, "What should I do today?"
[0411] Step 2:
[0412] The device sends voice data to the server.
[0413] The device acquires voice data and sends it to the emotion engine.
[0414] Step 3:
[0415] The server analyzes the voice data, and the emotion engine recognizes the user's emotions.
[0416] The server identifies whether it is "stressed" or "depressed."
[0417] Step 4:
[0418] The server sends appropriate responses or reminders to the user based on the emotion recognition results.
[0419] For example, when a user is feeling down, you can send them an encouraging message.
[0420] Step 5:
[0421] The device displays or plays emotion-based responses and reminders to the user via audio.
[0422] A message such as, "Why don't you take a short break to make yourself feel better?" is displayed or played audibly.
[0423] The above is a detailed breakdown of the system's processing steps, including the emotion engine.
[0424] (Example 2)
[0425] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0426] For elderly people, managing their daily schedules and lacking communication are problems that hinder their independent and safe living. Furthermore, conventional reminder and schedule management systems only send uniform notifications without considering the user's emotional state, resulting in a lack of psychological support. To address these issues, there is a need for personalized support that appropriately analyzes the user's emotional state and adapts to their individual circumstances.
[0427] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0428] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for inputting voice data into a generation AI model and utilizing the processing results, means for inputting voice data into an emotion engine and analyzing the emotional state, and means for adjusting the content or timing of reminder notifications based on the analysis results. This makes it possible to provide personalized schedule management and psychological support tailored to the emotional state of elderly people.
[0429] "User" refers to an individual who uses this system.
[0430] "Means for receiving voice input" refers to a device or method for collecting voice uttered by a user and inputting it into a system.
[0431] "Means of converting voice input to text" refers to a device or algorithm that analyzes collected voice data and converts it into textual information.
[0432] "Means for analyzing user schedule information and registering it in a database" refers to a device or method that analyzes a user's schedule based on text data and stores that information in a database.
[0433] "Means for generating and sending reminder notifications to the user's device" refers to a device or method that creates a notification based on registered schedule information and sends it to the user's device.
[0434] A "generative AI model" refers to a model that uses artificial intelligence technology to analyze speech and text data and generate appropriate information.
[0435] An "emotion engine" refers to an algorithm or device used to analyze a user's emotional state from voice data.
[0436] "Means for adjusting the content or timing of reminder notifications" refers to devices or methods for changing the text or sending time of notifications based on the results of the emotion engine's analysis.
[0437] A "server" refers to a computer system used for processing and analyzing data across an entire system.
[0438] "Terminal" refers to devices such as smartphones and tablets that are directly operated by the user.
[0439] This invention is an interactive life support system that helps elderly people live independently and safely. It accepts voice input from the user and provides personalized reminder notifications and schedule management using a generative AI model and emotion engine. A specific embodiment of the system is described below.
[0440] Server-side embodiment
[0441] User authentication and data storage
[0442] When a user logs into the app using a smartphone or tablet, the server accepts the user ID and password, verifies them against the database, and performs authentication. If authentication is successful, reusable user data and schedule information are saved to the database. For example, if a user logs in with the ID "abcd1234" and password "password," the server verifies this, and if authentication is successful, saves the user's new appointments and changes to the database.
[0443] Operation of Generative AI Models
[0444] The server receives voice input from the user and analyzes it using a generative AI model. The voice input is converted into text and matched to specific tasks or information. For example, if a user voice-inputs "I'm going for a walk at 3pm," the server passes this to the generative AI model, which converts it into text and registers "walk" as an event in the database.
[0445] Operation of the Emotion Engine
[0446] The server inputs voice data into an emotion engine to analyze the user's emotional state. This allows the server to adjust the content and timing of reminder notifications according to the user's emotions. For example, if a user says, "I'm not feeling well today," the server analyzes this and sets a reminder such as, "It's time to rest."
[0447] Terminal-side embodiment
[0448] Voice input function
[0449] The terminal receives voice input from the user and sends it to the server. As an initial process, it generates speech recognition data and sends it to the server for text conversion. For example, if the user says, "I will take my medicine at 2 pm," the terminal records this voice and sends it to the server.
[0450] Display reminders
[0451] The device receives reminder notifications sent from the server and displays them on the screen. Voice responses are also supported, allowing the user to check the notifications on the device and take the necessary actions. For example, if the device receives a reminder to "take your medicine at 2 PM," it will display this on the screen and also provide an audio notification.
[0452] Conversational interface
[0453] The device plays back responses from the generative AI model as audio, providing an environment where users can interact with the app naturally. For example, if a user asks, "What's the weather like today?", the device sends that audio to the server, and the generative AI model generates a response such as "It's sunny today." The device then plays this response back to the user as audio.
[0454] User-side embodiment
[0455] Inputting voice commands
[0456] Users use their smartphones to set daily schedules and ask questions using voice input. For example, they might speak instructions to their smartphone such as, "I'll go to the doctor at 10 AM tomorrow."
[0457] Notification confirmation
[0458] Users check reminder notifications displayed on their smartphones and act accordingly. For example, they might receive a notification saying, "It's 12 o'clock now. It's time for lunch."
[0459] Conversational communication
[0460] Users interact with the app via voice to configure and confirm settings. This helps alleviate feelings of loneliness for those living alone. For example, if asked, "What were your plans for today?", the device might respond, "I was planning to go shopping at 2 PM today."
[0461] Recognition and response to emotions
[0462] When a user gives a voice command, the device sends the voice to a server, where the emotion engine analyzes the user's emotional state. Based on the results, the system adjusts the content and timing of reminder notifications and generates comforting messages. For example, if the user says, "I'm not feeling very well," the server will generate a notification such as, "Please take some time to rest."
[0463] Example of a prompt
[0464] Please enter the following information by voice: "I will take my medicine at 8 AM tomorrow."
[0465] This invention enables the provision of personalized services that support the independent living of the elderly and take their emotions into consideration. This improves the elderly's sense of security and satisfaction, and enhances their quality of daily life.
[0466] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0467] Step 1:
[0468] The user performs voice input.
[0469] The user launches a smartphone application and uses voice input. For example, they might say, "I'm going to the doctor at 10 AM tomorrow." This input data is captured by the device in voice format.
[0470] Step 2:
[0471] The device records audio and sends it to the server.
[0472] The terminal receives the user's voice input and records it as audio data. This recorded data is sent to the server. Specifically, audio data is sent from the terminal to the server, and the server receives it. The input data is the user's voice data, and the output data is the audio data sent to the server.
[0473] Step 3:
[0474] The server generates audio data and analyzes it using an AI model.
[0475] The server inputs the received audio data into a generative AI model, which analyzes the audio and converts it into text data. The generative AI model converts the audio data into text and analyzes that text to identify the user's intent. For example, the output text might say, "I'm going to the doctor at 10 AM tomorrow." The input data is audio data, and the output data is the converted text data.
[0476] Step 4:
[0477] The server registers text data in the database.
[0478] The server analyzes user schedule information based on text data and registers it in the database. For example, information such as "I will go to the doctor at 10 AM tomorrow" is saved as a schedule in the database. The input data is text data, and the output data is the schedule information stored in the database.
[0479] Step 5:
[0480] The server analyzes the audio data using an emotion engine.
[0481] The server inputs audio data into the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes emotions from the tone and manner of speaking and outputs the results. For example, if the engine analyzes that the user is depressed, that information is output as emotion recognition data. The input data is audio data, and the output data is emotion recognition data.
[0482] Step 6:
[0483] The server generates and adjusts reminder notifications.
[0484] The server adjusts the content and timing of reminder notifications based on emotion recognition data. For example, if the server detects that the user is feeling down, the notification content changes to "Don't push yourself today, relax." The input data is emotion recognition data, and the output data is the adjusted reminder notification.
[0485] Step 7:
[0486] The server sends a reminder notification to the device.
[0487] The server sends a pre-arranged reminder notification to the device. For example, a reminder to "go to the doctor at 10 AM tomorrow" is sent to the device just before the scheduled time. The input data is the pre-arranged reminder notification, and the output data is the notification data sent to the device.
[0488] Step 8:
[0489] The device displays a reminder notification.
[0490] The device displays reminder notifications received from the server. For example, a notification saying "You have a doctor's appointment" is displayed on the screen and also announced audibly. The input data is the notification data received by the device, and the output data is the notification displayed to the user.
[0491] Through these steps, the system efficiently processes user voice input and utilizes a generative AI model and emotion engine to provide personalized support for daily life.
[0492] (Application Example 2)
[0493] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0494] This invention aims to solve problems related to support systems that enable elderly people to live independently and safely. While conventional technologies exist that convert voice input to text and provide mechanisms for schedule management and reminder notifications, they do not address the emotional state of the user or provide support for actual shopping. Elderly people often get lost or feel anxious when shopping in physical stores, and there is a need for a system that can alleviate these issues and allow them to shop in a relaxed state.
[0495] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0496] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for providing shopping assistance to the user based on the voice input, means including an emotion engine for analyzing the user's emotional state, and means for generating support messages based on the emotional state and sending them to the user's terminal. This makes it possible for elderly people to shop in physical stores without getting lost or feeling anxious, and to shop in a relaxed manner while receiving appropriate guidance and support.
[0497] A "voice input method" is a system for capturing a user's voice into a device such as a terminal and processing it.
[0498] A "text conversion method" is a mechanism for analyzing captured audio data and converting it into text-based data.
[0499] A "schedule information analysis tool" is a system that extracts and analyzes a user's schedule and tasks based on text data, and manages them accordingly.
[0500] A "database registration method" is a mechanism for registering analyzed schedule information into a database that centrally manages that information.
[0501] A "reminder notification generation method" is a mechanism for generating announcements and notifications for users at specific times and under specific conditions, based on registered schedule information.
[0502] A "reminder notification sending method" is a mechanism that sends generated notifications to the user's device so that the user can check them.
[0503] A "shopping assistant system" is a mechanism that supports users' shopping based on voice input, guiding them to necessary product information and locations.
[0504] An "emotion engine" is a technology that analyzes a user's voice data and evaluates and recognizes their emotional state.
[0505] A "support message generation mechanism" is a system for generating appropriate messages for users based on their analyzed emotional state.
[0506] A "support message sending mechanism" is a system that sends generated support messages to the user's device, allowing the user to receive them.
[0507] This invention relates to a system that supports elderly people in living independently and safely, and describes its specific embodiments. This system is primarily designed to support elderly people in their daily activities, such as shopping, through voice input. The main components of the system and their specific operation are described in detail below.
[0508] Server-side processing
[0509] User authentication and data storage
[0510] The server authenticates users when they log into the application. Upon successful authentication, it saves data such as the user's schedule and shopping list to a database. This information is managed and stored on the server whenever the user updates their schedule or tasks.
[0511] Operation of Generative AI Models
[0512] The server uses a generative AI model to analyze voice input from the user. When the user says "I'm looking for milk" by voice, the generative AI model converts that instruction into text and generates appropriate product information and directions to the store.
[0513] Operation of the Emotion Engine
[0514] The server uses an emotion engine to analyze the user's voice data to determine their emotional state. For example, if a user says, "I'm a little tired," the emotion engine recognizes this emotional state and generates an appropriate support message. This support message might include suggestions such as, "Let's take a break," or "I'll call a staff member."
[0515] Schedule notifications and support message sending
[0516] The server generates reminders and support messages based on the user's schedule information and emotional state, and sends them to the user's device. For example, if "milk" is included in the shopping list, when the user enters the nearest store, a message will be sent saying, "Milk is in the refrigerated section."
[0517] Terminal-side processing
[0518] Voice input function
[0519] User voice input is captured through the device's microphone. The captured voice is converted into text data by speech recognition software and sent to the server.
[0520] Display reminders
[0521] The device displays notifications and reminders sent from the server. For example, a reminder such as "Take your medicine at 3 PM" will be displayed on the user's smartphone or smart glasses.
[0522] Shopping assistant function
[0523] The device receives responses from the generative AI model and provides voice guidance to the user. For example, if the user asks, "Where is the milk section?", it will notify the user, "It's in the refrigerated section." It also provides voice support messages based on the analysis results of the emotion engine.
[0524] Hardware and software to use
[0525] This system uses the following hardware and software.
[0526] Hardware: Smartphone, smart glasses, microphone, speaker
[0527] Software: speech_recognition library (speech recognition), proprietary generative AI model, emotion engine
[0528] Specific example
[0529] In a real-world scenario, if a user goes to a shopping mall and uses voice input to say, "I'm looking for milk," the emotion engine detects that the user is feeling a little anxious and provides a message such as, "The milk is in the refrigerated section. Let's call a staff member nearby."
[0530] Examples of prompts for a generative AI model:
[0531] "When a user says 'I'm looking for milk,' analyze their response and the emotions they express to generate an appropriate response. If the user is anxious, include a reassuring message."
[0532] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0533] Step 1:
[0534] Acquiring voice input (user)
[0535] Users input voice commands through the microphone on their smartphone or smart glasses. This input is part of the user's shopping list and might say something like, "I'm looking for milk."
[0536] Input: Audio data
[0537] Output: Audio data (sent to the device)
[0538] Step 2:
[0539] Text conversion of audio data (on the device)
[0540] The device converts the acquired audio data into text data using a speech recognition library (for example, speech_recognition). The converted text becomes "I'm looking for milk."
[0541] Input: Audio data
[0542] Data processing: Convert speech to text using a speech recognition library.
[0543] Output: Text data
[0544] Step 3:
[0545] Sending text data (from terminal to server)
[0546] The terminal sends the converted text data to the server. The server receives this text data and begins analysis.
[0547] Input: Text data
[0548] Output: Text data sent to the server
[0549] Step 4:
[0550] User authentication and data storage (server)
[0551] The server performs user authentication before parsing the transmitted text data. If authentication is successful, the user's schedule and shopping list are saved to the database.
[0552] Input: User ID, password, submitted text data
[0553] Data processing: User authentication, saving to database.
[0554] Output: Authentication success message, saved user data
[0555] Step 5:
[0556] Analysis using a generative AI model (server)
[0557] The server uses a generative AI model to analyze text data and generate appropriate responses. For example, in response to the text "I'm looking for milk," the AI model generates the answer "It's in the refrigerated section."
[0558] Input: Text data
[0559] Data Processing: Text analysis and response generation using generative AI models.
[0560] Output: Generated response text
[0561] Step 6:
[0562] Emotional state analysis (server)
[0563] Along with the generated response text, the emotion engine analyzes the user's emotional state based on their voice data. For example, it can recognize that the user is feeling anxious and generate an appropriate support message.
[0564] Input: Audio data, generated response text
[0565] Data processing: Emotional analysis using an emotion engine
[0566] Output: Emotional state, support message
[0567] Step 7:
[0568] Sending responses and support messages (from server to terminal)
[0569] The server sends the generated response text and support message to the user's terminal. For example, a notification might be sent saying, "It's in the refrigerated section. I'll call a staff member nearby."
[0570] Input: Generated response text, support message
[0571] Output: Response text and support message sent to the terminal
[0572] Step 8:
[0573] Playback of response (terminal)
[0574] The terminal plays back the received response text and support message as audio. The user is notified, "It's in the refrigerated section. I'll call a staff member nearby," and voice assistance is provided.
[0575] Input: Response text, support message
[0576] Data processing: Converting text to speech using speech synthesis.
[0577] Output: Voice guidance and support messages
[0578] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0579] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0580] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0581] [Second Embodiment]
[0582] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0583] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0584] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0585] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0586] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0587] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0588] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0589] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0590] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0591] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0592] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0593] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0594] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. The system's program processing is described in detail below in natural language.
[0595] Server-side processing
[0596] User authentication and data storage
[0597] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[0598] Model operation of generative AI
[0599] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[0600] Schedule notification
[0601] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[0602] Terminal-side processing
[0603] Voice input function
[0604] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[0605] Display reminders
[0606] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[0607] Conversational interface
[0608] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[0609] User-side actions
[0610] Inputting voice commands
[0611] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[0612] Notification confirmation
[0613] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[0614] Conversational communication
[0615] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[0616] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. The simple operation via voice input and the provision of appropriate information by the generated AI create an efficient and secure living environment for the user.
[0617] The following describes the processing flow.
[0618] User authentication and data storage
[0619] Step 1:
[0620] The user launches the app on their smartphone.
[0621] The user taps the smartphone icon to open the app.
[0622] Step 2:
[0623] The device displays the login screen to the user.
[0624] The device displays a screen prompting the user to enter their user ID and password.
[0625] Step 3:
[0626] The user enters their user ID and password and taps the "Login" button.
[0627] The user enters the correct authentication information.
[0628] Step 4:
[0629] The device sends authentication information to the server.
[0630] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[0631] Step 5:
[0632] The server checks the authentication information received and searches the database for matching user information.
[0633] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[0634] Step 6:
[0635] The server sends the authentication result to the terminal.
[0636] The server returns the authentication success / failure result as a response.
[0637] Step 7:
[0638] The device displays a screen corresponding to the authentication result.
[0639] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[0640] Voice input and schedule registration
[0641] Step 1:
[0642] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[0643] The user taps the microphone button and enters their schedule by voice.
[0644] Step 2:
[0645] The device sends voice data to the server.
[0646] The device acquires voice data and sends it to the server.
[0647] Step 3:
[0648] The server receives the audio data and converts it into text using a generative AI model.
[0649] The server converts the audio data into text and then analyzes its content.
[0650] Step 4:
[0651] The server analyzes the text and extracts information related to the schedule.
[0652] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[0653] Step 5:
[0654] The server sends the registration results to the terminal.
[0655] The server notifies the terminal that the appointment registration was successful.
[0656] Step 6:
[0657] The device notifies the user of the registration result.
[0658] The device displays the message "The appointment has been successfully registered" to the user.
[0659] Reminder notifications and confirmations
[0660] Step 1:
[0661] The server sets the timing of reminders based on a schedule.
[0662] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[0663] Step 2:
[0664] When the server reaches the reminder time, it sends a reminder notification to the device.
[0665] The server will issue a reminder notification at the time set by the server.
[0666] Step 3:
[0667] The device receives a reminder notification.
[0668] The device receives the push notification and displays it on the screen.
[0669] Step 4:
[0670] The user checks the notification and takes action if necessary.
[0671] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[0672] Conversational interface
[0673] Step 1:
[0674] Users ask the app questions such as, "What's the weather like today?"
[0675] The user presses the microphone button to input their question by voice.
[0676] Step 2:
[0677] The device sends voice data to the server.
[0678] The device acquires voice data and sends it to the server.
[0679] Step 3:
[0680] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[0681] The server analyzes the audio data, retrieves weather information, and generates a response.
[0682] Step 4:
[0683] The server sends the generated response to the terminal.
[0684] The server sends a response to the terminal saying, "It's sunny today."
[0685] Step 5:
[0686] The device plays back the response it received as audio.
[0687] The device plays a voice message to the user saying, "It's sunny today."
[0688] The above is a detailed explanation of the program's processing flow, broken down into individual steps.
[0689] (Example 1)
[0690] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0691] In systems designed to support elderly individuals in living independent and safe daily lives, it is essential to efficiently analyze voice input, appropriately manage user schedules and provide reminder notifications, and enhance user convenience and peace of mind by offering interactive support. Furthermore, utilizing generative AI models to extract appropriate tasks from voice input and enabling natural dialogue with the user is also crucial.
[0692] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0693] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for operating a generative AI model and analyzing voice data to match it with specific tasks and information, and means for generating dialogue with the user based on the analysis results of the generative AI model. This makes it possible for elderly people to easily operate the system through voice input and to receive appropriate reminder notifications and interactive feedback.
[0694] "Means for receiving user voice input" refers to devices or systems for receiving information entered by a user via voice.
[0695] "Means for converting voice input to text" refers to methods and technologies for analyzing received voice data and converting it into text format.
[0696] "Means for analyzing user schedule information and registering it in a database" refers to a system or process that interprets information obtained as text and saves it in a database as the user's schedule or appointments.
[0697] "A means of generating and sending reminder notifications to the user's device" refers to a method for creating notifications based on registered schedule information and sending them to the user's device.
[0698] "A means of operating a generative AI model to analyze voice data and match it to specific tasks or information" refers to a technology that uses generative AI to process voice data and identify appropriate tasks or information.
[0699] "A means of generating user dialogue based on the analysis results of a generative AI model" refers to a system that generates and provides information for natural dialogue with the user based on the results of analysis by the generative AI.
[0700] This invention is an interactive lifestyle support system that helps elderly people live independently and safely, and provides its functions using a device such as a smartphone. This system accepts voice input from the user, analyzes it, and provides appropriate reminders and notifications, as well as conversational support utilizing a generative AI model.
[0701] Server-side processing
[0702] User Authentication
[0703] When a user logs into the app, the server authenticates them by comparing the entered user ID and password with the database. It also saves the data of successfully authenticated users to the database.
[0704] Specific operation: The user enters their user ID and password on their smartphone, and the server verifies this against the database to perform authentication.
[0705] Data storage
[0706] The server stores the schedule information and settings of authenticated users in a database. For example, if there are new appointments or changes to existing ones, this information is added to the database.
[0707] Specific operation: When a user voice-inputs "I want to change my dentist appointment to 3 PM tomorrow," the server saves this information to the database.
[0708] Operation of Generative AI Models
[0709] The server operates a generative AI model and analyzes the user's voice data. It extracts specific tasks and information from the voice input and registers them in a database.
[0710] Specific operation: When a user voice-inputs "I will take my medicine at 8 AM tomorrow," the generating AI model recognizes this as an appointment and saves it to the database.
[0711] Sending schedule notifications
[0712] The server generates reminder notifications based on the registered schedule information and sends them to the user's smartphone.
[0713] Specific action: A reminder notification to "take your medicine" is sent to the user's smartphone at 8:00 AM.
[0714] Terminal-side processing
[0715] Acquiring voice input
[0716] The device captures the user's voice and sends it to the server to be converted into text.
[0717] Specific operation: When the user uses voice input to say "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[0718] Display reminders
[0719] The device displays reminder notifications received from the server.
[0720] Specific action: A notification appears on the user's smartphone stating, "I'm going shopping at 2 PM."
[0721] Providing a conversational interface
[0722] The device plays the response from the generating AI as audio, allowing the user to interact with the app.
[0723] Specific operation: When the user voice-inputs "What's the weather like today?", the device plays the response "It's sunny today" from the server's AI generation system.
[0724] User-side processing
[0725] Inputting voice commands
[0726] Users use their smartphones to input their daily schedules and questions by voice.
[0727] Specific action: The user gives a voice command saying, "I will go to the doctor at 10 AM tomorrow."
[0728] Notification confirmation
[0729] Users check reminders and notifications displayed on their smartphones and take action based on them.
[0730] Specific actions: Check notifications such as "It's 12 o'clock now. It's lunchtime," and prepare the meal.
[0731] Conversational communication
[0732] Users can check and adjust their schedules by interacting with the app.
[0733] Specific action: When the user asks "What were your plans for today?", the app responds "I was planning to go shopping at 2pm today."
[0734] Examples of prompts for generative AI models
[0735] Example of a prompt
[0736] User voice input: "What's the weather like today?"
[0737] Generated AI model prompt: "The user wants to know today's weather. Please provide weather information such as sunny, cloudy, or rainy."
[0738] This system, with its interconnected components, provides support for elderly individuals to live independently through daily schedule management and dialogue. Its simple voice input operation and the provision of appropriate information by AI create an efficient and reassuring living environment for users.
[0739] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0740] Step 1:
[0741] Acquisition of user voice input
[0742] The user launches the app on their smartphone and enters voice commands. For example, they might voice-inform an appointment such as, "I'm going to the doctor tomorrow at 10 AM."
[0743] The terminal receives this audio and performs initial processing. The terminal receives the input as an audio signal and prepares it for conversion to text.
[0744] Input: Voice command (Example: "I will go to the doctor at 10 AM tomorrow")
[0745] Output: Audio data
[0746] Step 2:
[0747] Sending audio data to the server
[0748] The terminal uses a communication protocol to send the acquired voice data to the server. This transfers the user's voice information to the server.
[0749] Input: Audio data
[0750] Output: Sending audio data to the server
[0751] Specific operation: The device uses an internet connection to send voice data to the server. For example, it uses Wi-Fi or mobile data communication.
[0752] Step 3:
[0753] Voice input to text conversion
[0754] The server analyzes the received audio data and converts it into text. Speech recognition technology is used to convert the audio signal into linguistic information. Speech recognition APIs and software are used for this process.
[0755] Input: Audio data
[0756] Output: Text data (Example: "I'm going to the doctor at 10 AM tomorrow")
[0757] Specific operation: The server calls a speech recognition API to convert the audio data into text data. This API is often a cloud-based service.
[0758] Step 4:
[0759] Text data analysis and database registration
[0760] The server analyzes text data and extracts user schedule information. This information is then registered in a database. Natural language processing technology is used for the analysis, and the schedule information is added to the database as a new record.
[0761] Input: Text data
[0762] Output: Schedule information registered in the database
[0763] Specific operation: The server uses a natural language processing algorithm to extract date and time information from the text. Then, it stores the extracted data in a relational database.
[0764] Step 5:
[0765] Voice analysis of generative AI models
[0766] The server uses a generative AI model to analyze text data and match it with relevant tasks and information. The generative AI model understands the context from voice input and determines the appropriate action.
[0767] Input: Text data
[0768] Output: Matched task information or answers
[0769] Specific operation: The server runs the generated AI model, and the text "I will go to the doctor at 10 AM tomorrow" is matched as a scheduled task.
[0770] Step 6:
[0771] Generating and sending scheduled notifications
[0772] The server generates reminder notifications based on registered schedule information and sends these notifications to the user's device. Appropriate reminders are generated based on the specified time or conditions.
[0773] Input: Schedule information in the database
[0774] Output: Reminder notification to user's device
[0775] Specific operation: The server creates a reminder notification based on the schedule information and sends it to the user's smartphone using a push notification service. For example, a notification might be set saying, "I have a doctor's appointment tomorrow at 10 AM."
[0776] Step 7:
[0777] Display reminders
[0778] The device displays reminder notifications received from the server. The user checks the notification and takes action.
[0779] Input: Reminder notification from server
[0780] Output: Notification displayed on the smartphone screen
[0781] Specific action: The smartphone receives a push notification and displays a message on the screen saying, "Go to the doctor tomorrow at 10 AM." The user can tap it to view details or set an alarm.
[0782] Step 8:
[0783] Providing a conversational interface
[0784] The device plays back the responses from the generated AI as audio, allowing the user to communicate with the app in a conversational format.
[0785] Input: Generated AI response data from the server
[0786] Output: Generated AI response played back as audio.
[0787] Specific operation: When a user asks the device "What's the weather like today?", the device sends this to the server, and the AI-generated response "It's sunny today" is played aloud. This allows the user to have a natural conversation with the system.
[0788] As described above, through each processing step, the present invention provides a system that supports elderly people in living independently and safely, and realizes a convenient and secure living environment through voice input.
[0789] (Application Example 1)
[0790] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0791] To improve work efficiency within factories, there is a need for means to provide workers with timely and appropriate information and task reminders. Furthermore, an interactive user interface that is easy for elderly or inexperienced workers to use is required.
[0792] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0793] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for analyzing voice input from factory workers and matching it to specific tasks, and means for a robot to send push notifications to workers based on the work schedule. This makes it possible for workers to perform their work efficiently and obtain necessary information in a timely manner.
[0794] "User voice input" refers to voice commands that a user gives to the system through a voice input device such as a microphone.
[0795] "Means of converting to text" refers to software or hardware that has the function of converting voice-input data into text information.
[0796] "User schedule information" refers to data that records the user's planned activities and task schedules.
[0797] "Means of registering in the database" refers to the function of storing the analyzed schedule information in the database.
[0798] A "reminder notification" is a notification sent to a user based on a pre-set schedule.
[0799] "User's device" refers to electronic devices used by the user, such as smartphones, tablets, and personal computers.
[0800] A "factory worker" is an employee who performs physical or mechanical tasks within a factory.
[0801] "Means for analyzing voice input" refers to software or hardware that has the function of appropriately understanding and classifying received voice data.
[0802] "Means for matching to specific tasks" refers to a function that links information obtained through voice analysis to related tasks and work.
[0803] "Means of sending push notifications" refers to a function that sends information to a device in real time and displays a notification to the user.
[0804] Program generation
[0805] To implement this application, a program is needed that accepts user voice input and converts it into text. This program uses a speech recognition API (e.g., Google Cloud Speech-to-Text). After the voice data is converted to text, that text data is sent to a generating AI model (e.g., OpenAI GPT-4) to be matched with specific tasks or information.
[0806] Hardware and software usage
[0807] Hardware:
[0808] Microphone: Used to capture the user's voice input.
[0809] Speaker: Used to play the generated audio response.
[0810] Display (built into the robot): Used to display text information and reminders.
[0811] software:
[0812] Speech recognition APIs (e.g., Google Cloud Speech-to-Text): Used to convert speech data into text.
[0813] Generative AI models (e.g., OpenAI GPT-4): Used to analyze text data and match it to specific tasks or information.
[0814] Database management systems (e.g., MySQL): Used to store and manage schedule information.
[0815] Program processing
[0816] 1. Voice input:
[0817] The server captures the user's voice using the microphone and temporarily stores it locally.
[0818] For example, if a factory worker says, "Tell me the inventory of parts," the voice data is sent to the server and temporarily stored.
[0819] 2. Converting and transmitting audio data:
[0820] The speech recognition API is used to convert the audio data into text, which is then sent to a server where the generation AI model is running.
[0821] Example: The audio "Tell me the stock of parts" is converted to the text "Tell me the stock of parts".
[0822] 3. Processing of Generative AI Models:
[0823] The generating AI analyzes the received text data and extracts related tasks and information.
[0824] Example: Based on the instruction "Tell me the inventory status of the parts," the system retrieves the necessary information from the inventory database.
[0825] 4. Provision of information and notification:
[0826] It generates necessary information and reminders and sends them from the server to the robot.
[0827] For example, information such as "We have very little stock left of part A" is generated and communicated to the worker as a notification via the robot's display or speaker.
[0828] 5. Output:
[0829] The robot provides information to workers through voice and displays.
[0830] For example, if you ask "What's the next step?", you might get the answer "The next step is to install part D."
[0831] Examples of specific cases and prompt statements
[0832] As a concrete example, if a factory worker asks a robot, "When is my break time?", the robot will respond, "Today's break times are 10 AM and 3 PM." An example of the prompt in this case would be as follows:
[0833] "Please generate today's break times based on the voice instructions from the factory workers."
[0834] This allows workers to perform their tasks efficiently and obtain necessary information in a timely manner.
[0835] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0836] Step 1:
[0837] Accepts user voice input.
[0838] Input: The user speaks into the microphone and says, "Tell me the parts inventory."
[0839] Processing details: The device recognizes this voice input and temporarily saves it locally.
[0840] Output: An audio data file is generated and passed to the next processing step.
[0841] Step 2:
[0842] Convert audio data to text
[0843] Input: The audio data file generated in Step 1.
[0844] Processing details: Convert audio data into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text).
[0845] Output: The text data "Tell me the stock of parts" is generated.
[0846] Step 3:
[0847] Send text data to an AI model for generation.
[0848] Input: Text data generated in Step 2.
[0849] Processing details: Text data is sent to a server where an AI model (e.g., OpenAI GPT-4) is running.
[0850] Output: Text data is received by the generating AI model.
[0851] Step 4:
[0852] The generative AI model analyzes the text data.
[0853] Input: Text data received in Step 3.
[0854] Processing details: The generating AI model analyzes text data to identify tasks and information. In this case, it extracts the necessary information from the inventory database.
[0855] Output: Information such as "We have very little stock left of part A."
[0856] Step 5:
[0857] Create information and generate notifications
[0858] Input: Information obtained in Step 4.
[0859] Processing details: The server generates a notification message based on this information and sends it to the terminal.
[0860] Output: A notification message is generated stating, "We have very little stock left of part A."
[0861] Step 6:
[0862] Display and play audible notification messages.
[0863] Input: The notification message generated in Step 5.
[0864] Processing details: The device displays this notification message on its screen and plays it aloud through its speaker.
[0865] Output: The user is notified visually and audibly that "Part A is running low on stock."
[0866] Through the above processing steps, workers can efficiently receive information and act accordingly.
[0867] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0868] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it enables more personalized support. The system's program processing is described in detail below in natural language.
[0869] Server-side processing
[0870] User authentication and data storage
[0871] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[0872] Model operation of generative AI
[0873] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[0874] Operation of the Emotion Engine
[0875] The server operates an emotion engine that analyzes the user's emotional state from voice data. This allows it to adjust schedule information and reminder notifications based on the user's emotional state. For example, if it detects that the user is feeling down, it can adjust the timing and content of reminder notifications.
[0876] Schedule notification
[0877] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[0878] Terminal-side processing
[0879] Voice input function
[0880] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[0881] Display reminders
[0882] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[0883] Conversational interface
[0884] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[0885] User-side actions
[0886] Inputting voice commands
[0887] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[0888] Notification confirmation
[0889] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[0890] Conversational communication
[0891] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[0892] Recognition and response to emotions
[0893] When a user asks a question or gives a command by voice, the device sends the audio to the server, where the emotion engine analyzes the user's emotional state. For example, if the user says, "I'm not feeling very well," the server uses the emotion engine to analyze the user's emotional state and adjusts the content and timing of reminder notifications or generates comforting messages based on the results.
[0894] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. By combining it with an emotion engine, personalized support based on the user's emotional state becomes possible, further enhancing their sense of security and satisfaction.
[0895] The following describes the processing flow.
[0896] User authentication and data storage
[0897] Step 1:
[0898] The user launches the app on their smartphone.
[0899] The user taps the smartphone icon to open the app.
[0900] Step 2:
[0901] The device displays the login screen to the user.
[0902] The device displays a screen prompting the user to enter their user ID and password.
[0903] Step 3:
[0904] The user enters their user ID and password and taps the "Login" button.
[0905] The user enters the correct authentication information.
[0906] Step 4:
[0907] The device sends authentication information to the server.
[0908] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[0909] Step 5:
[0910] The server checks the authentication information received and searches the database for matching user information.
[0911] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[0912] Step 6:
[0913] The server sends the authentication result to the terminal.
[0914] The server returns the authentication success / failure result as a response.
[0915] Step 7:
[0916] The device displays a screen corresponding to the authentication result.
[0917] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[0918] Voice input and schedule registration
[0919] Step 1:
[0920] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[0921] The user taps the microphone button and enters their schedule by voice.
[0922] Step 2:
[0923] The device sends voice data to the server.
[0924] The device acquires voice data and sends it to the server.
[0925] Step 3:
[0926] The server receives the audio data and converts it into text using a generative AI model.
[0927] The server converts the audio data into text and then analyzes its content.
[0928] Step 4:
[0929] The server analyzes the text and extracts information related to the schedule.
[0930] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[0931] Step 5:
[0932] The server sends the registration results to the terminal.
[0933] The server notifies the terminal that the appointment registration was successful.
[0934] Step 6:
[0935] The device notifies the user of the registration result.
[0936] The device displays the message "The appointment has been successfully registered" to the user.
[0937] Reminder notifications and confirmations
[0938] Step 1:
[0939] The server sets the timing of reminders based on a schedule.
[0940] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[0941] Step 2:
[0942] When the server reaches the reminder time, it sends a reminder notification to the device.
[0943] The server will issue a reminder notification at the time set by the server.
[0944] Step 3:
[0945] The device receives a reminder notification.
[0946] The device receives the push notification and displays it on the screen.
[0947] Step 4:
[0948] The user checks the notification and takes action if necessary.
[0949] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[0950] Conversational interface
[0951] Step 1:
[0952] Users ask the app questions such as, "What's the weather like today?"
[0953] The user presses the microphone button to input their question by voice.
[0954] Step 2:
[0955] The device sends voice data to the server.
[0956] The device acquires voice data and sends it to the server.
[0957] Step 3:
[0958] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[0959] The server analyzes the audio data, retrieves weather information, and generates a response.
[0960] Step 4:
[0961] The server sends the generated response to the terminal.
[0962] The server sends a response to the terminal saying, "It's sunny today."
[0963] Step 5:
[0964] The device plays back the response it received as audio.
[0965] The device plays a voice message to the user saying, "It's sunny today."
[0966] Operation of the Emotion Engine
[0967] Step 1:
[0968] When a user voices a question or gives instructions, that voice data is sent to the server.
[0969] For example, the user might say, "What should I do today?"
[0970] Step 2:
[0971] The device sends voice data to the server.
[0972] The device acquires voice data and sends it to the emotion engine.
[0973] Step 3:
[0974] The server analyzes the voice data, and the emotion engine recognizes the user's emotions.
[0975] The server identifies whether it is "stressed" or "depressed."
[0976] Step 4:
[0977] The server sends appropriate responses or reminders to the user based on the emotion recognition results.
[0978] For example, when a user is feeling down, you can send them an encouraging message.
[0979] Step 5:
[0980] The device displays or plays emotion-based responses and reminders to the user via audio.
[0981] A message such as, "Why don't you take a short break to make yourself feel better?" is displayed or played audibly.
[0982] The above is a detailed breakdown of the system's processing steps, including the emotion engine.
[0983] (Example 2)
[0984] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0985] For elderly people, managing their daily schedules and lacking communication are problems that hinder their independent and safe living. Furthermore, conventional reminder and schedule management systems only send uniform notifications without considering the user's emotional state, resulting in a lack of psychological support. To address these issues, there is a need for personalized support that appropriately analyzes the user's emotional state and adapts to their individual circumstances.
[0986] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0987] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for inputting voice data into a generation AI model and utilizing the processing results, means for inputting voice data into an emotion engine and analyzing the emotional state, and means for adjusting the content or timing of reminder notifications based on the analysis results. This makes it possible to provide personalized schedule management and psychological support tailored to the emotional state of elderly people.
[0988] "User" refers to an individual who uses this system.
[0989] "Means for receiving voice input" refers to a device or method for collecting voice uttered by a user and inputting it into a system.
[0990] "Means of converting voice input to text" refers to a device or algorithm that analyzes collected voice data and converts it into textual information.
[0991] "Means for analyzing user schedule information and registering it in a database" refers to a device or method that analyzes a user's schedule based on text data and stores that information in a database.
[0992] "Means for generating and sending reminder notifications to the user's device" refers to a device or method that creates a notification based on registered schedule information and sends it to the user's device.
[0993] A "generative AI model" refers to a model that uses artificial intelligence technology to analyze speech and text data and generate appropriate information.
[0994] An "emotion engine" refers to an algorithm or device used to analyze a user's emotional state from voice data.
[0995] "Means for adjusting the content or timing of reminder notifications" refers to devices or methods for changing the text or sending time of notifications based on the results of the emotion engine's analysis.
[0996] A "server" refers to a computer system used for processing and analyzing data across an entire system.
[0997] "Terminal" refers to devices such as smartphones and tablets that are directly operated by the user.
[0998] This invention is an interactive life support system that helps elderly people live independently and safely. It accepts voice input from the user and provides personalized reminder notifications and schedule management using a generative AI model and emotion engine. A specific embodiment of the system is described below.
[0999] Server-side embodiment
[1000] User authentication and data storage
[1001] When a user logs into the app using a smartphone or tablet, the server accepts the user ID and password, verifies them against the database, and performs authentication. If authentication is successful, reusable user data and schedule information are saved to the database. For example, if a user logs in with the ID "abcd1234" and password "password," the server verifies this, and if authentication is successful, saves the user's new appointments and changes to the database.
[1002] Operation of Generative AI Models
[1003] The server receives voice input from the user and analyzes it using a generative AI model. The voice input is converted into text and matched to specific tasks or information. For example, if a user voice-inputs "I'm going for a walk at 3pm," the server passes this to the generative AI model, which converts it into text and registers "walk" as an event in the database.
[1004] Operation of the Emotion Engine
[1005] The server inputs voice data into an emotion engine to analyze the user's emotional state. This allows the server to adjust the content and timing of reminder notifications according to the user's emotions. For example, if a user says, "I'm not feeling well today," the server analyzes this and sets a reminder such as, "It's time to rest."
[1006] Terminal-side embodiment
[1007] Voice input function
[1008] The terminal receives voice input from the user and sends it to the server. As an initial process, it generates speech recognition data and sends it to the server for text conversion. For example, if the user says, "I will take my medicine at 2 pm," the terminal records this voice and sends it to the server.
[1009] Display reminders
[1010] The device receives reminder notifications sent from the server and displays them on the screen. Voice responses are also supported, allowing the user to check the notifications on the device and take the necessary actions. For example, if the device receives a reminder to "take your medicine at 2 PM," it will display this on the screen and also provide an audio notification.
[1011] Conversational interface
[1012] The device plays back responses from the generative AI model as audio, providing an environment where users can interact with the app naturally. For example, if a user asks, "What's the weather like today?", the device sends that audio to the server, and the generative AI model generates a response such as "It's sunny today." The device then plays this response back to the user as audio.
[1013] User-side embodiment
[1014] Inputting voice commands
[1015] Users use their smartphones to set daily schedules and ask questions using voice input. For example, they might speak instructions to their smartphone such as, "I'll go to the doctor at 10 AM tomorrow."
[1016] Notification confirmation
[1017] Users check reminder notifications displayed on their smartphones and act accordingly. For example, they might receive a notification saying, "It's 12 o'clock now. It's time for lunch."
[1018] Conversational communication
[1019] Users interact with the app via voice to configure and confirm settings. This helps alleviate feelings of loneliness for those living alone. For example, if asked, "What were your plans for today?", the device might respond, "I was planning to go shopping at 2 PM today."
[1020] Recognition and response to emotions
[1021] When a user gives a voice command, the device sends the voice to a server, where the emotion engine analyzes the user's emotional state. Based on the results, the system adjusts the content and timing of reminder notifications and generates comforting messages. For example, if the user says, "I'm not feeling very well," the server will generate a notification such as, "Please take some time to rest."
[1022] Example of a prompt
[1023] Please enter the following information by voice: "I will take my medicine at 8 AM tomorrow."
[1024] This invention enables the provision of personalized services that support the independent living of the elderly and take their emotions into consideration. This improves the elderly's sense of security and satisfaction, and enhances their quality of daily life.
[1025] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1026] Step 1:
[1027] The user performs voice input.
[1028] The user launches a smartphone application and uses voice input. For example, they might say, "I'm going to the doctor at 10 AM tomorrow." This input data is captured by the device in voice format.
[1029] Step 2:
[1030] The device records audio and sends it to the server.
[1031] The terminal receives the user's voice input and records it as audio data. This recorded data is sent to the server. Specifically, audio data is sent from the terminal to the server, and the server receives it. The input data is the user's voice data, and the output data is the audio data sent to the server.
[1032] Step 3:
[1033] The server generates audio data and analyzes it using an AI model.
[1034] The server inputs the received audio data into a generative AI model, which analyzes the audio and converts it into text data. The generative AI model converts the audio data into text and analyzes that text to identify the user's intent. For example, the output text might say, "I'm going to the doctor at 10 AM tomorrow." The input data is audio data, and the output data is the converted text data.
[1035] Step 4:
[1036] The server registers text data in the database.
[1037] The server analyzes user schedule information based on text data and registers it in the database. For example, information such as "I will go to the doctor at 10 AM tomorrow" is saved as a schedule in the database. The input data is text data, and the output data is the schedule information stored in the database.
[1038] Step 5:
[1039] The server analyzes the audio data using an emotion engine.
[1040] The server inputs audio data into the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes emotions from the tone and manner of speaking and outputs the results. For example, if the engine analyzes that the user is depressed, that information is output as emotion recognition data. The input data is audio data, and the output data is emotion recognition data.
[1041] Step 6:
[1042] The server generates and adjusts reminder notifications.
[1043] The server adjusts the content and timing of reminder notifications based on emotion recognition data. For example, if the server detects that the user is feeling down, the notification content changes to "Don't push yourself today, relax." The input data is emotion recognition data, and the output data is the adjusted reminder notification.
[1044] Step 7:
[1045] The server sends a reminder notification to the device.
[1046] The server sends a pre-arranged reminder notification to the device. For example, a reminder to "go to the doctor at 10 AM tomorrow" is sent to the device just before the scheduled time. The input data is the pre-arranged reminder notification, and the output data is the notification data sent to the device.
[1047] Step 8:
[1048] The device displays a reminder notification.
[1049] The device displays reminder notifications received from the server. For example, a notification saying "You have a doctor's appointment" is displayed on the screen and also announced audibly. The input data is the notification data received by the device, and the output data is the notification displayed to the user.
[1050] Through these steps, the system efficiently processes user voice input and utilizes a generative AI model and emotion engine to provide personalized support for daily life.
[1051] (Application Example 2)
[1052] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1053] This invention aims to solve problems related to support systems that enable elderly people to live independently and safely. While conventional technologies exist that convert voice input to text and provide mechanisms for schedule management and reminder notifications, they do not address the emotional state of the user or provide support for actual shopping. Elderly people often get lost or feel anxious when shopping in physical stores, and there is a need for a system that can alleviate these issues and allow them to shop in a relaxed state.
[1054] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1055] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for providing shopping assistance to the user based on the voice input, means including an emotion engine for analyzing the user's emotional state, and means for generating support messages based on the emotional state and sending them to the user's terminal. This makes it possible for elderly people to shop in physical stores without getting lost or feeling anxious, and to shop in a relaxed manner while receiving appropriate guidance and support.
[1056] A "voice input method" is a system for capturing a user's voice into a device such as a terminal and processing it.
[1057] A "text conversion method" is a mechanism for analyzing captured audio data and converting it into text-based data.
[1058] A "schedule information analysis tool" is a system that extracts and analyzes a user's schedule and tasks based on text data, and manages them accordingly.
[1059] A "database registration method" is a mechanism for registering analyzed schedule information into a database that centrally manages that information.
[1060] A "reminder notification generation method" is a mechanism for generating announcements and notifications for users at specific times and under specific conditions, based on registered schedule information.
[1061] A "reminder notification sending method" is a mechanism that sends generated notifications to the user's device so that the user can check them.
[1062] A "shopping assistant system" is a mechanism that supports users' shopping based on voice input, guiding them to necessary product information and locations.
[1063] An "emotion engine" is a technology that analyzes a user's voice data and evaluates and recognizes their emotional state.
[1064] A "support message generation mechanism" is a system for generating appropriate messages for users based on their analyzed emotional state.
[1065] A "support message sending mechanism" is a system that sends generated support messages to the user's device, allowing the user to receive them.
[1066] This invention relates to a system that supports elderly people in living independently and safely, and describes its specific embodiments. This system is primarily designed to support elderly people in their daily activities, such as shopping, through voice input. The main components of the system and their specific operation are described in detail below.
[1067] Server-side processing
[1068] User authentication and data storage
[1069] The server authenticates users when they log into the application. Upon successful authentication, it saves data such as the user's schedule and shopping list to a database. This information is managed and stored on the server whenever the user updates their schedule or tasks.
[1070] Operation of Generative AI Models
[1071] The server uses a generative AI model to analyze voice input from the user. When the user says "I'm looking for milk" by voice, the generative AI model converts that instruction into text and generates appropriate product information and directions to the store.
[1072] Operation of the Emotion Engine
[1073] The server uses an emotion engine to analyze the user's voice data to determine their emotional state. For example, if a user says, "I'm a little tired," the emotion engine recognizes this emotional state and generates an appropriate support message. This support message might include suggestions such as, "Let's take a break," or "I'll call a staff member."
[1074] Schedule notifications and support message sending
[1075] The server generates reminders and support messages based on the user's schedule information and emotional state, and sends them to the user's device. For example, if "milk" is included in the shopping list, when the user enters the nearest store, a message will be sent saying, "Milk is in the refrigerated section."
[1076] Terminal-side processing
[1077] Voice input function
[1078] User voice input is captured through the device's microphone. The captured voice is converted into text data by speech recognition software and sent to the server.
[1079] Display reminders
[1080] The device displays notifications and reminders sent from the server. For example, a reminder such as "Take your medicine at 3 PM" will be displayed on the user's smartphone or smart glasses.
[1081] Shopping assistant function
[1082] The device receives responses from the generative AI model and provides voice guidance to the user. For example, if the user asks, "Where is the milk section?", it will notify the user, "It's in the refrigerated section." It also provides voice support messages based on the analysis results of the emotion engine.
[1083] Hardware and software to use
[1084] This system uses the following hardware and software.
[1085] Hardware: Smartphone, smart glasses, microphone, speaker
[1086] Software: speech_recognition library (speech recognition), proprietary generative AI model, emotion engine
[1087] Specific example
[1088] In a real-world scenario, if a user goes to a shopping mall and uses voice input to say, "I'm looking for milk," the emotion engine detects that the user is feeling a little anxious and provides a message such as, "The milk is in the refrigerated section. Let's call a staff member nearby."
[1089] Examples of prompts for a generative AI model:
[1090] "When a user says 'I'm looking for milk,' analyze their response and the emotions they express to generate an appropriate response. If the user is anxious, include a reassuring message."
[1091] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1092] Step 1:
[1093] Acquiring voice input (user)
[1094] Users input voice commands through the microphone on their smartphone or smart glasses. This input is part of the user's shopping list and might say something like, "I'm looking for milk."
[1095] Input: Audio data
[1096] Output: Audio data (sent to the device)
[1097] Step 2:
[1098] Text conversion of audio data (on the device)
[1099] The device converts the acquired audio data into text data using a speech recognition library (for example, speech_recognition). The converted text becomes "I'm looking for milk."
[1100] Input: Audio data
[1101] Data processing: Convert speech to text using a speech recognition library.
[1102] Output: Text data
[1103] Step 3:
[1104] Sending text data (from terminal to server)
[1105] The terminal sends the converted text data to the server. The server receives this text data and begins analysis.
[1106] Input: Text data
[1107] Output: Text data sent to the server
[1108] Step 4:
[1109] User authentication and data storage (server)
[1110] The server performs user authentication before parsing the transmitted text data. If authentication is successful, the user's schedule and shopping list are saved to the database.
[1111] Input: User ID, password, submitted text data
[1112] Data processing: User authentication, saving to database.
[1113] Output: Authentication success message, saved user data
[1114] Step 5:
[1115] Analysis using a generative AI model (server)
[1116] The server uses a generative AI model to analyze text data and generate appropriate responses. For example, in response to the text "I'm looking for milk," the AI model generates the answer "It's in the refrigerated section."
[1117] Input: Text data
[1118] Data Processing: Text analysis and response generation using generative AI models.
[1119] Output: Generated response text
[1120] Step 6:
[1121] Emotional state analysis (server)
[1122] Along with the generated response text, the emotion engine analyzes the user's emotional state based on their voice data. For example, it can recognize that the user is feeling anxious and generate an appropriate support message.
[1123] Input: Audio data, generated response text
[1124] Data processing: Emotional analysis using an emotion engine
[1125] Output: Emotional state, support message
[1126] Step 7:
[1127] Sending responses and support messages (from server to terminal)
[1128] The server sends the generated response text and support message to the user's terminal. For example, a notification might be sent saying, "It's in the refrigerated section. I'll call a staff member nearby."
[1129] Input: Generated response text, support message
[1130] Output: Response text and support message sent to the terminal
[1131] Step 8:
[1132] Playback of response (terminal)
[1133] The terminal plays back the received response text and support message as audio. The user is notified, "It's in the refrigerated section. I'll call a staff member nearby," and voice assistance is provided.
[1134] Input: Response text, support message
[1135] Data processing: Converting text to speech using speech synthesis.
[1136] Output: Voice guidance and support messages
[1137] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1138] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1139] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1140] [Third Embodiment]
[1141] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1142] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1143] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1144] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1145] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1146] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1147] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1148] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1149] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1150] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1151] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1152] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1153] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. The system's program processing is described in detail below in natural language.
[1154] Server-side processing
[1155] User authentication and data storage
[1156] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[1157] Model operation of generative AI
[1158] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[1159] Schedule notification
[1160] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[1161] Terminal-side processing
[1162] Voice input function
[1163] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[1164] Display reminders
[1165] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[1166] Conversational interface
[1167] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[1168] User-side actions
[1169] Inputting voice commands
[1170] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[1171] Notification confirmation
[1172] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[1173] Conversational communication
[1174] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[1175] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. The simple operation via voice input and the provision of appropriate information by the generated AI create an efficient and secure living environment for the user.
[1176] The following describes the processing flow.
[1177] User authentication and data storage
[1178] Step 1:
[1179] The user launches the app on their smartphone.
[1180] The user taps the smartphone icon to open the app.
[1181] Step 2:
[1182] The device displays the login screen to the user.
[1183] The device displays a screen prompting the user to enter their user ID and password.
[1184] Step 3:
[1185] The user enters their user ID and password and taps the "Login" button.
[1186] The user enters the correct authentication information.
[1187] Step 4:
[1188] The device sends authentication information to the server.
[1189] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[1190] Step 5:
[1191] The server checks the authentication information received and searches the database for matching user information.
[1192] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[1193] Step 6:
[1194] The server sends the authentication result to the terminal.
[1195] The server returns the authentication success / failure result as a response.
[1196] Step 7:
[1197] The device displays a screen corresponding to the authentication result.
[1198] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[1199] Voice input and schedule registration
[1200] Step 1:
[1201] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[1202] The user taps the microphone button and enters their schedule by voice.
[1203] Step 2:
[1204] The device sends voice data to the server.
[1205] The device acquires voice data and sends it to the server.
[1206] Step 3:
[1207] The server receives the audio data and converts it into text using a generative AI model.
[1208] The server converts the audio data into text and then analyzes its content.
[1209] Step 4:
[1210] The server analyzes the text and extracts information related to the schedule.
[1211] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[1212] Step 5:
[1213] The server sends the registration results to the terminal.
[1214] The server notifies the terminal that the appointment registration was successful.
[1215] Step 6:
[1216] The device notifies the user of the registration result.
[1217] The device displays the message "The appointment has been successfully registered" to the user.
[1218] Reminder notifications and confirmations
[1219] Step 1:
[1220] The server sets the timing of reminders based on a schedule.
[1221] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[1222] Step 2:
[1223] When the server reaches the reminder time, it sends a reminder notification to the device.
[1224] The server will issue a reminder notification at the time set by the server.
[1225] Step 3:
[1226] The device receives a reminder notification.
[1227] The device receives the push notification and displays it on the screen.
[1228] Step 4:
[1229] The user checks the notification and takes action if necessary.
[1230] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[1231] Conversational interface
[1232] Step 1:
[1233] Users ask the app questions such as, "What's the weather like today?"
[1234] The user presses the microphone button to input their question by voice.
[1235] Step 2:
[1236] The device sends voice data to the server.
[1237] The device acquires voice data and sends it to the server.
[1238] Step 3:
[1239] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[1240] The server analyzes the audio data, retrieves weather information, and generates a response.
[1241] Step 4:
[1242] The server sends the generated response to the terminal.
[1243] The server sends a response to the terminal saying, "It's sunny today."
[1244] Step 5:
[1245] The device plays back the response it received as audio.
[1246] The device plays a voice message to the user saying, "It's sunny today."
[1247] The above is a detailed explanation of the program's processing flow, broken down into individual steps.
[1248] (Example 1)
[1249] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1250] In systems designed to support elderly individuals in living independent and safe daily lives, it is essential to efficiently analyze voice input, appropriately manage user schedules and provide reminder notifications, and enhance user convenience and peace of mind by offering interactive support. Furthermore, utilizing generative AI models to extract appropriate tasks from voice input and enabling natural dialogue with the user is also crucial.
[1251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1252] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for operating a generative AI model and analyzing voice data to match it with specific tasks and information, and means for generating dialogue with the user based on the analysis results of the generative AI model. This makes it possible for elderly people to easily operate the system through voice input and to receive appropriate reminder notifications and interactive feedback.
[1253] "Means for receiving user voice input" refers to devices or systems for receiving information entered by a user via voice.
[1254] "Means for converting voice input to text" refers to methods and technologies for analyzing received voice data and converting it into text format.
[1255] "Means for analyzing user schedule information and registering it in a database" refers to a system or process that interprets information obtained as text and saves it in a database as the user's schedule or appointments.
[1256] "A means of generating and sending reminder notifications to the user's device" refers to a method for creating notifications based on registered schedule information and sending them to the user's device.
[1257] "A means of operating a generative AI model to analyze voice data and match it to specific tasks or information" refers to a technology that uses generative AI to process voice data and identify appropriate tasks or information.
[1258] "A means of generating user dialogue based on the analysis results of a generative AI model" refers to a system that generates and provides information for natural dialogue with the user based on the results of analysis by the generative AI.
[1259] This invention is an interactive lifestyle support system that helps elderly people live independently and safely, and provides its functions using a device such as a smartphone. This system accepts voice input from the user, analyzes it, and provides appropriate reminders and notifications, as well as conversational support utilizing a generative AI model.
[1260] Server-side processing
[1261] User Authentication
[1262] When a user logs into the app, the server authenticates them by comparing the entered user ID and password with the database. It also saves the data of successfully authenticated users to the database.
[1263] Specific operation: The user enters their user ID and password on their smartphone, and the server verifies this against the database to perform authentication.
[1264] Data storage
[1265] The server stores the schedule information and settings of authenticated users in a database. For example, if there are new appointments or changes to existing ones, this information is added to the database.
[1266] Specific operation: When a user voice-inputs "I want to change my dentist appointment to 3 PM tomorrow," the server saves this information to the database.
[1267] Operation of Generative AI Models
[1268] The server operates a generative AI model and analyzes the user's voice data. It extracts specific tasks and information from the voice input and registers them in a database.
[1269] Specific operation: When a user voice-inputs "I will take my medicine at 8 AM tomorrow," the generating AI model recognizes this as an appointment and saves it to the database.
[1270] Sending schedule notifications
[1271] The server generates reminder notifications based on the registered schedule information and sends them to the user's smartphone.
[1272] Specific action: A reminder notification to "take your medicine" is sent to the user's smartphone at 8:00 AM.
[1273] Terminal-side processing
[1274] Acquiring voice input
[1275] The device captures the user's voice and sends it to the server to be converted into text.
[1276] Specific operation: When the user uses voice input to say "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[1277] Display reminders
[1278] The device displays reminder notifications received from the server.
[1279] Specific action: A notification appears on the user's smartphone stating, "I'm going shopping at 2 PM."
[1280] Providing a conversational interface
[1281] The device plays the response from the generating AI as audio, allowing the user to interact with the app.
[1282] Specific operation: When the user voice-inputs "What's the weather like today?", the device plays the response "It's sunny today" from the server's AI generation system.
[1283] User-side processing
[1284] Inputting voice commands
[1285] Users use their smartphones to input their daily schedules and questions by voice.
[1286] Specific action: The user gives a voice command saying, "I will go to the doctor at 10 AM tomorrow."
[1287] Notification confirmation
[1288] Users check reminders and notifications displayed on their smartphones and take action based on them.
[1289] Specific actions: Check notifications such as "It's 12 o'clock now. It's lunchtime," and prepare the meal.
[1290] Conversational communication
[1291] Users can check and adjust their schedules by interacting with the app.
[1292] Specific action: When the user asks "What were your plans for today?", the app responds "I was planning to go shopping at 2pm today."
[1293] Examples of prompts for generative AI models
[1294] Example of a prompt
[1295] User voice input: "What's the weather like today?"
[1296] Generated AI model prompt: "The user wants to know today's weather. Please provide weather information such as sunny, cloudy, or rainy."
[1297] This system, with its interconnected components, provides support for elderly individuals to live independently through daily schedule management and dialogue. Its simple voice input operation and the provision of appropriate information by AI create an efficient and reassuring living environment for users.
[1298] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1299] Step 1:
[1300] Acquisition of user voice input
[1301] The user launches the app on their smartphone and enters voice commands. For example, they might voice-inform an appointment such as, "I'm going to the doctor tomorrow at 10 AM."
[1302] The terminal receives this audio and performs initial processing. The terminal receives the input as an audio signal and prepares it for conversion to text.
[1303] Input: Voice command (Example: "I will go to the doctor at 10 AM tomorrow")
[1304] Output: Audio data
[1305] Step 2:
[1306] Sending audio data to the server
[1307] The terminal uses a communication protocol to send the acquired voice data to the server. This transfers the user's voice information to the server.
[1308] Input: Audio data
[1309] Output: Sending audio data to the server
[1310] Specific operation: The device uses an internet connection to send voice data to the server. For example, it uses Wi-Fi or mobile data communication.
[1311] Step 3:
[1312] Voice input to text conversion
[1313] The server analyzes the received audio data and converts it into text. Speech recognition technology is used to convert the audio signal into linguistic information. Speech recognition APIs and software are used for this process.
[1314] Input: Audio data
[1315] Output: Text data (Example: "I'm going to the doctor at 10 AM tomorrow")
[1316] Specific operation: The server calls a speech recognition API to convert the audio data into text data. This API is often a cloud-based service.
[1317] Step 4:
[1318] Text data analysis and database registration
[1319] The server analyzes text data and extracts user schedule information. This information is then registered in a database. Natural language processing technology is used for the analysis, and the schedule information is added to the database as a new record.
[1320] Input: Text data
[1321] Output: Schedule information registered in the database
[1322] Specific operation: The server uses a natural language processing algorithm to extract date and time information from the text. Then, it stores the extracted data in a relational database.
[1323] Step 5:
[1324] Voice analysis of generative AI models
[1325] The server uses a generative AI model to analyze text data and match it with relevant tasks and information. The generative AI model understands the context from voice input and determines the appropriate action.
[1326] Input: Text data
[1327] Output: Matched task information or answers
[1328] Specific operation: The server runs the generated AI model, and the text "I will go to the doctor at 10 AM tomorrow" is matched as a scheduled task.
[1329] Step 6:
[1330] Generating and sending scheduled notifications
[1331] The server generates reminder notifications based on registered schedule information and sends these notifications to the user's device. Appropriate reminders are generated based on the specified time or conditions.
[1332] Input: Schedule information in the database
[1333] Output: Reminder notification to user's device
[1334] Specific operation: The server creates a reminder notification based on the schedule information and sends it to the user's smartphone using a push notification service. For example, a notification might be set saying, "I have a doctor's appointment tomorrow at 10 AM."
[1335] Step 7:
[1336] Display reminders
[1337] The device displays reminder notifications received from the server. The user checks the notification and takes action.
[1338] Input: Reminder notification from server
[1339] Output: Notification displayed on the smartphone screen
[1340] Specific action: The smartphone receives a push notification and displays a message on the screen saying, "Go to the doctor tomorrow at 10 AM." The user can tap it to view details or set an alarm.
[1341] Step 8:
[1342] Providing a conversational interface
[1343] The device plays back the responses from the generated AI as audio, allowing the user to communicate with the app in a conversational format.
[1344] Input: Generated AI response data from the server
[1345] Output: Generated AI response played back as audio.
[1346] Specific operation: When a user asks the device "What's the weather like today?", the device sends this to the server, and the AI-generated response "It's sunny today" is played aloud. This allows the user to have a natural conversation with the system.
[1347] As described above, through each processing step, the present invention provides a system that supports elderly people in living independently and safely, and realizes a convenient and secure living environment through voice input.
[1348] (Application Example 1)
[1349] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1350] To improve work efficiency within factories, there is a need for means to provide workers with timely and appropriate information and task reminders. Furthermore, an interactive user interface that is easy for elderly or inexperienced workers to use is required.
[1351] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1352] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for analyzing voice input from factory workers and matching it to specific tasks, and means for a robot to send push notifications to workers based on the work schedule. This makes it possible for workers to perform their work efficiently and obtain necessary information in a timely manner.
[1353] "User voice input" refers to voice commands that a user gives to the system through a voice input device such as a microphone.
[1354] "Means of converting to text" refers to software or hardware that has the function of converting voice-input data into text information.
[1355] "User schedule information" refers to data that records the user's planned activities and task schedules.
[1356] "Means of registering in the database" refers to the function of storing the analyzed schedule information in the database.
[1357] A "reminder notification" is a notification sent to a user based on a pre-set schedule.
[1358] "User's device" refers to electronic devices used by the user, such as smartphones, tablets, and personal computers.
[1359] A "factory worker" is an employee who performs physical or mechanical tasks within a factory.
[1360] "Means for analyzing voice input" refers to software or hardware that has the function of appropriately understanding and classifying received voice data.
[1361] "Means for matching to specific tasks" refers to a function that links information obtained through voice analysis to related tasks and work.
[1362] "Means of sending push notifications" refers to a function that sends information to a device in real time and displays a notification to the user.
[1363] Program generation
[1364] To implement this application, a program is needed that accepts user voice input and converts it into text. This program uses a speech recognition API (e.g., Google Cloud Speech-to-Text). After the voice data is converted to text, that text data is sent to a generating AI model (e.g., OpenAI GPT-4) to be matched with specific tasks or information.
[1365] Hardware and software usage
[1366] Hardware:
[1367] Microphone: Used to capture the user's voice input.
[1368] Speaker: Used to play the generated audio response.
[1369] Display (built into the robot): Used to display text information and reminders.
[1370] software:
[1371] Speech recognition APIs (e.g., Google Cloud Speech-to-Text): Used to convert speech data into text.
[1372] Generative AI models (e.g., OpenAI GPT-4): Used to analyze text data and match it to specific tasks or information.
[1373] Database management systems (e.g., MySQL): Used to store and manage schedule information.
[1374] Program processing
[1375] 1. Voice input:
[1376] The server captures the user's voice using the microphone and temporarily stores it locally.
[1377] For example, if a factory worker says, "Tell me the inventory of parts," the voice data is sent to the server and temporarily stored.
[1378] 2. Converting and transmitting audio data:
[1379] The speech recognition API is used to convert the audio data into text, which is then sent to a server where the generation AI model is running.
[1380] Example: The audio "Tell me the stock of parts" is converted to the text "Tell me the stock of parts".
[1381] 3. Processing of Generative AI Models:
[1382] The generating AI analyzes the received text data and extracts related tasks and information.
[1383] Example: Based on the instruction "Tell me the inventory status of the parts," the system retrieves the necessary information from the inventory database.
[1384] 4. Provision of information and notification:
[1385] It generates necessary information and reminders and sends them from the server to the robot.
[1386] For example, information such as "We have very little stock left of part A" is generated and communicated to the worker as a notification via the robot's display or speaker.
[1387] 5. Output:
[1388] The robot provides information to workers through voice and displays.
[1389] For example, if you ask "What's the next step?", you might get the answer "The next step is to install part D."
[1390] Examples of specific cases and prompt statements
[1391] As a concrete example, if a factory worker asks a robot, "When is my break time?", the robot will respond, "Today's break times are 10 AM and 3 PM." An example of the prompt in this case would be as follows:
[1392] "Please generate today's break times based on the voice instructions from the factory workers."
[1393] This allows workers to perform their tasks efficiently and obtain necessary information in a timely manner.
[1394] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1395] Step 1:
[1396] Accepts user voice input.
[1397] Input: The user speaks into the microphone and says, "Tell me the parts inventory."
[1398] Processing details: The device recognizes this voice input and temporarily saves it locally.
[1399] Output: An audio data file is generated and passed to the next processing step.
[1400] Step 2:
[1401] Convert audio data to text
[1402] Input: The audio data file generated in Step 1.
[1403] Processing details: Convert audio data into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text).
[1404] Output: The text data "Tell me the stock of parts" is generated.
[1405] Step 3:
[1406] Send text data to an AI model for generation.
[1407] Input: Text data generated in Step 2.
[1408] Processing details: Text data is sent to a server where an AI model (e.g., OpenAI GPT-4) is running.
[1409] Output: Text data is received by the generating AI model.
[1410] Step 4:
[1411] The generative AI model analyzes the text data.
[1412] Input: Text data received in Step 3.
[1413] Processing details: The generating AI model analyzes text data to identify tasks and information. In this case, it extracts the necessary information from the inventory database.
[1414] Output: Information such as "We have very little stock left of part A."
[1415] Step 5:
[1416] Create information and generate notifications
[1417] Input: Information obtained in Step 4.
[1418] Processing details: The server generates a notification message based on this information and sends it to the terminal.
[1419] Output: A notification message is generated stating, "We have very little stock left of part A."
[1420] Step 6:
[1421] Display and play audible notification messages.
[1422] Input: The notification message generated in Step 5.
[1423] Processing details: The device displays this notification message on its screen and plays it aloud through its speaker.
[1424] Output: The user is notified visually and audibly that "Part A is running low on stock."
[1425] Through the above processing steps, workers can efficiently receive information and act accordingly.
[1426] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1427] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it enables more personalized support. The system's program processing is described in detail below in natural language.
[1428] Server-side processing
[1429] User authentication and data storage
[1430] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[1431] Model operation of generative AI
[1432] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[1433] Operation of the Emotion Engine
[1434] The server operates an emotion engine that analyzes the user's emotional state from voice data. This allows it to adjust schedule information and reminder notifications based on the user's emotional state. For example, if it detects that the user is feeling down, it can adjust the timing and content of reminder notifications.
[1435] Schedule notification
[1436] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[1437] Terminal-side processing
[1438] Voice input function
[1439] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[1440] Display reminders
[1441] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[1442] Conversational interface
[1443] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[1444] User-side actions
[1445] Inputting voice commands
[1446] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[1447] Notification confirmation
[1448] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[1449] Conversational communication
[1450] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[1451] Recognition and response to emotions
[1452] When a user asks a question or gives a command by voice, the device sends the audio to the server, where the emotion engine analyzes the user's emotional state. For example, if the user says, "I'm not feeling very well," the server uses the emotion engine to analyze the user's emotional state and adjusts the content and timing of reminder notifications or generates comforting messages based on the results.
[1453] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. By combining it with an emotion engine, personalized support based on the user's emotional state becomes possible, further enhancing their sense of security and satisfaction.
[1454] The following describes the processing flow.
[1455] User authentication and data storage
[1456] Step 1:
[1457] The user launches the app on their smartphone.
[1458] The user taps the smartphone icon to open the app.
[1459] Step 2:
[1460] The device displays the login screen to the user.
[1461] The device displays a screen prompting the user to enter their user ID and password.
[1462] Step 3:
[1463] The user enters their user ID and password and taps the "Login" button.
[1464] The user enters the correct authentication information.
[1465] Step 4:
[1466] The device sends authentication information to the server.
[1467] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[1468] Step 5:
[1469] The server checks the authentication information received and searches the database for matching user information.
[1470] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[1471] Step 6:
[1472] The server sends the authentication result to the terminal.
[1473] The server returns the authentication success / failure result as a response.
[1474] Step 7:
[1475] The device displays a screen corresponding to the authentication result.
[1476] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[1477] Voice input and schedule registration
[1478] Step 1:
[1479] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[1480] The user taps the microphone button and enters their schedule by voice.
[1481] Step 2:
[1482] The device sends voice data to the server.
[1483] The device acquires voice data and sends it to the server.
[1484] Step 3:
[1485] The server receives the audio data and converts it into text using a generative AI model.
[1486] The server converts the audio data into text and then analyzes its content.
[1487] Step 4:
[1488] The server analyzes the text and extracts information related to the schedule.
[1489] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[1490] Step 5:
[1491] The server sends the registration results to the terminal.
[1492] The server notifies the terminal that the appointment registration was successful.
[1493] Step 6:
[1494] The device notifies the user of the registration result.
[1495] The device displays the message "The appointment has been successfully registered" to the user.
[1496] Reminder notifications and confirmations
[1497] Step 1:
[1498] The server sets the timing of reminders based on a schedule.
[1499] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[1500] Step 2:
[1501] When the server reaches the reminder time, it sends a reminder notification to the device.
[1502] The server will issue a reminder notification at the time set by the server.
[1503] Step 3:
[1504] The device receives a reminder notification.
[1505] The device receives the push notification and displays it on the screen.
[1506] Step 4:
[1507] The user checks the notification and takes action if necessary.
[1508] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[1509] Conversational interface
[1510] Step 1:
[1511] Users ask the app questions such as, "What's the weather like today?"
[1512] The user presses the microphone button to input their question by voice.
[1513] Step 2:
[1514] The device sends voice data to the server.
[1515] The device acquires voice data and sends it to the server.
[1516] Step 3:
[1517] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[1518] The server analyzes the audio data, retrieves weather information, and generates a response.
[1519] Step 4:
[1520] The server sends the generated response to the terminal.
[1521] The server sends a response to the terminal saying, "It's sunny today."
[1522] Step 5:
[1523] The device plays back the response it received as audio.
[1524] The device plays a voice message to the user saying, "It's sunny today."
[1525] Operation of the Emotion Engine
[1526] Step 1:
[1527] When a user voices a question or gives instructions, that voice data is sent to the server.
[1528] For example, the user might say, "What should I do today?"
[1529] Step 2:
[1530] The device sends voice data to the server.
[1531] The device acquires voice data and sends it to the emotion engine.
[1532] Step 3:
[1533] The server analyzes the voice data, and the emotion engine recognizes the user's emotions.
[1534] The server identifies whether it is "stressed" or "depressed."
[1535] Step 4:
[1536] The server sends appropriate responses or reminders to the user based on the emotion recognition results.
[1537] For example, when a user is feeling down, you can send them an encouraging message.
[1538] Step 5:
[1539] The device displays or plays emotion-based responses and reminders to the user via audio.
[1540] A message such as, "Why don't you take a short break to make yourself feel better?" is displayed or played audibly.
[1541] The above is a detailed breakdown of the system's processing steps, including the emotion engine.
[1542] (Example 2)
[1543] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1544] For elderly people, managing their daily schedules and lacking communication are problems that hinder their independent and safe living. Furthermore, conventional reminder and schedule management systems only send uniform notifications without considering the user's emotional state, resulting in a lack of psychological support. To address these issues, there is a need for personalized support that appropriately analyzes the user's emotional state and adapts to their individual circumstances.
[1545] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1546] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for inputting voice data into a generation AI model and utilizing the processing results, means for inputting voice data into an emotion engine and analyzing the emotional state, and means for adjusting the content or timing of reminder notifications based on the analysis results. This makes it possible to provide personalized schedule management and psychological support tailored to the emotional state of elderly people.
[1547] "User" refers to an individual who uses this system.
[1548] "Means for receiving voice input" refers to a device or method for collecting voice uttered by a user and inputting it into a system.
[1549] "Means of converting voice input to text" refers to a device or algorithm that analyzes collected voice data and converts it into textual information.
[1550] "Means for analyzing user schedule information and registering it in a database" refers to a device or method that analyzes a user's schedule based on text data and stores that information in a database.
[1551] "Means for generating and sending reminder notifications to the user's device" refers to a device or method that creates a notification based on registered schedule information and sends it to the user's device.
[1552] A "generative AI model" refers to a model that uses artificial intelligence technology to analyze speech and text data and generate appropriate information.
[1553] An "emotion engine" refers to an algorithm or device used to analyze a user's emotional state from voice data.
[1554] "Means for adjusting the content or timing of reminder notifications" refers to devices or methods for changing the text or sending time of notifications based on the results of the emotion engine's analysis.
[1555] A "server" refers to a computer system used for processing and analyzing data across an entire system.
[1556] "Terminal" refers to devices such as smartphones and tablets that are directly operated by the user.
[1557] This invention is an interactive life support system that helps elderly people live independently and safely. It accepts voice input from the user and provides personalized reminder notifications and schedule management using a generative AI model and emotion engine. A specific embodiment of the system is described below.
[1558] Server-side embodiment
[1559] User authentication and data storage
[1560] When a user logs into the app using a smartphone or tablet, the server accepts the user ID and password, verifies them against the database, and performs authentication. If authentication is successful, reusable user data and schedule information are saved to the database. For example, if a user logs in with the ID "abcd1234" and password "password," the server verifies this, and if authentication is successful, saves the user's new appointments and changes to the database.
[1561] Operation of Generative AI Models
[1562] The server receives voice input from the user and analyzes it using a generative AI model. The voice input is converted into text and matched to specific tasks or information. For example, if a user voice-inputs "I'm going for a walk at 3pm," the server passes this to the generative AI model, which converts it into text and registers "walk" as an event in the database.
[1563] Operation of the Emotion Engine
[1564] The server inputs voice data into an emotion engine to analyze the user's emotional state. This allows the server to adjust the content and timing of reminder notifications according to the user's emotions. For example, if a user says, "I'm not feeling well today," the server analyzes this and sets a reminder such as, "It's time to rest."
[1565] Terminal-side embodiment
[1566] Voice input function
[1567] The terminal receives voice input from the user and sends it to the server. As an initial process, it generates speech recognition data and sends it to the server for text conversion. For example, if the user says, "I will take my medicine at 2 pm," the terminal records this voice and sends it to the server.
[1568] Display reminders
[1569] The device receives reminder notifications sent from the server and displays them on the screen. Voice responses are also supported, allowing the user to check the notifications on the device and take the necessary actions. For example, if the device receives a reminder to "take your medicine at 2 PM," it will display this on the screen and also provide an audio notification.
[1570] Conversational interface
[1571] The device plays back responses from the generative AI model as audio, providing an environment where users can interact with the app naturally. For example, if a user asks, "What's the weather like today?", the device sends that audio to the server, and the generative AI model generates a response such as "It's sunny today." The device then plays this response back to the user as audio.
[1572] User-side embodiment
[1573] Inputting voice commands
[1574] Users use their smartphones to set daily schedules and ask questions using voice input. For example, they might speak instructions to their smartphone such as, "I'll go to the doctor at 10 AM tomorrow."
[1575] Notification confirmation
[1576] Users check reminder notifications displayed on their smartphones and act accordingly. For example, they might receive a notification saying, "It's 12 o'clock now. It's time for lunch."
[1577] Conversational communication
[1578] Users interact with the app via voice to configure and confirm settings. This helps alleviate feelings of loneliness for those living alone. For example, if asked, "What were your plans for today?", the device might respond, "I was planning to go shopping at 2 PM today."
[1579] Recognition and response to emotions
[1580] When a user gives a voice command, the device sends the voice to a server, where the emotion engine analyzes the user's emotional state. Based on the results, the system adjusts the content and timing of reminder notifications and generates comforting messages. For example, if the user says, "I'm not feeling very well," the server will generate a notification such as, "Please take some time to rest."
[1581] Example of a prompt
[1582] Please enter the following information by voice: "I will take my medicine at 8 AM tomorrow."
[1583] This invention enables the provision of personalized services that support the independent living of the elderly and take their emotions into consideration. This improves the elderly's sense of security and satisfaction, and enhances their quality of daily life.
[1584] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1585] Step 1:
[1586] The user performs voice input.
[1587] The user launches a smartphone application and uses voice input. For example, they might say, "I'm going to the doctor at 10 AM tomorrow." This input data is captured by the device in voice format.
[1588] Step 2:
[1589] The device records audio and sends it to the server.
[1590] The terminal receives the user's voice input and records it as audio data. This recorded data is sent to the server. Specifically, audio data is sent from the terminal to the server, and the server receives it. The input data is the user's voice data, and the output data is the audio data sent to the server.
[1591] Step 3:
[1592] The server generates audio data and analyzes it using an AI model.
[1593] The server inputs the received audio data into a generative AI model, which analyzes the audio and converts it into text data. The generative AI model converts the audio data into text and analyzes that text to identify the user's intent. For example, the output text might say, "I'm going to the doctor at 10 AM tomorrow." The input data is audio data, and the output data is the converted text data.
[1594] Step 4:
[1595] The server registers text data in the database.
[1596] The server analyzes user schedule information based on text data and registers it in the database. For example, information such as "I will go to the doctor at 10 AM tomorrow" is saved as a schedule in the database. The input data is text data, and the output data is the schedule information stored in the database.
[1597] Step 5:
[1598] The server analyzes the audio data using an emotion engine.
[1599] The server inputs audio data into the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes emotions from the tone and manner of speaking and outputs the results. For example, if the engine analyzes that the user is depressed, that information is output as emotion recognition data. The input data is audio data, and the output data is emotion recognition data.
[1600] Step 6:
[1601] The server generates and adjusts reminder notifications.
[1602] The server adjusts the content and timing of reminder notifications based on emotion recognition data. For example, if the server detects that the user is feeling down, the notification content changes to "Don't push yourself today, relax." The input data is emotion recognition data, and the output data is the adjusted reminder notification.
[1603] Step 7:
[1604] The server sends a reminder notification to the device.
[1605] The server sends a pre-arranged reminder notification to the device. For example, a reminder to "go to the doctor at 10 AM tomorrow" is sent to the device just before the scheduled time. The input data is the pre-arranged reminder notification, and the output data is the notification data sent to the device.
[1606] Step 8:
[1607] The device displays a reminder notification.
[1608] The device displays reminder notifications received from the server. For example, a notification saying "You have a doctor's appointment" is displayed on the screen and also announced audibly. The input data is the notification data received by the device, and the output data is the notification displayed to the user.
[1609] Through these steps, the system efficiently processes user voice input and utilizes a generative AI model and emotion engine to provide personalized support for daily life.
[1610] (Application Example 2)
[1611] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1612] This invention aims to solve problems related to support systems that enable elderly people to live independently and safely. While conventional technologies exist that convert voice input to text and provide mechanisms for schedule management and reminder notifications, they do not address the emotional state of the user or provide support for actual shopping. Elderly people often get lost or feel anxious when shopping in physical stores, and there is a need for a system that can alleviate these issues and allow them to shop in a relaxed state.
[1613] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1614] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for providing shopping assistance to the user based on the voice input, means including an emotion engine for analyzing the user's emotional state, and means for generating support messages based on the emotional state and sending them to the user's terminal. This makes it possible for elderly people to shop in physical stores without getting lost or feeling anxious, and to shop in a relaxed manner while receiving appropriate guidance and support.
[1615] A "voice input method" is a system for capturing a user's voice into a device such as a terminal and processing it.
[1616] A "text conversion method" is a mechanism for analyzing captured audio data and converting it into text-based data.
[1617] A "schedule information analysis tool" is a system that extracts and analyzes a user's schedule and tasks based on text data, and manages them accordingly.
[1618] A "database registration method" is a mechanism for registering analyzed schedule information into a database that centrally manages that information.
[1619] A "reminder notification generation method" is a mechanism for generating announcements and notifications for users at specific times and under specific conditions, based on registered schedule information.
[1620] A "reminder notification sending method" is a mechanism that sends generated notifications to the user's device so that the user can check them.
[1621] A "shopping assistant system" is a mechanism that supports users' shopping based on voice input, guiding them to necessary product information and locations.
[1622] An "emotion engine" is a technology that analyzes a user's voice data and evaluates and recognizes their emotional state.
[1623] A "support message generation mechanism" is a system for generating appropriate messages for users based on their analyzed emotional state.
[1624] A "support message sending mechanism" is a system that sends generated support messages to the user's device, allowing the user to receive them.
[1625] This invention relates to a system that supports elderly people in living independently and safely, and describes its specific embodiments. This system is primarily designed to support elderly people in their daily activities, such as shopping, through voice input. The main components of the system and their specific operation are described in detail below.
[1626] Server-side processing
[1627] User authentication and data storage
[1628] The server authenticates users when they log into the application. Upon successful authentication, it saves data such as the user's schedule and shopping list to a database. This information is managed and stored on the server whenever the user updates their schedule or tasks.
[1629] Operation of Generative AI Models
[1630] The server uses a generative AI model to analyze voice input from the user. When the user says "I'm looking for milk" by voice, the generative AI model converts that instruction into text and generates appropriate product information and directions to the store.
[1631] Operation of the Emotion Engine
[1632] The server uses an emotion engine to analyze the user's voice data to determine their emotional state. For example, if a user says, "I'm a little tired," the emotion engine recognizes this emotional state and generates an appropriate support message. This support message might include suggestions such as, "Let's take a break," or "I'll call a staff member."
[1633] Schedule notifications and support message sending
[1634] The server generates reminders and support messages based on the user's schedule information and emotional state, and sends them to the user's device. For example, if "milk" is included in the shopping list, when the user enters the nearest store, a message will be sent saying, "Milk is in the refrigerated section."
[1635] Terminal-side processing
[1636] Voice input function
[1637] User voice input is captured through the device's microphone. The captured voice is converted into text data by speech recognition software and sent to the server.
[1638] Display reminders
[1639] The device displays notifications and reminders sent from the server. For example, a reminder such as "Take your medicine at 3 PM" will be displayed on the user's smartphone or smart glasses.
[1640] Shopping assistant function
[1641] The device receives responses from the generative AI model and provides voice guidance to the user. For example, if the user asks, "Where is the milk section?", it will notify the user, "It's in the refrigerated section." It also provides voice support messages based on the analysis results of the emotion engine.
[1642] Hardware and software to use
[1643] This system uses the following hardware and software.
[1644] Hardware: Smartphone, smart glasses, microphone, speaker
[1645] Software: speech_recognition library (speech recognition), proprietary generative AI model, emotion engine
[1646] Specific example
[1647] In a real-world scenario, if a user goes to a shopping mall and uses voice input to say, "I'm looking for milk," the emotion engine detects that the user is feeling a little anxious and provides a message such as, "The milk is in the refrigerated section. Let's call a staff member nearby."
[1648] Examples of prompts for a generative AI model:
[1649] "When a user says 'I'm looking for milk,' analyze their response and the emotions they express to generate an appropriate response. If the user is anxious, include a reassuring message."
[1650] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1651] Step 1:
[1652] Acquiring voice input (user)
[1653] Users input voice commands through the microphone on their smartphone or smart glasses. This input is part of the user's shopping list and might say something like, "I'm looking for milk."
[1654] Input: Audio data
[1655] Output: Audio data (sent to the device)
[1656] Step 2:
[1657] Text conversion of audio data (on the device)
[1658] The device converts the acquired audio data into text data using a speech recognition library (for example, speech_recognition). The converted text becomes "I'm looking for milk."
[1659] Input: Audio data
[1660] Data processing: Convert speech to text using a speech recognition library.
[1661] Output: Text data
[1662] Step 3:
[1663] Sending text data (from terminal to server)
[1664] The terminal sends the converted text data to the server. The server receives this text data and begins analysis.
[1665] Input: Text data
[1666] Output: Text data sent to the server
[1667] Step 4:
[1668] User authentication and data storage (server)
[1669] The server performs user authentication before parsing the transmitted text data. If authentication is successful, the user's schedule and shopping list are saved to the database.
[1670] Input: User ID, password, submitted text data
[1671] Data processing: User authentication, saving to database.
[1672] Output: Authentication success message, saved user data
[1673] Step 5:
[1674] Analysis using a generative AI model (server)
[1675] The server uses a generative AI model to analyze text data and generate appropriate responses. For example, in response to the text "I'm looking for milk," the AI model generates the answer "It's in the refrigerated section."
[1676] Input: Text data
[1677] Data Processing: Text analysis and response generation using generative AI models.
[1678] Output: Generated response text
[1679] Step 6:
[1680] Emotional state analysis (server)
[1681] Along with the generated response text, the emotion engine analyzes the user's emotional state based on their voice data. For example, it can recognize that the user is feeling anxious and generate an appropriate support message.
[1682] Input: Audio data, generated response text
[1683] Data processing: Emotional analysis using an emotion engine
[1684] Output: Emotional state, support message
[1685] Step 7:
[1686] Sending responses and support messages (from server to terminal)
[1687] The server sends the generated response text and support message to the user's terminal. For example, a notification might be sent saying, "It's in the refrigerated section. I'll call a staff member nearby."
[1688] Input: Generated response text, support message
[1689] Output: Response text and support message sent to the terminal
[1690] Step 8:
[1691] Playback of response (terminal)
[1692] The terminal plays back the received response text and support message as audio. The user is notified, "It's in the refrigerated section. I'll call a staff member nearby," and voice assistance is provided.
[1693] Input: Response text, support message
[1694] Data processing: Converting text to speech using speech synthesis.
[1695] Output: Voice guidance and support messages
[1696] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1697] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1698] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1699] [Fourth Embodiment]
[1700] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1701] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1702] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1703] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1704] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1705] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1706] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1707] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1708] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1709] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1710] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1711] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1712] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1713] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. The system's program processing is described in detail below in natural language.
[1714] Server-side processing
[1715] User authentication and data storage
[1716] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[1717] Model operation of generative AI
[1718] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[1719] Schedule notification
[1720] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[1721] Terminal-side processing
[1722] Voice input function
[1723] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[1724] Display reminders
[1725] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[1726] Conversational interface
[1727] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[1728] User-side actions
[1729] Inputting voice commands
[1730] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[1731] Notification confirmation
[1732] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[1733] Conversational communication
[1734] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[1735] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. The simple operation via voice input and the provision of appropriate information by the generated AI create an efficient and secure living environment for the user.
[1736] The following describes the processing flow.
[1737] User authentication and data storage
[1738] Step 1:
[1739] The user launches the app on their smartphone.
[1740] The user taps the smartphone icon to open the app.
[1741] Step 2:
[1742] The device displays the login screen to the user.
[1743] The device displays a screen prompting the user to enter their user ID and password.
[1744] Step 3:
[1745] The user enters their user ID and password and taps the "Login" button.
[1746] The user enters the correct authentication information.
[1747] Step 4:
[1748] The device sends authentication information to the server.
[1749] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[1750] Step 5:
[1751] The server checks the authentication information received and searches the database for matching user information.
[1752] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[1753] Step 6:
[1754] The server sends the authentication result to the terminal.
[1755] The server returns the authentication success / failure result as a response.
[1756] Step 7:
[1757] The device displays a screen corresponding to the authentication result.
[1758] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[1759] Voice input and schedule registration
[1760] Step 1:
[1761] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[1762] The user taps the microphone button and enters their schedule by voice.
[1763] Step 2:
[1764] The device sends voice data to the server.
[1765] The device acquires voice data and sends it to the server.
[1766] Step 3:
[1767] The server receives the audio data and converts it into text using a generative AI model.
[1768] The server converts the audio data into text and then analyzes its content.
[1769] Step 4:
[1770] The server analyzes the text and extracts information related to the schedule.
[1771] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[1772] Step 5:
[1773] The server sends the registration results to the terminal.
[1774] The server notifies the terminal that the appointment registration was successful.
[1775] Step 6:
[1776] The device notifies the user of the registration result.
[1777] The device displays the message "The appointment has been successfully registered" to the user.
[1778] Reminder notifications and confirmations
[1779] Step 1:
[1780] The server sets the timing of reminders based on a schedule.
[1781] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[1782] Step 2:
[1783] When the server reaches the reminder time, it sends a reminder notification to the device.
[1784] The server will issue a reminder notification at the time set by the server.
[1785] Step 3:
[1786] The device receives a reminder notification.
[1787] The device receives the push notification and displays it on the screen.
[1788] Step 4:
[1789] The user checks the notification and takes action if necessary.
[1790] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[1791] Conversational interface
[1792] Step 1:
[1793] Users ask the app questions such as, "What's the weather like today?"
[1794] The user presses the microphone button to input their question by voice.
[1795] Step 2:
[1796] The device sends voice data to the server.
[1797] The device acquires voice data and sends it to the server.
[1798] Step 3:
[1799] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[1800] The server analyzes the audio data, retrieves weather information, and generates a response.
[1801] Step 4:
[1802] The server sends the generated response to the terminal.
[1803] The server sends a response to the terminal saying, "It's sunny today."
[1804] Step 5:
[1805] The device plays back the response it received as audio.
[1806] The device plays a voice message to the user saying, "It's sunny today."
[1807] The above is a detailed explanation of the program's processing flow, broken down into individual steps.
[1808] (Example 1)
[1809] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1810] In systems designed to support elderly individuals in living independent and safe daily lives, it is essential to efficiently analyze voice input, appropriately manage user schedules and provide reminder notifications, and enhance user convenience and peace of mind by offering interactive support. Furthermore, utilizing generative AI models to extract appropriate tasks from voice input and enabling natural dialogue with the user is also crucial.
[1811] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1812] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for operating a generative AI model and analyzing voice data to match it with specific tasks and information, and means for generating dialogue with the user based on the analysis results of the generative AI model. This makes it possible for elderly people to easily operate the system through voice input and to receive appropriate reminder notifications and interactive feedback.
[1813] "Means for receiving user voice input" refers to devices or systems for receiving information entered by a user via voice.
[1814] "Means for converting voice input to text" refers to methods and technologies for analyzing received voice data and converting it into text format.
[1815] "Means for analyzing user schedule information and registering it in a database" refers to a system or process that interprets information obtained as text and saves it in a database as the user's schedule or appointments.
[1816] "A means of generating and sending reminder notifications to the user's device" refers to a method for creating notifications based on registered schedule information and sending them to the user's device.
[1817] "A means of operating a generative AI model to analyze voice data and match it to specific tasks or information" refers to a technology that uses generative AI to process voice data and identify appropriate tasks or information.
[1818] "A means of generating user dialogue based on the analysis results of a generative AI model" refers to a system that generates and provides information for natural dialogue with the user based on the results of analysis by the generative AI.
[1819] This invention is an interactive lifestyle support system that helps elderly people live independently and safely, and provides its functions using a device such as a smartphone. This system accepts voice input from the user, analyzes it, and provides appropriate reminders and notifications, as well as conversational support utilizing a generative AI model.
[1820] Server-side processing
[1821] User Authentication
[1822] When a user logs into the app, the server authenticates them by comparing the entered user ID and password with the database. It also saves the data of successfully authenticated users to the database.
[1823] Specific operation: The user enters their user ID and password on their smartphone, and the server verifies this against the database to perform authentication.
[1824] Data storage
[1825] The server stores the schedule information and settings of authenticated users in a database. For example, if there are new appointments or changes to existing ones, this information is added to the database.
[1826] Specific operation: When a user voice-inputs "I want to change my dentist appointment to 3 PM tomorrow," the server saves this information to the database.
[1827] Operation of Generative AI Models
[1828] The server operates a generative AI model and analyzes the user's voice data. It extracts specific tasks and information from the voice input and registers them in a database.
[1829] Specific operation: When a user voice-inputs "I will take my medicine at 8 AM tomorrow," the generating AI model recognizes this as an appointment and saves it to the database.
[1830] Sending schedule notifications
[1831] The server generates reminder notifications based on the registered schedule information and sends them to the user's smartphone.
[1832] Specific action: A reminder notification to "take your medicine" is sent to the user's smartphone at 8:00 AM.
[1833] Terminal-side processing
[1834] Acquiring voice input
[1835] The device captures the user's voice and sends it to the server to be converted into text.
[1836] Specific operation: When the user uses voice input to say "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[1837] Display reminders
[1838] The device displays reminder notifications received from the server.
[1839] Specific action: A notification appears on the user's smartphone stating, "I'm going shopping at 2 PM."
[1840] Providing a conversational interface
[1841] The device plays the response from the generating AI as audio, allowing the user to interact with the app.
[1842] Specific operation: When the user voice-inputs "What's the weather like today?", the device plays the response "It's sunny today" from the server's AI generation system.
[1843] User-side processing
[1844] Inputting voice commands
[1845] Users use their smartphones to input their daily schedules and questions by voice.
[1846] Specific action: The user gives a voice command saying, "I will go to the doctor at 10 AM tomorrow."
[1847] Notification confirmation
[1848] Users check reminders and notifications displayed on their smartphones and take action based on them.
[1849] Specific actions: Check notifications such as "It's 12 o'clock now. It's lunchtime," and prepare the meal.
[1850] Conversational communication
[1851] Users can check and adjust their schedules by interacting with the app.
[1852] Specific action: When the user asks "What were your plans for today?", the app responds "I was planning to go shopping at 2pm today."
[1853] Examples of prompts for generative AI models
[1854] Example of a prompt
[1855] User voice input: "What's the weather like today?"
[1856] Generated AI model prompt: "The user wants to know today's weather. Please provide weather information such as sunny, cloudy, or rainy."
[1857] This system, with its interconnected components, provides support for elderly individuals to live independently through daily schedule management and dialogue. Its simple voice input operation and the provision of appropriate information by AI create an efficient and reassuring living environment for users.
[1858] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1859] Step 1:
[1860] Acquisition of user voice input
[1861] The user launches the app on their smartphone and enters voice commands. For example, they might voice-inform an appointment such as, "I'm going to the doctor tomorrow at 10 AM."
[1862] The terminal receives this audio and performs initial processing. The terminal receives the input as an audio signal and prepares it for conversion to text.
[1863] Input: Voice command (Example: "I will go to the doctor at 10 AM tomorrow")
[1864] Output: Audio data
[1865] Step 2:
[1866] Sending audio data to the server
[1867] The terminal uses a communication protocol to send the acquired voice data to the server. This transfers the user's voice information to the server.
[1868] Input: Audio data
[1869] Output: Sending audio data to the server
[1870] Specific operation: The device uses an internet connection to send voice data to the server. For example, it uses Wi-Fi or mobile data communication.
[1871] Step 3:
[1872] Voice input to text conversion
[1873] The server analyzes the received audio data and converts it into text. Speech recognition technology is used to convert the audio signal into linguistic information. Speech recognition APIs and software are used for this process.
[1874] Input: Audio data
[1875] Output: Text data (Example: "I'm going to the doctor at 10 AM tomorrow")
[1876] Specific operation: The server calls a speech recognition API to convert the audio data into text data. This API is often a cloud-based service.
[1877] Step 4:
[1878] Text data analysis and database registration
[1879] The server analyzes text data and extracts user schedule information. This information is then registered in a database. Natural language processing technology is used for the analysis, and the schedule information is added to the database as a new record.
[1880] Input: Text data
[1881] Output: Schedule information registered in the database
[1882] Specific operation: The server uses a natural language processing algorithm to extract date and time information from the text. Then, it stores the extracted data in a relational database.
[1883] Step 5:
[1884] Voice analysis of generative AI models
[1885] The server uses a generative AI model to analyze text data and match it with relevant tasks and information. The generative AI model understands the context from voice input and determines the appropriate action.
[1886] Input: Text data
[1887] Output: Matched task information or answers
[1888] Specific operation: The server runs the generated AI model, and the text "I will go to the doctor at 10 AM tomorrow" is matched as a scheduled task.
[1889] Step 6:
[1890] Generating and sending scheduled notifications
[1891] The server generates reminder notifications based on registered schedule information and sends these notifications to the user's device. Appropriate reminders are generated based on the specified time or conditions.
[1892] Input: Schedule information in the database
[1893] Output: Reminder notification to user's device
[1894] Specific operation: The server creates a reminder notification based on the schedule information and sends it to the user's smartphone using a push notification service. For example, a notification might be set saying, "I have a doctor's appointment tomorrow at 10 AM."
[1895] Step 7:
[1896] Display reminders
[1897] The device displays reminder notifications received from the server. The user checks the notification and takes action.
[1898] Input: Reminder notification from server
[1899] Output: Notification displayed on the smartphone screen
[1900] Specific action: The smartphone receives a push notification and displays a message on the screen saying, "Go to the doctor tomorrow at 10 AM." The user can tap it to view details or set an alarm.
[1901] Step 8:
[1902] Providing a conversational interface
[1903] The device plays back the responses from the generated AI as audio, allowing the user to communicate with the app in a conversational format.
[1904] Input: Generated AI response data from the server
[1905] Output: Generated AI response played back as audio.
[1906] Specific operation: When a user asks the device "What's the weather like today?", the device sends this to the server, and the AI-generated response "It's sunny today" is played aloud. This allows the user to have a natural conversation with the system.
[1907] As described above, through each processing step, the present invention provides a system that supports elderly people in living independently and safely, and realizes a convenient and secure living environment through voice input.
[1908] (Application Example 1)
[1909] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1910] To improve work efficiency within factories, there is a need for means to provide workers with timely and appropriate information and task reminders. Furthermore, an interactive user interface that is easy for elderly or inexperienced workers to use is required.
[1911] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1912] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for analyzing voice input from factory workers and matching it to specific tasks, and means for a robot to send push notifications to workers based on the work schedule. This makes it possible for workers to perform their work efficiently and obtain necessary information in a timely manner.
[1913] "User voice input" refers to voice commands that a user gives to the system through a voice input device such as a microphone.
[1914] "Means of converting to text" refers to software or hardware that has the function of converting voice-input data into text information.
[1915] "User schedule information" refers to data that records the user's planned activities and task schedules.
[1916] "Means of registering in the database" refers to the function of storing the analyzed schedule information in the database.
[1917] A "reminder notification" is a notification sent to a user based on a pre-set schedule.
[1918] "User's device" refers to electronic devices used by the user, such as smartphones, tablets, and personal computers.
[1919] A "factory worker" is an employee who performs physical or mechanical tasks within a factory.
[1920] "Means for analyzing voice input" refers to software or hardware that has the function of appropriately understanding and classifying received voice data.
[1921] "Means for matching to specific tasks" refers to a function that links information obtained through voice analysis to related tasks and work.
[1922] "Means of sending push notifications" refers to a function that sends information to a device in real time and displays a notification to the user.
[1923] Program generation
[1924] To implement this application, a program is needed that accepts user voice input and converts it into text. This program uses a speech recognition API (e.g., Google Cloud Speech-to-Text). After the voice data is converted to text, that text data is sent to a generating AI model (e.g., OpenAI GPT-4) to be matched with specific tasks or information.
[1925] Hardware and software usage
[1926] Hardware:
[1927] Microphone: Used to capture the user's voice input.
[1928] Speaker: Used to play the generated audio response.
[1929] Display (built into the robot): Used to display text information and reminders.
[1930] software:
[1931] Speech recognition APIs (e.g., Google Cloud Speech-to-Text): Used to convert speech data into text.
[1932] Generative AI models (e.g., OpenAI GPT-4): Used to analyze text data and match it to specific tasks or information.
[1933] Database management systems (e.g., MySQL): Used to store and manage schedule information.
[1934] Program processing
[1935] 1. Voice input:
[1936] The server captures the user's voice using the microphone and temporarily stores it locally.
[1937] For example, if a factory worker says, "Tell me the inventory of parts," the voice data is sent to the server and temporarily stored.
[1938] 2. Converting and transmitting audio data:
[1939] The speech recognition API is used to convert the audio data into text, which is then sent to a server where the generation AI model is running.
[1940] Example: The audio "Tell me the stock of parts" is converted to the text "Tell me the stock of parts".
[1941] 3. Processing of Generative AI Models:
[1942] The generating AI analyzes the received text data and extracts related tasks and information.
[1943] Example: Based on the instruction "Tell me the inventory status of the parts," the system retrieves the necessary information from the inventory database.
[1944] 4. Provision of information and notification:
[1945] It generates necessary information and reminders and sends them from the server to the robot.
[1946] For example, information such as "We have very little stock left of part A" is generated and communicated to the worker as a notification via the robot's display or speaker.
[1947] 5. Output:
[1948] The robot provides information to workers through voice and displays.
[1949] For example, if you ask "What's the next step?", you might get the answer "The next step is to install part D."
[1950] Examples of specific cases and prompt statements
[1951] As a concrete example, if a factory worker asks a robot, "When is my break time?", the robot will respond, "Today's break times are 10 AM and 3 PM." An example of the prompt in this case would be as follows:
[1952] "Please generate today's break times based on the voice instructions from the factory workers."
[1953] This allows workers to perform their tasks efficiently and obtain necessary information in a timely manner.
[1954] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1955] Step 1:
[1956] Accepts user voice input.
[1957] Input: The user speaks into the microphone and says, "Tell me the parts inventory."
[1958] Processing details: The device recognizes this voice input and temporarily saves it locally.
[1959] Output: An audio data file is generated and passed to the next processing step.
[1960] Step 2:
[1961] Convert audio data to text
[1962] Input: The audio data file generated in Step 1.
[1963] Processing details: Convert audio data into text information using a speech recognition API (e.g., Google Cloud Speech-to-Text).
[1964] Output: The text data "Tell me the stock of parts" is generated.
[1965] Step 3:
[1966] Send text data to an AI model for generation.
[1967] Input: Text data generated in Step 2.
[1968] Processing details: Text data is sent to a server where an AI model (e.g., OpenAI GPT-4) is running.
[1969] Output: Text data is received by the generating AI model.
[1970] Step 4:
[1971] The generative AI model analyzes the text data.
[1972] Input: Text data received in Step 3.
[1973] Processing details: The generating AI model analyzes text data to identify tasks and information. In this case, it extracts the necessary information from the inventory database.
[1974] Output: Information such as "We have very little stock left of part A."
[1975] Step 5:
[1976] Create information and generate notifications
[1977] Input: Information obtained in Step 4.
[1978] Processing details: The server generates a notification message based on this information and sends it to the terminal.
[1979] Output: A notification message is generated stating, "We have very little stock left of part A."
[1980] Step 6:
[1981] Display and play audible notification messages.
[1982] Input: The notification message generated in Step 5.
[1983] Processing details: The device displays this notification message on its screen and plays it aloud through its speaker.
[1984] Output: The user is notified visually and audibly that "Part A is running low on stock."
[1985] Through the above processing steps, workers can efficiently receive information and act accordingly.
[1986] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1987] This invention is an interactive life support system that assists elderly people in living independently and safely, and provides its functions using a device such as a smartphone. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it enables more personalized support. The system's program processing is described in detail below in natural language.
[1988] Server-side processing
[1989] User authentication and data storage
[1990] The server authenticates users when they log into the app. If authentication is successful, the server saves the user's data. For example, a user opens the app on their smartphone and logs in. The server verifies the user ID and password and authenticates them. Then, if there are any new appointments or changes, that information is saved to the database.
[1991] Model operation of generative AI
[1992] The server operates a generative AI model and analyzes the input voice data. When a user makes a voice input, the server recognizes the input and matches it to a specific task or information. For example, if the voice input is "Take medicine at 8 AM tomorrow," the generative AI model recognizes this as an appointment and registers it in the database.
[1993] Operation of the Emotion Engine
[1994] The server operates an emotion engine that analyzes the user's emotional state from voice data. This allows it to adjust schedule information and reminder notifications based on the user's emotional state. For example, if it detects that the user is feeling down, it can adjust the timing and content of reminder notifications.
[1995] Schedule notification
[1996] The server sends push notifications to the user's smartphone based on the registered schedule. By sending reminders at the set time, it helps users not forget their appointments. For example, it might send a reminder to the user to "take your medicine" at 8 AM.
[1997] Terminal-side processing
[1998] Voice input function
[1999] The device acquires the user's voice and sends it to the server to be converted to text. The device performs initial processing for speech recognition. For example, if the user voice-inputs "I'm going shopping at 2pm," the device records the voice and sends it to the server.
[2000] Display reminders
[2001] The device displays notifications received from the server. The user checks the notifications and takes the necessary actions. For example, the user is shown a notification that says, "Go shopping at 2 PM."
[2002] Conversational interface
[2003] The device plays back the response from the generating AI as audio. This allows the user to interact with the app as if they were having a conversation. For example, if the user asks, "What's the weather like today?", the device sends that audio to the server and plays back the response from the generating AI, "It's sunny today."
[2004] User-side actions
[2005] Inputting voice commands
[2006] Users use their smartphones to input daily schedules and questions by voice. For example, they might give instructions like, "I'll go to the doctor tomorrow at 10 AM."
[2007] Notification confirmation
[2008] Users check reminders and notifications displayed on their smartphones and act accordingly. For example, a notification like "It's 12 o'clock now. It's lunchtime" helps them remember to eat.
[2009] Conversational communication
[2010] Users can configure and confirm settings by interacting with the app. This helps reduce feelings of loneliness, even for those living alone. For example, if you ask, "What were your plans for today?", the app will respond, "I was planning to go shopping at 2 PM today."
[2011] Recognition and response to emotions
[2012] When a user asks a question or gives a command by voice, the device sends the audio to the server, where the emotion engine analyzes the user's emotional state. For example, if the user says, "I'm not feeling very well," the server uses the emotion engine to analyze the user's emotional state and adjusts the content and timing of reminder notifications or generates comforting messages based on the results.
[2013] As described above, each component of the system works in conjunction to provide support for elderly people to live independently through daily schedule management and dialogue. By combining it with an emotion engine, personalized support based on the user's emotional state becomes possible, further enhancing their sense of security and satisfaction.
[2014] The following describes the processing flow.
[2015] User authentication and data storage
[2016] Step 1:
[2017] The user launches the app on their smartphone.
[2018] The user taps the smartphone icon to open the app.
[2019] Step 2:
[2020] The device displays the login screen to the user.
[2021] The device displays a screen prompting the user to enter their user ID and password.
[2022] Step 3:
[2023] The user enters their user ID and password and taps the "Login" button.
[2024] The user enters the correct authentication information.
[2025] Step 4:
[2026] The device sends authentication information to the server.
[2027] The terminal encrypts the entered user ID and password and sends them as an HTTP POST request.
[2028] Step 5:
[2029] The server checks the authentication information received and searches the database for matching user information.
[2030] The server searches the database for records that match the entered information and returns an authentication success flag if one exists.
[2031] Step 6:
[2032] The server sends the authentication result to the terminal.
[2033] The server returns the authentication success / failure result as a response.
[2034] Step 7:
[2035] The device displays a screen corresponding to the authentication result.
[2036] If authentication is successful, the user will be redirected to the main screen; otherwise, an error message will be displayed.
[2037] Voice input and schedule registration
[2038] Step 1:
[2039] The user gives a voice command saying, "I'm going to the doctor at 10 AM tomorrow."
[2040] The user taps the microphone button and enters their schedule by voice.
[2041] Step 2:
[2042] The device sends voice data to the server.
[2043] The device acquires voice data and sends it to the server.
[2044] Step 3:
[2045] The server receives the audio data and converts it into text using a generative AI model.
[2046] The server converts the audio data into text and then analyzes its content.
[2047] Step 4:
[2048] The server analyzes the text and extracts information related to the schedule.
[2049] The server registers the information "I will go to the doctor tomorrow at 10 AM" as an appointment in the database.
[2050] Step 5:
[2051] The server sends the registration results to the terminal.
[2052] The server notifies the terminal that the appointment registration was successful.
[2053] Step 6:
[2054] The device notifies the user of the registration result.
[2055] The device displays the message "The appointment has been successfully registered" to the user.
[2056] Reminder notifications and confirmations
[2057] Step 1:
[2058] The server sets the timing of reminders based on a schedule.
[2059] The server sets a reminder to "go to the doctor at 10 AM tomorrow" and saves it to the database.
[2060] Step 2:
[2061] When the server reaches the reminder time, it sends a reminder notification to the device.
[2062] The server will issue a reminder notification at the time set by the server.
[2063] Step 3:
[2064] The device receives a reminder notification.
[2065] The device receives the push notification and displays it on the screen.
[2066] Step 4:
[2067] The user checks the notification and takes action if necessary.
[2068] The user sees the notification on the screen that says "You have an appointment to see a doctor now" and begins to prepare.
[2069] Conversational interface
[2070] Step 1:
[2071] Users ask the app questions such as, "What's the weather like today?"
[2072] The user presses the microphone button to input their question by voice.
[2073] Step 2:
[2074] The device sends voice data to the server.
[2075] The device acquires voice data and sends it to the server.
[2076] Step 3:
[2077] The server receives the audio data, analyzes it using a generative AI model, and generates the corresponding text.
[2078] The server analyzes the audio data, retrieves weather information, and generates a response.
[2079] Step 4:
[2080] The server sends the generated response to the terminal.
[2081] The server sends a response to the terminal saying, "It's sunny today."
[2082] Step 5:
[2083] The device plays back the response it received as audio.
[2084] The device plays a voice message to the user saying, "It's sunny today."
[2085] Operation of the Emotion Engine
[2086] Step 1:
[2087] When a user voices a question or gives instructions, that voice data is sent to the server.
[2088] For example, the user might say, "What should I do today?"
[2089] Step 2:
[2090] The device sends voice data to the server.
[2091] The device acquires voice data and sends it to the emotion engine.
[2092] Step 3:
[2093] The server analyzes the voice data, and the emotion engine recognizes the user's emotions.
[2094] The server identifies whether it is "stressed" or "depressed."
[2095] Step 4:
[2096] The server sends appropriate responses or reminders to the user based on the emotion recognition results.
[2097] For example, when a user is feeling down, you can send them an encouraging message.
[2098] Step 5:
[2099] The device displays or plays emotion-based responses and reminders to the user via audio.
[2100] A message such as, "Why don't you take a short break to make yourself feel better?" is displayed or played audibly.
[2101] The above is a detailed breakdown of the system's processing steps, including the emotion engine.
[2102] (Example 2)
[2103] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2104] For elderly people, managing their daily schedules and lacking communication are problems that hinder their independent and safe living. Furthermore, conventional reminder and schedule management systems only send uniform notifications without considering the user's emotional state, resulting in a lack of psychological support. To address these issues, there is a need for personalized support that appropriately analyzes the user's emotional state and adapts to their individual circumstances.
[2105] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[2106] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for inputting voice data into a generation AI model and utilizing the processing results, means for inputting voice data into an emotion engine and analyzing the emotional state, and means for adjusting the content or timing of reminder notifications based on the analysis results. This makes it possible to provide personalized schedule management and psychological support tailored to the emotional state of elderly people.
[2107] "User" refers to an individual who uses this system.
[2108] "Means for receiving voice input" refers to a device or method for collecting voice uttered by a user and inputting it into a system.
[2109] "Means of converting voice input to text" refers to a device or algorithm that analyzes collected voice data and converts it into textual information.
[2110] "Means for analyzing user schedule information and registering it in a database" refers to a device or method that analyzes a user's schedule based on text data and stores that information in a database.
[2111] "Means for generating and sending reminder notifications to the user's device" refers to a device or method that creates a notification based on registered schedule information and sends it to the user's device.
[2112] A "generative AI model" refers to a model that uses artificial intelligence technology to analyze speech and text data and generate appropriate information.
[2113] An "emotion engine" refers to an algorithm or device used to analyze a user's emotional state from voice data.
[2114] "Means for adjusting the content or timing of reminder notifications" refers to devices or methods for changing the text or sending time of notifications based on the results of the emotion engine's analysis.
[2115] A "server" refers to a computer system used for processing and analyzing data across an entire system.
[2116] "Terminal" refers to devices such as smartphones and tablets that are directly operated by the user.
[2117] This invention is an interactive life support system that helps elderly people live independently and safely. It accepts voice input from the user and provides personalized reminder notifications and schedule management using a generative AI model and emotion engine. A specific embodiment of the system is described below.
[2118] Server-side embodiment
[2119] User authentication and data storage
[2120] When a user logs into the app using a smartphone or tablet, the server accepts the user ID and password, verifies them against the database, and performs authentication. If authentication is successful, reusable user data and schedule information are saved to the database. For example, if a user logs in with the ID "abcd1234" and password "password," the server verifies this, and if authentication is successful, saves the user's new appointments and changes to the database.
[2121] Operation of Generative AI Models
[2122] The server receives voice input from the user and analyzes it using a generative AI model. The voice input is converted into text and matched to specific tasks or information. For example, if a user voice-inputs "I'm going for a walk at 3pm," the server passes this to the generative AI model, which converts it into text and registers "walk" as an event in the database.
[2123] Operation of the Emotion Engine
[2124] The server inputs voice data into an emotion engine to analyze the user's emotional state. This allows the server to adjust the content and timing of reminder notifications according to the user's emotions. For example, if a user says, "I'm not feeling well today," the server analyzes this and sets a reminder such as, "It's time to rest."
[2125] Terminal-side embodiment
[2126] Voice input function
[2127] The terminal receives voice input from the user and sends it to the server. As an initial process, it generates speech recognition data and sends it to the server for text conversion. For example, if the user says, "I will take my medicine at 2 pm," the terminal records this voice and sends it to the server.
[2128] Display reminders
[2129] The device receives reminder notifications sent from the server and displays them on the screen. Voice responses are also supported, allowing the user to check the notifications on the device and take the necessary actions. For example, if the device receives a reminder to "take your medicine at 2 PM," it will display this on the screen and also provide an audio notification.
[2130] Conversational interface
[2131] The device plays back responses from the generative AI model as audio, providing an environment where users can interact with the app naturally. For example, if a user asks, "What's the weather like today?", the device sends that audio to the server, and the generative AI model generates a response such as "It's sunny today." The device then plays this response back to the user as audio.
[2132] User-side embodiment
[2133] Inputting voice commands
[2134] Users use their smartphones to set daily schedules and ask questions using voice input. For example, they might speak instructions to their smartphone such as, "I'll go to the doctor at 10 AM tomorrow."
[2135] Notification confirmation
[2136] Users check reminder notifications displayed on their smartphones and act accordingly. For example, they might receive a notification saying, "It's 12 o'clock now. It's time for lunch."
[2137] Conversational communication
[2138] Users interact with the app via voice to configure and confirm settings. This helps alleviate feelings of loneliness for those living alone. For example, if asked, "What were your plans for today?", the device might respond, "I was planning to go shopping at 2 PM today."
[2139] Recognition and response to emotions
[2140] When a user gives a voice command, the device sends the voice to a server, where the emotion engine analyzes the user's emotional state. Based on the results, the system adjusts the content and timing of reminder notifications and generates comforting messages. For example, if the user says, "I'm not feeling very well," the server will generate a notification such as, "Please take some time to rest."
[2141] Example of a prompt
[2142] Please enter the following information by voice: "I will take my medicine at 8 AM tomorrow."
[2143] This invention enables the provision of personalized services that support the independent living of the elderly and take their emotions into consideration. This improves the elderly's sense of security and satisfaction, and enhances their quality of daily life.
[2144] The flow of the specific processing in Example 2 will be explained using Figure 13.
[2145] Step 1:
[2146] The user performs voice input.
[2147] The user launches a smartphone application and uses voice input. For example, they might say, "I'm going to the doctor at 10 AM tomorrow." This input data is captured by the device in voice format.
[2148] Step 2:
[2149] The device records audio and sends it to the server.
[2150] The terminal receives the user's voice input and records it as audio data. This recorded data is sent to the server. Specifically, audio data is sent from the terminal to the server, and the server receives it. The input data is the user's voice data, and the output data is the audio data sent to the server.
[2151] Step 3:
[2152] The server generates audio data and analyzes it using an AI model.
[2153] The server inputs the received audio data into a generative AI model, which analyzes the audio and converts it into text data. The generative AI model converts the audio data into text and analyzes that text to identify the user's intent. For example, the output text might say, "I'm going to the doctor at 10 AM tomorrow." The input data is audio data, and the output data is the converted text data.
[2154] Step 4:
[2155] The server registers text data in the database.
[2156] The server analyzes user schedule information based on text data and registers it in the database. For example, information such as "I will go to the doctor at 10 AM tomorrow" is saved as a schedule in the database. The input data is text data, and the output data is the schedule information stored in the database.
[2157] Step 5:
[2158] The server analyzes the audio data using an emotion engine.
[2159] The server inputs audio data into the emotion engine, which analyzes the user's emotional state. The emotion engine analyzes emotions from the tone and manner of speaking and outputs the results. For example, if the engine analyzes that the user is depressed, that information is output as emotion recognition data. The input data is audio data, and the output data is emotion recognition data.
[2160] Step 6:
[2161] The server generates and adjusts reminder notifications.
[2162] The server adjusts the content and timing of reminder notifications based on emotion recognition data. For example, if the server detects that the user is feeling down, the notification content changes to "Don't push yourself today, relax." The input data is emotion recognition data, and the output data is the adjusted reminder notification.
[2163] Step 7:
[2164] The server sends a reminder notification to the device.
[2165] The server sends a pre-arranged reminder notification to the device. For example, a reminder to "go to the doctor at 10 AM tomorrow" is sent to the device just before the scheduled time. The input data is the pre-arranged reminder notification, and the output data is the notification data sent to the device.
[2166] Step 8:
[2167] The device displays a reminder notification.
[2168] The device displays reminder notifications received from the server. For example, a notification saying "You have a doctor's appointment" is displayed on the screen and also announced audibly. The input data is the notification data received by the device, and the output data is the notification displayed to the user.
[2169] Through these steps, the system efficiently processes user voice input and utilizes a generative AI model and emotion engine to provide personalized support for daily life.
[2170] (Application Example 2)
[2171] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2172] This invention aims to solve problems related to support systems that enable elderly people to live independently and safely. While conventional technologies exist that convert voice input to text and provide mechanisms for schedule management and reminder notifications, they do not address the emotional state of the user or provide support for actual shopping. Elderly people often get lost or feel anxious when shopping in physical stores, and there is a need for a system that can alleviate these issues and allow them to shop in a relaxed state.
[2173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[2174] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the user's schedule information based on the text and registering it in a database, means for generating reminder notifications based on the schedule information and sending them to the user's terminal, means for providing shopping assistance to the user based on the voice input, means including an emotion engine for analyzing the user's emotional state, and means for generating support messages based on the emotional state and sending them to the user's terminal. This makes it possible for elderly people to shop in physical stores without getting lost or feeling anxious, and to shop in a relaxed manner while receiving appropriate guidance and support.
[2175] A "voice input method" is a system for capturing a user's voice into a device such as a terminal and processing it.
[2176] A "text conversion method" is a mechanism for analyzing captured audio data and converting it into text-based data.
[2177] A "schedule information analysis tool" is a system that extracts and analyzes a user's schedule and tasks based on text data, and manages them accordingly.
[2178] A "database registration method" is a mechanism for registering analyzed schedule information into a database that centrally manages that information.
[2179] A "reminder notification generation method" is a mechanism for generating announcements and notifications for users at specific times and under specific conditions, based on registered schedule information.
[2180] A "reminder notification sending method" is a mechanism that sends generated notifications to the user's device so that the user can check them.
[2181] A "shopping assistant system" is a mechanism that supports users' shopping based on voice input, guiding them to necessary product information and locations.
[2182] An "emotion engine" is a technology that analyzes a user's voice data and evaluates and recognizes their emotional state.
[2183] A "support message generation mechanism" is a system for generating appropriate messages for users based on their analyzed emotional state.
[2184] A "support message sending mechanism" is a system that sends generated support messages to the user's device, allowing the user to receive them.
[2185] This invention relates to a system that supports elderly people in living independently and safely, and describes its specific embodiments. This system is primarily designed to support elderly people in their daily activities, such as shopping, through voice input. The main components of the system and their specific operation are described in detail below.
[2186] Server-side processing
[2187] User authentication and data storage
[2188] The server authenticates users when they log into the application. Upon successful authentication, it saves data such as the user's schedule and shopping list to a database. This information is managed and stored on the server whenever the user updates their schedule or tasks.
[2189] Operation of Generative AI Models
[2190] The server uses a generative AI model to analyze voice input from the user. When the user says "I'm looking for milk" by voice, the generative AI model converts that instruction into text and generates appropriate product information and directions to the store.
[2191] Operation of the Emotion Engine
[2192] The server uses an emotion engine to analyze the user's voice data to determine their emotional state. For example, if a user says, "I'm a little tired," the emotion engine recognizes this emotional state and generates an appropriate support message. This support message might include suggestions such as, "Let's take a break," or "I'll call a staff member."
[2193] Schedule notifications and support message sending
[2194] The server generates reminders and support messages based on the user's schedule information and emotional state, and sends them to the user's device. For example, if "milk" is included in the shopping list, when the user enters the nearest store, a message will be sent saying, "Milk is in the refrigerated section."
[2195] Terminal-side processing
[2196] Voice input function
[2197] User voice input is captured through the device's microphone. The captured voice is converted into text data by speech recognition software and sent to the server.
[2198] Display reminders
[2199] The device displays notifications and reminders sent from the server. For example, a reminder such as "Take your medicine at 3 PM" will be displayed on the user's smartphone or smart glasses.
[2200] Shopping assistant function
[2201] The device receives responses from the generative AI model and provides voice guidance to the user. For example, if the user asks, "Where is the milk section?", it will notify the user, "It's in the refrigerated section." It also provides voice support messages based on the analysis results of the emotion engine.
[2202] Hardware and software to use
[2203] This system uses the following hardware and software.
[2204] Hardware: Smartphone, smart glasses, microphone, speaker
[2205] Software: speech_recognition library (speech recognition), proprietary generative AI model, emotion engine
[2206] Specific example
[2207] In a real-world scenario, if a user goes to a shopping mall and uses voice input to say, "I'm looking for milk," the emotion engine detects that the user is feeling a little anxious and provides a message such as, "The milk is in the refrigerated section. Let's call a staff member nearby."
[2208] Examples of prompts for a generative AI model:
[2209] "When a user says 'I'm looking for milk,' analyze their response and the emotions they express to generate an appropriate response. If the user is anxious, include a reassuring message."
[2210] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2211] Step 1:
[2212] Acquiring voice input (user)
[2213] Users input voice commands through the microphone on their smartphone or smart glasses. This input is part of the user's shopping list and might say something like, "I'm looking for milk."
[2214] Input: Audio data
[2215] Output: Audio data (sent to the device)
[2216] Step 2:
[2217] Text conversion of audio data (on the device)
[2218] The device converts the acquired audio data into text data using a speech recognition library (for example, speech_recognition). The converted text becomes "I'm looking for milk."
[2219] Input: Audio data
[2220] Data processing: Convert speech to text using a speech recognition library.
[2221] Output: Text data
[2222] Step 3:
[2223] Sending text data (from terminal to server)
[2224] The terminal sends the converted text data to the server. The server receives this text data and begins analysis.
[2225] Input: Text data
[2226] Output: Text data sent to the server
[2227] Step 4:
[2228] User authentication and data storage (server)
[2229] The server performs user authentication before parsing the transmitted text data. If authentication is successful, the user's schedule and shopping list are saved to the database.
[2230] Input: User ID, password, submitted text data
[2231] Data processing: User authentication, saving to database.
[2232] Output: Authentication success message, saved user data
[2233] Step 5:
[2234] Analysis using a generative AI model (server)
[2235] The server uses a generative AI model to analyze text data and generate appropriate responses. For example, in response to the text "I'm looking for milk," the AI model generates the answer "It's in the refrigerated section."
[2236] Input: Text data
[2237] Data Processing: Text analysis and response generation using generative AI models.
[2238] Output: Generated response text
[2239] Step 6:
[2240] Emotional state analysis (server)
[2241] Along with the generated response text, the emotion engine analyzes the user's emotional state based on their voice data. For example, it can recognize that the user is feeling anxious and generate an appropriate support message.
[2242] Input: Audio data, generated response text
[2243] Data processing: Emotional analysis using an emotion engine
[2244] Output: Emotional state, support message
[2245] Step 7:
[2246] Sending responses and support messages (from server to terminal)
[2247] The server sends the generated response text and support message to the user's terminal. For example, a notification might be sent saying, "It's in the refrigerated section. I'll call a staff member nearby."
[2248] Input: Generated response text, support message
[2249] Output: Response text and support message sent to the terminal
[2250] Step 8:
[2251] Playback of response (terminal)
[2252] The terminal plays back the received response text and support message as audio. The user is notified, "It's in the refrigerated section. I'll call a staff member nearby," and voice assistance is provided.
[2253] Input: Response text, support message
[2254] Data processing: Converting text to speech using speech synthesis.
[2255] Output: Voice guidance and support messages
[2256] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2257] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2258] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[2259] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2260] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[2261] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[2262] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[2263] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[2264] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[2265] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2266] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2267] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2268] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[2269] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2270] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[2271] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[2272] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[2273] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[2274] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[2275] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[2276] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[2277] The following is further disclosed regarding the embodiments described above.
[2278] (Claim 1)
[2279] A means of receiving user voice input,
[2280] means for converting the aforementioned voice input into text,
[2281] A means for analyzing the user's schedule information based on the aforementioned text and registering it in a database,
[2282] A means for generating a reminder notification based on the aforementioned schedule information and sending it to the user's terminal,
[2283] A system that includes this.
[2284] (Claim 2)
[2285] The system according to claim 1, further comprising means for comprehensively analyzing the aforementioned voice input and providing additional information to the user in a conversational format.
[2286] (Claim 3)
[2287] The system according to claim 1, further comprising means for automatically adjusting the reminder notification based on a specific time or condition.
[2288] "Example 1"
[2289] (Claim 1)
[2290] A means of receiving user voice input,
[2291] means for converting the aforementioned voice input into text,
[2292] A means for analyzing the user's schedule information based on the aforementioned text and registering it in a database,
[2293] A means for generating a reminder notification based on the aforementioned schedule information and sending it to the user's terminal,
[2294] A method for operating a generative AI model to analyze voice data and match it to specific tasks or information,
[2295] A means for generating user interaction based on the analysis results of the aforementioned generated AI model,
[2296] A system that includes this.
[2297] (Claim 2)
[2298] The system according to claim 1, further comprising means for comprehensively analyzing the aforementioned voice input and providing additional information to the user in a conversational format.
[2299] (Claim 3)
[2300] The system according to claim 1, further comprising means for automatically adjusting the reminder notification based on a specific time or condition.
[2301] "Application Example 1"
[2302] (Claim 1)
[2303] A means of receiving user voice input,
[2304] means for converting the aforementioned voice input into text,
[2305] A means for analyzing the user's schedule information based on the aforementioned text and registering it in a database,
[2306] A means for generating a reminder notification based on the aforementioned schedule information and sending it to the user's terminal,
[2307] A method for analyzing voice input from factory workers and matching it to specific tasks,
[2308] A means by which a robot sends push notifications to workers based on the work schedule,
[2309] A system that includes this.
[2310] (Claim 2)
[2311] The system according to claim 1, further comprising means for comprehensively analyzing the aforementioned voice input and providing additional information to the user in a conversational format, and for displaying and playing notifications received by the user's terminal on a display or by sound.
[2312] (Claim 3)
[2313] The system according to claim 1, further comprising means for automatically adjusting the reminder notification based on a specific time or condition and for using a generative AI model to extract specific tasks or information.
[2314] "Example 2 of combining an emotion engine"
[2315] (Claim 1)
[2316] A means of receiving user voice input,
[2317] means for converting the aforementioned voice input into text,
[2318] A means for analyzing the user's schedule information based on the aforementioned text and registering it in a database,
[2319] A means for generating a reminder notification based on the aforementioned schedule information and sending it to the user's terminal,
[2320] A means of inputting the aforementioned audio data into a generating AI model and utilizing the processing results,
[2321] A means for inputting the aforementioned audio data into an emotion engine and analyzing the emotional state,
[2322] A means of adjusting the content or timing of reminder notifications based on the analysis results,
[2323] A system that includes this.
[2324] (Claim 2)
[2325] The system according to claim 1, further comprising means for comprehensively analyzing the aforementioned voice input and providing additional information to the user in a conversational format.
[2326] (Claim 3)
[2327] The system according to claim 1, further comprising means for automatically adjusting the content or timing of reminder notifications based on the aforementioned emotional state.
[2328] "Application example 2 of combining emotional engines"
[2329] (Claim 1)
[2330] A means of receiving user voice input,
[2331] means for converting the aforementioned voice input into text,
[2332] A means for analyzing the user's schedule information based on the aforementioned text and registering it in a database,
[2333] A means for generating a reminder notification based on the aforementioned schedule information and sending it to the user's terminal,
[2334] A means of providing a user's shopping assistant based on voice input,
[2335] A means including an emotion engine for analyzing the user's emotional state,
[2336] A means for generating support messages based on emotional state and sending them to the user's device,
[2337] A system that includes this.
[2338] (Claim 2)
[2339] A means for comprehensively analyzing the aforementioned voice input and providing additional information to the user in a dialogue format, and
[2340] The system according to claim 1, further comprising means for managing a user's shopping list and providing product information.
[2341] (Claim 3)
[2342] The system according to claim 1, further comprising means for automatically adjusting the reminder notifications and support messages based on specific times or conditions. [Explanation of Symbols]
[2343] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving user voice input, means for converting the aforementioned voice input into text, A means for analyzing the user's schedule information based on the aforementioned text and registering it in a database, A means for generating a reminder notification based on the aforementioned schedule information and sending it to the user's terminal, A system that includes this.
2. The system according to claim 1, further comprising means for comprehensively analyzing the aforementioned voice input and providing additional information to the user in a dialogue format.
3. The system according to claim 1, further comprising means for automatically adjusting the reminder notification based on a specific time or condition.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A