System
A system integrating voice input, natural language processing, and GPS-based route calculation addresses caregiver burdens by offering efficient daily support and guidance, enhancing independent living for the elderly and disabled.
Patent Information
- Application Number
- JP2024129290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2026-02-18
AI Technical Summary
The increasing burden on caregivers and those receiving care in an aging society, particularly in areas like daily conversation support, cooking, and route guidance, is significant, exacerbated by caregiver shortages and rising care costs, necessitating a system that provides efficient and comfortable support.
A system that includes voice input conversion to text, natural language processing for response generation, health condition and ingredient analysis for cooking guidance, and GPS-based route calculation, integrated to provide real-time assistance via voice or on-screen displays.
Reduces the burden on caregivers and enables those receiving care to live independently by providing efficient daily support and guidance.
Smart Images

Figure 2026026869000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In an aging society, the burden on both caregivers and those receiving care is increasing. In particular, areas such as daily conversation support, cooking and meal assistance based on health status, and route guidance during travel require a lot of time and effort, placing a heavy burden on caregivers and the elderly. Furthermore, the shortage of caregivers and rising care costs are becoming serious problems. It is necessary to solve these issues and provide a comfortable and efficient lifestyle for all users, including caregivers, the elderly, and people with disabilities. [Means for solving the problem]
[0005] The present invention solves the above problems by providing a system including the following means.
[0006] A system including: means for acquiring a user's voice input and converting the voice data into text data; means for analyzing the text data and generating an appropriate response using a natural language processing engine; and means for converting the generated response into speech and providing it to the user.
[0007] The system further includes a means for acquiring user input data and transmitting health condition and ingredient information to a server, a means for generating cooking methods and recipes based on the user's health condition and ingredient data, and a means for providing the generated cooking instructions to the user by voice or on-screen display.
[0008] The system also includes a means for obtaining the user's current location using GPS and transmitting destination information to a server, a means for calculating the optimal route based on real-time traffic information, and a means for providing the calculated route information to the user by voice or on-screen display.
[0009] By combining these measures, the burden on caregivers can be significantly reduced, and those receiving care can live comfortable, independent lives.
[0010] "User" refers to the end users of this system, such as elderly people, people with disabilities, and caregivers.
[0011] "Voice input" refers to voice data uttered by a user, which is input information that the system converts into text data.
[0012] "Character data" is the result of converting voice input into character information, and is digital text data that the system uses for analysis.
[0013] A "natural language processing engine" is a software mechanism that analyzes text data obtained from a user, understands the context, and generates an appropriate response.
[0014] A "generated response" is appropriate response information analyzed and generated by a natural language processing engine.
[0015] "Health Status" refers to information that indicates the user's current physical and medical condition and vital signs.
[0016] "Ingredient information" is information about the types, quantities, and nutritional values of ingredients available to the user.
[0017] A "cooking method" is a cooking procedure or process that the system generates based on the user's health condition and ingredient information.
[0018] A "recipe" is a document or data that shows how to make a dish and the ingredients in it, in detail based on the cooking method.
[0019] "GPS" stands for Global Positioning System, a satellite positioning technology that obtains and provides a user's current location in real time.
[0020] "Real-time traffic information" refers to information that provides information on the operation status of transportation services and road congestion status in real time.
[0021] An "optimal route" is a route calculated based on real-time traffic information and GPS data that will allow a user to travel to their destination most efficiently.
[0022] A "server" is a computer system that receives and analyzes data from users and performs processes such as natural language processing and route calculation.
[0023] A "terminal" is a device that is directly operated by a user, and is a device that receives voice input, transmits data to a server, provides responses, and so on. [Brief explanation of the drawings]
[0024] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0025] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0026] First, the terms used in the following description will be explained.
[0027] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0028] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0029] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0030] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0031] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0032] [First embodiment]
[0033] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0034] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0035] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0036] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0037] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0038] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0039] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0040] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0041] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0043] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0045] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable life. Specific embodiments of the system are described below.
[0046] This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[0047] 1. Conversation support system
[0048] Overview of conversation support
[0049] The server hosts a natural language processing (NLP) engine that understands voice input and generates appropriate responses to assist the user in the conversation.
[0050] Program processing
[0051] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0052] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0053] 3. Server: The received text data is analyzed using a natural language processing engine and an appropriate response is generated, such as "Today's weather is sunny."
[0054] 4. Server: Sends the generated response data to the terminal.
[0055] 5. Terminal: The received response data is converted into voice using a voice synthesis function and provided to the user.
[0056] Specific examples
[0057] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question in real time and responds, based on the current weather information, with, "Today's weather is sunny." This response is provided to the user by voice from the device.
[0058] 2. Cooking and meal follow-up system
[0059] Cooking and Meal Follow-up Overview
[0060] This is a system that optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking.
[0061] Program processing
[0062] 1. Terminal: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[0063] 2. Server: Based on the received data, it generates cooking methods and recipes that take health into consideration.
[0064] 3. Server: Sends the generated cooking instructions to the terminal.
[0065] 4. Terminal: Provides the received cooking instructions to the user via voice or screen display, and also sends specific operating instructions to the robot assistant as needed.
[0066] Specific examples
[0067] If a user requests a low-sugar dessert, the system will provide the optimal recipe, taking into account the user's health status (e.g., whether they are prone to diabetes), and even automatically prepare part of the dessert with the help of a robotic assistant if needed.
[0068] 3. Route Guidance and Public Transportation Integration System
[0069] Overview of route guidance during travel
[0070] This system presents the optimal route based on the user's current location and destination. It provides the optimal route by including real-time traffic information.
[0071] Program processing
[0072] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[0073] 2. Server: Obtains real-time traffic information, calculates the optimal route, and sends the results to the device.
[0074] 3. Terminal: Provides the received route information to the user via voice and screen display. Tracks the user's current location in real time while moving and updates the route as needed.
[0075] Specific examples
[0076] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[0077] By integrating these functions, the system aims to reduce the burden of caregiving and enable those receiving care to live more independently.
[0078] The processing flow will be explained below.
[0079] Conversation support processing flow
[0080] Step 1:
[0081] Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0082] Step 2:
[0083] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0084] Step 3:
[0085] Server: Analyzes the received text data using a natural language processing engine and generates an appropriate response based on its content, such as "Today's weather is sunny."
[0086] Step 4:
[0087] Server: Sends the generated response data to the terminal.
[0088] Step 5:
[0089] Terminal: The received response data is converted into voice using a speech synthesis function and provided to the user. The response is provided at a speed and voice quality that is easy for the user to hear.
[0090] Cooking and meal follow-up process flow
[0091] Step 1:
[0092] Device: The user voices or texts a request such as "I want to make a diet lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[0093] Step 2:
[0094] Server: Based on the received data, it generates cooking methods and recipes that take health conditions into consideration.
[0095] Step 3:
[0096] Server: Sends the generated cooking instructions to the terminal.
[0097] Step 4:
[0098] Terminal: Provides received cooking instructions to the user via voice or screen display, and also sends signals to instruct the robot assistant on cooking operations as needed.
[0099] Step 5:
[0100] Robot: Based on instructions from a terminal, the robot will begin specific cooking tasks, such as automatically chopping vegetables and putting food in a pot.
[0101] Processing flow of the system for linking route guidance and public transportation during travel
[0102] Step 1:
[0103] Device: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[0104] Step 2:
[0105] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[0106] Step 3:
[0107] Server: Sends the calculated route information to the terminal.
[0108] Step 4:
[0109] Terminal: Provides the received route information to the user via voice and screen display, providing specific guidance such as "turn right at the next traffic light."
[0110] Step 5:
[0111] Device: GPS information is constantly updated while the user is moving, tracking the user's current location in real time. If the user deviates from the route, the device recalculates the optimal route and provides new route guidance.
[0112] Step 6:
[0113] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[0114] Through specific processing at each step, this system aims to provide real-time support to users and reduce the overall burden of caregiving.
[0115] Example 1
[0116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0117] In an aging society, reducing the burden of caregiving and providing a comfortable lifestyle are major challenges. In particular, advanced technology is required to provide consistent support for elderly people in conversation, eating, and mobility. However, existing systems have difficulty integrating multiple functions and providing them in a unified manner, limiting their practicality and effectiveness in the field. Therefore, there is a need to provide a comprehensive support system that allows elderly people to live independently.
[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0119] In this invention, the server includes means for receiving a user's voice input and converting the voice data into text data, means for transmitting the text data to the server, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for transmitting the generated response to a terminal, means for converting the generated response into text and providing it to the user, means for receiving user request data and transmitting health status and ingredient data to the server, means for analyzing the user's health status and ingredient data and generating cooking methods and recipes, means for transmitting the generated cooking instructions to the terminal, means for providing the generated cooking instructions to the user by voice or on-screen display, means for obtaining the user's current location using GPS and transmitting destination information to the server, means for calculating an optimal route based on real-time traffic information, means for transmitting the calculated route information to the terminal, means for providing the calculated route information to the user by voice or on-screen display, and means for tracking the user's current location and updating the route during travel. This provides users with conversation assistance, cooking and meal follow-up, and route guidance during travel in an integrated manner, enabling elderly people to live independently and comfortably.
[0120] "Voice input" refers to the act of a user providing information to a system using voice.
[0121] "Voice data" refers to a digital recording of a user's voice input.
[0122] "Character data" is data in text format converted by voice recognition.
[0123] A "server" is a computer system that provides functionality such as processing and storing data and generating responses.
[0124] A "terminal" is a device that allows a user to interact with the system through an interface, and includes smartphones, tablets, PCs, etc.
[0125] A "natural language processing engine" is software that analyzes text data, understands its meaning, and generates responses.
[0126] A "response" is a reply or instruction generated by a server based on user input.
[0127] "Health status data" is data that includes information related to the user's health.
[0128] "Ingredient data" is data that includes information about ingredients owned by the user.
[0129] A "cooking method" is a procedure for preparing a dish using specific ingredients.
[0130] A "recipe" is a set of instructions that lists the ingredients and steps needed to make a particular dish.
[0131] "GPS" is a Global Positioning System for obtaining geographical location information.
[0132] "Real-time traffic information" means up-to-date data on current traffic conditions.
[0133] An "optimal route" is the most efficient route to reach a destination.
[0134] "Location tracking" refers to the real-time monitoring of a moving user's current location.
[0135] "Screen display" refers to the visual presentation of information on a terminal display.
[0136] MODE FOR CARRYING OUT THE INVENTION
[0137] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable lifestyle. This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[0138] Conversation support system
[0139] In a conversation support system, a server hosts a natural language processing (NLP) engine that understands the user's voice input and generates appropriate responses. This system uses the following hardware and software:
[0140] Hardware
[0141] Devices: PC, smartphone, tablet, etc.
[0142] Devices with built-in microphones
[0143] software
[0144] Speech recognition engine (e.g., Google Speech-to-Text, Amazon Transcribe)
[0145] Natural language processing engines (e.g., Google Cloud Natural Language API, OpenAI's GPT-3)
[0146] Text-to-speech software (e.g., Amazon Polly, Google Text-to-Speech)
[0147] Specific examples
[0148] When a user asks, "What's the weather like today?", the system captures the voice with the device's microphone, and the speech recognition engine converts it into text data. The converted text data is sent to the server and analyzed by the natural language processing engine. As a result, a response such as "The weather is sunny today" is generated and sent to the device. Finally, the device converts the received response data into speech using speech synthesis software and provides it to the user.
[0149] Prompt Sentence Examples
[0150] "When a user says, 'What's the weather like today?' generate an appropriate response."
[0151] Cooking and meal follow-up system
[0152] The cooking and meal follow-up system optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking. This system uses the following hardware and software:
[0153] Hardware
[0154] Smartphones and tablets
[0155] Smart appliances (smart ovens, smart refrigerators)
[0156] software
[0157] Health management app
[0158] Ingredient Management Software
[0159] Recipe Generation Engine
[0160] Robot Assistant Control Software
[0161] Specific examples
[0162] If a user requests, "I want to make a low-sugar dessert," the system will consider the user's health condition (for example, whether they have a tendency toward diabetes) and provide the optimal recipe. The smartphone receives the request via voice or text and sends it to the server. The server generates the optimal recipe based on the received data and pre-registered health data and sends it to the smartphone. Furthermore, if necessary, the robot assistant will automatically prepare part of the dessert.
[0163] Prompt Sentence Examples
[0164] "If a user requests, 'I want to make a dessert with less sugar,' generate a recipe that takes their health into consideration."
[0165] Route guidance and public transport linkage system
[0166] The route guidance and public transport integration system for travel is a system that presents the optimal travel route based on the user's current location and destination. It provides the optimal route by including real-time traffic information. This system uses the following hardware and software.
[0167] Hardware
[0168] GPS-enabled devices (smartphones, tablets)
[0169] software
[0170] Real-time traffic information API (e.g., Google Maps API, HERE API)
[0171] Route Calculation Algorithm
[0172] Location Tracking Software
[0173] Specific examples
[0174] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." The smartphone obtains the user's current location using GPS and sends a request to the server. The server calculates a route based on real-time traffic information and sends it to the smartphone. The smartphone then provides the calculated route to the user by voice or on-screen display, and tracks the user's current location in real time while traveling, updating the route as needed.
[0175] Prompt Sentence Examples
[0176] "When a user requests 'Take me to the station,' generate a plan that provides the optimal route."
[0177] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0178] Conversation support system
[0179] Program processing
[0180] Step 1:
[0181] The user gives a voice input of "What's the weather like today?" The device captures this voice input using the built-in microphone. The input is the user's voice data, and the output is the captured voice data. The device collects voice data in real time.
[0182] Step 2:
[0183] The device sends the captured voice data to a voice recognition engine and converts it into text data. Specifically, the voice data is sent to a voice recognition service on the cloud, and text data is obtained. The input here is voice data, and the output is converted text data. The device performs the operation of converting voice data into text data.
[0184] Step 3:
[0185] The terminal sends the converted character data to the server. The data is sent securely using a secure protocol (such as HTTPS). The input here is character data, and the output is the character data sent to the server. The terminal performs the data sending operation.
[0186] Step 4:
[0187] The server inputs the received text data into a natural language processing engine (NLP engine) for analysis. The NLP engine analyzes the text data and generates an appropriate response. The input here is text data, and the output is the generated response data. The server performs the operations of analyzing the text data and generating a response.
[0188] Step 5:
[0189] The server sends the generated response data to the terminal. Again, communication is performed using a secure protocol. The input here is the response data, and the output is the response data sent to the terminal. The server performs the data transmission operation.
[0190] Step 6:
[0191] The terminal converts the received response data into speech using a speech synthesis engine. The speech synthesis engine is used to convert text data into speech data, which is then provided to the user through a speaker. The input here is the response data, and the output is the generated speech data. The terminal then performs the operation of playing back the speech data.
[0192] Cooking and meal follow-up system
[0193] Program processing
[0194] Step 1:
[0195] The user inputs a request into the device by voice or text, such as "I want to make a dietary lunch." The device captures this input. The input is the user's request data, and the output is the captured request data. The device performs the operation of collecting voice and text data.
[0196] Step 2:
[0197] The terminal sends the request data, pre-registered health condition data, and ingredient information to the server. The data is sent using a secure protocol. The input is the request data and the attached health condition data and ingredient information, and the output is the data sent to the server. The terminal performs the data transmission operation.
[0198] Step 3:
[0199] The server analyzes the received data and generates appropriate cooking methods and recipes that take health status into consideration. The input here is the request data, health status data, and ingredient information, and the output is the generated recipe data. The server performs the data analysis and recipe generation operations.
[0200] Step 4:
[0201] The server sends the generated recipe data to the terminal. Communication is again performed using a secure protocol. The input is the recipe data, and the output is the recipe data sent to the terminal. The server performs the data transmission operation.
[0202] Step 5:
[0203] The terminal provides the received recipe data to the user by voice or screen display. The input here is recipe data, and the output is recipe information provided to the user. The terminal performs the operations of displaying information and playing voice.
[0204] Step 6:
[0205] If necessary, the terminal sends specific cooking operation instructions to the robot assistant. The input is cooking operation instruction data, and the output is instruction data sent to the robot assistant. The terminal performs the operation of sending instructions.
[0206] Route guidance and public transport linkage system
[0207] Program processing
[0208] Step 1:
[0209] The user inputs a request into the terminal by voice or text, such as "Take me to the station." The terminal captures this input. The input is the user's request data, and the output is the captured request data. The terminal performs the operation of collecting voice and text data.
[0210] Step 2:
[0211] The device uses the built-in GPS to obtain the user's current location. The input here is GPS data, and the output is the obtained current location data. The device performs the operation of collecting location data.
[0212] Step 3:
[0213] The device sends the request data and current location data to the server. The data is sent using a secure protocol. The input is the request data and current location data, and the output is the data sent to the server. The device performs the data sending operation.
[0214] Step 4:
[0215] The server obtains real-time traffic information and calculates the optimal route. The input here is the current location data and real-time traffic information, and the output is the calculated route data. The server performs data analysis and route calculation.
[0216] Step 5:
[0217] The server sends the calculated route data to the terminal. Communication is again performed using a secure protocol. The input is the route data, and the output is the route data sent to the terminal. The server performs the data transmission operation.
[0218] Step 6:
[0219] The terminal provides the received route data to the user by voice or screen display. The input here is the route data, and the output is the route information provided to the user. The terminal performs the operations of displaying information and playing voice.
[0220] Step 7:
[0221] While moving, the device uses GPS to track the user's current location and updates the route as needed. The input here is GPS data and real-time traffic information, and the output is updated route data. The device performs location tracking and route update operations.
[0222] (Application example 1)
[0223] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0224] This invention relates to a system for supporting and ensuring the safety of specific groups, such as the elderly, in their daily lives. In particular, the system aims to reduce the burden on caregivers and support the independent living of those receiving care by utilizing the user's voice input and location information to provide comprehensive support tailored to a wide range of needs, including conversation assistance, cooking and meal follow-up, and emergency response. There is also a need for a system that can respond quickly and appropriately in emergencies.
[0225] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0226] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for acquiring the user's current location, means for transmitting the user's location information to the server, and means for analyzing the user's voice in an emergency and transmitting an alert together with the location information, thereby supporting the user's daily life and enabling a prompt and appropriate response, particularly in an emergency.
[0227] "Voice input" refers to data obtained by a device capturing words spoken by a user.
[0228] "Voice data" is data that digitally represents voice input obtained from a user.
[0229] "Character data" is data obtained by analyzing voice data and converting the content into text format.
[0230] A "natural language processing engine" is a program or system that analyzes text data and generates appropriate responses.
[0231] "Response generation" refers to the process of using a natural language processing engine to create an appropriate response from text data.
[0232] "Speech synthesis" is the technique of converting the generated response back into speech form.
[0233] "Current location of user" refers to the geographic location where the user is currently located.
[0234] "GPS" is a satellite positioning system that identifies a user's current location.
[0235] "Real-time traffic information" is dynamic information that shows the current traffic situation.
[0236] An "optimal route" refers to the most efficient route from the current location to the destination.
[0237] "Emergency speech analysis" is the process of identifying specific speech inputs that indicate an emergency.
[0238] "Transmitting location information" means sending data indicating the user's current location to a server or other receiving device.
[0239] "Sending an alert" means communicating a warning message based on a specific condition.
[0240] This invention is a system for supporting specific groups, such as the elderly, in their daily lives and ensuring their safety, and is realized by each of the entities, the server, the terminal, and the user, fulfilling their specific roles. The following describes in detail the mode for carrying out this invention.
[0241] Embodiment of conversation support system
[0242] The server hosts a natural language processing engine (NLP engine). The device captures the user's voice input and converts this voice data into text data. The converted text data is sent to the server, where the NLP engine generates an appropriate response. The generated response is sent back to the device, where it is converted into speech and provided to the user. This system allows users to obtain a variety of information through natural conversation.
[0243] Hardware / Software: A microphone is used for voice input, the speech_recognition library is used for voice recognition, an NLP engine is used for natural language processing, and the pyttsx3 library is used for speech synthesis.
[0244] Example: When a user asks, "What's the weather like today?", the system responds by saying, "The weather is sunny today."
[0245] Cooking and meal follow-up system
[0246] The device receives input data from the user regarding their health condition and ingredients, and sends this data to the server. Based on the received data, the server generates cooking methods and recipes optimized for the user's health condition. The generated cooking instructions are provided to the user via the device via voice or screen display. This allows users to easily cook in a way that takes their health into consideration.
[0247] Hardware and software: Smartphones and tablets are used to acquire and transmit data. AI models are used to generate recipes, and existing speech synthesis libraries and display technologies are used for voice and screen display.
[0248] Example: When a user requests to make a low-sugar dessert, the system provides the optimal recipe based on the user's health status.
[0249] Route guidance and public transport linkage system for travel
[0250] The device acquires the user's current location and sends destination information to the server. The server calculates the optimal route based on real-time traffic information and sends that information to the device. The device then provides the calculated route information to the user via voice and on-screen display. In addition, the device has the ability to analyze the user's voice in an emergency and send an alert along with location information.
[0251] Hardware and software: GPS devices, real-time traffic information platforms, and speech recognition and speech synthesis libraries are used.
[0252] Example: If a user yells "Help!", the system interprets the audio as an emergency alert and sends information, including the user's current location, to emergency contacts.
[0253] Prompt Sentence Examples
[0254] "How's the weather today?"
[0255] "I want to make a dessert with less sugar."
[0256] "help me!"
[0257] In this way, the present invention utilizes the user's voice and location information to provide a wide range of support for daily life, thereby reducing the burden on elderly people and their caregivers.
[0258] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0259] Processing steps of the conversation support system
[0260] Step 1:
[0261] The terminal obtains the user's voice input.
[0262] Specifically, when a user says, "What's the weather like today?", the device's microphone captures the voice. The input is voice data.
[0263] Step 2:
[0264] The terminal converts the captured voice data into text data.
[0265] Specifically, it uses a speech recognition library (e.g., speech_recognition) to convert voice data into text format. The input is voice data, and the output is text data.
[0266] Step 3:
[0267] The terminal transmits the converted character data to the server.
[0268] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[0269] Step 4:
[0270] The server analyzes the received text data using a natural language processing engine.
[0271] Specifically, an NLP engine (for example, a custom NLP server) analyzes the text data and generates an appropriate response to the question, "What's the weather like today?" The input is the text data, and the output is the analysis result and response data.
[0272] Step 5:
[0273] The server sends the generated response data to the terminal.
[0274] Specifically, the generated response data is sent to the terminal via an HTTP request. The input is the response data, and the output is data sent to the terminal.
[0275] Step 6:
[0276] The terminal converts the received response data into voice and provides it to the user.
[0277] Specifically, it uses a speech synthesis library (e.g., pyttsx3) to convert text-based response data into speech and provides it to the user through a speaker. The input is the response data, and the output is speech output.
[0278] Emergency response system processing steps
[0279] Step 1:
[0280] The terminal obtains the user's voice input.
[0281] Specifically, when a user shouts "Help!", the device's microphone captures the voice. The input is voice data.
[0282] Step 2:
[0283] The terminal converts the captured voice data into text data.
[0284] Specifically, it uses a speech recognition library to convert voice data into text format. The input is voice data and the output is text data.
[0285] Step 3:
[0286] The terminal transmits the converted character data to the server.
[0287] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[0288] Step 4:
[0289] The server analyzes the text data and identifies it as an urgent message.
[0290] Specifically, the NLP engine recognizes the emergency message "Help!" and prepares an appropriate response. The input is text data, and the output is emergency alert data.
[0291] Step 5:
[0292] The device obtains the user's current location using GPS.
[0293] Specifically, the GPS module of the device acquires the current location information. The input is a signal from the GPS device, and the output is location information data.
[0294] Step 6:
[0295] The server sends location information and emergency alerts to emergency contacts.
[0296] Specifically, location information and an emergency message are sent to emergency contacts (e.g., family members or emergency services) via an HTTP request. The input is the emergency alert data and location data, and the output is the transmission of the data to the contacts.
[0297] Step 7:
[0298] The terminal provides a voice message to reassure the user.
[0299] Specifically, it uses a speech synthesis library to generate a message such as "Rescue is currently being called. Please rest assured," and provides it to the user through a speaker. The input is the data for the reassurance message, and the output is voice output.
[0300] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0301] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live a comfortable life. Specific embodiments of this invention are described below.
[0302] In this system, each of the server, terminal, and user plays a role in recognizing the user's emotions and providing optimal responses and support to the user's mental and physical needs.
[0303] 1. Conversation support system
[0304] Overview of conversation support
[0305] The server hosts an emotion engine and a natural language processing (NLP) engine, which recognizes emotions from the user's voice input and generates appropriate responses to provide to the user.
[0306] Program processing
[0307] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0308] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0309] 3. Server: Analyzes the received text data using an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).
[0310] 4. Server: Uses a natural language processing engine to generate an appropriate response based on context and sentiment, such as "The weather is sunny today."
[0311] 5. Server: Sends a tone based on the generated response data and emotional state to the terminal.
[0312] 6. Terminal: The received response data is converted into voice with an appropriate tone using a voice synthesis function and provided to the user.
[0313] Specific examples
[0314] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question, and if it recognizes that the user is feeling "depressed," it responds by adding encouraging words such as, "Cheer up, it's sunny today, why don't you go outside?"
[0315] 2. Cooking and meal follow-up system
[0316] Cooking and Meal Follow-up Overview
[0317] This system provides optimal cooking instructions based on the user's health and emotional state.
[0318] Program processing
[0319] 1. Device: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status, ingredient information, and emotional data, is sent to the server.
[0320] 2. Server: Based on the received data, it generates cooking methods and recipes that take into account the health and emotional state of the user.
[0321] 3. Server: Sends the generated cooking instructions and related emotional support information to the terminal.
[0322] 4. Terminal: Provides the received cooking instructions and emotional support information (e.g., suggestions for ingredients with a relaxing effect when the user is feeling stressed) to the user via voice or on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[0323] Specific examples
[0324] If a user requests "I want to make a healthy breakfast," the system will provide a relaxing recipe using oatmeal based on the user's health status and feelings of "tension." Furthermore, if a robotic assistant is needed, it will automate some of the oatmeal cooking process.
[0325] 3. Route Guidance and Public Transportation Integration System
[0326] Overview of route guidance during travel
[0327] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[0328] Program processing
[0329] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[0330] 2. Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination, taking into account the emotional state.
[0331] 3. Server: Sends the calculated route information and emotional support information (for example, calming voice guidance if the user is feeling anxious) to the device.
[0332] 4. Device: Provides the user with the received route information and emotional support information via voice and screen display. It also tracks the user's current location in real time while moving and updates the route as needed.
[0333] Specific examples
[0334] If the system detects that the user is "nervous" while traveling from home to the station, it calculates the shortest and safest route based on the latest traffic information and provides specific instructions in a gentle tone, such as "Turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[0335] As described above, the AI system combined with the emotion engine provides services that take into account the user's mental state, aiming to reduce the burden of caregiving and improve the user's quality of life.
[0336] The processing flow will be explained below.
[0337] Conversation support processing flow
[0338] Step 1:
[0339] User: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0340] Step 2:
[0341] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0342] Step 3:
[0343] Server: Analyzes the user's emotional state using the emotion engine and sends the results of this analysis to the natural language processing engine.
[0344] Step 4:
[0345] Server: Uses a natural language processing engine to generate appropriate responses based on context and sentiment.
[0346] Step 5:
[0347] Server: Sends tone information based on the generated response data and emotional state to the terminal.
[0348] Step 6:
[0349] Terminal: The terminal converts the received response data into speech with an appropriate tone using a speech synthesis function and provides it to the user.
[0350] Cooking and meal follow-up process flow
[0351] Step 1:
[0352] User: The user inputs a request by voice or text, such as "I want to make a dietary lunch." The device then sends this information, along with previously entered health status and ingredient information, to the server.
[0353] Step 2:
[0354] Server: Based on the received information about the user's health condition and ingredients, the emotion engine analyzes the user's emotional state.
[0355] Step 3:
[0356] Server: Generates cooking methods and recipes that take into account health and emotional states.
[0357] Step 4:
[0358] Server: Sends the generated cooking instructions and emotional support information to the terminal.
[0359] Step 5:
[0360] Terminal: Provides the received cooking instructions and emotional support information to the user via voice and on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[0361] Step 6:
[0362] Robot: Based on instructions from a terminal, the robot begins specific cooking tasks. For example, it automatically performs basic operations such as chopping vegetables and putting food in a pot.
[0363] Processing flow of the system for linking route guidance and public transportation during travel
[0364] Step 1:
[0365] User: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[0366] Step 2:
[0367] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[0368] Step 3:
[0369] Server: Based on the received user's emotional state, the server adjusts the guidance method (tone and pace) for the optimal route.
[0370] Step 4:
[0371] Server: Sends the calculated route information and emotional support information to the terminal.
[0372] Step 5:
[0373] Terminal: Provides the received route information and emotional support information to the user via voice and on-screen display. For example, it provides specific guidance such as "Turn right at the next traffic light," but changes the tone depending on the user's emotional state.
[0374] Step 6:
[0375] Device: GPS information is constantly updated while moving, tracking the user's current location in real time. Each time, the server recalculates the optimal route based on the user's emotional state and provides new route guidance.
[0376] Step 7:
[0377] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[0378] This detailed processing flow allows the system to provide more personalized assistance while taking into account the user's emotional state, thereby reducing the burden of caregiving and enabling users to live a comfortable and efficient life.
[0379] Example 2
[0380] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0381] In recent years, the burden of caregiving has increased with the progress of the aging society. There is also a demand for support to maintain users' mental and physical health. However, conventional AI systems do not take into account the user's emotional state, limiting the user experience. Furthermore, even in health management and mobility assistance, they can only make suggestions that ignore the user's emotional state, making it difficult to provide accurate support.
[0382] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the user's emotional state from the text data using an emotion engine, means for generating an appropriate response based on the context and the emotional state using a natural language processing engine, and means for converting the generated response into voice and providing it to the user. This makes it possible to provide appropriate responses and support according to the user's emotional state.
[0383] "Audio input" refers to an audio signal emitted by a user through an input device such as a microphone.
[0384] "Audio Data" means data in digital form obtained from audio input.
[0385] "Character data" refers to text-format data obtained by analyzing voice data using voice recognition technology.
[0386] An "emotion engine" is software or hardware that analyzes text data and estimates the user's emotional state.
[0387] An "emotional state" is a mental state that a user is feeling, and examples include happiness, sadness, anger, etc.
[0388] A "natural language processing engine" is software that analyzes text data, understands meaning and context, and generates appropriate responses.
[0389] A "response" is a reply or instruction to the user generated by a natural language processing engine.
[0390] "Speech synthesis" is a technology that converts text data into audio data and turns it into audio that can be played through speakers, etc.
[0391] "Health status" refers to information about the user's current physical health, including information about illnesses, symptoms, and physical conditions.
[0392] "Ingredient information" refers to data relating to the types and amounts of ingredients available to the user.
[0393] A "cooking method" is the procedure or process of preparing a dish using ingredients.
[0394] A "recipe" is a document that describes the ingredients and steps for making a particular dish.
[0395] A "location information system" is a system that obtains a user's current location using technology such as GPS.
[0396] "Real-time traffic information" refers to information about current traffic conditions and transportation options.
[0397] An "optimal route" is the most efficient route for a user to travel from their current location to their destination.
[0398] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live more comfortable lives. Specific embodiments of this invention will be described in detail below.
[0399] 1. Conversation support system
[0400] System Configuration
[0401] In this system, users ask questions or make requests by voice, and the server processes them and provides a voice response. The main components are as follows:
[0402] Terminal: A device that captures the user's voice and sends the voice data to the server. It uses the Google Speech-to-Text API for voice recognition.
[0403] Server: Acquires text data and analyzes it using an emotion engine (e.g., IBM Watson Tone Analyzer). OpenAI's GPT-3 natural language processing engine is used.
[0404] Speech synthesis engine: Converts the generated text response into speech using Amazon Polly.
[0405] Specific examples
[0406] When a user asks, "What's the weather like today?", the voice data captured by the device is converted into text data and sent to the server. The server uses an emotion engine to recognize the emotional state of "interested" and generates an appropriate response using a natural language processing engine. For example, it might respond, "The weather is sunny today." This is then converted into speech using a speech synthesis engine and provided to the user via the device.
[0407] Prompt Sentence Examples
[0408] User: "What's the weather like today?"
[0409] system:
[0410] 1. The server converts the speech into text and analyzes it using an emotion engine.
[0411] 2. If the emotion is perceived as "interested," generate an appropriate response.
[0412] 2. Cooking and meal follow-up system
[0413] System Configuration
[0414] This system provides optimal cooking instructions based on the user's health and emotional state.
[0415] Device: Captures user requests and sends them to the server. The voice recognition function uses the Google Speech-to-Text API.
[0416] Server: Analyzes health status, ingredient information, and emotional state to generate cooking methods and recipes. IBM Watson Tone Analyzer is used as the emotion engine, and OpenAI's GPT-3 is used to generate recipes.
[0417] Robot assistant: A device that automates cooking operations as needed.
[0418] Specific examples
[0419] If a user requests "I want to make a dietary lunch," the device sends this information to the server. The server uses its emotion engine to recognize that the user is "feeling stressed" and generates a cooking method using ingredients that have a relaxing effect. The generated recipe is for oatmeal and salad and is provided to the user via voice and on-screen display. If necessary, some of the cooking can be automated by a robotic assistant.
[0420] Prompt Sentence Examples
[0421] User: "I want to make a diet lunch."
[0422] system:
[0423] 1. The server receives the request and user data and generates an appropriate recipe.
[0424] 2. The recipe uses oatmeal and provides audio instructions.
[0425] 3. Route Guidance and Public Transportation Integration System
[0426] System Configuration
[0427] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[0428] Terminal: Receives the user's request, obtains the current location using GPS, and sends it to the server.
[0429] Server: Obtains real-time traffic information and calculates the optimal route. IBM Watson Tone Analyzer is used as the emotion engine, and Google Maps API is used to obtain traffic information.
[0430] Speech synthesis engine: Converts route directions into speech.
[0431] Specific examples
[0432] When a user requests "Guide me to the station," the device obtains the user's current location, sets "station" as the destination, and sends the information to the server. The server then uses its emotion engine to recognize the user as "stressed" and calculates the optimal route to reassure the user. For example, the server uses a speech synthesis engine to convert specific instructions, such as "Turn right at the next traffic light," into voice and provides the information to the user via the device.
[0433] Prompt Sentence Examples
[0434] User: "Take me to the station."
[0435] system:
[0436] 1. The server calculates a route based on the current location and destination, and provides guidance according to the user's emotional state.
[0437] 2. If they are feeling anxious, guide them with a reassuring route and a gentle tone.
[0438] As described above, the present invention can provide support in various situations while taking into consideration the emotional state of the user, reduce the burden of caregiving, and improve the quality of life of the user.
[0439] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0440] Conversation support system processing flow
[0441] Step 1:
[0442] Description: The user provides voice input.
[0443] Specific action: The user speaks to the device, "What's the weather like today?"
[0444] Input: User's voice input.
[0445] Output: The captured audio data.
[0446] Step 2:
[0447] Description: The device receives voice data and converts it into text data using voice recognition.
[0448] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[0449] Input: Captured audio data.
[0450] Output: The converted text data (e.g., "What's the weather like today?").
[0451] Step 3:
[0452] Description: The terminal sends the converted character data to the server.
[0453] Specific operation: The terminal sends an HTTP request to the server and sends character data.
[0454] Input: The converted character data.
[0455] Output: Character data sent to the server.
[0456] Step 4:
[0457] Description: The server receives text data and analyzes it with the emotion engine.
[0458] Specific operation: The server uses IBM Watson Tone Analyzer to analyze the user's emotional state (e.g., "interested") from the text data.
[0459] Input: The character data sent.
[0460] Output: Parsed emotional state.
[0461] Step 5:
[0462] Description: The server uses a natural language processing engine to generate appropriate responses based on context and emotional state.
[0463] Specific operation: The server calls OpenAI's GPT-3 and generates an appropriate response sentence using the text data and emotional state as input (e.g., "The weather is sunny today.").
[0464] Input: Text data, emotional state.
[0465] Output: The generated response sentence.
[0466] Step 6:
[0467] Description: The server sends the generated response to the terminal.
[0468] Specific operation: The server sends an HTTP response to the terminal and sends the generated response text.
[0469] Input: The generated response sentence.
[0470] Output: The response sent to the terminal.
[0471] Step 7:
[0472] Description: The response received by the terminal is converted into speech using the speech synthesis function.
[0473] Specific operation: The device uses Amazon Polly to convert the response into voice data.
[0474] Input: The response received.
[0475] Output: The converted audio data.
[0476] Step 8:
[0477] Description: The device plays the converted audio data to the user.
[0478] Specific operation: Play audio data through the device speaker.
[0479] Input: The converted audio data.
[0480] Output: A spoken response to the user.
[0481] Cooking and meal follow-up system processing flow
[0482] Step 1:
[0483] Description: A user enters a cooking request.
[0484] Specific operation: The user speaks to the device, saying, "I want to make a dietary lunch."
[0485] Input: User's voice input.
[0486] Output: The captured audio data.
[0487] Step 2:
[0488] Description: The device receives voice data and converts it into text data using voice recognition.
[0489] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[0490] Input: Captured audio data.
[0491] Output: The converted text (e.g. "I want to make a dietary lunch").
[0492] Step 3:
[0493] Description: The device acquires health status, food ingredient information, and emotional state and sends them to the server.
[0494] Specific operation: The device obtains pre-registered health status, food information, and emotional state, and sends an HTTP request to the server.
[0495] Input: Text data, health status, food information, emotional state.
[0496] Output: The data sent to the server.
[0497] Step 4:
[0498] Description: The server generates cooking instructions and emotional support information based on the data received.
[0499] Specific operation: Using an emotion engine (e.g., IBM Watson Tone Analyzer) and a natural language processing engine (e.g., OpenAI's GPT-3), it generates appropriate cooking methods, recipes, and emotional support information.
[0500] Input: health status, food information, emotional state.
[0501] Output: Generated cooking instructions and emotional support information.
[0502] Step 5:
[0503] Description: The server sends the generated cooking instructions and emotional support information to the terminal.
[0504] Specific operation: The server sends an HTTP response to the terminal, conveying instructions and information.
[0505] Input: Generated cooking instructions and emotional support information.
[0506] Output: Cooking instructions and emotional support information sent to the device.
[0507] Step 6:
[0508] Description: The device provides the user with cooking instructions and emotional support information received.
[0509] How it works: The device uses speech synthesis to convert instructions into voice, provides information on the screen, and issues instructions to the robot assistant as needed.
[0510] Input: cooking instructions and emotional support information.
[0511] Output: Voice and visual instructions, instruction signals to the robot assistant.
[0512] Processing flow of the system for linking route guidance and public transportation during travel
[0513] Step 1:
[0514] Description: User requests directions.
[0515] Specific operation: The user speaks to the terminal, saying, "Take me to the station."
[0516] Input: User's voice input.
[0517] Output: The captured audio data.
[0518] Step 2:
[0519] Description: The device receives voice data and converts it into text data using voice recognition.
[0520] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[0521] Input: Captured audio data.
[0522] Output: The converted text data (e.g. "Take me to the station").
[0523] Step 3:
[0524] Description: The device obtains its current location using GPS and sends destination information and emotional state to the server.
[0525] Specific operation: The device acquires GPS data and sends an HTTP request to the server along with the user's emotional state information.
[0526] Input: Text data, current location information, destination information, emotional state.
[0527] Output: The data sent to the server.
[0528] Step 4:
[0529] Description: The server calculates the optimal route based on real-time traffic information and emotional state.
[0530] Specific operation: The server uses the Google Maps API to calculate the optimal route, taking into account the traffic information obtained and the emotional state analyzed by the emotion engine.
[0531] Input: current location information, destination information, emotional state.
[0532] Output: Calculated optimal route information.
[0533] Step 5:
[0534] Description: The server sends the calculated route information and emotional support information to the terminal.
[0535] Specific operation: The server sends an HTTP response to the device, conveying route information and emotional support information.
[0536] Input: Calculated optimal route information, emotional support information.
[0537] Output: Route information and emotional support information sent to the device.
[0538] Step 6:
[0539] Description: Provides the user with route information and emotional support information received by the terminal.
[0540] Specific operation: The device uses a speech synthesis function to convert route guidance into voice and also provides route information on the screen.
[0541] Input: Received route information, emotional support information.
[0542] Output: Route guidance by voice and on-screen display.
[0543] Step 7:
[0544] Description: The device tracks the user's location in real time and updates the route as needed.
[0545] How it works: The device uses GPS to continuously obtain its current location while moving, and resends the information to the server as needed to receive new route guidance.
[0546] Input: current location information, route information.
[0547] Output: Updated route guidance information.
[0548] This is the specific processing flow of this system. At each step, the system converts voice data, analyzes emotions, processes natural language, calculates routes, and performs other operations based on the user's input, providing appropriate responses and support.
[0549] (Application example 2)
[0550] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0551] Conventional security and guidance systems tend to respond in a uniform manner without considering the user's emotional state, making it impossible to provide optimal alerts and guidance to users. Furthermore, even in situations where users feel anxious or scared, appropriate responses are often delayed, reducing the user's sense of security. Therefore, there is a need for a system that provides effective responses and a sense of security that takes the user's emotional state into account.
[0552] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0553] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for capturing the user's face and recognizing emotions using an emotion engine, and means for providing appropriate security alerts and guidance based on the recognized emotions. This enables optimal responses that take the user's emotional state into consideration, thereby improving the user's sense of security.
[0554] "Voice input" refers to data that is input by the user using voice.
[0555] "Voice data" refers to data that represents in digital form the voice input by the user.
[0556] "Character data" refers to text-format data obtained by converting voice data into characters.
[0557] A "natural language processing engine" is a processing engine that analyzes text data, understands its meaning, and generates an appropriate response.
[0558] An "emotion engine" is an engine that has the function of determining emotions from the user's voice, facial expressions, etc.
[0559] "Response generation" is the process of creating appropriate responses or guidance based on the user's input data and their emotional state.
[0560] "Speech synthesis" is a technology that outputs text data as speech.
[0561] "Emotion recognition" is the process of determining a user's emotions from their facial expressions, voice, etc.
[0562] A "security alert" is a notification that warns or warns users to ensure their safety.
[0563] "Guidance" refers to instructions that instruct the user on appropriate actions or responses.
[0564] "Capturing a user's face" means obtaining an image of the user's face using a device such as a camera.
[0565] A "location information system" is a system that obtains a user's current location using GPS or other means.
[0566] "Real-time traffic information" refers to data that obtains the latest information on traffic conditions in real time.
[0567] An "optimal route" is a route that allows a user to reach a destination most efficiently.
[0568] The system of the present invention is an AI system that combines an emotion engine and a natural language processing engine to provide optimal support for the user's mental and physical needs. The system allows the server, terminal, and user to play their respective roles, improving user security.
[0569] Hardware and Software
[0570] Hardware
[0571] Smart glasses (a device with a general camera function)
[0572] Smartphone (Android or iOS device)
[0573] Head-mounted display (general VR device)
[0574] software
[0575] Emotion engine (e.g. Affectiva SDK)
[0576] Natural language processing engine (e.g. Google Cloud Natural Language API)
[0577] Image and voice recognition tools (e.g., Google Cloud Vision API, Speech-to-Text API)
[0578] Speech synthesis tools (e.g., Google Cloud Text-to-Speech API)
[0579] Data processing and calculation
[0580] server
[0581] The server performs the following processes. First, it receives the user's voice input and facial image from the device. The voice data is converted into text data using a voice recognition function, and the text data and facial image are sent in parallel to the emotion engine. The emotion engine analyzes the user's emotional state and provides this data to the natural language processing engine. The natural language processing engine generates an appropriate response based on the text data and emotional data. Finally, the generated response data is converted into voice data using a voice synthesis tool and sent to the device.
[0582] Terminal
[0583] The device performs the following processes: First, it captures voice input and facial images from the user and sends them to the server. Then it receives response data returned from the server and outputs voice to the user at the appropriate time. If the user feels anxious, it provides security alerts and guidance based on emotion recognition.
[0584] Specific examples
[0585] scenario
[0586] When a user is walking down a street at night, the AI detects that they are feeling anxious.
[0587] Response: The smart glasses say, "User, it's okay. There's a station nearby. Would you like to contact the fastest security company right away?"
[0588] Prompt Sentence Examples
[0589] Design an AI assistant with an emotion engine that recognizes the user's emotional state and provides appropriate security alerts and guidance. For a specific scenario, generate a reassuring message when the user is feeling anxious. The devices used could be smart glasses, a smartphone, or a head-mounted display.
[0590] To implement the present invention, devices and software work together to analyze and process user information in real time, ensuring the user's safety and security, and providing security alerts and guidance tailored to the user's emotional state, improving the user's quality of life.
[0591] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0592] Step 1:
[0593] Input: User's voice input, face image
[0594] Operation: The device captures the user's voice input and acquires a facial image using the camera function. The voice data and facial image data are sent to the server.
[0595] Output: Audio data, facial image data
[0596] Step 2:
[0597] Input: Audio data
[0598] Operation: The server converts the received voice data into text data using a speech recognition function. The speech is converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API).
[0599] Output: Character data
[0600] Step 3:
[0601] Input: Facial image data
[0602] How it works: The server sends facial image data to an emotion engine (e.g., Affectiva SDK) that analyzes the user's emotional state. The emotion engine recognizes the user's emotion (e.g., anxiety, anger, joy) from the facial image and sends the result back to the server.
[0603] Output: Emotion data
[0604] Step 4:
[0605] Input: Text data, emotion data
[0606] How it works: The server inputs text data and emotion data into a natural language processing engine (e.g., Google Cloud Natural Language API) to generate an appropriate response. The natural language processing engine understands the context based on the user's speech and emotion and creates the optimal response.
[0607] Output: Response data
[0608] Step 5:
[0609] Input: Response data
[0610] How it works: The server inputs the generated response data into a speech synthesis tool (e.g., Google Cloud Text-to-Speech API) to convert the text to speech. The speech synthesis tool then generates speech in the appropriate tone.
[0611] Output: Voice response data
[0612] Step 6:
[0613] Input: Voice response data
[0614] Operation: The device receives the voice response data from the server and provides it to the user as a voice. The device outputs the voice response using a speaker, giving the user a sense of security.
[0615] Output: The user receives a spoken response
[0616] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0617] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0618] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0619] [Second embodiment]
[0620] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0621] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0622] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0623] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0624] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0625] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0626] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0627] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0628] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0629] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0630] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0631] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0632] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable life. Specific embodiments of the system are described below.
[0633] This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[0634] 1. Conversation support system
[0635] Overview of conversation support
[0636] The server hosts a natural language processing (NLP) engine that understands voice input and generates appropriate responses to assist the user in the conversation.
[0637] Program processing
[0638] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0639] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0640] 3. Server: The received text data is analyzed using a natural language processing engine and an appropriate response is generated, such as "Today's weather is sunny."
[0641] 4. Server: Sends the generated response data to the terminal.
[0642] 5. Terminal: The received response data is converted into voice using a voice synthesis function and provided to the user.
[0643] Specific examples
[0644] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question in real time and responds, based on the current weather information, with, "Today's weather is sunny." This response is provided to the user by voice from the device.
[0645] 2. Cooking and meal follow-up system
[0646] Cooking and Meal Follow-up Overview
[0647] This is a system that optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking.
[0648] Program processing
[0649] 1. Terminal: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[0650] 2. Server: Based on the received data, it generates cooking methods and recipes that take health into consideration.
[0651] 3. Server: Sends the generated cooking instructions to the terminal.
[0652] 4. Terminal: Provides the received cooking instructions to the user via voice or screen display, and also sends specific operating instructions to the robot assistant as needed.
[0653] Specific examples
[0654] If a user requests a low-sugar dessert, the system will provide the optimal recipe, taking into account the user's health status (e.g., whether they are prone to diabetes), and even automatically prepare part of the dessert with the help of a robotic assistant if needed.
[0655] 3. Route Guidance and Public Transportation Integration System
[0656] Overview of route guidance during travel
[0657] This system presents the optimal route based on the user's current location and destination. It provides the optimal route by including real-time traffic information.
[0658] Program processing
[0659] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[0660] 2. Server: Obtains real-time traffic information, calculates the optimal route, and sends the results to the device.
[0661] 3. Terminal: Provides the received route information to the user via voice and screen display. Tracks the user's current location in real time while moving and updates the route as needed.
[0662] Specific examples
[0663] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[0664] By integrating these functions, the system aims to reduce the burden of caregiving and enable those receiving care to live more independently.
[0665] The processing flow will be explained below.
[0666] Conversation support processing flow
[0667] Step 1:
[0668] Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0669] Step 2:
[0670] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0671] Step 3:
[0672] Server: Analyzes the received text data using a natural language processing engine and generates an appropriate response based on its content, such as "Today's weather is sunny."
[0673] Step 4:
[0674] Server: Sends the generated response data to the terminal.
[0675] Step 5:
[0676] Terminal: The received response data is converted into voice using a speech synthesis function and provided to the user. The response is provided at a speed and voice quality that is easy for the user to hear.
[0677] Cooking and meal follow-up process flow
[0678] Step 1:
[0679] Device: The user voices or texts a request such as "I want to make a diet lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[0680] Step 2:
[0681] Server: Based on the received data, it generates cooking methods and recipes that take health conditions into consideration.
[0682] Step 3:
[0683] Server: Sends the generated cooking instructions to the terminal.
[0684] Step 4:
[0685] Terminal: Provides received cooking instructions to the user via voice or screen display, and also sends signals to instruct the robot assistant on cooking operations as needed.
[0686] Step 5:
[0687] Robot: Based on instructions from a terminal, the robot will begin specific cooking tasks, such as automatically chopping vegetables and putting food in a pot.
[0688] Processing flow of the system for linking route guidance and public transportation during travel
[0689] Step 1:
[0690] Device: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[0691] Step 2:
[0692] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[0693] Step 3:
[0694] Server: Sends the calculated route information to the terminal.
[0695] Step 4:
[0696] Terminal: Provides the received route information to the user via voice and screen display, providing specific guidance such as "turn right at the next traffic light."
[0697] Step 5:
[0698] Device: GPS information is constantly updated while the user is moving, tracking the user's current location in real time. If the user deviates from the route, the device recalculates the optimal route and provides new route guidance.
[0699] Step 6:
[0700] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[0701] Through specific processing at each step, this system aims to provide real-time support to users and reduce the overall burden of caregiving.
[0702] Example 1
[0703] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0704] In an aging society, reducing the burden of caregiving and providing a comfortable lifestyle are major challenges. In particular, advanced technology is required to provide consistent support for elderly people in conversation, eating, and mobility. However, existing systems have difficulty integrating multiple functions and providing them in a unified manner, limiting their practicality and effectiveness in the field. Therefore, there is a need to provide a comprehensive support system that allows elderly people to live independently.
[0705] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0706] In this invention, the server includes means for receiving a user's voice input and converting the voice data into text data, means for transmitting the text data to the server, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for transmitting the generated response to a terminal, means for converting the generated response into text and providing it to the user, means for receiving user request data and transmitting health status and ingredient data to the server, means for analyzing the user's health status and ingredient data and generating cooking methods and recipes, means for transmitting the generated cooking instructions to the terminal, means for providing the generated cooking instructions to the user by voice or on-screen display, means for obtaining the user's current location using GPS and transmitting destination information to the server, means for calculating an optimal route based on real-time traffic information, means for transmitting the calculated route information to the terminal, means for providing the calculated route information to the user by voice or on-screen display, and means for tracking the user's current location and updating the route during travel. This provides users with conversation assistance, cooking and meal follow-up, and route guidance during travel in an integrated manner, enabling elderly people to live independently and comfortably.
[0707] "Voice input" refers to the act of a user providing information to a system using voice.
[0708] "Voice data" refers to a digital recording of a user's voice input.
[0709] "Character data" is data in text format converted by voice recognition.
[0710] A "server" is a computer system that provides functionality such as processing and storing data and generating responses.
[0711] A "terminal" is a device that allows a user to interact with the system through an interface, and includes smartphones, tablets, PCs, etc.
[0712] A "natural language processing engine" is software that analyzes text data, understands its meaning, and generates responses.
[0713] A "response" is a reply or instruction generated by a server based on user input.
[0714] "Health status data" is data that includes information related to the user's health.
[0715] "Ingredient data" is data that includes information about ingredients owned by the user.
[0716] A "cooking method" is a procedure for preparing a dish using specific ingredients.
[0717] A "recipe" is a set of instructions that lists the ingredients and steps needed to make a particular dish.
[0718] "GPS" is a Global Positioning System for obtaining geographical location information.
[0719] "Real-time traffic information" means up-to-date data on current traffic conditions.
[0720] An "optimal route" is the most efficient route to reach a destination.
[0721] "Location tracking" refers to the real-time monitoring of a moving user's current location.
[0722] "Screen display" refers to the visual presentation of information on a terminal display.
[0723] MODE FOR CARRYING OUT THE INVENTION
[0724] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable lifestyle. This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[0725] Conversation support system
[0726] In a conversation support system, a server hosts a natural language processing (NLP) engine that understands the user's voice input and generates appropriate responses. This system uses the following hardware and software:
[0727] Hardware
[0728] Devices: PC, smartphone, tablet, etc.
[0729] Devices with built-in microphones
[0730] software
[0731] Speech recognition engine (e.g., Google Speech-to-Text, Amazon Transcribe)
[0732] Natural language processing engines (e.g., Google Cloud Natural Language API, OpenAI's GPT-3)
[0733] Text-to-speech software (e.g., Amazon Polly, Google Text-to-Speech)
[0734] Specific examples
[0735] When a user asks, "What's the weather like today?", the system captures the voice with the device's microphone, and the speech recognition engine converts it into text data. The converted text data is sent to the server and analyzed by the natural language processing engine. As a result, a response such as "The weather is sunny today" is generated and sent to the device. Finally, the device converts the received response data into speech using speech synthesis software and provides it to the user.
[0736] Prompt Sentence Examples
[0737] "When a user says, 'What's the weather like today?' generate an appropriate response."
[0738] Cooking and meal follow-up system
[0739] The cooking and meal follow-up system optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking. This system uses the following hardware and software:
[0740] Hardware
[0741] Smartphones and tablets
[0742] Smart appliances (smart ovens, smart refrigerators)
[0743] software
[0744] Health management app
[0745] Ingredient Management Software
[0746] Recipe Generation Engine
[0747] Robot Assistant Control Software
[0748] Specific examples
[0749] If a user requests, "I want to make a low-sugar dessert," the system will consider the user's health condition (for example, whether they have a tendency toward diabetes) and provide the optimal recipe. The smartphone receives the request via voice or text and sends it to the server. The server generates the optimal recipe based on the received data and pre-registered health data and sends it to the smartphone. Furthermore, if necessary, the robot assistant will automatically prepare part of the dessert.
[0750] Prompt Sentence Examples
[0751] "If a user requests, 'I want to make a dessert with less sugar,' generate a recipe that takes their health into consideration."
[0752] Route guidance and public transport linkage system
[0753] The route guidance and public transport integration system for travel is a system that presents the optimal travel route based on the user's current location and destination. It provides the optimal route by including real-time traffic information. This system uses the following hardware and software.
[0754] Hardware
[0755] GPS-enabled devices (smartphones, tablets)
[0756] software
[0757] Real-time traffic information API (e.g., Google Maps API, HERE API)
[0758] Route Calculation Algorithm
[0759] Location Tracking Software
[0760] Specific examples
[0761] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." The smartphone obtains the user's current location using GPS and sends a request to the server. The server calculates a route based on real-time traffic information and sends it to the smartphone. The smartphone then provides the calculated route to the user by voice or on-screen display, and tracks the user's current location in real time while traveling, updating the route as needed.
[0762] Prompt Sentence Examples
[0763] "When a user requests 'Take me to the station,' generate a plan that provides the optimal route."
[0764] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0765] Conversation support system
[0766] Program processing
[0767] Step 1:
[0768] The user gives a voice input of "What's the weather like today?" The device captures this voice input using the built-in microphone. The input is the user's voice data, and the output is the captured voice data. The device collects voice data in real time.
[0769] Step 2:
[0770] The device sends the captured voice data to a voice recognition engine and converts it into text data. Specifically, the voice data is sent to a voice recognition service on the cloud, and text data is obtained. The input here is voice data, and the output is converted text data. The device performs the operation of converting voice data into text data.
[0771] Step 3:
[0772] The terminal sends the converted character data to the server. The data is sent securely using a secure protocol (such as HTTPS). The input here is character data, and the output is the character data sent to the server. The terminal performs the data sending operation.
[0773] Step 4:
[0774] The server inputs the received text data into a natural language processing engine (NLP engine) for analysis. The NLP engine analyzes the text data and generates an appropriate response. The input here is text data, and the output is the generated response data. The server performs the operations of analyzing the text data and generating a response.
[0775] Step 5:
[0776] The server sends the generated response data to the terminal. Again, communication is performed using a secure protocol. The input here is the response data, and the output is the response data sent to the terminal. The server performs the data transmission operation.
[0777] Step 6:
[0778] The terminal converts the received response data into speech using a speech synthesis engine. The speech synthesis engine is used to convert text data into speech data, which is then provided to the user through a speaker. The input here is the response data, and the output is the generated speech data. The terminal then performs the operation of playing back the speech data.
[0779] Cooking and meal follow-up system
[0780] Program processing
[0781] Step 1:
[0782] The user inputs a request into the device by voice or text, such as "I want to make a dietary lunch." The device captures this input. The input is the user's request data, and the output is the captured request data. The device performs the operation of collecting voice and text data.
[0783] Step 2:
[0784] The terminal sends the request data, pre-registered health condition data, and ingredient information to the server. The data is sent using a secure protocol. The input is the request data and the attached health condition data and ingredient information, and the output is the data sent to the server. The terminal performs the data transmission operation.
[0785] Step 3:
[0786] The server analyzes the received data and generates appropriate cooking methods and recipes that take health status into consideration. The input here is the request data, health status data, and ingredient information, and the output is the generated recipe data. The server performs the data analysis and recipe generation operations.
[0787] Step 4:
[0788] The server sends the generated recipe data to the terminal. Communication is again performed using a secure protocol. The input is the recipe data, and the output is the recipe data sent to the terminal. The server performs the data transmission operation.
[0789] Step 5:
[0790] The terminal provides the received recipe data to the user by voice or screen display. The input here is recipe data, and the output is recipe information provided to the user. The terminal performs the operations of displaying information and playing voice.
[0791] Step 6:
[0792] If necessary, the terminal sends specific cooking operation instructions to the robot assistant. The input is cooking operation instruction data, and the output is instruction data sent to the robot assistant. The terminal performs the operation of sending instructions.
[0793] Route guidance and public transport linkage system
[0794] Program processing
[0795] Step 1:
[0796] The user inputs a request into the terminal by voice or text, such as "Take me to the station." The terminal captures this input. The input is the user's request data, and the output is the captured request data. The terminal performs the operation of collecting voice and text data.
[0797] Step 2:
[0798] The device uses the built-in GPS to obtain the user's current location. The input here is GPS data, and the output is the obtained current location data. The device performs the operation of collecting location data.
[0799] Step 3:
[0800] The device sends the request data and current location data to the server. The data is sent using a secure protocol. The input is the request data and current location data, and the output is the data sent to the server. The device performs the data sending operation.
[0801] Step 4:
[0802] The server obtains real-time traffic information and calculates the optimal route. The input here is the current location data and real-time traffic information, and the output is the calculated route data. The server performs data analysis and route calculation.
[0803] Step 5:
[0804] The server sends the calculated route data to the terminal. Communication is again performed using a secure protocol. The input is the route data, and the output is the route data sent to the terminal. The server performs the data transmission operation.
[0805] Step 6:
[0806] The terminal provides the received route data to the user by voice or screen display. The input here is the route data, and the output is the route information provided to the user. The terminal performs the operations of displaying information and playing voice.
[0807] Step 7:
[0808] While moving, the device uses GPS to track the user's current location and updates the route as needed. The input here is GPS data and real-time traffic information, and the output is updated route data. The device performs location tracking and route update operations.
[0809] (Application example 1)
[0810] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0811] This invention relates to a system for supporting and ensuring the safety of specific groups, such as the elderly, in their daily lives. In particular, the system aims to reduce the burden on caregivers and support the independent living of those receiving care by utilizing the user's voice input and location information to provide comprehensive support tailored to a wide range of needs, including conversation assistance, cooking and meal follow-up, and emergency response. There is also a need for a system that can respond quickly and appropriately in emergencies.
[0812] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0813] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for acquiring the user's current location, means for transmitting the user's location information to the server, and means for analyzing the user's voice in an emergency and transmitting an alert together with the location information, thereby supporting the user's daily life and enabling a prompt and appropriate response, particularly in an emergency.
[0814] "Voice input" refers to data obtained by a device capturing words spoken by a user.
[0815] "Voice data" is data that digitally represents voice input obtained from a user.
[0816] "Character data" is data obtained by analyzing voice data and converting the content into text format.
[0817] A "natural language processing engine" is a program or system that analyzes text data and generates appropriate responses.
[0818] "Response generation" refers to the process of using a natural language processing engine to create an appropriate response from text data.
[0819] "Speech synthesis" is the technique of converting the generated response back into speech form.
[0820] "Current location of user" refers to the geographic location where the user is currently located.
[0821] "GPS" is a satellite positioning system that identifies a user's current location.
[0822] "Real-time traffic information" is dynamic information that shows the current traffic situation.
[0823] An "optimal route" refers to the most efficient route from the current location to the destination.
[0824] "Emergency speech analysis" is the process of identifying specific speech inputs that indicate an emergency.
[0825] "Transmitting location information" means sending data indicating the user's current location to a server or other receiving device.
[0826] "Sending an alert" means communicating a warning message based on a specific condition.
[0827] This invention is a system for supporting specific groups, such as the elderly, in their daily lives and ensuring their safety, and is realized by each of the entities, the server, the terminal, and the user, fulfilling their specific roles. The following describes in detail the mode for carrying out this invention.
[0828] Embodiment of conversation support system
[0829] The server hosts a natural language processing engine (NLP engine). The device captures the user's voice input and converts this voice data into text data. The converted text data is sent to the server, where the NLP engine generates an appropriate response. The generated response is sent back to the device, where it is converted into speech and provided to the user. This system allows users to obtain a variety of information through natural conversation.
[0830] Hardware / Software: A microphone is used for voice input, the speech_recognition library is used for voice recognition, an NLP engine is used for natural language processing, and the pyttsx3 library is used for speech synthesis.
[0831] Example: When a user asks, "What's the weather like today?", the system responds by saying, "The weather is sunny today."
[0832] Cooking and meal follow-up system
[0833] The device receives input data from the user regarding their health condition and ingredients, and sends this data to the server. Based on the received data, the server generates cooking methods and recipes optimized for the user's health condition. The generated cooking instructions are provided to the user via the device via voice or screen display. This allows users to easily cook in a way that takes their health into consideration.
[0834] Hardware and software: Smartphones and tablets are used to acquire and transmit data. AI models are used to generate recipes, and existing speech synthesis libraries and display technologies are used for voice and screen display.
[0835] Example: When a user requests to make a low-sugar dessert, the system provides the optimal recipe based on the user's health status.
[0836] Route guidance and public transport linkage system for travel
[0837] The device acquires the user's current location and sends destination information to the server. The server calculates the optimal route based on real-time traffic information and sends that information to the device. The device then provides the calculated route information to the user via voice and on-screen display. In addition, the device has the ability to analyze the user's voice in an emergency and send an alert along with location information.
[0838] Hardware and software: GPS devices, real-time traffic information platforms, and speech recognition and speech synthesis libraries are used.
[0839] Example: If a user yells "Help!", the system interprets the audio as an emergency alert and sends information, including the user's current location, to emergency contacts.
[0840] Prompt Sentence Examples
[0841] "How's the weather today?"
[0842] "I want to make a dessert with less sugar."
[0843] "help me!"
[0844] In this way, the present invention utilizes the user's voice and location information to provide a wide range of support for daily life, thereby reducing the burden on elderly people and their caregivers.
[0845] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0846] Processing steps of the conversation support system
[0847] Step 1:
[0848] The terminal obtains the user's voice input.
[0849] Specifically, when a user says, "What's the weather like today?", the device's microphone captures the voice. The input is voice data.
[0850] Step 2:
[0851] The terminal converts the captured voice data into text data.
[0852] Specifically, it uses a speech recognition library (e.g., speech_recognition) to convert voice data into text format. The input is voice data, and the output is text data.
[0853] Step 3:
[0854] The terminal transmits the converted character data to the server.
[0855] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[0856] Step 4:
[0857] The server analyzes the received text data using a natural language processing engine.
[0858] Specifically, an NLP engine (for example, a custom NLP server) analyzes the text data and generates an appropriate response to the question, "What's the weather like today?" The input is the text data, and the output is the analysis result and response data.
[0859] Step 5:
[0860] The server sends the generated response data to the terminal.
[0861] Specifically, the generated response data is sent to the terminal via an HTTP request. The input is the response data, and the output is data sent to the terminal.
[0862] Step 6:
[0863] The terminal converts the received response data into voice and provides it to the user.
[0864] Specifically, it uses a speech synthesis library (e.g., pyttsx3) to convert text-based response data into speech and provides it to the user through a speaker. The input is the response data, and the output is speech output.
[0865] Emergency response system processing steps
[0866] Step 1:
[0867] The terminal obtains the user's voice input.
[0868] Specifically, when a user shouts "Help!", the device's microphone captures the voice. The input is voice data.
[0869] Step 2:
[0870] The terminal converts the captured voice data into text data.
[0871] Specifically, it uses a speech recognition library to convert voice data into text format. The input is voice data and the output is text data.
[0872] Step 3:
[0873] The terminal transmits the converted character data to the server.
[0874] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[0875] Step 4:
[0876] The server analyzes the text data and identifies it as an urgent message.
[0877] Specifically, the NLP engine recognizes the emergency message "Help!" and prepares an appropriate response. The input is text data, and the output is emergency alert data.
[0878] Step 5:
[0879] The device obtains the user's current location using GPS.
[0880] Specifically, the GPS module of the device acquires the current location information. The input is a signal from the GPS device, and the output is location information data.
[0881] Step 6:
[0882] The server sends location information and emergency alerts to emergency contacts.
[0883] Specifically, location information and an emergency message are sent to emergency contacts (e.g., family members or emergency services) via an HTTP request. The input is the emergency alert data and location data, and the output is the transmission of the data to the contacts.
[0884] Step 7:
[0885] The terminal provides a voice message to reassure the user.
[0886] Specifically, it uses a speech synthesis library to generate a message such as "Rescue is currently being called. Please rest assured," and provides it to the user through a speaker. The input is the data for the reassurance message, and the output is voice output.
[0887] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0888] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live a comfortable life. Specific embodiments of this invention are described below.
[0889] In this system, each of the server, terminal, and user plays a role in recognizing the user's emotions and providing optimal responses and support to the user's mental and physical needs.
[0890] 1. Conversation support system
[0891] Overview of conversation support
[0892] The server hosts an emotion engine and a natural language processing (NLP) engine, which recognizes emotions from the user's voice input and generates appropriate responses to provide to the user.
[0893] Program processing
[0894] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0895] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0896] 3. Server: Analyzes the received text data using an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).
[0897] 4. Server: Uses a natural language processing engine to generate an appropriate response based on context and sentiment, such as "The weather is sunny today."
[0898] 5. Server: Sends a tone based on the generated response data and emotional state to the terminal.
[0899] 6. Terminal: The received response data is converted into voice with an appropriate tone using a voice synthesis function and provided to the user.
[0900] Specific examples
[0901] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question, and if it recognizes that the user is feeling "depressed," it responds by adding encouraging words such as, "Cheer up, it's sunny today, why don't you go outside?"
[0902] 2. Cooking and meal follow-up system
[0903] Cooking and Meal Follow-up Overview
[0904] This system provides optimal cooking instructions based on the user's health and emotional state.
[0905] Program processing
[0906] 1. Device: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status, ingredient information, and emotional data, is sent to the server.
[0907] 2. Server: Based on the received data, it generates cooking methods and recipes that take into account the health and emotional state of the user.
[0908] 3. Server: Sends the generated cooking instructions and related emotional support information to the terminal.
[0909] 4. Terminal: Provides the received cooking instructions and emotional support information (e.g., suggestions for ingredients with a relaxing effect when the user is feeling stressed) to the user via voice or on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[0910] Specific examples
[0911] If a user requests "I want to make a healthy breakfast," the system will provide a relaxing recipe using oatmeal based on the user's health status and feelings of "tension." Furthermore, if a robotic assistant is needed, it will automate some of the oatmeal cooking process.
[0912] 3. Route Guidance and Public Transportation Integration System
[0913] Overview of route guidance during travel
[0914] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[0915] Program processing
[0916] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[0917] 2. Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination, taking into account the emotional state.
[0918] 3. Server: Sends the calculated route information and emotional support information (for example, calming voice guidance if the user is feeling anxious) to the device.
[0919] 4. Device: Provides the user with the received route information and emotional support information via voice and screen display. It also tracks the user's current location in real time while moving and updates the route as needed.
[0920] Specific examples
[0921] If the system detects that the user is "nervous" while traveling from home to the station, it calculates the shortest and safest route based on the latest traffic information and provides specific instructions in a gentle tone, such as "Turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[0922] As described above, the AI system combined with the emotion engine provides services that take into account the user's mental state, aiming to reduce the burden of caregiving and improve the user's quality of life.
[0923] The processing flow will be explained below.
[0924] Conversation support processing flow
[0925] Step 1:
[0926] User: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[0927] Step 2:
[0928] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[0929] Step 3:
[0930] Server: Analyzes the user's emotional state using the emotion engine and sends the results of this analysis to the natural language processing engine.
[0931] Step 4:
[0932] Server: Uses a natural language processing engine to generate appropriate responses based on context and sentiment.
[0933] Step 5:
[0934] Server: Sends tone information based on the generated response data and emotional state to the terminal.
[0935] Step 6:
[0936] Terminal: The terminal converts the received response data into speech with an appropriate tone using a speech synthesis function and provides it to the user.
[0937] Cooking and meal follow-up process flow
[0938] Step 1:
[0939] User: The user inputs a request by voice or text, such as "I want to make a dietary lunch." The device then sends this information, along with previously entered health status and ingredient information, to the server.
[0940] Step 2:
[0941] Server: Based on the received information about the user's health condition and ingredients, the emotion engine analyzes the user's emotional state.
[0942] Step 3:
[0943] Server: Generates cooking methods and recipes that take into account health and emotional states.
[0944] Step 4:
[0945] Server: Sends the generated cooking instructions and emotional support information to the terminal.
[0946] Step 5:
[0947] Terminal: Provides the received cooking instructions and emotional support information to the user via voice and on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[0948] Step 6:
[0949] Robot: Based on instructions from a terminal, the robot begins specific cooking tasks. For example, it automatically performs basic operations such as chopping vegetables and putting food in a pot.
[0950] Processing flow of the system for linking route guidance and public transportation during travel
[0951] Step 1:
[0952] User: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[0953] Step 2:
[0954] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[0955] Step 3:
[0956] Server: Based on the received user's emotional state, the server adjusts the guidance method (tone and pace) for the optimal route.
[0957] Step 4:
[0958] Server: Sends the calculated route information and emotional support information to the terminal.
[0959] Step 5:
[0960] Terminal: Provides the received route information and emotional support information to the user via voice and on-screen display. For example, it provides specific guidance such as "Turn right at the next traffic light," but changes the tone depending on the user's emotional state.
[0961] Step 6:
[0962] Device: GPS information is constantly updated while moving, tracking the user's current location in real time. Each time, the server recalculates the optimal route based on the user's emotional state and provides new route guidance.
[0963] Step 7:
[0964] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[0965] This detailed processing flow allows the system to provide more personalized assistance while taking into account the user's emotional state, thereby reducing the burden of caregiving and enabling users to live a comfortable and efficient life.
[0966] Example 2
[0967] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0968] In recent years, the burden of caregiving has increased with the progress of the aging society. There is also a demand for support to maintain users' mental and physical health. However, conventional AI systems do not take into account the user's emotional state, limiting the user experience. Furthermore, even in health management and mobility assistance, they can only make suggestions that ignore the user's emotional state, making it difficult to provide accurate support.
[0969] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the user's emotional state from the text data using an emotion engine, means for generating an appropriate response based on the context and the emotional state using a natural language processing engine, and means for converting the generated response into voice and providing it to the user. This makes it possible to provide appropriate responses and support according to the user's emotional state.
[0970] "Audio input" refers to an audio signal emitted by a user through an input device such as a microphone.
[0971] "Audio Data" means data in digital form obtained from audio input.
[0972] "Character data" refers to text-format data obtained by analyzing voice data using voice recognition technology.
[0973] An "emotion engine" is software or hardware that analyzes text data and estimates the user's emotional state.
[0974] An "emotional state" is a mental state that a user is feeling, and examples include happiness, sadness, anger, etc.
[0975] A "natural language processing engine" is software that analyzes text data, understands meaning and context, and generates appropriate responses.
[0976] A "response" is a reply or instruction to the user generated by a natural language processing engine.
[0977] "Speech synthesis" is a technology that converts text data into audio data and turns it into audio that can be played through speakers, etc.
[0978] "Health status" refers to information about the user's current physical health, including information about illnesses, symptoms, and physical conditions.
[0979] "Ingredient information" refers to data relating to the types and amounts of ingredients available to the user.
[0980] A "cooking method" is the procedure or process of preparing a dish using ingredients.
[0981] A "recipe" is a document that describes the ingredients and steps for making a particular dish.
[0982] A "location information system" is a system that obtains a user's current location using technology such as GPS.
[0983] "Real-time traffic information" refers to information about current traffic conditions and transportation options.
[0984] An "optimal route" is the most efficient route for a user to travel from their current location to their destination.
[0985] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live more comfortable lives. Specific embodiments of this invention will be described in detail below.
[0986] 1. Conversation support system
[0987] System Configuration
[0988] In this system, users ask questions or make requests by voice, and the server processes them and provides a voice response. The main components are as follows:
[0989] Terminal: A device that captures the user's voice and sends the voice data to the server. It uses the Google Speech-to-Text API for voice recognition.
[0990] Server: Acquires text data and analyzes it using an emotion engine (e.g., IBM Watson Tone Analyzer). OpenAI's GPT-3 natural language processing engine is used.
[0991] Speech synthesis engine: Converts the generated text response into speech using Amazon Polly.
[0992] Specific examples
[0993] When a user asks, "What's the weather like today?", the voice data captured by the device is converted into text data and sent to the server. The server uses an emotion engine to recognize the emotional state of "interested" and generates an appropriate response using a natural language processing engine. For example, it might respond, "The weather is sunny today." This is then converted into speech using a speech synthesis engine and provided to the user via the device.
[0994] Prompt Sentence Examples
[0995] User: "What's the weather like today?"
[0996] system:
[0997] 1. The server converts the speech into text and analyzes it using an emotion engine.
[0998] 2. If the emotion is perceived as "interested," generate an appropriate response.
[0999] 2. Cooking and meal follow-up system
[1000] System Configuration
[1001] This system provides optimal cooking instructions based on the user's health and emotional state.
[1002] Device: Captures user requests and sends them to the server. The voice recognition function uses the Google Speech-to-Text API.
[1003] Server: Analyzes health status, ingredient information, and emotional state to generate cooking methods and recipes. IBM Watson Tone Analyzer is used as the emotion engine, and OpenAI's GPT-3 is used to generate recipes.
[1004] Robot assistant: A device that automates cooking operations as needed.
[1005] Specific examples
[1006] If a user requests "I want to make a dietary lunch," the device sends this information to the server. The server uses its emotion engine to recognize that the user is "feeling stressed" and generates a cooking method using ingredients that have a relaxing effect. The generated recipe is for oatmeal and salad and is provided to the user via voice and on-screen display. If necessary, some of the cooking can be automated by a robotic assistant.
[1007] Prompt Sentence Examples
[1008] User: "I want to make a diet lunch."
[1009] system:
[1010] 1. The server receives the request and user data and generates an appropriate recipe.
[1011] 2. The recipe uses oatmeal and provides audio instructions.
[1012] 3. Route Guidance and Public Transportation Integration System
[1013] System Configuration
[1014] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[1015] Terminal: Receives the user's request, obtains the current location using GPS, and sends it to the server.
[1016] Server: Obtains real-time traffic information and calculates the optimal route. IBM Watson Tone Analyzer is used as the emotion engine, and Google Maps API is used to obtain traffic information.
[1017] Speech synthesis engine: Converts route directions into speech.
[1018] Specific examples
[1019] When a user requests "Guide me to the station," the device obtains the user's current location, sets "station" as the destination, and sends the information to the server. The server then uses its emotion engine to recognize the user as "stressed" and calculates the optimal route to reassure the user. For example, the server uses a speech synthesis engine to convert specific instructions, such as "Turn right at the next traffic light," into voice and provides the information to the user via the device.
[1020] Prompt Sentence Examples
[1021] User: "Take me to the station."
[1022] system:
[1023] 1. The server calculates a route based on the current location and destination, and provides guidance according to the user's emotional state.
[1024] 2. If they are feeling anxious, guide them with a reassuring route and a gentle tone.
[1025] As described above, the present invention can provide support in various situations while taking into consideration the emotional state of the user, reduce the burden of caregiving, and improve the quality of life of the user.
[1026] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1027] Conversation support system processing flow
[1028] Step 1:
[1029] Description: The user provides voice input.
[1030] Specific action: The user speaks to the device, "What's the weather like today?"
[1031] Input: User's voice input.
[1032] Output: The captured audio data.
[1033] Step 2:
[1034] Description: The device receives voice data and converts it into text data using voice recognition.
[1035] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[1036] Input: Captured audio data.
[1037] Output: The converted text data (e.g., "What's the weather like today?").
[1038] Step 3:
[1039] Description: The terminal sends the converted character data to the server.
[1040] Specific operation: The terminal sends an HTTP request to the server and sends character data.
[1041] Input: The converted character data.
[1042] Output: Character data sent to the server.
[1043] Step 4:
[1044] Description: The server receives text data and analyzes it with the emotion engine.
[1045] Specific operation: The server uses IBM Watson Tone Analyzer to analyze the user's emotional state (e.g., "interested") from the text data.
[1046] Input: The character data sent.
[1047] Output: Parsed emotional state.
[1048] Step 5:
[1049] Description: The server uses a natural language processing engine to generate appropriate responses based on context and emotional state.
[1050] Specific operation: The server calls OpenAI's GPT-3 and generates an appropriate response sentence using the text data and emotional state as input (e.g., "The weather is sunny today.").
[1051] Input: Text data, emotional state.
[1052] Output: The generated response sentence.
[1053] Step 6:
[1054] Description: The server sends the generated response to the terminal.
[1055] Specific operation: The server sends an HTTP response to the terminal and sends the generated response text.
[1056] Input: The generated response sentence.
[1057] Output: The response sent to the terminal.
[1058] Step 7:
[1059] Description: The response received by the terminal is converted into speech using the speech synthesis function.
[1060] Specific operation: The device uses Amazon Polly to convert the response into voice data.
[1061] Input: The response received.
[1062] Output: The converted audio data.
[1063] Step 8:
[1064] Description: The device plays the converted audio data to the user.
[1065] Specific operation: Play audio data through the device speaker.
[1066] Input: The converted audio data.
[1067] Output: A spoken response to the user.
[1068] Cooking and meal follow-up system processing flow
[1069] Step 1:
[1070] Description: A user enters a cooking request.
[1071] Specific operation: The user speaks to the device, saying, "I want to make a dietary lunch."
[1072] Input: User's voice input.
[1073] Output: The captured audio data.
[1074] Step 2:
[1075] Description: The device receives voice data and converts it into text data using voice recognition.
[1076] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[1077] Input: Captured audio data.
[1078] Output: The converted text (e.g. "I want to make a dietary lunch").
[1079] Step 3:
[1080] Description: The device acquires health status, food ingredient information, and emotional state and sends them to the server.
[1081] Specific operation: The device obtains pre-registered health status, food information, and emotional state, and sends an HTTP request to the server.
[1082] Input: Text data, health status, food information, emotional state.
[1083] Output: The data sent to the server.
[1084] Step 4:
[1085] Description: The server generates cooking instructions and emotional support information based on the data received.
[1086] Specific operation: Using an emotion engine (e.g., IBM Watson Tone Analyzer) and a natural language processing engine (e.g., OpenAI's GPT-3), it generates appropriate cooking methods, recipes, and emotional support information.
[1087] Input: health status, food information, emotional state.
[1088] Output: Generated cooking instructions and emotional support information.
[1089] Step 5:
[1090] Description: The server sends the generated cooking instructions and emotional support information to the terminal.
[1091] Specific operation: The server sends an HTTP response to the terminal, conveying instructions and information.
[1092] Input: Generated cooking instructions and emotional support information.
[1093] Output: Cooking instructions and emotional support information sent to the device.
[1094] Step 6:
[1095] Description: The device provides the user with cooking instructions and emotional support information received.
[1096] How it works: The device uses speech synthesis to convert instructions into voice, provides information on the screen, and issues instructions to the robot assistant as needed.
[1097] Input: cooking instructions and emotional support information.
[1098] Output: Voice and visual instructions, instruction signals to the robot assistant.
[1099] Processing flow of the system for linking route guidance and public transportation during travel
[1100] Step 1:
[1101] Description: User requests directions.
[1102] Specific operation: The user speaks to the terminal, saying, "Take me to the station."
[1103] Input: User's voice input.
[1104] Output: The captured audio data.
[1105] Step 2:
[1106] Description: The device receives voice data and converts it into text data using voice recognition.
[1107] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[1108] Input: Captured audio data.
[1109] Output: The converted text data (e.g. "Take me to the station").
[1110] Step 3:
[1111] Description: The device obtains its current location using GPS and sends destination information and emotional state to the server.
[1112] Specific operation: The device acquires GPS data and sends an HTTP request to the server along with the user's emotional state information.
[1113] Input: Text data, current location information, destination information, emotional state.
[1114] Output: The data sent to the server.
[1115] Step 4:
[1116] Description: The server calculates the optimal route based on real-time traffic information and emotional state.
[1117] Specific operation: The server uses the Google Maps API to calculate the optimal route, taking into account the traffic information obtained and the emotional state analyzed by the emotion engine.
[1118] Input: current location information, destination information, emotional state.
[1119] Output: Calculated optimal route information.
[1120] Step 5:
[1121] Description: The server sends the calculated route information and emotional support information to the terminal.
[1122] Specific operation: The server sends an HTTP response to the device, conveying route information and emotional support information.
[1123] Input: Calculated optimal route information, emotional support information.
[1124] Output: Route information and emotional support information sent to the device.
[1125] Step 6:
[1126] Description: Provides the user with route information and emotional support information received by the terminal.
[1127] Specific operation: The device uses a speech synthesis function to convert route guidance into voice and also provides route information on the screen.
[1128] Input: Received route information, emotional support information.
[1129] Output: Route guidance by voice and on-screen display.
[1130] Step 7:
[1131] Description: The device tracks the user's location in real time and updates the route as needed.
[1132] How it works: The device uses GPS to continuously obtain its current location while moving, and resends the information to the server as needed to receive new route guidance.
[1133] Input: current location information, route information.
[1134] Output: Updated route guidance information.
[1135] This is the specific processing flow of this system. At each step, the system converts voice data, analyzes emotions, processes natural language, calculates routes, and performs other operations based on the user's input, providing appropriate responses and support.
[1136] (Application example 2)
[1137] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1138] Conventional security and guidance systems tend to respond in a uniform manner without considering the user's emotional state, making it impossible to provide optimal alerts and guidance to users. Furthermore, even in situations where users feel anxious or scared, appropriate responses are often delayed, reducing the user's sense of security. Therefore, there is a need for a system that provides effective responses and a sense of security that takes the user's emotional state into account.
[1139] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1140] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for capturing the user's face and recognizing emotions using an emotion engine, and means for providing appropriate security alerts and guidance based on the recognized emotions. This enables optimal responses that take the user's emotional state into consideration, thereby improving the user's sense of security.
[1141] "Voice input" refers to data that is input by the user using voice.
[1142] "Voice data" refers to data that represents in digital form the voice input by the user.
[1143] "Character data" refers to text-format data obtained by converting voice data into characters.
[1144] A "natural language processing engine" is a processing engine that analyzes text data, understands its meaning, and generates an appropriate response.
[1145] An "emotion engine" is an engine that has the function of determining emotions from the user's voice, facial expressions, etc.
[1146] "Response generation" is the process of creating appropriate responses or guidance based on the user's input data and their emotional state.
[1147] "Speech synthesis" is a technology that outputs text data as speech.
[1148] "Emotion recognition" is the process of determining a user's emotions from their facial expressions, voice, etc.
[1149] A "security alert" is a notification that warns or warns users to ensure their safety.
[1150] "Guidance" refers to instructions that instruct the user on appropriate actions or responses.
[1151] "Capturing a user's face" means obtaining an image of the user's face using a device such as a camera.
[1152] A "location information system" is a system that obtains a user's current location using GPS or other means.
[1153] "Real-time traffic information" refers to data that obtains the latest information on traffic conditions in real time.
[1154] An "optimal route" is a route that allows a user to reach a destination most efficiently.
[1155] The system of the present invention is an AI system that combines an emotion engine and a natural language processing engine to provide optimal support for the user's mental and physical needs. The system allows the server, terminal, and user to play their respective roles, improving user security.
[1156] Hardware and Software
[1157] Hardware
[1158] Smart glasses (a device with a general camera function)
[1159] Smartphone (Android or iOS device)
[1160] Head-mounted display (general VR device)
[1161] software
[1162] Emotion engine (e.g. Affectiva SDK)
[1163] Natural language processing engine (e.g. Google Cloud Natural Language API)
[1164] Image and voice recognition tools (e.g., Google Cloud Vision API, Speech-to-Text API)
[1165] Speech synthesis tools (e.g., Google Cloud Text-to-Speech API)
[1166] Data processing and calculation
[1167] server
[1168] The server performs the following processes. First, it receives the user's voice input and facial image from the device. The voice data is converted into text data using a voice recognition function, and the text data and facial image are sent in parallel to the emotion engine. The emotion engine analyzes the user's emotional state and provides this data to the natural language processing engine. The natural language processing engine generates an appropriate response based on the text data and emotional data. Finally, the generated response data is converted into voice data using a voice synthesis tool and sent to the device.
[1169] Terminal
[1170] The device performs the following processes: First, it captures voice input and facial images from the user and sends them to the server. Then it receives response data returned from the server and outputs voice to the user at the appropriate time. If the user feels anxious, it provides security alerts and guidance based on emotion recognition.
[1171] Specific examples
[1172] scenario
[1173] When a user is walking down a street at night, the AI detects that they are feeling anxious.
[1174] Response: The smart glasses say, "User, it's okay. There's a station nearby. Would you like to contact the fastest security company right away?"
[1175] Prompt Sentence Examples
[1176] Design an AI assistant with an emotion engine that recognizes the user's emotional state and provides appropriate security alerts and guidance. For a specific scenario, generate a reassuring message when the user is feeling anxious. The devices used could be smart glasses, a smartphone, or a head-mounted display.
[1177] To implement the present invention, devices and software work together to analyze and process user information in real time, ensuring the user's safety and security, and providing security alerts and guidance tailored to the user's emotional state, improving the user's quality of life.
[1178] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1179] Step 1:
[1180] Input: User's voice input, face image
[1181] Operation: The device captures the user's voice input and acquires a facial image using the camera function. The voice data and facial image data are sent to the server.
[1182] Output: Audio data, facial image data
[1183] Step 2:
[1184] Input: Audio data
[1185] Operation: The server converts the received voice data into text data using a speech recognition function. The speech is converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API).
[1186] Output: Character data
[1187] Step 3:
[1188] Input: Facial image data
[1189] How it works: The server sends facial image data to an emotion engine (e.g., Affectiva SDK) that analyzes the user's emotional state. The emotion engine recognizes the user's emotion (e.g., anxiety, anger, joy) from the facial image and sends the result back to the server.
[1190] Output: Emotion data
[1191] Step 4:
[1192] Input: Text data, emotion data
[1193] How it works: The server inputs text data and emotion data into a natural language processing engine (e.g., Google Cloud Natural Language API) to generate an appropriate response. The natural language processing engine understands the context based on the user's speech and emotion and creates the optimal response.
[1194] Output: Response data
[1195] Step 5:
[1196] Input: Response data
[1197] How it works: The server inputs the generated response data into a speech synthesis tool (e.g., Google Cloud Text-to-Speech API) to convert the text to speech. The speech synthesis tool then generates speech in the appropriate tone.
[1198] Output: Voice response data
[1199] Step 6:
[1200] Input: Voice response data
[1201] Operation: The device receives the voice response data from the server and provides it to the user as a voice. The device outputs the voice response using a speaker, giving the user a sense of security.
[1202] Output: The user receives a spoken response
[1203] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1204] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1205] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1206] [Third embodiment]
[1207] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1208] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1209] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1210] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1211] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1212] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1213] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1214] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1215] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1216] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1217] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1218] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1219] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable life. Specific embodiments of the system are described below.
[1220] This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[1221] 1. Conversation support system
[1222] Overview of conversation support
[1223] The server hosts a natural language processing (NLP) engine that understands voice input and generates appropriate responses to assist the user in the conversation.
[1224] Program processing
[1225] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[1226] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[1227] 3. Server: The received text data is analyzed using a natural language processing engine and an appropriate response is generated, such as "Today's weather is sunny."
[1228] 4. Server: Sends the generated response data to the terminal.
[1229] 5. Terminal: The received response data is converted into voice using a voice synthesis function and provided to the user.
[1230] Specific examples
[1231] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question in real time and responds, based on the current weather information, with, "Today's weather is sunny." This response is provided to the user by voice from the device.
[1232] 2. Cooking and meal follow-up system
[1233] Cooking and Meal Follow-up Overview
[1234] This is a system that optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking.
[1235] Program processing
[1236] 1. Terminal: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[1237] 2. Server: Based on the received data, it generates cooking methods and recipes that take health into consideration.
[1238] 3. Server: Sends the generated cooking instructions to the terminal.
[1239] 4. Terminal: Provides the received cooking instructions to the user via voice or screen display, and also sends specific operating instructions to the robot assistant as needed.
[1240] Specific examples
[1241] If a user requests a low-sugar dessert, the system will provide the optimal recipe, taking into account the user's health status (e.g., whether they are prone to diabetes), and even automatically prepare part of the dessert with the help of a robotic assistant if needed.
[1242] 3. Route Guidance and Public Transportation Integration System
[1243] Overview of route guidance during travel
[1244] This system presents the optimal route based on the user's current location and destination. It provides the optimal route by including real-time traffic information.
[1245] Program processing
[1246] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[1247] 2. Server: Obtains real-time traffic information, calculates the optimal route, and sends the results to the device.
[1248] 3. Terminal: Provides the received route information to the user via voice and screen display. Tracks the user's current location in real time while moving and updates the route as needed.
[1249] Specific examples
[1250] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[1251] By integrating these functions, the system aims to reduce the burden of caregiving and enable those receiving care to live more independently.
[1252] The processing flow will be explained below.
[1253] Conversation support processing flow
[1254] Step 1:
[1255] Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[1256] Step 2:
[1257] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[1258] Step 3:
[1259] Server: Analyzes the received text data using a natural language processing engine and generates an appropriate response based on its content, such as "Today's weather is sunny."
[1260] Step 4:
[1261] Server: Sends the generated response data to the terminal.
[1262] Step 5:
[1263] Terminal: The received response data is converted into voice using a speech synthesis function and provided to the user. The response is provided at a speed and voice quality that is easy for the user to hear.
[1264] Cooking and meal follow-up process flow
[1265] Step 1:
[1266] Device: The user voices or texts a request such as "I want to make a diet lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[1267] Step 2:
[1268] Server: Based on the received data, it generates cooking methods and recipes that take health conditions into consideration.
[1269] Step 3:
[1270] Server: Sends the generated cooking instructions to the terminal.
[1271] Step 4:
[1272] Terminal: Provides received cooking instructions to the user via voice or screen display, and also sends signals to instruct the robot assistant on cooking operations as needed.
[1273] Step 5:
[1274] Robot: Based on instructions from a terminal, the robot will begin specific cooking tasks, such as automatically chopping vegetables and putting food in a pot.
[1275] Processing flow of the system for linking route guidance and public transportation during travel
[1276] Step 1:
[1277] Device: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[1278] Step 2:
[1279] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[1280] Step 3:
[1281] Server: Sends the calculated route information to the terminal.
[1282] Step 4:
[1283] Terminal: Provides the received route information to the user via voice and screen display, providing specific guidance such as "turn right at the next traffic light."
[1284] Step 5:
[1285] Device: GPS information is constantly updated while the user is moving, tracking the user's current location in real time. If the user deviates from the route, the device recalculates the optimal route and provides new route guidance.
[1286] Step 6:
[1287] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[1288] Through specific processing at each step, this system aims to provide real-time support to users and reduce the overall burden of caregiving.
[1289] Example 1
[1290] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1291] In an aging society, reducing the burden of caregiving and providing a comfortable lifestyle are major challenges. In particular, advanced technology is required to provide consistent support for elderly people in conversation, eating, and mobility. However, existing systems have difficulty integrating multiple functions and providing them in a unified manner, limiting their practicality and effectiveness in the field. Therefore, there is a need to provide a comprehensive support system that allows elderly people to live independently.
[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1293] In this invention, the server includes means for receiving a user's voice input and converting the voice data into text data, means for transmitting the text data to the server, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for transmitting the generated response to a terminal, means for converting the generated response into text and providing it to the user, means for receiving user request data and transmitting health status and ingredient data to the server, means for analyzing the user's health status and ingredient data and generating cooking methods and recipes, means for transmitting the generated cooking instructions to the terminal, means for providing the generated cooking instructions to the user by voice or on-screen display, means for obtaining the user's current location using GPS and transmitting destination information to the server, means for calculating an optimal route based on real-time traffic information, means for transmitting the calculated route information to the terminal, means for providing the calculated route information to the user by voice or on-screen display, and means for tracking the user's current location and updating the route during travel. This provides users with conversation assistance, cooking and meal follow-up, and route guidance during travel in an integrated manner, enabling elderly people to live independently and comfortably.
[1294] "Voice input" refers to the act of a user providing information to a system using voice.
[1295] "Voice data" refers to a digital recording of a user's voice input.
[1296] "Character data" is data in text format converted by voice recognition.
[1297] A "server" is a computer system that provides functionality such as processing and storing data and generating responses.
[1298] A "terminal" is a device that allows a user to interact with the system through an interface, and includes smartphones, tablets, PCs, etc.
[1299] A "natural language processing engine" is software that analyzes text data, understands its meaning, and generates responses.
[1300] A "response" is a reply or instruction generated by a server based on user input.
[1301] "Health status data" is data that includes information related to the user's health.
[1302] "Ingredient data" is data that includes information about ingredients owned by the user.
[1303] A "cooking method" is a procedure for preparing a dish using specific ingredients.
[1304] A "recipe" is a set of instructions that lists the ingredients and steps needed to make a particular dish.
[1305] "GPS" is a Global Positioning System for obtaining geographical location information.
[1306] "Real-time traffic information" means up-to-date data on current traffic conditions.
[1307] An "optimal route" is the most efficient route to reach a destination.
[1308] "Location tracking" refers to the real-time monitoring of a moving user's current location.
[1309] "Screen display" refers to the visual presentation of information on a terminal display.
[1310] MODE FOR CARRYING OUT THE INVENTION
[1311] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable lifestyle. This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[1312] Conversation support system
[1313] In a conversation support system, a server hosts a natural language processing (NLP) engine that understands the user's voice input and generates appropriate responses. This system uses the following hardware and software:
[1314] Hardware
[1315] Devices: PC, smartphone, tablet, etc.
[1316] Devices with built-in microphones
[1317] software
[1318] Speech recognition engine (e.g., Google Speech-to-Text, Amazon Transcribe)
[1319] Natural language processing engines (e.g., Google Cloud Natural Language API, OpenAI's GPT-3)
[1320] Text-to-speech software (e.g., Amazon Polly, Google Text-to-Speech)
[1321] Specific examples
[1322] When a user asks, "What's the weather like today?", the system captures the voice with the device's microphone, and the speech recognition engine converts it into text data. The converted text data is sent to the server and analyzed by the natural language processing engine. As a result, a response such as "The weather is sunny today" is generated and sent to the device. Finally, the device converts the received response data into speech using speech synthesis software and provides it to the user.
[1323] Prompt Sentence Examples
[1324] "When a user says, 'What's the weather like today?' generate an appropriate response."
[1325] Cooking and meal follow-up system
[1326] The cooking and meal follow-up system optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking. This system uses the following hardware and software:
[1327] Hardware
[1328] Smartphones and tablets
[1329] Smart appliances (smart ovens, smart refrigerators)
[1330] software
[1331] Health management app
[1332] Ingredient Management Software
[1333] Recipe Generation Engine
[1334] Robot Assistant Control Software
[1335] Specific examples
[1336] If a user requests, "I want to make a low-sugar dessert," the system will consider the user's health condition (for example, whether they have a tendency toward diabetes) and provide the optimal recipe. The smartphone receives the request via voice or text and sends it to the server. The server generates the optimal recipe based on the received data and pre-registered health data and sends it to the smartphone. Furthermore, if necessary, the robot assistant will automatically prepare part of the dessert.
[1337] Prompt Sentence Examples
[1338] "If a user requests, 'I want to make a dessert with less sugar,' generate a recipe that takes their health into consideration."
[1339] Route guidance and public transport linkage system
[1340] The route guidance and public transport integration system for travel is a system that presents the optimal travel route based on the user's current location and destination. It provides the optimal route by including real-time traffic information. This system uses the following hardware and software.
[1341] Hardware
[1342] GPS-enabled devices (smartphones, tablets)
[1343] software
[1344] Real-time traffic information API (e.g., Google Maps API, HERE API)
[1345] Route Calculation Algorithm
[1346] Location Tracking Software
[1347] Specific examples
[1348] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." The smartphone obtains the user's current location using GPS and sends a request to the server. The server calculates a route based on real-time traffic information and sends it to the smartphone. The smartphone then provides the calculated route to the user by voice or on-screen display, and tracks the user's current location in real time while traveling, updating the route as needed.
[1349] Prompt Sentence Examples
[1350] "When a user requests 'Take me to the station,' generate a plan that provides the optimal route."
[1351] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1352] Conversation support system
[1353] Program processing
[1354] Step 1:
[1355] The user gives a voice input of "What's the weather like today?" The device captures this voice input using the built-in microphone. The input is the user's voice data, and the output is the captured voice data. The device collects voice data in real time.
[1356] Step 2:
[1357] The device sends the captured voice data to a voice recognition engine and converts it into text data. Specifically, the voice data is sent to a voice recognition service on the cloud, and text data is obtained. The input here is voice data, and the output is converted text data. The device performs the operation of converting voice data into text data.
[1358] Step 3:
[1359] The terminal sends the converted character data to the server. The data is sent securely using a secure protocol (such as HTTPS). The input here is character data, and the output is the character data sent to the server. The terminal performs the data sending operation.
[1360] Step 4:
[1361] The server inputs the received text data into a natural language processing engine (NLP engine) for analysis. The NLP engine analyzes the text data and generates an appropriate response. The input here is text data, and the output is the generated response data. The server performs the operations of analyzing the text data and generating a response.
[1362] Step 5:
[1363] The server sends the generated response data to the terminal. Again, communication is performed using a secure protocol. The input here is the response data, and the output is the response data sent to the terminal. The server performs the data transmission operation.
[1364] Step 6:
[1365] The terminal converts the received response data into speech using a speech synthesis engine. The speech synthesis engine is used to convert text data into speech data, which is then provided to the user through a speaker. The input here is the response data, and the output is the generated speech data. The terminal then performs the operation of playing back the speech data.
[1366] Cooking and meal follow-up system
[1367] Program processing
[1368] Step 1:
[1369] The user inputs a request into the device by voice or text, such as "I want to make a dietary lunch." The device captures this input. The input is the user's request data, and the output is the captured request data. The device performs the operation of collecting voice and text data.
[1370] Step 2:
[1371] The terminal sends the request data, pre-registered health condition data, and ingredient information to the server. The data is sent using a secure protocol. The input is the request data and the attached health condition data and ingredient information, and the output is the data sent to the server. The terminal performs the data transmission operation.
[1372] Step 3:
[1373] The server analyzes the received data and generates appropriate cooking methods and recipes that take health status into consideration. The input here is the request data, health status data, and ingredient information, and the output is the generated recipe data. The server performs the data analysis and recipe generation operations.
[1374] Step 4:
[1375] The server sends the generated recipe data to the terminal. Communication is again performed using a secure protocol. The input is the recipe data, and the output is the recipe data sent to the terminal. The server performs the data transmission operation.
[1376] Step 5:
[1377] The terminal provides the received recipe data to the user by voice or screen display. The input here is recipe data, and the output is recipe information provided to the user. The terminal performs the operations of displaying information and playing voice.
[1378] Step 6:
[1379] If necessary, the terminal sends specific cooking operation instructions to the robot assistant. The input is cooking operation instruction data, and the output is instruction data sent to the robot assistant. The terminal performs the operation of sending instructions.
[1380] Route guidance and public transport linkage system
[1381] Program processing
[1382] Step 1:
[1383] The user inputs a request into the terminal by voice or text, such as "Take me to the station." The terminal captures this input. The input is the user's request data, and the output is the captured request data. The terminal performs the operation of collecting voice and text data.
[1384] Step 2:
[1385] The device uses the built-in GPS to obtain the user's current location. The input here is GPS data, and the output is the obtained current location data. The device performs the operation of collecting location data.
[1386] Step 3:
[1387] The device sends the request data and current location data to the server. The data is sent using a secure protocol. The input is the request data and current location data, and the output is the data sent to the server. The device performs the data sending operation.
[1388] Step 4:
[1389] The server obtains real-time traffic information and calculates the optimal route. The input here is the current location data and real-time traffic information, and the output is the calculated route data. The server performs data analysis and route calculation.
[1390] Step 5:
[1391] The server sends the calculated route data to the terminal. Communication is again performed using a secure protocol. The input is the route data, and the output is the route data sent to the terminal. The server performs the data transmission operation.
[1392] Step 6:
[1393] The terminal provides the received route data to the user by voice or screen display. The input here is the route data, and the output is the route information provided to the user. The terminal performs the operations of displaying information and playing voice.
[1394] Step 7:
[1395] While moving, the device uses GPS to track the user's current location and updates the route as needed. The input here is GPS data and real-time traffic information, and the output is updated route data. The device performs location tracking and route update operations.
[1396] (Application example 1)
[1397] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1398] This invention relates to a system for supporting and ensuring the safety of specific groups, such as the elderly, in their daily lives. In particular, the system aims to reduce the burden on caregivers and support the independent living of those receiving care by utilizing the user's voice input and location information to provide comprehensive support tailored to a wide range of needs, including conversation assistance, cooking and meal follow-up, and emergency response. There is also a need for a system that can respond quickly and appropriately in emergencies.
[1399] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1400] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for acquiring the user's current location, means for transmitting the user's location information to the server, and means for analyzing the user's voice in an emergency and transmitting an alert together with the location information, thereby supporting the user's daily life and enabling a prompt and appropriate response, particularly in an emergency.
[1401] "Voice input" refers to data obtained by a device capturing words spoken by a user.
[1402] "Voice data" is data that digitally represents voice input obtained from a user.
[1403] "Character data" is data obtained by analyzing voice data and converting the content into text format.
[1404] A "natural language processing engine" is a program or system that analyzes text data and generates appropriate responses.
[1405] "Response generation" refers to the process of using a natural language processing engine to create an appropriate response from text data.
[1406] "Speech synthesis" is the technique of converting the generated response back into speech form.
[1407] "Current location of user" refers to the geographic location where the user is currently located.
[1408] "GPS" is a satellite positioning system that identifies a user's current location.
[1409] "Real-time traffic information" is dynamic information that shows the current traffic situation.
[1410] An "optimal route" refers to the most efficient route from the current location to the destination.
[1411] "Emergency speech analysis" is the process of identifying specific speech inputs that indicate an emergency.
[1412] "Transmitting location information" means sending data indicating the user's current location to a server or other receiving device.
[1413] "Sending an alert" means communicating a warning message based on a specific condition.
[1414] This invention is a system for supporting specific groups, such as the elderly, in their daily lives and ensuring their safety, and is realized by each of the entities, the server, the terminal, and the user, fulfilling their specific roles. The following describes in detail the mode for carrying out this invention.
[1415] Embodiment of conversation support system
[1416] The server hosts a natural language processing engine (NLP engine). The device captures the user's voice input and converts this voice data into text data. The converted text data is sent to the server, where the NLP engine generates an appropriate response. The generated response is sent back to the device, where it is converted into speech and provided to the user. This system allows users to obtain a variety of information through natural conversation.
[1417] Hardware / Software: A microphone is used for voice input, the speech_recognition library is used for voice recognition, an NLP engine is used for natural language processing, and the pyttsx3 library is used for speech synthesis.
[1418] Example: When a user asks, "What's the weather like today?", the system responds by saying, "The weather is sunny today."
[1419] Cooking and meal follow-up system
[1420] The device receives input data from the user regarding their health condition and ingredients, and sends this data to the server. Based on the received data, the server generates cooking methods and recipes optimized for the user's health condition. The generated cooking instructions are provided to the user via the device via voice or screen display. This allows users to easily cook in a way that takes their health into consideration.
[1421] Hardware and software: Smartphones and tablets are used to acquire and transmit data. AI models are used to generate recipes, and existing speech synthesis libraries and display technologies are used for voice and screen display.
[1422] Example: When a user requests to make a low-sugar dessert, the system provides the optimal recipe based on the user's health status.
[1423] Route guidance and public transport linkage system for travel
[1424] The device acquires the user's current location and sends destination information to the server. The server calculates the optimal route based on real-time traffic information and sends that information to the device. The device then provides the calculated route information to the user via voice and on-screen display. In addition, the device has the ability to analyze the user's voice in an emergency and send an alert along with location information.
[1425] Hardware and software: GPS devices, real-time traffic information platforms, and speech recognition and speech synthesis libraries are used.
[1426] Example: If a user yells "Help!", the system interprets the audio as an emergency alert and sends information, including the user's current location, to emergency contacts.
[1427] Prompt Sentence Examples
[1428] "How's the weather today?"
[1429] "I want to make a dessert with less sugar."
[1430] "help me!"
[1431] In this way, the present invention utilizes the user's voice and location information to provide a wide range of support for daily life, thereby reducing the burden on elderly people and their caregivers.
[1432] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1433] Processing steps of the conversation support system
[1434] Step 1:
[1435] The terminal obtains the user's voice input.
[1436] Specifically, when a user says, "What's the weather like today?", the device's microphone captures the voice. The input is voice data.
[1437] Step 2:
[1438] The terminal converts the captured voice data into text data.
[1439] Specifically, it uses a speech recognition library (e.g., speech_recognition) to convert voice data into text format. The input is voice data, and the output is text data.
[1440] Step 3:
[1441] The terminal transmits the converted character data to the server.
[1442] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[1443] Step 4:
[1444] The server analyzes the received text data using a natural language processing engine.
[1445] Specifically, an NLP engine (for example, a custom NLP server) analyzes the text data and generates an appropriate response to the question, "What's the weather like today?" The input is the text data, and the output is the analysis result and response data.
[1446] Step 5:
[1447] The server sends the generated response data to the terminal.
[1448] Specifically, the generated response data is sent to the terminal via an HTTP request. The input is the response data, and the output is data sent to the terminal.
[1449] Step 6:
[1450] The terminal converts the received response data into voice and provides it to the user.
[1451] Specifically, it uses a speech synthesis library (e.g., pyttsx3) to convert text-based response data into speech and provides it to the user through a speaker. The input is the response data, and the output is speech output.
[1452] Emergency response system processing steps
[1453] Step 1:
[1454] The terminal obtains the user's voice input.
[1455] Specifically, when a user shouts "Help!", the device's microphone captures the voice. The input is voice data.
[1456] Step 2:
[1457] The terminal converts the captured voice data into text data.
[1458] Specifically, it uses a speech recognition library to convert voice data into text format. The input is voice data and the output is text data.
[1459] Step 3:
[1460] The terminal transmits the converted character data to the server.
[1461] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[1462] Step 4:
[1463] The server analyzes the text data and identifies it as an urgent message.
[1464] Specifically, the NLP engine recognizes the emergency message "Help!" and prepares an appropriate response. The input is text data, and the output is emergency alert data.
[1465] Step 5:
[1466] The device obtains the user's current location using GPS.
[1467] Specifically, the GPS module of the device acquires the current location information. The input is a signal from the GPS device, and the output is location information data.
[1468] Step 6:
[1469] The server sends location information and emergency alerts to emergency contacts.
[1470] Specifically, location information and an emergency message are sent to emergency contacts (e.g., family members or emergency services) via an HTTP request. The input is the emergency alert data and location data, and the output is the transmission of the data to the contacts.
[1471] Step 7:
[1472] The terminal provides a voice message to reassure the user.
[1473] Specifically, it uses a speech synthesis library to generate a message such as "Rescue is currently being called. Please rest assured," and provides it to the user through a speaker. The input is the data for the reassurance message, and the output is voice output.
[1474] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1475] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live a comfortable life. Specific embodiments of this invention are described below.
[1476] In this system, each of the server, terminal, and user plays a role in recognizing the user's emotions and providing optimal responses and support to the user's mental and physical needs.
[1477] 1. Conversation support system
[1478] Overview of conversation support
[1479] The server hosts an emotion engine and a natural language processing (NLP) engine, which recognizes emotions from the user's voice input and generates appropriate responses to provide to the user.
[1480] Program processing
[1481] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[1482] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[1483] 3. Server: Analyzes the received text data using an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).
[1484] 4. Server: Uses a natural language processing engine to generate an appropriate response based on context and sentiment, such as "The weather is sunny today."
[1485] 5. Server: Sends a tone based on the generated response data and emotional state to the terminal.
[1486] 6. Terminal: The received response data is converted into voice with an appropriate tone using a voice synthesis function and provided to the user.
[1487] Specific examples
[1488] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question, and if it recognizes that the user is feeling "depressed," it responds by adding encouraging words such as, "Cheer up, it's sunny today, why don't you go outside?"
[1489] 2. Cooking and meal follow-up system
[1490] Cooking and Meal Follow-up Overview
[1491] This system provides optimal cooking instructions based on the user's health and emotional state.
[1492] Program processing
[1493] 1. Device: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status, ingredient information, and emotional data, is sent to the server.
[1494] 2. Server: Based on the received data, it generates cooking methods and recipes that take into account the health and emotional state of the user.
[1495] 3. Server: Sends the generated cooking instructions and related emotional support information to the terminal.
[1496] 4. Terminal: Provides the received cooking instructions and emotional support information (e.g., suggestions for ingredients with a relaxing effect when the user is feeling stressed) to the user via voice or on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[1497] Specific examples
[1498] If a user requests "I want to make a healthy breakfast," the system will provide a relaxing recipe using oatmeal based on the user's health status and feelings of "tension." Furthermore, if a robotic assistant is needed, it will automate some of the oatmeal cooking process.
[1499] 3. Route Guidance and Public Transportation Integration System
[1500] Overview of route guidance during travel
[1501] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[1502] Program processing
[1503] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[1504] 2. Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination, taking into account the emotional state.
[1505] 3. Server: Sends the calculated route information and emotional support information (for example, calming voice guidance if the user is feeling anxious) to the device.
[1506] 4. Device: Provides the user with the received route information and emotional support information via voice and screen display. It also tracks the user's current location in real time while moving and updates the route as needed.
[1507] Specific examples
[1508] If the system detects that the user is "nervous" while traveling from home to the station, it calculates the shortest and safest route based on the latest traffic information and provides specific instructions in a gentle tone, such as "Turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[1509] As described above, the AI system combined with the emotion engine provides services that take into account the user's mental state, aiming to reduce the burden of caregiving and improve the user's quality of life.
[1510] The processing flow will be explained below.
[1511] Conversation support processing flow
[1512] Step 1:
[1513] User: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[1514] Step 2:
[1515] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[1516] Step 3:
[1517] Server: Analyzes the user's emotional state using the emotion engine and sends the results of this analysis to the natural language processing engine.
[1518] Step 4:
[1519] Server: Uses a natural language processing engine to generate appropriate responses based on context and sentiment.
[1520] Step 5:
[1521] Server: Sends tone information based on the generated response data and emotional state to the terminal.
[1522] Step 6:
[1523] Terminal: The terminal converts the received response data into speech with an appropriate tone using a speech synthesis function and provides it to the user.
[1524] Cooking and meal follow-up process flow
[1525] Step 1:
[1526] User: The user inputs a request by voice or text, such as "I want to make a dietary lunch." The device then sends this information, along with previously entered health status and ingredient information, to the server.
[1527] Step 2:
[1528] Server: Based on the received information about the user's health condition and ingredients, the emotion engine analyzes the user's emotional state.
[1529] Step 3:
[1530] Server: Generates cooking methods and recipes that take into account health and emotional states.
[1531] Step 4:
[1532] Server: Sends the generated cooking instructions and emotional support information to the terminal.
[1533] Step 5:
[1534] Terminal: Provides the received cooking instructions and emotional support information to the user via voice and on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[1535] Step 6:
[1536] Robot: Based on instructions from a terminal, the robot begins specific cooking tasks. For example, it automatically performs basic operations such as chopping vegetables and putting food in a pot.
[1537] Processing flow of the system for linking route guidance and public transportation during travel
[1538] Step 1:
[1539] User: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[1540] Step 2:
[1541] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[1542] Step 3:
[1543] Server: Based on the received user's emotional state, the server adjusts the guidance method (tone and pace) for the optimal route.
[1544] Step 4:
[1545] Server: Sends the calculated route information and emotional support information to the terminal.
[1546] Step 5:
[1547] Terminal: Provides the received route information and emotional support information to the user via voice and on-screen display. For example, it provides specific guidance such as "Turn right at the next traffic light," but changes the tone depending on the user's emotional state.
[1548] Step 6:
[1549] Device: GPS information is constantly updated while moving, tracking the user's current location in real time. Each time, the server recalculates the optimal route based on the user's emotional state and provides new route guidance.
[1550] Step 7:
[1551] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[1552] This detailed processing flow allows the system to provide more personalized assistance while taking into account the user's emotional state, thereby reducing the burden of caregiving and enabling users to live a comfortable and efficient life.
[1553] Example 2
[1554] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1555] In recent years, the burden of caregiving has increased with the progress of the aging society. There is also a demand for support to maintain users' mental and physical health. However, conventional AI systems do not take into account the user's emotional state, limiting the user experience. Furthermore, even in health management and mobility assistance, they can only make suggestions that ignore the user's emotional state, making it difficult to provide accurate support.
[1556] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the user's emotional state from the text data using an emotion engine, means for generating an appropriate response based on the context and the emotional state using a natural language processing engine, and means for converting the generated response into voice and providing it to the user. This makes it possible to provide appropriate responses and support according to the user's emotional state.
[1557] "Audio input" refers to an audio signal emitted by a user through an input device such as a microphone.
[1558] "Audio Data" means data in digital form obtained from audio input.
[1559] "Character data" refers to text-format data obtained by analyzing voice data using voice recognition technology.
[1560] An "emotion engine" is software or hardware that analyzes text data and estimates the user's emotional state.
[1561] An "emotional state" is a mental state that a user is feeling, and examples include happiness, sadness, anger, etc.
[1562] A "natural language processing engine" is software that analyzes text data, understands meaning and context, and generates appropriate responses.
[1563] A "response" is a reply or instruction to the user generated by a natural language processing engine.
[1564] "Speech synthesis" is a technology that converts text data into audio data and turns it into audio that can be played through speakers, etc.
[1565] "Health status" refers to information about the user's current physical health, including information about illnesses, symptoms, and physical conditions.
[1566] "Ingredient information" refers to data relating to the types and amounts of ingredients available to the user.
[1567] A "cooking method" is the procedure or process of preparing a dish using ingredients.
[1568] A "recipe" is a document that describes the ingredients and steps for making a particular dish.
[1569] A "location information system" is a system that obtains a user's current location using technology such as GPS.
[1570] "Real-time traffic information" refers to information about current traffic conditions and transportation options.
[1571] An "optimal route" is the most efficient route for a user to travel from their current location to their destination.
[1572] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live more comfortable lives. Specific embodiments of this invention will be described in detail below.
[1573] 1. Conversation support system
[1574] System Configuration
[1575] In this system, users ask questions or make requests by voice, and the server processes them and provides a voice response. The main components are as follows:
[1576] Terminal: A device that captures the user's voice and sends the voice data to the server. It uses the Google Speech-to-Text API for voice recognition.
[1577] Server: Acquires text data and analyzes it using an emotion engine (e.g., IBM Watson Tone Analyzer). OpenAI's GPT-3 natural language processing engine is used.
[1578] Speech synthesis engine: Converts the generated text response into speech using Amazon Polly.
[1579] Specific examples
[1580] When a user asks, "What's the weather like today?", the voice data captured by the device is converted into text data and sent to the server. The server uses an emotion engine to recognize the emotional state of "interested" and generates an appropriate response using a natural language processing engine. For example, it might respond, "The weather is sunny today." This is then converted into speech using a speech synthesis engine and provided to the user via the device.
[1581] Prompt Sentence Examples
[1582] User: "What's the weather like today?"
[1583] system:
[1584] 1. The server converts the speech into text and analyzes it using an emotion engine.
[1585] 2. If the emotion is perceived as "interested," generate an appropriate response.
[1586] 2. Cooking and meal follow-up system
[1587] System Configuration
[1588] This system provides optimal cooking instructions based on the user's health and emotional state.
[1589] Device: Captures user requests and sends them to the server. The voice recognition function uses the Google Speech-to-Text API.
[1590] Server: Analyzes health status, ingredient information, and emotional state to generate cooking methods and recipes. IBM Watson Tone Analyzer is used as the emotion engine, and OpenAI's GPT-3 is used to generate recipes.
[1591] Robot assistant: A device that automates cooking operations as needed.
[1592] Specific examples
[1593] If a user requests "I want to make a dietary lunch," the device sends this information to the server. The server uses its emotion engine to recognize that the user is "feeling stressed" and generates a cooking method using ingredients that have a relaxing effect. The generated recipe is for oatmeal and salad and is provided to the user via voice and on-screen display. If necessary, some of the cooking can be automated by a robotic assistant.
[1594] Prompt Sentence Examples
[1595] User: "I want to make a diet lunch."
[1596] system:
[1597] 1. The server receives the request and user data and generates an appropriate recipe.
[1598] 2. The recipe uses oatmeal and provides audio instructions.
[1599] 3. Route Guidance and Public Transportation Integration System
[1600] System Configuration
[1601] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[1602] Terminal: Receives the user's request, obtains the current location using GPS, and sends it to the server.
[1603] Server: Obtains real-time traffic information and calculates the optimal route. IBM Watson Tone Analyzer is used as the emotion engine, and Google Maps API is used to obtain traffic information.
[1604] Speech synthesis engine: Converts route directions into speech.
[1605] Specific examples
[1606] When a user requests "Guide me to the station," the device obtains the user's current location, sets "station" as the destination, and sends the information to the server. The server then uses its emotion engine to recognize the user as "stressed" and calculates the optimal route to reassure the user. For example, the server uses a speech synthesis engine to convert specific instructions, such as "Turn right at the next traffic light," into voice and provides the information to the user via the device.
[1607] Prompt Sentence Examples
[1608] User: "Take me to the station."
[1609] system:
[1610] 1. The server calculates a route based on the current location and destination, and provides guidance according to the user's emotional state.
[1611] 2. If they are feeling anxious, guide them with a reassuring route and a gentle tone.
[1612] As described above, the present invention can provide support in various situations while taking into consideration the emotional state of the user, reduce the burden of caregiving, and improve the quality of life of the user.
[1613] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1614] Conversation support system processing flow
[1615] Step 1:
[1616] Description: The user provides voice input.
[1617] Specific action: The user speaks to the device, "What's the weather like today?"
[1618] Input: User's voice input.
[1619] Output: The captured audio data.
[1620] Step 2:
[1621] Description: The device receives voice data and converts it into text data using voice recognition.
[1622] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[1623] Input: Captured audio data.
[1624] Output: The converted text data (e.g., "What's the weather like today?").
[1625] Step 3:
[1626] Description: The terminal sends the converted character data to the server.
[1627] Specific operation: The terminal sends an HTTP request to the server and sends character data.
[1628] Input: The converted character data.
[1629] Output: Character data sent to the server.
[1630] Step 4:
[1631] Description: The server receives text data and analyzes it with the emotion engine.
[1632] Specific operation: The server uses IBM Watson Tone Analyzer to analyze the user's emotional state (e.g., "interested") from the text data.
[1633] Input: The character data sent.
[1634] Output: Parsed emotional state.
[1635] Step 5:
[1636] Description: The server uses a natural language processing engine to generate appropriate responses based on context and emotional state.
[1637] Specific operation: The server calls OpenAI's GPT-3 and generates an appropriate response sentence using the text data and emotional state as input (e.g., "The weather is sunny today.").
[1638] Input: Text data, emotional state.
[1639] Output: The generated response sentence.
[1640] Step 6:
[1641] Description: The server sends the generated response to the terminal.
[1642] Specific operation: The server sends an HTTP response to the terminal and sends the generated response text.
[1643] Input: The generated response sentence.
[1644] Output: The response sent to the terminal.
[1645] Step 7:
[1646] Description: The response received by the terminal is converted into speech using the speech synthesis function.
[1647] Specific operation: The device uses Amazon Polly to convert the response into voice data.
[1648] Input: The response received.
[1649] Output: The converted audio data.
[1650] Step 8:
[1651] Description: The device plays the converted audio data to the user.
[1652] Specific operation: Play audio data through the device speaker.
[1653] Input: The converted audio data.
[1654] Output: A spoken response to the user.
[1655] Cooking and meal follow-up system processing flow
[1656] Step 1:
[1657] Description: A user enters a cooking request.
[1658] Specific operation: The user speaks to the device, saying, "I want to make a dietary lunch."
[1659] Input: User's voice input.
[1660] Output: The captured audio data.
[1661] Step 2:
[1662] Description: The device receives voice data and converts it into text data using voice recognition.
[1663] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[1664] Input: Captured audio data.
[1665] Output: The converted text (e.g. "I want to make a dietary lunch").
[1666] Step 3:
[1667] Description: The device acquires health status, food ingredient information, and emotional state and sends them to the server.
[1668] Specific operation: The device obtains pre-registered health status, food information, and emotional state, and sends an HTTP request to the server.
[1669] Input: Text data, health status, food information, emotional state.
[1670] Output: The data sent to the server.
[1671] Step 4:
[1672] Description: The server generates cooking instructions and emotional support information based on the data received.
[1673] Specific operation: Using an emotion engine (e.g., IBM Watson Tone Analyzer) and a natural language processing engine (e.g., OpenAI's GPT-3), it generates appropriate cooking methods, recipes, and emotional support information.
[1674] Input: health status, food information, emotional state.
[1675] Output: Generated cooking instructions and emotional support information.
[1676] Step 5:
[1677] Description: The server sends the generated cooking instructions and emotional support information to the terminal.
[1678] Specific operation: The server sends an HTTP response to the terminal, conveying instructions and information.
[1679] Input: Generated cooking instructions and emotional support information.
[1680] Output: Cooking instructions and emotional support information sent to the device.
[1681] Step 6:
[1682] Description: The device provides the user with cooking instructions and emotional support information received.
[1683] How it works: The device uses speech synthesis to convert instructions into voice, provides information on the screen, and issues instructions to the robot assistant as needed.
[1684] Input: cooking instructions and emotional support information.
[1685] Output: Voice and visual instructions, instruction signals to the robot assistant.
[1686] Processing flow of the system for linking route guidance and public transportation during travel
[1687] Step 1:
[1688] Description: User requests directions.
[1689] Specific operation: The user speaks to the terminal, saying, "Take me to the station."
[1690] Input: User's voice input.
[1691] Output: The captured audio data.
[1692] Step 2:
[1693] Description: The device receives voice data and converts it into text data using voice recognition.
[1694] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[1695] Input: Captured audio data.
[1696] Output: The converted text data (e.g. "Take me to the station").
[1697] Step 3:
[1698] Description: The device obtains its current location using GPS and sends destination information and emotional state to the server.
[1699] Specific operation: The device acquires GPS data and sends an HTTP request to the server along with the user's emotional state information.
[1700] Input: Text data, current location information, destination information, emotional state.
[1701] Output: The data sent to the server.
[1702] Step 4:
[1703] Description: The server calculates the optimal route based on real-time traffic information and emotional state.
[1704] Specific operation: The server uses the Google Maps API to calculate the optimal route, taking into account the traffic information obtained and the emotional state analyzed by the emotion engine.
[1705] Input: current location information, destination information, emotional state.
[1706] Output: Calculated optimal route information.
[1707] Step 5:
[1708] Description: The server sends the calculated route information and emotional support information to the terminal.
[1709] Specific operation: The server sends an HTTP response to the device, conveying route information and emotional support information.
[1710] Input: Calculated optimal route information, emotional support information.
[1711] Output: Route information and emotional support information sent to the device.
[1712] Step 6:
[1713] Description: Provides the user with route information and emotional support information received by the terminal.
[1714] Specific operation: The device uses a speech synthesis function to convert route guidance into voice and also provides route information on the screen.
[1715] Input: Received route information, emotional support information.
[1716] Output: Route guidance by voice and on-screen display.
[1717] Step 7:
[1718] Description: The device tracks the user's location in real time and updates the route as needed.
[1719] How it works: The device uses GPS to continuously obtain its current location while moving, and resends the information to the server as needed to receive new route guidance.
[1720] Input: current location information, route information.
[1721] Output: Updated route guidance information.
[1722] This is the specific processing flow of this system. At each step, the system converts voice data, analyzes emotions, processes natural language, calculates routes, and performs other operations based on the user's input, providing appropriate responses and support.
[1723] (Application example 2)
[1724] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1725] Conventional security and guidance systems tend to respond in a uniform manner without considering the user's emotional state, making it impossible to provide optimal alerts and guidance to users. Furthermore, even in situations where users feel anxious or scared, appropriate responses are often delayed, reducing the user's sense of security. Therefore, there is a need for a system that provides effective responses and a sense of security that takes the user's emotional state into account.
[1726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1727] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for capturing the user's face and recognizing emotions using an emotion engine, and means for providing appropriate security alerts and guidance based on the recognized emotions. This enables optimal responses that take the user's emotional state into consideration, thereby improving the user's sense of security.
[1728] "Voice input" refers to data that is input by the user using voice.
[1729] "Voice data" refers to data that represents in digital form the voice input by the user.
[1730] "Character data" refers to text-format data obtained by converting voice data into characters.
[1731] A "natural language processing engine" is a processing engine that analyzes text data, understands its meaning, and generates an appropriate response.
[1732] An "emotion engine" is an engine that has the function of determining emotions from the user's voice, facial expressions, etc.
[1733] "Response generation" is the process of creating appropriate responses or guidance based on the user's input data and their emotional state.
[1734] "Speech synthesis" is a technology that outputs text data as speech.
[1735] "Emotion recognition" is the process of determining a user's emotions from their facial expressions, voice, etc.
[1736] A "security alert" is a notification that warns or warns users to ensure their safety.
[1737] "Guidance" refers to instructions that instruct the user on appropriate actions or responses.
[1738] "Capturing a user's face" means obtaining an image of the user's face using a device such as a camera.
[1739] A "location information system" is a system that obtains a user's current location using GPS or other means.
[1740] "Real-time traffic information" refers to data that obtains the latest information on traffic conditions in real time.
[1741] An "optimal route" is a route that allows a user to reach a destination most efficiently.
[1742] The system of the present invention is an AI system that combines an emotion engine and a natural language processing engine to provide optimal support for the user's mental and physical needs. The system allows the server, terminal, and user to play their respective roles, improving user security.
[1743] Hardware and Software
[1744] Hardware
[1745] Smart glasses (a device with a general camera function)
[1746] Smartphone (Android or iOS device)
[1747] Head-mounted display (general VR device)
[1748] software
[1749] Emotion engine (e.g. Affectiva SDK)
[1750] Natural language processing engine (e.g. Google Cloud Natural Language API)
[1751] Image and voice recognition tools (e.g., Google Cloud Vision API, Speech-to-Text API)
[1752] Speech synthesis tools (e.g., Google Cloud Text-to-Speech API)
[1753] Data processing and calculation
[1754] server
[1755] The server performs the following processes. First, it receives the user's voice input and facial image from the device. The voice data is converted into text data using a voice recognition function, and the text data and facial image are sent in parallel to the emotion engine. The emotion engine analyzes the user's emotional state and provides this data to the natural language processing engine. The natural language processing engine generates an appropriate response based on the text data and emotional data. Finally, the generated response data is converted into voice data using a voice synthesis tool and sent to the device.
[1756] Terminal
[1757] The device performs the following processes: First, it captures voice input and facial images from the user and sends them to the server. Then it receives response data returned from the server and outputs voice to the user at the appropriate time. If the user feels anxious, it provides security alerts and guidance based on emotion recognition.
[1758] Specific examples
[1759] scenario
[1760] When a user is walking down a street at night, the AI detects that they are feeling anxious.
[1761] Response: The smart glasses say, "User, it's okay. There's a station nearby. Would you like to contact the fastest security company right away?"
[1762] Prompt Sentence Examples
[1763] Design an AI assistant with an emotion engine that recognizes the user's emotional state and provides appropriate security alerts and guidance. For a specific scenario, generate a reassuring message when the user is feeling anxious. The devices used could be smart glasses, a smartphone, or a head-mounted display.
[1764] To implement the present invention, devices and software work together to analyze and process user information in real time, ensuring the user's safety and security, and providing security alerts and guidance tailored to the user's emotional state, improving the user's quality of life.
[1765] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1766] Step 1:
[1767] Input: User's voice input, face image
[1768] Operation: The device captures the user's voice input and acquires a facial image using the camera function. The voice data and facial image data are sent to the server.
[1769] Output: Audio data, facial image data
[1770] Step 2:
[1771] Input: Audio data
[1772] Operation: The server converts the received voice data into text data using a speech recognition function. The speech is converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text API).
[1773] Output: Character data
[1774] Step 3:
[1775] Input: Facial image data
[1776] How it works: The server sends facial image data to an emotion engine (e.g., Affectiva SDK) that analyzes the user's emotional state. The emotion engine recognizes the user's emotion (e.g., anxiety, anger, joy) from the facial image and sends the result back to the server.
[1777] Output: Emotion data
[1778] Step 4:
[1779] Input: Text data, emotion data
[1780] How it works: The server inputs text data and emotion data into a natural language processing engine (e.g., Google Cloud Natural Language API) to generate an appropriate response. The natural language processing engine understands the context based on the user's speech and emotion and creates the optimal response.
[1781] Output: Response data
[1782] Step 5:
[1783] Input: Response data
[1784] How it works: The server inputs the generated response data into a speech synthesis tool (e.g., Google Cloud Text-to-Speech API) to convert the text to speech. The speech synthesis tool then generates speech in the appropriate tone.
[1785] Output: Voice response data
[1786] Step 6:
[1787] Input: Voice response data
[1788] Operation: The device receives the voice response data from the server and provides it to the user as a voice. The device outputs the voice response using a speaker, giving the user a sense of security.
[1789] Output: The user receives a spoken response
[1790] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1791] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1792] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1793] [Fourth embodiment]
[1794] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1795] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1796] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1797] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1798] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1799] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1800] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1801] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1802] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1803] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1804] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1805] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1806] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1807] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable life. Specific embodiments of the system are described below.
[1808] This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[1809] 1. Conversation support system
[1810] Overview of conversation support
[1811] The server hosts a natural language processing (NLP) engine that understands voice input and generates appropriate responses to assist the user in the conversation.
[1812] Program processing
[1813] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[1814] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[1815] 3. Server: The received text data is analyzed using a natural language processing engine and an appropriate response is generated, such as "Today's weather is sunny."
[1816] 4. Server: Sends the generated response data to the terminal.
[1817] 5. Terminal: The received response data is converted into voice using a voice synthesis function and provided to the user.
[1818] Specific examples
[1819] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question in real time and responds, based on the current weather information, with, "Today's weather is sunny." This response is provided to the user by voice from the device.
[1820] 2. Cooking and meal follow-up system
[1821] Cooking and Meal Follow-up Overview
[1822] This is a system that optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking.
[1823] Program processing
[1824] 1. Terminal: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[1825] 2. Server: Based on the received data, it generates cooking methods and recipes that take health into consideration.
[1826] 3. Server: Sends the generated cooking instructions to the terminal.
[1827] 4. Terminal: Provides the received cooking instructions to the user via voice or screen display, and also sends specific operating instructions to the robot assistant as needed.
[1828] Specific examples
[1829] If a user requests a low-sugar dessert, the system will provide the optimal recipe, taking into account the user's health status (e.g., whether they are prone to diabetes), and even automatically prepare part of the dessert with the help of a robotic assistant if needed.
[1830] 3. Route Guidance and Public Transportation Integration System
[1831] Overview of route guidance during travel
[1832] This system presents the optimal route based on the user's current location and destination. It provides the optimal route by including real-time traffic information.
[1833] Program processing
[1834] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[1835] 2. Server: Obtains real-time traffic information, calculates the optimal route, and sends the results to the device.
[1836] 3. Terminal: Provides the received route information to the user via voice and screen display. Tracks the user's current location in real time while moving and updates the route as needed.
[1837] Specific examples
[1838] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[1839] By integrating these functions, the system aims to reduce the burden of caregiving and enable those receiving care to live more independently.
[1840] The processing flow will be explained below.
[1841] Conversation support processing flow
[1842] Step 1:
[1843] Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[1844] Step 2:
[1845] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[1846] Step 3:
[1847] Server: Analyzes the received text data using a natural language processing engine and generates an appropriate response based on its content, such as "Today's weather is sunny."
[1848] Step 4:
[1849] Server: Sends the generated response data to the terminal.
[1850] Step 5:
[1851] Terminal: The received response data is converted into voice using a speech synthesis function and provided to the user. The response is provided at a speed and voice quality that is easy for the user to hear.
[1852] Cooking and meal follow-up process flow
[1853] Step 1:
[1854] Device: The user voices or texts a request such as "I want to make a diet lunch." This information, along with previously entered health status and ingredient information, is sent to the server.
[1855] Step 2:
[1856] Server: Based on the received data, it generates cooking methods and recipes that take health conditions into consideration.
[1857] Step 3:
[1858] Server: Sends the generated cooking instructions to the terminal.
[1859] Step 4:
[1860] Terminal: Provides received cooking instructions to the user via voice or screen display, and also sends signals to instruct the robot assistant on cooking operations as needed.
[1861] Step 5:
[1862] Robot: Based on instructions from a terminal, the robot will begin specific cooking tasks, such as automatically chopping vegetables and putting food in a pot.
[1863] Processing flow of the system for linking route guidance and public transportation during travel
[1864] Step 1:
[1865] Device: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[1866] Step 2:
[1867] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[1868] Step 3:
[1869] Server: Sends the calculated route information to the terminal.
[1870] Step 4:
[1871] Terminal: Provides the received route information to the user via voice and screen display, providing specific guidance such as "turn right at the next traffic light."
[1872] Step 5:
[1873] Device: GPS information is constantly updated while the user is moving, tracking the user's current location in real time. If the user deviates from the route, the device recalculates the optimal route and provides new route guidance.
[1874] Step 6:
[1875] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[1876] Through specific processing at each step, this system aims to provide real-time support to users and reduce the overall burden of caregiving.
[1877] Example 1
[1878] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1879] In an aging society, reducing the burden of caregiving and providing a comfortable lifestyle are major challenges. In particular, advanced technology is required to provide consistent support for elderly people in conversation, eating, and mobility. However, existing systems have difficulty integrating multiple functions and providing them in a unified manner, limiting their practicality and effectiveness in the field. Therefore, there is a need to provide a comprehensive support system that allows elderly people to live independently.
[1880] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1881] In this invention, the server includes means for receiving a user's voice input and converting the voice data into text data, means for transmitting the text data to the server, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for transmitting the generated response to a terminal, means for converting the generated response into text and providing it to the user, means for receiving user request data and transmitting health status and ingredient data to the server, means for analyzing the user's health status and ingredient data and generating cooking methods and recipes, means for transmitting the generated cooking instructions to the terminal, means for providing the generated cooking instructions to the user by voice or on-screen display, means for obtaining the user's current location using GPS and transmitting destination information to the server, means for calculating an optimal route based on real-time traffic information, means for transmitting the calculated route information to the terminal, means for providing the calculated route information to the user by voice or on-screen display, and means for tracking the user's current location and updating the route during travel. This provides users with conversation assistance, cooking and meal follow-up, and route guidance during travel in an integrated manner, enabling elderly people to live independently and comfortably.
[1882] "Voice input" refers to the act of a user providing information to a system using voice.
[1883] "Voice data" refers to a digital recording of a user's voice input.
[1884] "Character data" is data in text format converted by voice recognition.
[1885] A "server" is a computer system that provides functionality such as processing and storing data and generating responses.
[1886] A "terminal" is a device that allows a user to interact with the system through an interface, and includes smartphones, tablets, PCs, etc.
[1887] A "natural language processing engine" is software that analyzes text data, understands its meaning, and generates responses.
[1888] A "response" is a reply or instruction generated by a server based on user input.
[1889] "Health status data" is data that includes information related to the user's health.
[1890] "Ingredient data" is data that includes information about ingredients owned by the user.
[1891] A "cooking method" is a procedure for preparing a dish using specific ingredients.
[1892] A "recipe" is a set of instructions that lists the ingredients and steps needed to make a particular dish.
[1893] "GPS" is a Global Positioning System for obtaining geographical location information.
[1894] "Real-time traffic information" means up-to-date data on current traffic conditions.
[1895] An "optimal route" is the most efficient route to reach a destination.
[1896] "Location tracking" refers to the real-time monitoring of a moving user's current location.
[1897] "Screen display" refers to the visual presentation of information on a terminal display.
[1898] MODE FOR CARRYING OUT THE INVENTION
[1899] This invention relates to a system that utilizes AI and multiple digital technologies to reduce the burden of nursing care in an aging society and provide a comfortable lifestyle. This system is realized by each of the entities, "server," "terminal," and "user," each playing a specific role.
[1900] Conversation support system
[1901] In a conversation support system, a server hosts a natural language processing (NLP) engine that understands the user's voice input and generates appropriate responses. This system uses the following hardware and software:
[1902] Hardware
[1903] Devices: PC, smartphone, tablet, etc.
[1904] Devices with built-in microphones
[1905] software
[1906] Speech recognition engine (e.g., Google Speech-to-Text, Amazon Transcribe)
[1907] Natural language processing engines (e.g., Google Cloud Natural Language API, OpenAI's GPT-3)
[1908] Text-to-speech software (e.g., Amazon Polly, Google Text-to-Speech)
[1909] Specific examples
[1910] When a user asks, "What's the weather like today?", the system captures the voice with the device's microphone, and the speech recognition engine converts it into text data. The converted text data is sent to the server and analyzed by the natural language processing engine. As a result, a response such as "The weather is sunny today" is generated and sent to the device. Finally, the device converts the received response data into speech using speech synthesis software and provides it to the user.
[1911] Prompt Sentence Examples
[1912] "When a user says, 'What's the weather like today?' generate an appropriate response."
[1913] Cooking and meal follow-up system
[1914] The cooking and meal follow-up system optimizes cooking instructions based on the user's health condition and reduces the effort required for cooking. This system uses the following hardware and software:
[1915] Hardware
[1916] Smartphones and tablets
[1917] Smart appliances (smart ovens, smart refrigerators)
[1918] software
[1919] Health management app
[1920] Ingredient Management Software
[1921] Recipe Generation Engine
[1922] Robot Assistant Control Software
[1923] Specific examples
[1924] If a user requests, "I want to make a low-sugar dessert," the system will consider the user's health condition (for example, whether they have a tendency toward diabetes) and provide the optimal recipe. The smartphone receives the request via voice or text and sends it to the server. The server generates the optimal recipe based on the received data and pre-registered health data and sends it to the smartphone. Furthermore, if necessary, the robot assistant will automatically prepare part of the dessert.
[1925] Prompt Sentence Examples
[1926] "If a user requests, 'I want to make a dessert with less sugar,' generate a recipe that takes their health into consideration."
[1927] Route guidance and public transport linkage system
[1928] The route guidance and public transport integration system for travel is a system that presents the optimal travel route based on the user's current location and destination. It provides the optimal route by including real-time traffic information. This system uses the following hardware and software.
[1929] Hardware
[1930] GPS-enabled devices (smartphones, tablets)
[1931] software
[1932] Real-time traffic information API (e.g., Google Maps API, HERE API)
[1933] Route Calculation Algorithm
[1934] Location Tracking Software
[1935] Specific examples
[1936] When a user wants to travel from their home to the station, the system calculates the shortest and most optimal route based on the latest traffic information and provides the user with specific instructions such as "turn right at the next traffic light." The smartphone obtains the user's current location using GPS and sends a request to the server. The server calculates a route based on real-time traffic information and sends it to the smartphone. The smartphone then provides the calculated route to the user by voice or on-screen display, and tracks the user's current location in real time while traveling, updating the route as needed.
[1937] Prompt Sentence Examples
[1938] "When a user requests 'Take me to the station,' generate a plan that provides the optimal route."
[1939] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1940] Conversation support system
[1941] Program processing
[1942] Step 1:
[1943] The user gives a voice input of "What's the weather like today?" The device captures this voice input using the built-in microphone. The input is the user's voice data, and the output is the captured voice data. The device collects voice data in real time.
[1944] Step 2:
[1945] The device sends the captured voice data to a voice recognition engine and converts it into text data. Specifically, the voice data is sent to a voice recognition service on the cloud, and text data is obtained. The input here is voice data, and the output is converted text data. The device performs the operation of converting voice data into text data.
[1946] Step 3:
[1947] The terminal sends the converted character data to the server. The data is sent securely using a secure protocol (such as HTTPS). The input here is character data, and the output is the character data sent to the server. The terminal performs the data sending operation.
[1948] Step 4:
[1949] The server inputs the received text data into a natural language processing engine (NLP engine) for analysis. The NLP engine analyzes the text data and generates an appropriate response. The input here is text data, and the output is the generated response data. The server performs the operations of analyzing the text data and generating a response.
[1950] Step 5:
[1951] The server sends the generated response data to the terminal. Again, communication is performed using a secure protocol. The input here is the response data, and the output is the response data sent to the terminal. The server performs the data transmission operation.
[1952] Step 6:
[1953] The terminal converts the received response data into speech using a speech synthesis engine. The speech synthesis engine is used to convert text data into speech data, which is then provided to the user through a speaker. The input here is the response data, and the output is the generated speech data. The terminal then performs the operation of playing back the speech data.
[1954] Cooking and meal follow-up system
[1955] Program processing
[1956] Step 1:
[1957] The user inputs a request into the device by voice or text, such as "I want to make a dietary lunch." The device captures this input. The input is the user's request data, and the output is the captured request data. The device performs the operation of collecting voice and text data.
[1958] Step 2:
[1959] The terminal sends the request data, pre-registered health condition data, and ingredient information to the server. The data is sent using a secure protocol. The input is the request data and the attached health condition data and ingredient information, and the output is the data sent to the server. The terminal performs the data transmission operation.
[1960] Step 3:
[1961] The server analyzes the received data and generates appropriate cooking methods and recipes that take health status into consideration. The input here is the request data, health status data, and ingredient information, and the output is the generated recipe data. The server performs the data analysis and recipe generation operations.
[1962] Step 4:
[1963] The server sends the generated recipe data to the terminal. Communication is again performed using a secure protocol. The input is the recipe data, and the output is the recipe data sent to the terminal. The server performs the data transmission operation.
[1964] Step 5:
[1965] The terminal provides the received recipe data to the user by voice or screen display. The input here is recipe data, and the output is recipe information provided to the user. The terminal performs the operations of displaying information and playing voice.
[1966] Step 6:
[1967] If necessary, the terminal sends specific cooking operation instructions to the robot assistant. The input is cooking operation instruction data, and the output is instruction data sent to the robot assistant. The terminal performs the operation of sending instructions.
[1968] Route guidance and public transport linkage system
[1969] Program processing
[1970] Step 1:
[1971] The user inputs a request into the terminal by voice or text, such as "Take me to the station." The terminal captures this input. The input is the user's request data, and the output is the captured request data. The terminal performs the operation of collecting voice and text data.
[1972] Step 2:
[1973] The device uses the built-in GPS to obtain the user's current location. The input here is GPS data, and the output is the obtained current location data. The device performs the operation of collecting location data.
[1974] Step 3:
[1975] The device sends the request data and current location data to the server. The data is sent using a secure protocol. The input is the request data and current location data, and the output is the data sent to the server. The device performs the data sending operation.
[1976] Step 4:
[1977] The server obtains real-time traffic information and calculates the optimal route. The input here is the current location data and real-time traffic information, and the output is the calculated route data. The server performs data analysis and route calculation.
[1978] Step 5:
[1979] The server sends the calculated route data to the terminal. Communication is again performed using a secure protocol. The input is the route data, and the output is the route data sent to the terminal. The server performs the data transmission operation.
[1980] Step 6:
[1981] The terminal provides the received route data to the user by voice or screen display. The input here is the route data, and the output is the route information provided to the user. The terminal performs the operations of displaying information and playing voice.
[1982] Step 7:
[1983] While moving, the device uses GPS to track the user's current location and updates the route as needed. The input here is GPS data and real-time traffic information, and the output is updated route data. The device performs location tracking and route update operations.
[1984] (Application example 1)
[1985] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1986] This invention relates to a system for supporting and ensuring the safety of specific groups, such as the elderly, in their daily lives. In particular, the system aims to reduce the burden on caregivers and support the independent living of those receiving care by utilizing the user's voice input and location information to provide comprehensive support tailored to a wide range of needs, including conversation assistance, cooking and meal follow-up, and emergency response. There is also a need for a system that can respond quickly and appropriately in emergencies.
[1987] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1988] In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the text data and generating an appropriate response using a natural language processing engine, means for converting the generated response into speech and providing it to the user, means for acquiring the user's current location, means for transmitting the user's location information to the server, and means for analyzing the user's voice in an emergency and transmitting an alert together with the location information, thereby supporting the user's daily life and enabling a prompt and appropriate response, particularly in an emergency.
[1989] "Voice input" refers to data obtained by a device capturing words spoken by a user.
[1990] "Voice data" is data that digitally represents voice input obtained from a user.
[1991] "Character data" is data obtained by analyzing voice data and converting the content into text format.
[1992] A "natural language processing engine" is a program or system that analyzes text data and generates appropriate responses.
[1993] "Response generation" refers to the process of using a natural language processing engine to create an appropriate response from text data.
[1994] "Speech synthesis" is the technique of converting the generated response back into speech form.
[1995] "Current location of user" refers to the geographic location where the user is currently located.
[1996] "GPS" is a satellite positioning system that identifies a user's current location.
[1997] "Real-time traffic information" is dynamic information that shows the current traffic situation.
[1998] An "optimal route" refers to the most efficient route from the current location to the destination.
[1999] "Emergency speech analysis" is the process of identifying specific speech inputs that indicate an emergency.
[2000] "Transmitting location information" means sending data indicating the user's current location to a server or other receiving device.
[2001] "Sending an alert" means communicating a warning message based on a specific condition.
[2002] This invention is a system for supporting specific groups, such as the elderly, in their daily lives and ensuring their safety, and is realized by each of the entities, the server, the terminal, and the user, fulfilling their specific roles. The following describes in detail the mode for carrying out this invention.
[2003] Embodiment of conversation support system
[2004] The server hosts a natural language processing engine (NLP engine). The device captures the user's voice input and converts this voice data into text data. The converted text data is sent to the server, where the NLP engine generates an appropriate response. The generated response is sent back to the device, where it is converted into speech and provided to the user. This system allows users to obtain a variety of information through natural conversation.
[2005] Hardware / Software: A microphone is used for voice input, the speech_recognition library is used for voice recognition, an NLP engine is used for natural language processing, and the pyttsx3 library is used for speech synthesis.
[2006] Example: When a user asks, "What's the weather like today?", the system responds by saying, "The weather is sunny today."
[2007] Cooking and meal follow-up system
[2008] The device receives input data from the user regarding their health condition and ingredients, and sends this data to the server. Based on the received data, the server generates cooking methods and recipes optimized for the user's health condition. The generated cooking instructions are provided to the user via the device via voice or screen display. This allows users to easily cook in a way that takes their health into consideration.
[2009] Hardware and software: Smartphones and tablets are used to acquire and transmit data. AI models are used to generate recipes, and existing speech synthesis libraries and display technologies are used for voice and screen display.
[2010] Example: When a user requests to make a low-sugar dessert, the system provides the optimal recipe based on the user's health status.
[2011] Route guidance and public transport linkage system for travel
[2012] The device acquires the user's current location and sends destination information to the server. The server calculates the optimal route based on real-time traffic information and sends that information to the device. The device then provides the calculated route information to the user via voice and on-screen display. In addition, the device has the ability to analyze the user's voice in an emergency and send an alert along with location information.
[2013] Hardware and software: GPS devices, real-time traffic information platforms, and speech recognition and speech synthesis libraries are used.
[2014] Example: If a user yells "Help!", the system interprets the audio as an emergency alert and sends information, including the user's current location, to emergency contacts.
[2015] Prompt Sentence Examples
[2016] "How's the weather today?"
[2017] "I want to make a dessert with less sugar."
[2018] "help me!"
[2019] In this way, the present invention utilizes the user's voice and location information to provide a wide range of support for daily life, thereby reducing the burden on elderly people and their caregivers.
[2020] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2021] Processing steps of the conversation support system
[2022] Step 1:
[2023] The terminal obtains the user's voice input.
[2024] Specifically, when a user says, "What's the weather like today?", the device's microphone captures the voice. The input is voice data.
[2025] Step 2:
[2026] The terminal converts the captured voice data into text data.
[2027] Specifically, it uses a speech recognition library (e.g., speech_recognition) to convert voice data into text format. The input is voice data, and the output is text data.
[2028] Step 3:
[2029] The terminal transmits the converted character data to the server.
[2030] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[2031] Step 4:
[2032] The server analyzes the received text data using a natural language processing engine.
[2033] Specifically, an NLP engine (for example, a custom NLP server) analyzes the text data and generates an appropriate response to the question, "What's the weather like today?" The input is the text data, and the output is the analysis result and response data.
[2034] Step 5:
[2035] The server sends the generated response data to the terminal.
[2036] Specifically, the generated response data is sent to the terminal via an HTTP request. The input is the response data, and the output is data sent to the terminal.
[2037] Step 6:
[2038] The terminal converts the received response data into voice and provides it to the user.
[2039] Specifically, it uses a speech synthesis library (e.g., pyttsx3) to convert text-based response data into speech and provides it to the user through a speaker. The input is the response data, and the output is speech output.
[2040] Emergency response system processing steps
[2041] Step 1:
[2042] The terminal obtains the user's voice input.
[2043] Specifically, when a user shouts "Help!", the device's microphone captures the voice. The input is voice data.
[2044] Step 2:
[2045] The terminal converts the captured voice data into text data.
[2046] Specifically, it uses a speech recognition library to convert voice data into text format. The input is voice data and the output is text data.
[2047] Step 3:
[2048] The terminal transmits the converted character data to the server.
[2049] Specifically, text data is sent to the server via an HTTP request. The input is character data, and the output is data sent to the server.
[2050] Step 4:
[2051] The server analyzes the text data and identifies it as an urgent message.
[2052] Specifically, the NLP engine recognizes the emergency message "Help!" and prepares an appropriate response. The input is text data, and the output is emergency alert data.
[2053] Step 5:
[2054] The device obtains the user's current location using GPS.
[2055] Specifically, the GPS module of the device acquires the current location information. The input is a signal from the GPS device, and the output is location information data.
[2056] Step 6:
[2057] The server sends location information and emergency alerts to emergency contacts.
[2058] Specifically, location information and an emergency message are sent to emergency contacts (e.g., family members or emergency services) via an HTTP request. The input is the emergency alert data and location data, and the output is the transmission of the data to the contacts.
[2059] Step 7:
[2060] The terminal provides a voice message to reassure the user.
[2061] Specifically, it uses a speech synthesis library to generate a message such as "Rescue is currently being called. Please rest assured," and provides it to the user through a speaker. The input is the data for the reassurance message, and the output is voice output.
[2062] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2063] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live a comfortable life. Specific embodiments of this invention are described below.
[2064] In this system, each of the server, terminal, and user plays a role in recognizing the user's emotions and providing optimal responses and support to the user's mental and physical needs.
[2065] 1. Conversation support system
[2066] Overview of conversation support
[2067] The server hosts an emotion engine and a natural language processing (NLP) engine, which recognizes emotions from the user's voice input and generates appropriate responses to provide to the user.
[2068] Program processing
[2069] 1. Device: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[2070] 2. Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[2071] 3. Server: Analyzes the received text data using an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger).
[2072] 4. Server: Uses a natural language processing engine to generate an appropriate response based on context and sentiment, such as "The weather is sunny today."
[2073] 5. Server: Sends a tone based on the generated response data and emotional state to the terminal.
[2074] 6. Terminal: The received response data is converted into voice with an appropriate tone using a voice synthesis function and provided to the user.
[2075] Specific examples
[2076] When a user asks, "What's the weather like today?", the system analyzes the meaning of the question, and if it recognizes that the user is feeling "depressed," it responds by adding encouraging words such as, "Cheer up, it's sunny today, why don't you go outside?"
[2077] 2. Cooking and meal follow-up system
[2078] Cooking and Meal Follow-up Overview
[2079] This system provides optimal cooking instructions based on the user's health and emotional state.
[2080] Program processing
[2081] 1. Device: The user inputs a request such as "I want to make a dietary lunch." This information, along with previously entered health status, ingredient information, and emotional data, is sent to the server.
[2082] 2. Server: Based on the received data, it generates cooking methods and recipes that take into account the health and emotional state of the user.
[2083] 3. Server: Sends the generated cooking instructions and related emotional support information to the terminal.
[2084] 4. Terminal: Provides the received cooking instructions and emotional support information (e.g., suggestions for ingredients with a relaxing effect when the user is feeling stressed) to the user via voice or on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[2085] Specific examples
[2086] If a user requests "I want to make a healthy breakfast," the system will provide a relaxing recipe using oatmeal based on the user's health status and feelings of "tension." Furthermore, if a robotic assistant is needed, it will automate some of the oatmeal cooking process.
[2087] 3. Route Guidance and Public Transportation Integration System
[2088] Overview of route guidance during travel
[2089] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[2090] Program processing
[2091] 1. Device: The user inputs a request such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends the information to the server.
[2092] 2. Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination, taking into account the emotional state.
[2093] 3. Server: Sends the calculated route information and emotional support information (for example, calming voice guidance if the user is feeling anxious) to the device.
[2094] 4. Device: Provides the user with the received route information and emotional support information via voice and screen display. It also tracks the user's current location in real time while moving and updates the route as needed.
[2095] Specific examples
[2096] If the system detects that the user is "nervous" while traveling from home to the station, it calculates the shortest and safest route based on the latest traffic information and provides specific instructions in a gentle tone, such as "Turn right at the next traffic light." If the user deviates from the instructions, the server immediately recalculates and provides new route guidance.
[2097] As described above, the AI system combined with the emotion engine provides services that take into account the user's mental state, aiming to reduce the burden of caregiving and improve the user's quality of life.
[2098] The processing flow will be explained below.
[2099] Conversation support processing flow
[2100] Step 1:
[2101] User: The user speaks, "What's the weather like today?" This speech is captured in real time by the device.
[2102] Step 2:
[2103] Terminal: The captured voice data is converted into text data using a voice recognition function, and the converted results are sent to the server.
[2104] Step 3:
[2105] Server: Analyzes the user's emotional state using the emotion engine and sends the results of this analysis to the natural language processing engine.
[2106] Step 4:
[2107] Server: Uses a natural language processing engine to generate appropriate responses based on context and sentiment.
[2108] Step 5:
[2109] Server: Sends tone information based on the generated response data and emotional state to the terminal.
[2110] Step 6:
[2111] Terminal: The terminal converts the received response data into speech with an appropriate tone using a speech synthesis function and provides it to the user.
[2112] Cooking and meal follow-up process flow
[2113] Step 1:
[2114] User: The user inputs a request by voice or text, such as "I want to make a dietary lunch." The device then sends this information, along with previously entered health status and ingredient information, to the server.
[2115] Step 2:
[2116] Server: Based on the received information about the user's health condition and ingredients, the emotion engine analyzes the user's emotional state.
[2117] Step 3:
[2118] Server: Generates cooking methods and recipes that take into account health and emotional states.
[2119] Step 4:
[2120] Server: Sends the generated cooking instructions and emotional support information to the terminal.
[2121] Step 5:
[2122] Terminal: Provides the received cooking instructions and emotional support information to the user via voice and on-screen display. It also sends signals to instruct the robot assistant on cooking operations as needed.
[2123] Step 6:
[2124] Robot: Based on instructions from a terminal, the robot begins specific cooking tasks. For example, it automatically performs basic operations such as chopping vegetables and putting food in a pot.
[2125] Processing flow of the system for linking route guidance and public transportation during travel
[2126] Step 1:
[2127] User: The user inputs a request by voice or text, such as "Take me to the station." The device obtains its current location using GPS, sets "station" as the destination, and sends this information to the server.
[2128] Step 2:
[2129] Server: Obtains real-time traffic information and calculates the optimal route from the current location to the destination.
[2130] Step 3:
[2131] Server: Based on the received user's emotional state, the server adjusts the guidance method (tone and pace) for the optimal route.
[2132] Step 4:
[2133] Server: Sends the calculated route information and emotional support information to the terminal.
[2134] Step 5:
[2135] Terminal: Provides the received route information and emotional support information to the user via voice and on-screen display. For example, it provides specific guidance such as "Turn right at the next traffic light," but changes the tone depending on the user's emotional state.
[2136] Step 6:
[2137] Device: GPS information is constantly updated while moving, tracking the user's current location in real time. Each time, the server recalculates the optimal route based on the user's emotional state and provides new route guidance.
[2138] Step 7:
[2139] User: Follow the directions and reach the destination. During this time, the system will constantly update and provide optimal guidance.
[2140] This detailed processing flow allows the system to provide more personalized assistance while taking into account the user's emotional state, thereby reducing the burden of caregiving and enabling users to live a comfortable and efficient life.
[2141] Example 2
[2142] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2143] In recent years, the burden of caregiving has increased with the progress of the aging society. There is also a demand for support to maintain users' mental and physical health. However, conventional AI systems do not take into account the user's emotional state, limiting the user experience. Furthermore, even in health management and mobility assistance, they can only make suggestions that ignore the user's emotional state, making it difficult to provide accurate support.
[2144] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for acquiring a user's voice input and converting the voice data into text data, means for analyzing the user's emotional state from the text data using an emotion engine, means for generating an appropriate response based on the context and the emotional state using a natural language processing engine, and means for converting the generated response into voice and providing it to the user. This makes it possible to provide appropriate responses and support according to the user's emotional state.
[2145] "Audio input" refers to an audio signal emitted by a user through an input device such as a microphone.
[2146] "Audio Data" means data in digital form obtained from audio input.
[2147] "Character data" refers to text-format data obtained by analyzing voice data using voice recognition technology.
[2148] An "emotion engine" is software or hardware that analyzes text data and estimates the user's emotional state.
[2149] An "emotional state" is a mental state that a user is feeling, and examples include happiness, sadness, anger, etc.
[2150] A "natural language processing engine" is software that analyzes text data, understands meaning and context, and generates appropriate responses.
[2151] A "response" is a reply or instruction to the user generated by a natural language processing engine.
[2152] "Speech synthesis" is a technology that converts text data into audio data and turns it into audio that can be played through speakers, etc.
[2153] "Health status" refers to information about the user's current physical health, including information about illnesses, symptoms, and physical conditions.
[2154] "Ingredient information" refers to data relating to the types and amounts of ingredients available to the user.
[2155] A "cooking method" is the procedure or process of preparing a dish using ingredients.
[2156] A "recipe" is a document that describes the ingredients and steps for making a particular dish.
[2157] A "location information system" is a system that obtains a user's current location using technology such as GPS.
[2158] "Real-time traffic information" refers to information about current traffic conditions and transportation options.
[2159] An "optimal route" is the most efficient route for a user to travel from their current location to their destination.
[2160] This invention utilizes an AI system combined with an emotion engine to reduce the burden of caregiving and help users live more comfortable lives. Specific embodiments of this invention will be described in detail below.
[2161] 1. Conversation support system
[2162] System Configuration
[2163] In this system, users ask questions or make requests by voice, and the server processes them and provides a voice response. The main components are as follows:
[2164] Terminal: A device that captures the user's voice and sends the voice data to the server. It uses the Google Speech-to-Text API for voice recognition.
[2165] Server: Acquires text data and analyzes it using an emotion engine (e.g., IBM Watson Tone Analyzer). OpenAI's GPT-3 natural language processing engine is used.
[2166] Speech synthesis engine: Converts the generated text response into speech using Amazon Polly.
[2167] Specific examples
[2168] When a user asks, "What's the weather like today?", the voice data captured by the device is converted into text data and sent to the server. The server uses an emotion engine to recognize the emotional state of "interested" and generates an appropriate response using a natural language processing engine. For example, it might respond, "The weather is sunny today." This is then converted into speech using a speech synthesis engine and provided to the user via the device.
[2169] Prompt Sentence Examples
[2170] User: "What's the weather like today?"
[2171] system:
[2172] 1. The server converts the speech into text and analyzes it using an emotion engine.
[2173] 2. If the emotion is perceived as "interested," generate an appropriate response.
[2174] 2. Cooking and meal follow-up system
[2175] System Configuration
[2176] This system provides optimal cooking instructions based on the user's health and emotional state.
[2177] Device: Captures user requests and sends them to the server. The voice recognition function uses the Google Speech-to-Text API.
[2178] Server: Analyzes health status, ingredient information, and emotional state to generate cooking methods and recipes. IBM Watson Tone Analyzer is used as the emotion engine, and OpenAI's GPT-3 is used to generate recipes.
[2179] Robot assistant: A device that automates cooking operations as needed.
[2180] Specific examples
[2181] If a user requests "I want to make a dietary lunch," the device sends this information to the server. The server uses its emotion engine to recognize that the user is "feeling stressed" and generates a cooking method using ingredients that have a relaxing effect. The generated recipe is for oatmeal and salad and is provided to the user via voice and on-screen display. If necessary, some of the cooking can be automated by a robotic assistant.
[2182] Prompt Sentence Examples
[2183] User: "I want to make a diet lunch."
[2184] system:
[2185] 1. The server receives the request and user data and generates an appropriate recipe.
[2186] 2. The recipe uses oatmeal and provides audio instructions.
[2187] 3. Route Guidance and Public Transportation Integration System
[2188] System Configuration
[2189] This system presents the optimal route based on the user's current location and destination, taking into account their emotional state.
[2190] Terminal: Receives the user's request, obtains the current location using GPS, and sends it to the server.
[2191] Server: Obtains real-time traffic information and calculates the optimal route. IBM Watson Tone Analyzer is used as the emotion engine, and Google Maps API is used to obtain traffic information.
[2192] Speech synthesis engine: Converts route directions into speech.
[2193] Specific examples
[2194] When a user requests "Guide me to the station," the device obtains the user's current location, sets "station" as the destination, and sends the information to the server. The server then uses its emotion engine to recognize the user as "stressed" and calculates the optimal route to reassure the user. For example, the server uses a speech synthesis engine to convert specific instructions, such as "Turn right at the next traffic light," into voice and provides the information to the user via the device.
[2195] Prompt Sentence Examples
[2196] User: "Take me to the station."
[2197] system:
[2198] 1. The server calculates a route based on the current location and destination, and provides guidance according to the user's emotional state.
[2199] 2. If they are feeling anxious, guide them with a reassuring route and a gentle tone.
[2200] As described above, the present invention can provide support in various situations while taking into consideration the emotional state of the user, reduce the burden of caregiving, and improve the quality of life of the user.
[2201] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2202] Conversation support system processing flow
[2203] Step 1:
[2204] Description: The user provides voice input.
[2205] Specific action: The user speaks to the device, "What's the weather like today?"
[2206] Input: User's voice input.
[2207] Output: The captured audio data.
[2208] Step 2:
[2209] Description: The device receives voice data and converts it into text data using voice recognition.
[2210] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[2211] Input: Captured audio data.
[2212] Output: The converted text data (e.g., "What's the weather like today?").
[2213] Step 3:
[2214] Description: The terminal sends the converted character data to the server.
[2215] Specific operation: The terminal sends an HTTP request to the server and sends character data.
[2216] Input: The converted character data.
[2217] Output: Character data sent to the server.
[2218] Step 4:
[2219] Description: The server receives text data and analyzes it with the emotion engine.
[2220] Specific operation: The server uses IBM Watson Tone Analyzer to analyze the user's emotional state (e.g., "interested") from the text data.
[2221] Input: The character data sent.
[2222] Output: Parsed emotional state.
[2223] Step 5:
[2224] Description: The server uses a natural language processing engine to generate appropriate responses based on context and emotional state.
[2225] Specific operation: The server calls OpenAI's GPT-3 and generates an appropriate response sentence using the text data and emotional state as input (e.g., "The weather is sunny today.").
[2226] Input: Text data, emotional state.
[2227] Output: The generated response sentence.
[2228] Step 6:
[2229] Description: The server sends the generated response to the terminal.
[2230] Specific operation: The server sends an HTTP response to the terminal and sends the generated response text.
[2231] Input: The generated response sentence.
[2232] Output: The response sent to the terminal.
[2233] Step 7:
[2234] Description: The response received by the terminal is converted into speech using the speech synthesis function.
[2235] Specific operation: The device uses Amazon Polly to convert the response into voice data.
[2236] Input: The response received.
[2237] Output: The converted audio data.
[2238] Step 8:
[2239] Description: The device plays the converted audio data to the user.
[2240] Specific operation: Play audio data through the device speaker.
[2241] Input: The converted audio data.
[2242] Output: A spoken response to the user.
[2243] Cooking and meal follow-up system processing flow
[2244] Step 1:
[2245] Description: A user enters a cooking request.
[2246] Specific operation: The user speaks to the device, saying, "I want to make a dietary lunch."
[2247] Input: User's voice input.
[2248] Output: The captured audio data.
[2249] Step 2:
[2250] Description: The device receives voice data and converts it into text data using voice recognition.
[2251] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[2252] Input: Captured audio data.
[2253] Output: The converted text (e.g. "I want to make a dietary lunch").
[2254] Step 3:
[2255] Description: The device acquires health status, food ingredient information, and emotional state and sends them to the server.
[2256] Specific operation: The device obtains pre-registered health status, food information, and emotional state, and sends an HTTP request to the server.
[2257] Input: Text data, health status, food information, emotional state.
[2258] Output: The data sent to the server.
[2259] Step 4:
[2260] Description: The server generates cooking instructions and emotional support information based on the data received.
[2261] Specific operation: Using an emotion engine (e.g., IBM Watson Tone Analyzer) and a natural language processing engine (e.g., OpenAI's GPT-3), it generates appropriate cooking methods, recipes, and emotional support information.
[2262] Input: health status, food information, emotional state.
[2263] Output: Generated cooking instructions and emotional support information.
[2264] Step 5:
[2265] Description: The server sends the generated cooking instructions and emotional support information to the terminal.
[2266] Specific operation: The server sends an HTTP response to the terminal, conveying instructions and information.
[2267] Input: Generated cooking instructions and emotional support information.
[2268] Output: Cooking instructions and emotional support information sent to the device.
[2269] Step 6:
[2270] Description: The device provides the user with cooking instructions and emotional support information received.
[2271] How it works: The device uses speech synthesis to convert instructions into voice, provides information on the screen, and issues instructions to the robot assistant as needed.
[2272] Input: cooking instructions and emotional support information.
[2273] Output: Voice and visual instructions, instruction signals to the robot assistant.
[2274] Processing flow of the system for linking route guidance and public transportation during travel
[2275] Step 1:
[2276] Description: User requests directions.
[2277] Specific operation: The user speaks to the terminal, saying, "Take me to the station."
[2278] Input: User's voice input.
[2279] Output: The captured audio data.
[2280] Step 2:
[2281] Description: The device receives voice data and converts it into text data using voice recognition.
[2282] Specific operation: The device calls the Google Speech-to-Text API and processes the audio data for speech recognition.
[2283] Input: Captured audio data.
[2284] Output: The converted text data (e.g. "Take me to the station").
[2285] Step 3:
[2286] Description: The device obtains its current location using GPS and sends destination information and emotional state to the server.
[2287] Specific operation: The device acquires GPS data and sends an HTTP request to the server along with the user's emotional state information.
[2288] Input: Text data, current location information, destination information, emotional state.
[2289] Output: The data sent to the server.
[2290] Step 4:
[2291] Description: The server calculates the optimal route based on real-time traffic information and emotional state.
[2292] Specific operation: The server uses the Google Maps API to calculate the optimal rout...
Claims
1. means for receiving a user's voice input and converting the voice data into text data; means for analyzing the text data and generating an appropriate response using a natural language processing engine; The system includes means for converting the generated response into speech and providing it to the user.
2. A means for acquiring user input data and transmitting health status and food ingredient information to a server; A means for generating cooking methods and recipes based on the user's health condition and ingredient data; The system according to claim 1, further comprising means for providing the generated cooking instructions to the user by voice or by display on a screen.
3. A means for acquiring the user's current location by GPS and transmitting destination information to a server; means for calculating an optimal route based on real-time traffic information; 2. The system according to claim 1, further comprising means for providing the calculated route information to the user by voice or by display on a screen.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A