system

A generative AI-based system addresses the challenges of elderly care by providing personalized conversational support and care assistance, enhancing mental and physical well-being through daily interactions and health-adapted advice.

JP2026064647APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The super-aging population and caregiver shortage exacerbate the challenges of providing appropriate care for the elderly, leading to loneliness and deteriorating mental and physical health due to limited daily conversation and timely care opportunities.

Method used

A system utilizing generative AI for conversational support and care assistance, including user authentication, voice input conversion to text, data transmission to a server, conversation content generation, and voice playback, with health status management and advice generation based on collected data, to provide personalized responses and advice.

Benefits of technology

The system effectively alleviates loneliness, maintains cognitive function, and improves caregiving quality by enabling daily conversations and care assistance tailored to the elderly's needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064647000001_ABST
    Figure 2026064647000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for performing user authentication, A means of managing user profile information, A means of converting voice input to text, A means of sending the converted text data to the server, A means of generating conversation content based on received text data, A means of converting the generated conversation content into speech, A means of playing it back as audio, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, the super-aging population is progressing, and the loneliness of the elderly and the risk of dementia have become social problems. Also, in the field of caregiving, the population decline and caregiver shortage are worsening, making it difficult to provide appropriate care for the elderly. Therefore, there is a need for an effective and efficient elderly support system to maintain the mental and physical health of the elderly and improve the quality of caregiving.

Means for Solving the Problems

[0005] This invention provides a system utilizing generative AI for providing conversational support and care assistance to the elderly. This system includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to a server, means for generating conversation content based on the received text data, means for converting the generated conversation content to voice, and means for playing it back as voice. Furthermore, by including means for managing the user's health status, means for collecting data on the user's physical condition and concerns, means for generating appropriate advice based on the collected data, means for converting the generated advice to voice, and means for playing it back as voice, the mental and physical care of the elderly can be efficiently provided. By storing the user's profile information and health status in a database and referring to past health history, it is possible to provide more personalized responses and advice.

[0006] User authentication is a method of verifying a user's identity when they access a system.

[0007] "Profile information" refers to data that includes a user's personal information and specific characteristics.

[0008] "Voice input" is a method of inputting the voice spoken by the user into the system.

[0009] "Converting to text" is the process of changing audio data into written information.

[0010] A "server" is a remote computer device used for storing and processing data.

[0011] "Conversation content" refers to the text data of the response that the system generates in response to user input.

[0012] "Converting to speech" is the process of changing text data into audio data.

[0013] "Playback" is an operation to let the user hear the generated voice data.

[0014] "Health condition" is information regarding the user's physical condition and medical situation.

[0015] "Database" is a storage system for storing user information and history.

[0016] "Advice" is a proposal or recommendation to assist the user in problem-solving and health management.

[0017] "Collect" is a process of gathering information.

Brief Description of Drawings

[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10]Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the language used in the following description will be described.

[0021] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0022] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. The aim of this system is to alleviate the user's feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support.

[0040] System Configuration

[0041] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using an AI model.

[0042] Basic operation

[0043] User authentication and profile information management

[0044] User:

[0045] Users log in to the system by using voice commands or touch controls on the device.

[0046] Terminal:

[0047] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[0048] Send authentication information to the server.

[0049] server:

[0050] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[0051] If authentication is successful, the user's profile information will be sent to the device.

[0052] Everyday conversation

[0053] User:

[0054] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[0055] Terminal:

[0056] The device converts voice input into text and sends the text data to the server.

[0057] server:

[0058] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[0059] The generated conversation content is sent to the terminal as text data.

[0060] Terminal:

[0061] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0062] Caregiving assistance

[0063] User:

[0064] The user says, "My lower back has been a little sore since yesterday."

[0065] Terminal:

[0066] The device converts voice input into text and sends the text data to the server.

[0067] server:

[0068] The server compares the received text data with the database and refers to the user's past health history.

[0069] It generates appropriate advice and sends text data to the device saying, "Keep your lower back warm. Use a bath towel to cool it down if necessary."

[0070] Terminal:

[0071] The device converts the received text data into speech and plays it back to the user.

[0072] Specific example

[0073] The following will explain this with specific examples.

[0074] Example 1: Everyday conversation

[0075] The user launches the smartphone app and performs fingerprint authentication.

[0076] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[0077] The device receives this, converts the audio to text, and sends it to the server.

[0078] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[0079] The device converts this into audio and plays it back.

[0080] Example 2: Caregiving advice

[0081] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[0082] The device converts the audio to text and sends it to the server.

[0083] The server references the user's health history and generates appropriate advice.

[0084] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0085] The device converts the audio to sound and plays it back.

[0086] As described above, this system utilizes generative AI to efficiently perform daily conversations and provide care assistance for the elderly. This system is expected to alleviate feelings of loneliness among the elderly, maintain cognitive function, and improve the quality of care.

[0087] The following describes the processing flow.

[0088] User authentication and profile information management

[0089] User

[0090] Step 1:

[0091] Launch the smartphone app.

[0092] terminal

[0093] Step 2:

[0094] It displays a screen for fingerprint authentication, facial recognition, or entry of a username and password.

[0095] User

[0096] Step 3:

[0097] You can use fingerprint authentication, facial recognition, or enter your username and password.

[0098] terminal

[0099] Step 4:

[0100] Send authentication information to the server.

[0101] server

[0102] Step 5:

[0103] The received authentication information is compared with the database to authenticate the user.

[0104] Step 6:

[0105] If authentication is successful, the user's profile information is retrieved from the database and sent to the device.

[0106] terminal

[0107] Step 7:

[0108] The received profile information is cached to prepare for the next process.

[0109] Everyday conversation

[0110] User

[0111] Step 1:

[0112] I say to the device, "Good morning, it's a beautiful day today."

[0113] terminal

[0114] Step 2:

[0115] It receives voice input and records it as audio data.

[0116] Step 3:

[0117] Converting audio data into text data (speech recognition processing).

[0118] Step 4:

[0119] Send the converted text data to the server.

[0120] server

[0121] Step 5:

[0122] The received text data is input into an AI model to generate appropriate conversation content (natural language processing).

[0123] Step 6:

[0124] The generated conversation content is sent to the terminal as text data.

[0125] terminal

[0126] Step 7:

[0127] Converts received text data into audio data (speech synthesis).

[0128] Step 8:

[0129] The generated audio data is played for the user.

[0130] Caregiving assistance

[0131] User

[0132] Step 1:

[0133] I spoke to the device, saying, "My lower back has been a little sore since yesterday."

[0134] terminal

[0135] Step 2:

[0136] It receives voice input and records it as audio data.

[0137] Step 3:

[0138] Converting audio data into text data (speech recognition processing).

[0139] Step 4:

[0140] Send the converted text data to the server.

[0141] server

[0142] Step 5:

[0143] The received text data is compared against a database to retrieve past health history.

[0144] Step 6:

[0145] Based on the acquired health history, appropriate advice is generated (health analysis and advice generation).

[0146] Step 7:

[0147] The generated advice is sent to the device as text data.

[0148] terminal

[0149] Step 8:

[0150] Converts received text data into audio data (speech synthesis).

[0151] Step 9:

[0152] The generated advice will be played back as audio.

[0153] The above is a detailed explanation of the program's processing flow.

[0154] (Example 1)

[0155] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0156] In modern society, there is a need to alleviate loneliness among the elderly, maintain cognitive function, and provide appropriate care support. However, many elderly people have limited opportunities for daily conversation and timely care, putting them at high risk of deteriorating mental and physical health. To solve this problem, a more effective and efficient system is needed.

[0157] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0158] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content using a generative AI model based on the received text data, means for converting the generated conversation content to speech, and means for playing it back as speech. This makes it possible for elderly people to receive everyday conversations and care advice without feeling lonely. Furthermore, by combining means for managing the user's health status, means for collecting data on the user's physical condition and problems, means for generating appropriate advice using a generative AI model based on the collected data, means for converting the generated advice to speech, and means for playing it back as speech, it becomes possible to provide appropriate responses based on the user's past health history.

[0159] "User authentication" is the process of verifying that the user accessing the system is indeed that specific user.

[0160] "Profile information" refers to data that includes personal information and attributes about the user.

[0161] "Means for converting voice input to text" refers to devices or software that use speech recognition technology to convert a user's speech into text format.

[0162] "Means of sending converted text data to the server" refers to the process of sending text data generated on the terminal to the server via the network.

[0163] A "generative AI model" is an algorithm or software that uses natural language processing and machine learning to generate appropriate responses from input text data.

[0164] "Means for generating conversation content" refers to the process of generating appropriate text for interacting with the user using a generation AI model based on received text data.

[0165] "Means for converting generated conversation content into speech" refers to devices or software that convert the text of generated conversation content into speech format using speech synthesis technology.

[0166] "Means of playback as audio" refers to the process of playing audio data to the user through audio equipment such as speakers.

[0167] "Means for managing health status" refers to devices and software that record and monitor a user's health information and periodically evaluate that status.

[0168] "Means of collecting data on health and problems" refers to the process of collecting information on health and problems through user feedback and questions.

[0169] "Means for generating appropriate advice" refers to the process of generating appropriate advice and suggestions for users using a generative AI model based on collected data and past history.

[0170] "A means of referencing past health history based on saved data" refers to the process of searching for a user's health information stored in a database and referencing their past health status and history.

[0171] Modes for carrying out the invention

[0172] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. This system aims to alleviate feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support by implementing user authentication, daily conversation, and care assistance functions.

[0173] System Configuration

[0174] The system consists of user terminals and servers that process data.

[0175] A terminal is a device that performs voice input and output, and includes smartphones, tablets, and smart speakers. A server is a computer device that performs natural language processing using AI models, and examples include cloud servers and dedicated servers.

[0176] Hardware and software usage

[0177] The system uses the following specific hardware and software.

[0178] Speech recognition software for converting voice input to text (e.g., Google® Cloud Speech-to-Text, IBM Watson® Speech to Text)

[0179] Generative AI models for generating conversation content based on text data (e.g., OpenAI® GPT-3®)

[0180] Text-to-speech software for converting text data into speech (e.g., Amazon Polly, Google Cloud Text-to-Speech)

[0181] Authentication technologies for user authentication (e.g., fingerprint recognition, facial recognition, username and password)

[0182] User authentication and profile information management

[0183] The user attempts to log in by using voice or touch controls on the device.

[0184] The device accepts fingerprint authentication, facial recognition, and username and password input, and sends the authentication information to the server.

[0185] The server compares the received authentication information with the database to determine whether authentication was successful. If authentication is successful, the user's profile information is sent to the device.

[0186] Everyday conversation

[0187] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[0188] The device converts voice input into text and sends that text data to the server.

[0189] The server uses an AI model to generate appropriate conversation content based on the received text data. The generated conversation content is then sent to the terminal as text data.

[0190] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0191] Specific examples of prompt statements are as follows:

[0192] User: Good morning, it's a beautiful day today.

[0193] AI Model: Good morning. It's sunny today. Shall we go for a walk?

[0194] Caregiving assistance

[0195] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[0196] The device converts voice input into text and sends that text data to the server.

[0197] The server uses the received text data to compare and refer to the user's past health history in a database. Then, using a generative AI model, it generates appropriate advice and sends text data to the terminal such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0198] The device converts the received text data into speech and plays it back to the user.

[0199] Specific examples of prompt statements are as follows:

[0200] User: My lower back has been a little sore since yesterday.

[0201] AI Model: Keep your lower back warm. Use a bath towel to cool it down if necessary.

[0202] The above describes the embodiments for carrying out the present invention. This system provides assistance with daily conversation and care for the elderly, contributing to the alleviation of feelings of loneliness, the maintenance and improvement of cognitive function, and the improvement of the quality of care.

[0203] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0204] User Authentication

[0205] Processing flow

[0206] Step 1:

[0207] The user attempts to log in by using voice or touch controls on the device.

[0208] Input: User voice or touch operation

[0209] Output: Login attempt data

[0210] Step 2:

[0211] The device accepts fingerprint authentication, facial recognition, and username and password input.

[0212] Input: User's fingerprint data, facial data, username and password

[0213] Output: Authentication information data

[0214] Step 3:

[0215] The terminal sends the entered authentication information to the server.

[0216] Input: Authentication information data

[0217] Output: Authentication request to the server

[0218] Step 4:

[0219] The server compares the received authentication information with the database to determine whether the authentication was successful.

[0220] Input: Authentication information data

[0221] Output: Authentication success / failure result

[0222] Step 5:

[0223] If authentication is successful, the server sends the user's profile information to the device.

[0224] Input: Authentication success / failure result

[0225] Output: Profile information data

[0226] Everyday conversation

[0227] Processing flow

[0228] Step 1:

[0229] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[0230] Input: User voice input

[0231] Output: Audio data

[0232] Step 2:

[0233] The device converts voice input into text and sends that text data to the server.

[0234] Input: Audio data

[0235] Output: Text data, request to send to the server

[0236] Step 3:

[0237] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[0238] Input: Text data

[0239] Output: Generated conversation text

[0240] Step 4:

[0241] The server sends the generated conversation content to the terminal as text data.

[0242] Input: Generated conversation text

[0243] Output: Conversation text data, request to send to terminal

[0244] Step 5:

[0245] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0246] Input: Conversation text data

[0247] Output: Audio data

[0248] Caregiving assistance

[0249] Processing flow

[0250] Step 1:

[0251] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[0252] Input: User voice input

[0253] Output: Audio data

[0254] Step 2:

[0255] The device converts voice input into text and sends that text data to the server.

[0256] Input: Audio data

[0257] Output: Text data, request to send to the server

[0258] Step 3:

[0259] The server uses the received text data to compare and reference the user's past health history from the database.

[0260] Input: Text data

[0261] Output: Health history data

[0262] Step 4:

[0263] The server uses a generative AI model to generate appropriate advice and sends the text data to the terminal.

[0264] Input: Health history data

[0265] Output: Generated advice text data, request to send to terminal

[0266] Step 5:

[0267] The device converts the received text data into speech and plays back, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0268] Input: Advice text data

[0269] Output: Audio data

[0270] (Application Example 1)

[0271] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0272] Elderly people often find food delivery services difficult to use due to their complex operation. Furthermore, they struggle to select nutritionally balanced meals, hindering proper dietary management. This highlights the growing need for food delivery applications that improve the quality of life for the elderly.

[0273] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0274] In this invention, the server includes means for generating food delivery order content using a generative AI model, means for transmitting the generated order content to a food delivery system, and means for providing the generated nutritional advice. This makes it possible for elderly people to easily order food delivery by voice and also receive appropriate nutritional advice.

[0275] User authentication is the process of identifying users who are accessing a system and verifying that they are authorized users.

[0276] "Profile information" refers to personal information and settings information about a user, and is data necessary for the system to provide services tailored to the user.

[0277] "Voice input" is the process by which a user speaks to the system through a microphone or similar device, and the system receives voice data in response.

[0278] "Converting to text" is the process of converting audio data into text data.

[0279] The "generative AI model" refers to an algorithm or program that uses artificial intelligence to generate conversation content or order content based on user input.

[0280] The "food delivery order content" refers to the details of food and drinks specified by the user when using the food delivery service.

[0281] The "food delivery system" refers to the entire service for delivering food to the orderer and its operation system.

[0282] The "nutritional advice" refers to professional advice to enable the user to choose a diet that takes health into consideration.

[0283] "Conversion to voice" refers to the process of converting character data into voice data so that it can be played back to the user.

[0284] "Playback as voice" refers to the process of playing the converted voice data to the user through a speaker or the like.

[0285] The "past health history" refers to past data and records related to the user's health status, which are used to generate current advice.

[0286] The present invention is a system that simplifies operations and simultaneously provides nutritional advice when elderly people use the food delivery service. This system includes a series of processes for voice input, conversation generation using a generative AI model, order content generation, and provision of nutritional advice.

[0287] System Configuration

[0288] The system is composed of a terminal used by the user and a server that performs data processing. The terminal is a device that performs voice input and voice output, and the server is a computer device that performs natural language processing using a generative AI model.

[0289] User authentication and profile information management

[0290] User:

[0291] The user launches the food delivery app and logs into the system using fingerprint authentication or similar methods.

[0292] Terminal:

[0293] The device accepts fingerprint authentication, facial recognition, or username and password input for user authentication.

[0294] Send authentication information to the server.

[0295] server:

[0296] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[0297] If authentication is successful, the user's profile information will be sent to the device.

[0298] Food delivery order generation

[0299] User:

[0300] The user speaks to the device and says, "I want to eat taco rice today."

[0301] Terminal:

[0302] The device converts voice input into text and sends the text data to the server.

[0303] server:

[0304] The server receives the text data and sends it as a prompt to the AI ​​model, which then generates an appropriate food delivery order.

[0305] Send the generated order content to the food delivery system.

[0306] Terminal:

[0307] The terminal converts the order confirmation content received from the server into voice and plays it for the user.

[0308] Nutritional advice

[0309] User:

[0310] The user says to the terminal, "I've been concerned about my nutritional balance lately."

[0311] Terminal:

[0312] The terminal converts the voice input into text and sends the text data to the server.

[0313] Server:

[0314] The server refers to the user's past health history and uses a generated AI model to generate appropriate nutritional advice.

[0315] Generate a message saying, "Takoyaki rice has a balanced nutrition. If you want to take additional vitamin C, let's add a salad," and send it to the terminal.

[0316] Terminal:

[0317] The terminal converts the received nutritional advice into voice and plays it for the user.

[0318] Specific example

[0319] Example of a food delivery order:

[0320] Usage example: The user says, "I want to eat takoyaki rice today."

[0321] AI: "So, you'd like taco rice. Would you like it spicy or mild?"

[0322] The user replied, "Please make it mild."

[0323] AI: "Thank you. I'll order the taco rice (mild) then."

[0324] Examples of nutritional advice:

[0325] Example of use: The user says, "Lately, I've been worried about my nutritional balance."

[0326] AI: "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[0327] Example of a prompt

[0328] "I want to eat taco rice today."

[0329] "Lately, I've been worried about my nutritional balance."

[0330] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0331] Step 1:

[0332] User: The user launches the food delivery app and logs into the system using fingerprint authentication or by entering a password. During this process, the user's fingerprint information and password are entered into the system.

[0333] Step 2:

[0334] Terminal: The terminal obtains user authentication information through on-screen input forms or fingerprint sensors and sends it to the server. In this case, the entered authentication information is the data sent to the server.

[0335] Step 3:

[0336] Server: The server verifies the received authentication information and compares it with the profile information in the database. If the comparison is successful, it sends the user's profile information back to the terminal. It verifies that the authentication information is correct and outputs the user's profile data.

[0337] Step 4:

[0338] User: After successful authentication, the user speaks into the device saying, "I want to eat taco rice today."

[0339] Step 5:

[0340] Terminal: The terminal performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process. Speech recognition software (e.g., Google Speech Recognition API) is used to convert the speech input into text.

[0341] Step 6:

[0342] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[0343] Step 7:

[0344] Server: The server receives text data and sends it to the generative AI model as a prompt. The prompt sent here becomes the input to the AI ​​model. The generative AI model (e.g., OpenAI GPT-3) is used to generate appropriate food delivery order details.

[0345] Step 8:

[0346] Server: Retrieves order details generated by the generation AI model and sends them to the food delivery system. The generated order details become the output data.

[0347] Step 9:

[0348] Terminal: Converts the order confirmation received from the server into audio. This converted audio data becomes the input data for the next process.

[0349] Step 10:

[0350] Terminal: The terminal plays back the converted order details and provides feedback to the user. It uses a voice output device to play back the text-to-speech conversion.

[0351] Step 11:

[0352] User: The user speaks to the device, saying, "Lately, I've been worried about my nutritional balance."

[0353] Step 12:

[0354] Terminal: Performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process.

[0355] Step 13:

[0356] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[0357] Step 14:

[0358] Server: The server receives text data, references the user's health history, generates prompts about nutritional balance, and sends them to an AI model. The user's health history data and prompts serve as input data.

[0359] Step 15:

[0360] Server: Retrieves nutritional advice generated by the generation AI model and sends it to the terminal as text data. The generated nutritional advice becomes the output data.

[0361] Step 16:

[0362] Terminal: Converts text data of nutritional advice received from the server into speech. This converted speech data becomes input data for the next process.

[0363] Step 17:

[0364] Device: The device plays nutritional advice converted into audio and provides feedback to the user. It uses an audio output device to play the nutritional advice.

[0365] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0366] This invention is a system that combines a generative AI, which is used to provide conversational support and care assistance for the elderly, with an emotion engine that recognizes the user's emotions. The aim of this system is to provide more effective mental care for the elderly, by offering conversations and care advice tailored to their emotional state.

[0367] System Configuration

[0368] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, while the server is a computer that uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more personalized responses.

[0369] Basic operation

[0370] User authentication and profile information management

[0371] User

[0372] Users log in to the system by using voice commands or touch controls on the device.

[0373] terminal

[0374] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[0375] Send authentication information to the server.

[0376] server

[0377] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[0378] If authentication is successful, the user's profile information will be sent to the device.

[0379] Everyday conversation

[0380] User

[0381] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[0382] terminal

[0383] The device receives voice input and records it as voice data.

[0384] The audio data is converted into text data (speech recognition processing) and sent to the server.

[0385] server

[0386] The server inputs the received text data into an AI model and generates appropriate conversation content (natural language processing).

[0387] The generated conversation content is sent to the device as text data.

[0388] terminal

[0389] The device converts the received text data into audio data (speech synthesis) and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0390] emotion recognition

[0391] User

[0392] You can talk to the device about everyday topics, your health, or any problems you're facing.

[0393] terminal

[0394] It receives voice input and records it as audio data.

[0395] The system analyzes the user's emotional state from voice data using an emotion engine and sends the data to the server.

[0396] server

[0397] The server uses an emotion engine to analyze the received audio data and determine the user's emotional state.

[0398] Based on the determined emotional state, the system generates conversation content and advice.

[0399] Caregiving assistance

[0400] User

[0401] "My lower back has been a little sore since yesterday," he said.

[0402] terminal

[0403] It receives voice input and records it as audio data.

[0404] The audio data is converted into text data (speech recognition processing) and sent to the server.

[0405] server

[0406] Text data is compared against a database to retrieve past health history.

[0407] Based on health history and emotional state, the system generates appropriate advice and sends a message to the device saying, "Please keep your lower back warm. If necessary, cool it with a bath towel."

[0408] terminal

[0409] Text data is converted into audio data (speech synthesis), and then played back to the user.

[0410] Specific example

[0411] Example 1: Everyday conversation

[0412] The user launches the smartphone app and performs fingerprint authentication.

[0413] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[0414] The device receives this, converts the audio to text, and sends it to the server.

[0415] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[0416] The device converts this into audio and plays it back.

[0417] Example 2: Caregiving advice

[0418] The user says, "My lower back has been a little sore since yesterday."

[0419] The device converts the audio to text and sends it to the server.

[0420] The server references the user's health history and generates appropriate advice.

[0421] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0422] The device converts the audio to sound and plays it back.

[0423] Example 3: Emotion recognition

[0424] The user says, "I'm feeling a little down today."

[0425] The device converts speech to text and then uses an emotion engine to analyze the emotional state.

[0426] Based on the analysis results, the server generates a conversation message such as, "Why don't you listen to your favorite music to change your mood?" and sends it to the terminal.

[0427] The device converts this into audio and plays it back.

[0428] In this way, this system combines generative AI and an emotion engine to achieve more personalized responses for the elderly. As a result, it becomes possible to provide efficient and effective mental and physical care for the elderly.

[0429] The following describes the processing flow.

[0430] User authentication and profile information management

[0431] User

[0432] Step 1: Launch the smartphone app.

[0433] terminal

[0434] Step 2: Display a screen for fingerprint authentication, facial recognition, or username and password entry.

[0435] User

[0436] Step 3: Perform fingerprint authentication or facial recognition, or enter your username and password.

[0437] terminal

[0438] Step 4: Send authentication information to the server.

[0439] server

[0440] Step 5: The received authentication information is compared with the database to authenticate the user.

[0441] Step 6: If authentication is successful, retrieve the user's profile information from the database and send it to the device.

[0442] terminal

[0443] Step 7: Cache the received profile information and prepare for the next step.

[0444] Everyday conversation

[0445] User

[0446] Step 1: Speak to your device and say, "Good morning, it's a beautiful day today."

[0447] terminal

[0448] Step 2: Receive voice input and record it as audio data.

[0449] Step 3: Convert the audio data to text (speech recognition process).

[0450] Step 4: Send the converted text data to the server.

[0451] server

[0452] Step 5: Input the received text data into the AI ​​model and generate appropriate conversation content (natural language processing).

[0453] Step 6: Send the generated conversation content as text data to the device.

[0454] terminal

[0455] Step 7: Convert the received text data into audio data (speech synthesis).

[0456] Step 8: Play the generated audio data to the user.

[0457] emotion recognition

[0458] User

[0459] Step 1: Speak into the device about everyday topics, your health, or any problems you're experiencing.

[0460] terminal

[0461] Step 2: Receive voice input and record it as audio data.

[0462] Step 3: Analyze the audio data with the emotion engine to determine the emotional state.

[0463] Step 4: Send the analyzed emotional state as text data to the server.

[0464] server

[0465] Step 5: Generate conversation content and advice based on the received text data and emotional state.

[0466] Step 6: Send the generated conversation content and advice to the device.

[0467] Caregiving assistance

[0468] User

[0469] Step 1: Start by saying, "My lower back has been a little sore since yesterday."

[0470] terminal

[0471] Step 2: Receive voice input and record it as audio data.

[0472] Step 3: Convert the audio data to text (speech recognition process).

[0473] Step 4: Send the converted text data to the server.

[0474] server

[0475] Step 5: Match the text data to the database to retrieve past health history.

[0476] Step 6: Generate appropriate advice based on health history and emotional state.

[0477] Step 7: Send the generated advice as text data to the device.

[0478] terminal

[0479] Step 8: Convert the received text data into audio data (speech synthesis).

[0480] Step 9: Play the generated advice as audio.

[0481] The above is a detailed explanation of the processing flow of the system that combines the emotion engine.

[0482] (Example 2)

[0483] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0484] There is a need for systems that provide mental and physical care to the elderly in a more effective and personalized way. In particular, it is necessary to accurately understand the emotional state of the elderly when they talk about their daily lives or health, and to provide appropriate responses and advice. However, current systems have shortcomings in emotional recognition, making it difficult to respond to each individual's situation.

[0485] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0486] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content based on the received text data, means for converting the generated conversation content to voice, means for analyzing the user's voice data and determining their emotional state, means for generating conversation content and advice based on the determined emotional state, and means for playing it back as voice. This makes it possible to provide personalized conversations and care advice tailored to the emotional state of elderly people.

[0487] "User authentication" refers to the process of verifying a user's identity in order to access a system.

[0488] "Profile information" refers to data that includes personal information and specific characteristics about a user.

[0489] "Voice input" refers to audio data used to capture a user's speech into the system.

[0490] "Text conversion" refers to the process of converting audio data into text data.

[0491] A "server" refers to a computer device that processes and stores data.

[0492] "Conversation content generation" refers to the process of creating appropriate responses or messages based on input data.

[0493] "Speech conversion" refers to the process of converting text data into speech data.

[0494] "Audio playback" refers to the act of making the generated audio data playable for the user.

[0495] "Emotional state" refers to the emotional state of a user as analyzed from their speech patterns and content.

[0496] "Health status" refers to data that indicates the user's physical and mental condition.

[0497] "Advice generation" refers to the process of creating appropriate advice based on the user's situation and data.

[0498] "Health history" refers to records of a user's past health condition.

[0499] A "database" refers to a system for effectively storing, managing, and retrieving information.

[0500] An "emotion engine" refers to software that analyzes a user's voice data to determine their emotional state.

[0501] This invention is a system that combines user authentication, speech recognition, a generative AI model, and an emotion recognition engine to improve the mental and physical care of the elderly. Specifically, the user communicates by voice using a terminal, the voice is analyzed by a server, and appropriate responses and advice are generated. The specific method for implementing this system is described below.

[0502] User authentication and profile information management

[0503] Users log in to the device using fingerprint authentication, facial recognition, or by entering a username and password. This allows the system to recognize who the individual is and provide personalized services.

[0504] The device receives the authentication information entered by the user and sends it to the server. Specifically, a standard facial recognition camera can be used for facial authentication, and a fingerprint scanner can be used for fingerprint authentication. The login information is encrypted and securely transmitted to the server.

[0505] The server compares the transmitted authentication information with the user's profile information stored in the database. If authentication is successful, the server sends the user's profile information to the device. This allows the device to provide services tailored to that user.

[0506] Speech recognition and conversation content generation

[0507] Users can converse with the device using their voice. For example, they can say, "Good morning, it's a nice day today."

[0508] The device receives voice input and converts the voice data into text using speech recognition software such as the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[0509] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate appropriate conversation content based on the received text data. The generated text data is then sent to the terminal.

[0510] The device converts the generated conversation into speech using speech synthesis software such as Amazon Polly and plays it back to the user. A concrete example of such a response might be, "Good morning. It's sunny today. Shall we go for a walk?"

[0511] emotion recognition

[0512] Users can talk to the device about their physical condition and emotions. For example, they might say, "I'm feeling a little down today."

[0513] The device records the input voice data and analyzes the emotional state using emotion recognition software such as Microsoft® Azure® Emotion API. The results are then sent to the server.

[0514] The server uses a generative AI model based on the analysis results to generate conversation content and produce appropriate responses that match the user's emotional state. For example, a possible response might be, "Why don't you listen to your favorite music to change your mood?"

[0515] The device converts the generated text data into speech and plays it back to the user.

[0516] Caregiving assistance

[0517] Users can talk about specific health issues they have. For example, they might say, "My lower back has been a little sore since yesterday."

[0518] The device records the audio data and converts it into text data using the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[0519] The server retrieves the user's past health history from the database based on the received text data. Then, it uses a generative AI model to generate appropriate advice, such as a message like, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0520] The device converts the generated text data into speech using speech synthesis software and plays it back to the user.

[0521] Specific example

[0522] Example of a prompt:

[0523] Everyday conversation: "Good morning, it's a beautiful day today."

[0524] Emotion recognition: "I'm feeling a little down today."

[0525] Caregiving advice: "My lower back has been a little sore since yesterday."

[0526] Thus, the present invention integrates a series of processes, including user authentication, voice recognition, emotion recognition, conversation generation, and health management, to realize more personalized responses for the elderly. This makes it possible to provide efficient and effective mental and physical care for the elderly.

[0527] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0528] Program processing flow

[0529] User authentication and profile information management

[0530] Step 1:

[0531] The user authenticates using their fingerprint or facial recognition, or enters a username and password.

[0532] Input: Fingerprint data, facial image, username and password

[0533] Output: Authentication information

[0534] Specific action: The user places their finger on the fingerprint sensor of their smartphone.

[0535] Step 2:

[0536] The terminal receives the entered authentication information, encrypts it, and sends it to the server.

[0537] Input: Authentication information

[0538] Output: Encrypted authentication information

[0539] Specific operation: The device encrypts the received fingerprint data using AES (Advanced Encryption Standard).

[0540] Step 3:

[0541] The server compares the received authentication information with the registered information in the database.

[0542] Input: Encrypted credentials

[0543] Output: Authentication results and profile information

[0544] Specific operation: The server performs user authentication by comparing the user's profile information with that in the database. If authentication is successful, the server then retrieves the profile information.

[0545] Step 4:

[0546] If the server successfully authenticates, it will send the profile information to the device.

[0547] Input: Profile Information

[0548] Output: Profile information sent to the device

[0549] Specific operation: The server encrypts the profile information it has acquired and sends it to the terminal.

[0550] Everyday conversation

[0551] Step 1:

[0552] The user speaks to the device, saying, "Good morning, it's a nice day today."

[0553] Input: Audio data

[0554] Output: Audio data recorded by the device's microphone

[0555] Specific action: The user speaks into the smart speaker.

[0556] Step 2:

[0557] The device receives voice input and uses the Google Cloud Speech-to-Text API to convert the voice data into text.

[0558] Input: Audio data

[0559] Output: Text data

[0560] Specific operation: The device sends voice data to the API and retrieves the text data "It's a nice day today."

[0561] Step 3:

[0562] The terminal sends the converted text data to the server.

[0563] Input: Text data

[0564] Output: Text data sent to the server

[0565] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[0566] Step 4:

[0567] The server inputs the received text data into a generating AI model (e.g., GPT-3) to generate appropriate conversation content.

[0568] Input: Text data

[0569] Output: Generated conversation content

[0570] Specific operation: The server inputs the text "It's a nice day today" into GPT-3 and generates the response "Good morning. It's sunny today. Shall we go for a walk?".

[0571] Step 5:

[0572] The server sends the generated conversation content to the terminal.

[0573] Input: Generated conversation content

[0574] Output: Text data sent to the terminal

[0575] Specific operation: The server sends the generated text data to the terminal.

[0576] Step 6:

[0577] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[0578] Input: Text data of the generated conversation

[0579] Output: Audio data

[0580] Specific action: The device sends the text message "Good morning. It's sunny today. Shall we go for a walk?" to Amazon Polly and plays the audio data.

[0581] emotion recognition

[0582] Step 1:

[0583] The user speaks to the device about their physical condition and emotions.

[0584] Input: Audio data

[0585] Output: Audio data recorded by the device's microphone

[0586] Specific action: The user says, "I'm feeling a little down today."

[0587] Step 2:

[0588] The device records voice data and analyzes the emotional state using the Microsoft Azure Emotion API.

[0589] Input: Audio data

[0590] Output: Emotion analysis results

[0591] Specific operation: The device sends voice data to the API to obtain the emotional state (e.g., "feeling down").

[0592] Step 3:

[0593] The terminal sends the analysis results to the server.

[0594] Input: Sentiment analysis results

[0595] Output: Sentiment analysis results sent to the server

[0596] Specific operation: The device sends the emotion analysis results to the server using the HTTPS protocol.

[0597] Step 4:

[0598] The server generates conversation content using an AI model based on the emotion analysis results it receives.

[0599] Input: Sentiment analysis results

[0600] Output: Generated conversation content

[0601] Specific operation: The server inputs the emotional state of "feeling down" into GPT-3 and generates the response, "Why don't you try listening to your favorite music to cheer yourself up?"

[0602] Step 5:

[0603] The server sends the generated conversation content to the terminal.

[0604] Input: Generated conversation content

[0605] Output: Text data sent to the terminal

[0606] Specific operation: The server sends the generated text data to the terminal.

[0607] Step 6:

[0608] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[0609] Input: Text data of the generated conversation

[0610] Output: Audio data

[0611] Specific action: The device sends a text message to Amazon Polly saying, "Why not listen to your favorite music to change your mood?" and plays the audio data.

[0612] Caregiving assistance

[0613] Step 1:

[0614] The user speaks to the device about their health problems.

[0615] Input: Audio data

[0616] Output: Audio data recorded by the device's microphone

[0617] Specific action: The user says, "My lower back has been a little sore since yesterday."

[0618] Step 2:

[0619] The device records the voice data and converts it into text data using the Google Cloud Speech-to-Text API.

[0620] Input: Audio data

[0621] Output: Text data

[0622] Specific operation: The device sends voice data to the API and retrieves text data saying, "My lower back has been a little sore since yesterday."

[0623] Step 3:

[0624] The terminal sends the converted text data to the server.

[0625] Input: Text data

[0626] Output: Text data sent to the server

[0627] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[0628] Step 4:

[0629] The server retrieves the user's past health history from the database based on text data.

[0630] Input: Text data

[0631] Output: Health history data

[0632] Specific operation: The server queries the database using the text "My lower back has been a little sore since yesterday" to retrieve past health history.

[0633] Step 5:

[0634] The server uses the generated AI model to produce appropriate advice.

[0635] Input: Health history data

[0636] Output: Generated advice

[0637] Specific operation: The server inputs health history data into GPT-3 and generates advice such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0638] Step 6:

[0639] The server sends the generated advice to the terminal.

[0640] Input: Generated advice

[0641] Output: Text data sent to the terminal

[0642] Specific operation: The server sends the generated text data to the terminal.

[0643] Step 7:

[0644] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[0645] Input: Text data of the generated advice

[0646] Output: Audio data

[0647] Specific action: The device sends a text message to Amazon Polly saying, "Please keep your lower back warm. Use a bath towel to cool it down if necessary," and plays an audio message.

[0648] In this way, by processing data based on specific input data at each step, a system is realized that provides personalized conversations and care advice to the elderly.

[0649] (Application Example 2)

[0650] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0651] In systems designed to support the safe and secure lives of the elderly, it is crucial not only to provide conversation and care assistance, but also to understand the user's emotional state in real time and provide appropriate support and advice tailored to their feelings. Furthermore, when using autonomous vehicles, it is especially important to alleviate the anxiety and tension felt by the elderly and support safe travel. Conventional technologies have difficulty adequately reflecting the user's emotions and state, making the enhancement of mental care a challenge.

[0652] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state, means for adjusting the conversation content based on the analysis results, and means for providing safety information for the autonomous vehicle. As a result, the user's emotional state can be grasped in real time, and personalized conversations and support can be provided accordingly, thereby increasing the elderly's sense of security and reducing anxiety and tension during travel. Furthermore, this makes it possible to provide an environment in which the elderly can travel with peace of mind even when using an autonomous vehicle.

[0653] "Means of user authentication" refers to methods used to verify the identity of a user when they access a system.

[0654] "Means of managing user profile information" refers to a method of centrally managing users' personal information and settings.

[0655] "Methods for converting voice input to text" refer to technologies that convert a user's voice data into text data.

[0656] "Means for sending converted text data to the server" refers to a method for converting voice input into text data and then sending that data to the server.

[0657] "Methods for generating conversation content based on received text data" refers to technologies that create appropriate dialogue content based on received text data.

[0658] "Means of converting generated conversation content into speech" refers to technology that converts the created dialogue content back into speech data.

[0659] "Means of playback as audio" refers to methods for making the generated audio data listenable to the user.

[0660] "Means for analyzing a user's emotional state" refers to technologies for analyzing a user's emotions and psychological state in real time.

[0661] "Means of adjusting conversation content based on analysis results" refers to technologies that change the content of a dialogue based on analyzed emotional data.

[0662] "Means for managing a user's health status" refers to methods for collecting and managing a user's health data.

[0663] "Means for collecting data on users' physical condition and problems" refers to technologies for collecting information on users' physical condition and problems.

[0664] "Means of generating appropriate advice based on collected data" refers to methods of creating appropriate advice based on acquired data.

[0665] "Means for providing safety information on autonomous vehicles" refers to technologies that provide users with safety information related to autonomous vehicles.

[0666] "Means for storing user profile information and health status in a database" refers to methods for storing users' personal information and health information in a database.

[0667] "A means of referring to past health history based on stored data" refers to a technology that allows users to check their past health status based on data stored in a database.

[0668] "Means of providing real-time support based on the user's emotional state" refers to technologies that provide immediate support and advice based on emotional data analyzed in real time.

[0669] This invention is a system that analyzes a user's emotional state in real time and provides conversation content and safety information for autonomous vehicles accordingly. The system consists of a terminal used by the user and a server that processes data. The specific implementation method is described below.

[0670] System Configuration

[0671] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, and functions as an interface for smartphones and autonomous vehicles. The server uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more appropriate responses.

[0672] Basic operation

[0673] User authentication and profile information management

[0674] Users log in to the system using their device via voice or touch controls. The device accepts fingerprint or facial recognition, or username and password input, and sends the authentication information to the server. The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information. If authentication is successful, the server sends the user's profile information to the device.

[0675] Everyday conversation

[0676] When a user speaks to the device, saying "Good morning, it's a nice day today," the device receives the voice input and records it as audio data. It then converts the audio data to text and sends it to the server. The server inputs the received text data into an AI model and generates appropriate conversation content. The generated conversation content is sent to the device as text data, which then converts it back into audio data and plays it back.

[0677] Emotional recognition and regulation

[0678] The device, upon receiving the user's everyday conversations and statements about their physical condition, analyzes the audio data using an emotion engine to determine their emotional state. It then sends this emotional state information to the server. Based on the received emotional information, the server appropriately adjusts the conversation content and support information to provide a personalized response.

[0679] Caregiving assistance and health management

[0680] For example, if a user says, "My lower back has been a little sore since yesterday," the device converts the audio to text and sends it to the server. The server refers to the user's health history and emotional state and generates appropriate advice. The generated advice is sent to the device as a message such as, "Keep your lower back warm. Use a bath towel to cool it if necessary," and the device converts this back to audio and plays it.

[0681] Applications in autonomous vehicles

[0682] If a user feels "a little uneasy" while using an autonomous vehicle, the terminal converts the voice input into text and analyzes it with an emotion engine. Based on the analysis results, the server provides real-time support such as, "It's okay, the car is driving safely. Please relax. Would you like some tea?" Also, if the user says, "I hear a strange noise," the system checks the vehicle's status and provides reassuring feedback in real time.

[0683] Hardware and software to be used

[0684] Hardware: Smartphone, microphone, speaker, autonomous vehicle interface

[0685] Software: Python, speech recognition library (speech_recognition), speech synthesis library (gTTS), natural language processing library (transformers)

[0686] Specific examples and prompt statements

[0687] Example 1: Anxiety

[0688] If the user says, "I'm a little worried," the system will generate a response such as, "We've determined your emotion to be negative, but it's okay. The car is driving safely, so please relax."

[0689] Specific example 2: Problem

[0690] If you say, "I hear a strange noise," and your emotion is "NEGATIVE," the response you'll get is, "We're checking the car's condition. There doesn't seem to be any problem, so please don't worry."

[0691] Example of a prompt:

[0692] "Generate a reassuring response regarding vehicle safety for elderly individuals who are expressing anxiety."

[0693] "Please generate relaxing responses to the negative emotions expressed by older adults."

[0694] By constructing a concrete system in this way, it is possible to enhance the mental care of the elderly and provide an environment in which they can use autonomous vehicles with peace of mind.

[0695] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0696] Step 1:

[0697] User Login

[0698] Input: The terminal receives the user's voice, fingerprint, facial recognition, or username and password.

[0699] Operation:

[0700] The terminal converts the received authentication information into text data and sends it to the server.

[0701] Output: Authentication information is sent to the server as text data.

[0702] Step 2:

[0703] Verification of authentication information and retrieval of profile information

[0704] Input: The server authenticates the user based on the authentication information received from the terminal.

[0705] Operation:

[0706] The server compares the user information with the user information in the database, and if authentication is successful, retrieves the user's profile information.

[0707] Output: Authentication results and profile information are sent to the device.

[0708] Step 3:

[0709] Processing everyday conversations

[0710] Input: The user speaks to the device (e.g., "Good morning, it's a nice day today").

[0711] Operation:

[0712] The device records voice input as data and converts the voice data into text data.

[0713] Send the converted text data to the server.

[0714] Output: The converted text data is sent to the server.

[0715] Step 4:

[0716] Generating conversation content

[0717] Input: The server inputs the received text data into the natural language processing AI.

[0718] Operation:

[0719] The server performs natural language processing to generate appropriate conversation content.

[0720] The conversation content is generated as text data and sent to the device.

[0721] Output: The generated conversation content is sent as text data to the terminal.

[0722] Step 5:

[0723] Speech conversion and playback of conversation content

[0724] Input: The terminal receives text data of the conversation received from the server.

[0725] Operation:

[0726] The device converts the received text data into audio data.

[0727] Play audio data and let the user listen.

[0728] Output: The generated audio data is played.

[0729] Step 6:

[0730] emotion recognition

[0731] Input: The user talks about everyday topics or their health.

[0732] Operation:

[0733] The device records voice input and feeds the voice data into the emotion engine.

[0734] The emotion engine analyzes the audio data and determines the emotional state.

[0735] Send emotional state information to the server.

[0736] Output: The analyzed emotional state information is sent to the server.

[0737] Step 7:

[0738] Emotion-based response generation

[0739] Input: The server generates a response based on the user's emotional state information and the content of the previous conversation.

[0740] Operation:

[0741] The server adjusts the conversation based on sentiment data and generates customized responses.

[0742] The response is sent to the terminal as text data.

[0743] Output: Text data containing the customized response is sent to the terminal.

[0744] Step 8:

[0745] Caregiving advice

[0746] Input: The user talks about their physical condition or any problems they are experiencing (e.g., "My back has been hurting since yesterday").

[0747] Operation:

[0748] The device converts the audio to text and sends it to the server.

[0749] The server references the user's health history and emotional state to generate appropriate advice.

[0750] The advice is sent to the device, converted to audio, and then played back.

[0751] Output: The generated audio data is played back to the user.

[0752] Step 9:

[0753] Support for autonomous vehicles

[0754] Input: When the user feels anxious while traveling (e.g., "I feel a little anxious").

[0755] Operation:

[0756] The device converts voice input into text and analyzes it using an emotion engine.

[0757] The server generates a reassuring response based on emotional state information.

[0758] The response is converted into audio data and played back to the user.

[0759] Output: The generated audio data is played back inside the autonomous vehicle to reassure the user.

[0760] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0761] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0762] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0763] [Second Embodiment]

[0764] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0765] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0766] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0767] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0768] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0769] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0770] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0771] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0772] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0773] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0774] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0775] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0776] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. The aim of this system is to alleviate the user's feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support.

[0777] System Configuration

[0778] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using an AI model.

[0779] Basic operation

[0780] User authentication and profile information management

[0781] User:

[0782] Users log in to the system by using voice commands or touch controls on the device.

[0783] Terminal:

[0784] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[0785] Send authentication information to the server.

[0786] server:

[0787] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[0788] If authentication is successful, the user's profile information will be sent to the device.

[0789] Everyday conversation

[0790] User:

[0791] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[0792] Terminal:

[0793] The device converts voice input into text and sends the text data to the server.

[0794] server:

[0795] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[0796] The generated conversation content is sent to the terminal as text data.

[0797] Terminal:

[0798] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0799] Caregiving assistance

[0800] User:

[0801] The user says, "My lower back has been a little sore since yesterday."

[0802] Terminal:

[0803] The device converts voice input into text and sends the text data to the server.

[0804] server:

[0805] The server compares the received text data with the database and refers to the user's past health history.

[0806] It generates appropriate advice and sends text data to the device saying, "Keep your lower back warm. Use a bath towel to cool it down if necessary."

[0807] Terminal:

[0808] The device converts the received text data into speech and plays it back to the user.

[0809] Specific example

[0810] The following will explain this with specific examples.

[0811] Example 1: Everyday conversation

[0812] The user launches the smartphone app and performs fingerprint authentication.

[0813] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[0814] The device receives this, converts the audio to text, and sends it to the server.

[0815] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[0816] The device converts this into audio and plays it back.

[0817] Example 2: Caregiving advice

[0818] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[0819] The device converts the audio to text and sends it to the server.

[0820] The server references the user's health history and generates appropriate advice.

[0821] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0822] The device converts the audio to sound and plays it back.

[0823] As described above, this system utilizes generative AI to efficiently perform daily conversations and provide care assistance for the elderly. This system is expected to alleviate feelings of loneliness among the elderly, maintain cognitive function, and improve the quality of care.

[0824] The following describes the processing flow.

[0825] User authentication and profile information management

[0826] User

[0827] Step 1:

[0828] Launch the smartphone app.

[0829] terminal

[0830] Step 2:

[0831] It displays a screen for fingerprint authentication, facial recognition, or entry of a username and password.

[0832] User

[0833] Step 3:

[0834] You can use fingerprint authentication, facial recognition, or enter your username and password.

[0835] terminal

[0836] Step 4:

[0837] Send authentication information to the server.

[0838] server

[0839] Step 5:

[0840] The received authentication information is compared with the database to authenticate the user.

[0841] Step 6:

[0842] If authentication is successful, the user's profile information is retrieved from the database and sent to the device.

[0843] terminal

[0844] Step 7:

[0845] The received profile information is cached to prepare for the next process.

[0846] Everyday conversation

[0847] User

[0848] Step 1:

[0849] I say to the device, "Good morning, it's a beautiful day today."

[0850] terminal

[0851] Step 2:

[0852] It receives voice input and records it as audio data.

[0853] Step 3:

[0854] Converting audio data into text data (speech recognition processing).

[0855] Step 4:

[0856] Send the converted text data to the server.

[0857] server

[0858] Step 5:

[0859] The received text data is input into an AI model to generate appropriate conversation content (natural language processing).

[0860] Step 6:

[0861] The generated conversation content is sent to the terminal as text data.

[0862] terminal

[0863] Step 7:

[0864] Converts received text data into audio data (speech synthesis).

[0865] Step 8:

[0866] The generated audio data is played for the user.

[0867] Caregiving assistance

[0868] User

[0869] Step 1:

[0870] I spoke to the device, saying, "My lower back has been a little sore since yesterday."

[0871] terminal

[0872] Step 2:

[0873] It receives voice input and records it as audio data.

[0874] Step 3:

[0875] Converting audio data into text data (speech recognition processing).

[0876] Step 4:

[0877] Send the converted text data to the server.

[0878] server

[0879] Step 5:

[0880] The received text data is compared against a database to retrieve past health history.

[0881] Step 6:

[0882] Based on the acquired health history, appropriate advice is generated (health analysis and advice generation).

[0883] Step 7:

[0884] The generated advice is sent to the device as text data.

[0885] terminal

[0886] Step 8:

[0887] Converts received text data into audio data (speech synthesis).

[0888] Step 9:

[0889] The generated advice will be played back as audio.

[0890] The above is a detailed explanation of the program's processing flow.

[0891] (Example 1)

[0892] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0893] In modern society, there is a need to alleviate loneliness among the elderly, maintain cognitive function, and provide appropriate care support. However, many elderly people have limited opportunities for daily conversation and timely care, putting them at high risk of deteriorating mental and physical health. To solve this problem, a more effective and efficient system is needed.

[0894] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0895] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content using a generative AI model based on the received text data, means for converting the generated conversation content to speech, and means for playing it back as speech. This makes it possible for elderly people to receive everyday conversations and care advice without feeling lonely. Furthermore, by combining means for managing the user's health status, means for collecting data on the user's physical condition and problems, means for generating appropriate advice using a generative AI model based on the collected data, means for converting the generated advice to speech, and means for playing it back as speech, it becomes possible to provide appropriate responses based on the user's past health history.

[0896] "User authentication" is the process of verifying that the user accessing the system is indeed that specific user.

[0897] "Profile information" refers to data that includes personal information and attributes about the user.

[0898] "Means for converting voice input to text" refers to devices or software that use speech recognition technology to convert a user's speech into text format.

[0899] "Means of sending converted text data to the server" refers to the process of sending text data generated on the terminal to the server via the network.

[0900] A "generative AI model" is an algorithm or software that uses natural language processing and machine learning to generate appropriate responses from input text data.

[0901] "Means for generating conversation content" refers to the process of generating appropriate text for interacting with the user using a generation AI model based on received text data.

[0902] "Means for converting generated conversation content into speech" refers to devices or software that convert the text of generated conversation content into speech format using speech synthesis technology.

[0903] "Means of playback as audio" refers to the process of playing audio data to the user through audio equipment such as speakers.

[0904] "Means for managing health status" refers to devices and software that record and monitor a user's health information and periodically evaluate that status.

[0905] "Means of collecting data on health and problems" refers to the process of collecting information on health and problems through user feedback and questions.

[0906] "Means for generating appropriate advice" refers to the process of generating appropriate advice and suggestions for users using a generative AI model based on collected data and past history.

[0907] "A means of referencing past health history based on saved data" refers to the process of searching for a user's health information stored in a database and referencing their past health status and history.

[0908] Modes for carrying out the invention

[0909] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. This system aims to alleviate feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support by implementing user authentication, daily conversation, and care assistance functions.

[0910] System Configuration

[0911] The system consists of user terminals and servers that process data.

[0912] A terminal is a device that performs voice input and output, and includes smartphones, tablets, and smart speakers. A server is a computer device that performs natural language processing using AI models, and examples include cloud servers and dedicated servers.

[0913] Hardware and software usage

[0914] The system uses the following specific hardware and software.

[0915] Speech recognition software for converting voice input to text (e.g., Google Cloud Speech-to-Text, IBM Watson Speech to Text)

[0916] Generative AI models for generating conversation content based on text data (e.g., OpenAI GPT-3)

[0917] Text-to-speech software for converting text data into speech (e.g., Amazon Polly, Google Cloud Text-to-Speech)

[0918] Authentication technologies for user authentication (e.g., fingerprint recognition, facial recognition, username and password)

[0919] User authentication and profile information management

[0920] The user attempts to log in by using voice or touch controls on the device.

[0921] The device accepts fingerprint authentication, facial recognition, and username and password input, and sends the authentication information to the server.

[0922] The server compares the received authentication information with the database to determine whether authentication was successful. If authentication is successful, the user's profile information is sent to the device.

[0923] Everyday conversation

[0924] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[0925] The device converts voice input into text and sends that text data to the server.

[0926] The server uses an AI model to generate appropriate conversation content based on the received text data. The generated conversation content is then sent to the terminal as text data.

[0927] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0928] Specific examples of prompt statements are as follows:

[0929] User: Good morning, it's a beautiful day today.

[0930] AI Model: Good morning. It's sunny today. Shall we go for a walk?

[0931] Caregiving assistance

[0932] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[0933] The device converts voice input into text and sends that text data to the server.

[0934] The server uses the received text data to compare and refer to the user's past health history in a database. Then, using a generative AI model, it generates appropriate advice and sends text data to the terminal such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[0935] The device converts the received text data into speech and plays it back to the user.

[0936] Specific examples of prompt statements are as follows:

[0937] User: My lower back has been a little sore since yesterday.

[0938] AI Model: Keep your lower back warm. Use a bath towel to cool it down if necessary.

[0939] The above describes the embodiments for carrying out the present invention. This system provides assistance with daily conversation and care for the elderly, contributing to the alleviation of feelings of loneliness, the maintenance and improvement of cognitive function, and the improvement of the quality of care.

[0940] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0941] User Authentication

[0942] Processing flow

[0943] Step 1:

[0944] The user attempts to log in by using voice or touch controls on the device.

[0945] Input: User voice or touch operation

[0946] Output: Login attempt data

[0947] Step 2:

[0948] The device accepts fingerprint authentication, facial recognition, and username and password input.

[0949] Input: User's fingerprint data, facial data, username and password

[0950] Output: Authentication information data

[0951] Step 3:

[0952] The terminal sends the entered authentication information to the server.

[0953] Input: Authentication information data

[0954] Output: Authentication request to the server

[0955] Step 4:

[0956] The server compares the received authentication information with the database to determine whether the authentication was successful.

[0957] Input: Authentication information data

[0958] Output: Authentication success / failure result

[0959] Step 5:

[0960] If authentication is successful, the server sends the user's profile information to the device.

[0961] Input: Authentication success / failure result

[0962] Output: Profile information data

[0963] Everyday conversation

[0964] Processing flow

[0965] Step 1:

[0966] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[0967] Input: User voice input

[0968] Output: Audio data

[0969] Step 2:

[0970] The device converts voice input into text and sends that text data to the server.

[0971] Input: Audio data

[0972] Output: Text data, request to send to the server

[0973] Step 3:

[0974] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[0975] Input: Text data

[0976] Output: Generated conversation text

[0977] Step 4:

[0978] The server sends the generated conversation content to the terminal as text data.

[0979] Input: Generated conversation text

[0980] Output: Conversation text data, request to send to terminal

[0981] Step 5:

[0982] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[0983] Input: Conversation text data

[0984] Output: Audio data

[0985] Caregiving assistance

[0986] Processing flow

[0987] Step 1:

[0988] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[0989] Input: User voice input

[0990] Output: Audio data

[0991] Step 2:

[0992] The device converts voice input into text and sends that text data to the server.

[0993] Input: Audio data

[0994] Output: Text data, request to send to the server

[0995] Step 3:

[0996] The server uses the received text data to compare and reference the user's past health history from the database.

[0997] Input: Text data

[0998] Output: Health history data

[0999] Step 4:

[1000] The server uses a generative AI model to generate appropriate advice and sends the text data to the terminal.

[1001] Input: Health history data

[1002] Output: Generated advice text data, request to send to terminal

[1003] Step 5:

[1004] The device converts the received text data into speech and plays back, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1005] Input: Advice text data

[1006] Output: Audio data

[1007] (Application Example 1)

[1008] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1009] Elderly people often find food delivery services difficult to use due to their complex operation. Furthermore, they struggle to select nutritionally balanced meals, hindering proper dietary management. This highlights the growing need for food delivery applications that improve the quality of life for the elderly.

[1010] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1011] In this invention, the server includes means for generating food delivery order content using a generative AI model, means for transmitting the generated order content to a food delivery system, and means for providing the generated nutritional advice. This makes it possible for elderly people to easily order food delivery by voice and also receive appropriate nutritional advice.

[1012] User authentication is the process of identifying users who are accessing a system and verifying that they are authorized users.

[1013] "Profile information" refers to personal information and settings information about a user, and is data necessary for the system to provide services tailored to the user.

[1014] "Voice input" is the process by which a user speaks to the system through a microphone or similar device, and the system receives voice data in response.

[1015] "Converting to text" is the process of converting audio data into text data.

[1016] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate conversation content or order details based on user input.

[1017] "Food delivery order details" refers to the details of the food and drinks that a user specifies when using a food delivery service.

[1018] A "food delivery system" refers to the entire service and its operating system for delivering food to customers who order it.

[1019] "Nutritional advice" refers to expert advice designed to help users choose healthy meals.

[1020] "Converting to speech" is the process of converting text data into audio data, making it playable for the user.

[1021] "Playback as audio" refers to the process of playing the converted audio data to the user through a speaker or other device.

[1022] "Past health history" refers to past data and records regarding the user's health status, which are used to generate current advice.

[1023] This invention provides a system that simplifies the operation of food delivery services for elderly people and simultaneously provides nutritional advice. This system includes a series of processes for voice input, conversation generation using a generative AI model, order content generation, and provision of nutritional advice.

[1024] System Configuration

[1025] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using a generative AI model.

[1026] User authentication and profile information management

[1027] User:

[1028] The user launches the food delivery app and logs into the system using fingerprint authentication or similar methods.

[1029] Terminal:

[1030] The device accepts fingerprint authentication, facial recognition, or username and password input for user authentication.

[1031] Send authentication information to the server.

[1032] server:

[1033] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[1034] If authentication is successful, the user's profile information will be sent to the device.

[1035] Food delivery order generation

[1036] User:

[1037] The user speaks to the device and says, "I want to eat taco rice today."

[1038] Terminal:

[1039] The device converts voice input into text and sends the text data to the server.

[1040] server:

[1041] The server receives the text data and sends it as a prompt to the AI ​​model, which then generates an appropriate food delivery order.

[1042] The generated order details are sent to the food delivery system.

[1043] Terminal:

[1044] The terminal converts the order confirmation received from the server into audio and plays it back to the user.

[1045] Nutritional advice

[1046] User:

[1047] The user says to the device, "Lately, I've been worried about my nutritional balance."

[1048] Terminal:

[1049] The device converts voice input into text and sends the text data to the server.

[1050] server:

[1051] The server references the user's past health history and uses a generative AI model to generate appropriate nutritional advice.

[1052] The system generates and sends a message to the device saying, "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[1053] Terminal:

[1054] The device converts the received nutritional advice into audio and plays it back to the user.

[1055] Specific example

[1056] Example of a food delivery order:

[1057] Example usage: The user says, "I want to eat taco rice today."

[1058] AI: "So, you'd like taco rice. Would you like it spicy or mild?"

[1059] The user replied, "Please make it mild."

[1060] AI: "Thank you. I'll order the taco rice (mild) then."

[1061] Examples of nutritional advice:

[1062] Example of use: The user says, "Lately, I've been worried about my nutritional balance."

[1063] AI: "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[1064] Example of a prompt

[1065] "I want to eat taco rice today."

[1066] "Lately, I've been worried about nutritional balance."

[1067] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1068] Step 1:

[1069] User: The user launches the food delivery app and logs into the system using fingerprint authentication or by entering a password. During this process, the user's fingerprint information and password are entered into the system.

[1070] Step 2:

[1071] Terminal: The terminal obtains user authentication information through on-screen input forms or fingerprint sensors and sends it to the server. In this case, the entered authentication information is the data sent to the server.

[1072] Step 3:

[1073] Server: The server verifies the received authentication information and compares it with the profile information in the database. If the comparison is successful, it sends the user's profile information back to the terminal. It verifies that the authentication information is correct and outputs the user's profile data.

[1074] Step 4:

[1075] User: After successful authentication, the user speaks into the device saying, "I want to eat taco rice today."

[1076] Step 5:

[1077] Terminal: The terminal performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process. Speech recognition software (e.g., Google Speech Recognition API) is used to convert the speech input into text.

[1078] Step 6:

[1079] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[1080] Step 7:

[1081] Server: The server receives text data and sends it to the generative AI model as a prompt. This prompt becomes the input to the AI ​​model. The generative AI model (e.g., OpenAI GPT-3) is used to generate appropriate food delivery order details.

[1082] Step 8:

[1083] Server: Retrieves order details generated by the generation AI model and sends them to the food delivery system. The generated order details become the output data.

[1084] Step 9:

[1085] Terminal: Converts the order confirmation received from the server into audio. This converted audio data becomes the input data for the next process.

[1086] Step 10:

[1087] Terminal: The terminal plays back the converted order details and provides feedback to the user. It uses a voice output device to play back the text-to-speech conversion.

[1088] Step 11:

[1089] User: The user speaks to the device, saying, "Lately, I've been worried about my nutritional balance."

[1090] Step 12:

[1091] Terminal: Performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process.

[1092] Step 13:

[1093] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[1094] Step 14:

[1095] Server: The server receives text data, references the user's health history, generates prompts about nutritional balance, and sends them to an AI model. The user's health history data and prompts serve as input data.

[1096] Step 15:

[1097] Server: Retrieves nutritional advice generated by the generation AI model and sends it to the terminal as text data. The generated nutritional advice becomes the output data.

[1098] Step 16:

[1099] Terminal: Converts text data of nutritional advice received from the server into speech. This converted speech data becomes input data for the next process.

[1100] Step 17:

[1101] Device: The device plays nutritional advice converted into audio and provides feedback to the user. It uses an audio output device to play the nutritional advice.

[1102] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1103] This invention is a system that combines a generative AI, which is used to provide conversational support and care assistance for the elderly, with an emotion engine that recognizes the user's emotions. The aim of this system is to provide more effective mental care for the elderly, by offering conversations and care advice tailored to their emotional state.

[1104] System Configuration

[1105] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, while the server is a computer that uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more personalized responses.

[1106] Basic operation

[1107] User authentication and profile information management

[1108] User

[1109] Users log in to the system by using voice commands or touch controls on the device.

[1110] terminal

[1111] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[1112] Send authentication information to the server.

[1113] server

[1114] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[1115] If authentication is successful, the user's profile information will be sent to the device.

[1116] Everyday conversation

[1117] User

[1118] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[1119] terminal

[1120] The device receives voice input and records it as voice data.

[1121] The audio data is converted into text data (speech recognition processing) and sent to the server.

[1122] server

[1123] The server inputs the received text data into an AI model and generates appropriate conversation content (natural language processing).

[1124] The generated conversation content is sent to the terminal as text data.

[1125] terminal

[1126] The device converts the received text data into audio data (speech synthesis) and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[1127] emotion recognition

[1128] User

[1129] You can talk to the device about everyday topics, your health, or any problems you're facing.

[1130] terminal

[1131] It receives voice input and records it as audio data.

[1132] The system analyzes the user's emotional state from voice data using an emotion engine and sends the data to the server.

[1133] server

[1134] The server uses an emotion engine to analyze the received audio data and determine the user's emotional state.

[1135] Based on the determined emotional state, the system generates conversation content and advice.

[1136] Caregiving assistance

[1137] User

[1138] "My lower back has been a little sore since yesterday," he said.

[1139] terminal

[1140] It receives voice input and records it as audio data.

[1141] The audio data is converted into text data (speech recognition processing) and sent to the server.

[1142] server

[1143] Text data is compared against a database to retrieve past health history.

[1144] Based on health history and emotional state, the system generates appropriate advice and sends a message to the device saying, "Please keep your lower back warm. If necessary, cool it with a bath towel."

[1145] terminal

[1146] Text data is converted into audio data (speech synthesis), and then played back to the user.

[1147] Specific example

[1148] Example 1: Everyday conversation

[1149] The user launches the smartphone app and performs fingerprint authentication.

[1150] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[1151] The device receives this, converts the audio to text, and sends it to the server.

[1152] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[1153] The device converts this into audio and plays it back.

[1154] Example 2: Caregiving advice

[1155] The user says, "My lower back has been a little sore since yesterday."

[1156] The device converts the audio to text and sends it to the server.

[1157] The server references the user's health history and generates appropriate advice.

[1158] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1159] The device converts the audio to sound and plays it back.

[1160] Example 3: Emotion recognition

[1161] The user says, "I'm feeling a little down today."

[1162] The device converts speech to text and then uses an emotion engine to analyze the emotional state.

[1163] Based on the analysis results, the server generates a conversation message such as, "Why don't you listen to your favorite music to change your mood?" and sends it to the terminal.

[1164] The device converts this into audio and plays it back.

[1165] In this way, this system combines generative AI and an emotion engine to achieve more personalized responses for the elderly. As a result, it becomes possible to provide efficient and effective mental and physical care for the elderly.

[1166] The following describes the processing flow.

[1167] User authentication and profile information management

[1168] User

[1169] Step 1: Launch the smartphone app.

[1170] terminal

[1171] Step 2: Display a screen for fingerprint authentication, facial recognition, or username and password entry.

[1172] User

[1173] Step 3: Perform fingerprint authentication or facial recognition, or enter your username and password.

[1174] terminal

[1175] Step 4: Send authentication information to the server.

[1176] server

[1177] Step 5: The received authentication information is compared with the database to authenticate the user.

[1178] Step 6: If authentication is successful, retrieve the user's profile information from the database and send it to the device.

[1179] terminal

[1180] Step 7: Cache the received profile information and prepare for the next step.

[1181] Everyday conversation

[1182] User

[1183] Step 1: Speak to your device and say, "Good morning, it's a beautiful day today."

[1184] terminal

[1185] Step 2: Receive voice input and record it as audio data.

[1186] Step 3: Convert the audio data to text (speech recognition process).

[1187] Step 4: Send the converted text data to the server.

[1188] server

[1189] Step 5: Input the received text data into the AI ​​model and generate appropriate conversation content (natural language processing).

[1190] Step 6: Send the generated conversation content as text data to the device.

[1191] terminal

[1192] Step 7: Convert the received text data into audio data (speech synthesis).

[1193] Step 8: Play the generated audio data to the user.

[1194] emotion recognition

[1195] User

[1196] Step 1: Speak into the device about everyday topics, your health, or any problems you're experiencing.

[1197] terminal

[1198] Step 2: Receive voice input and record it as audio data.

[1199] Step 3: Analyze the audio data with the emotion engine to determine the emotional state.

[1200] Step 4: Send the analyzed emotional state as text data to the server.

[1201] server

[1202] Step 5: Generate conversation content and advice based on the received text data and emotional state.

[1203] Step 6: Send the generated conversation content and advice to the device.

[1204] Caregiving assistance

[1205] User

[1206] Step 1: Start by saying, "My lower back has been a little sore since yesterday."

[1207] terminal

[1208] Step 2: Receive voice input and record it as audio data.

[1209] Step 3: Convert the audio data to text (speech recognition process).

[1210] Step 4: Send the converted text data to the server.

[1211] server

[1212] Step 5: Match the text data to the database to retrieve past health history.

[1213] Step 6: Generate appropriate advice based on health history and emotional state.

[1214] Step 7: Send the generated advice as text data to the device.

[1215] terminal

[1216] Step 8: Convert the received text data into audio data (speech synthesis).

[1217] Step 9: Play the generated advice as audio.

[1218] The above is a detailed explanation of the processing flow of the system that combines the emotion engine.

[1219] (Example 2)

[1220] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[1221] There is a need for systems that provide mental and physical care to the elderly in a more effective and personalized way. In particular, it is necessary to accurately understand the emotional state of the elderly when they talk about their daily lives or health, and to provide appropriate responses and advice. However, current systems have shortcomings in emotional recognition, making it difficult to respond to each individual's situation.

[1222] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1223] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content based on the received text data, means for converting the generated conversation content to voice, means for analyzing the user's voice data and determining their emotional state, means for generating conversation content and advice based on the determined emotional state, and means for playing it back as voice. This makes it possible to provide personalized conversations and care advice tailored to the emotional state of elderly people.

[1224] "User authentication" refers to the process of verifying a user's identity in order to access a system.

[1225] "Profile information" refers to data that includes personal information and specific characteristics about a user.

[1226] "Voice input" refers to audio data used to capture a user's speech into the system.

[1227] "Text conversion" refers to the process of converting audio data into text data.

[1228] A "server" refers to a computer device that processes and stores data.

[1229] "Conversation content generation" refers to the process of creating appropriate responses or messages based on input data.

[1230] "Speech conversion" refers to the process of converting text data into speech data.

[1231] "Audio playback" refers to the act of making the generated audio data playable for the user.

[1232] "Emotional state" refers to the emotional state of a user as analyzed from their speech patterns and content.

[1233] "Health status" refers to data that indicates the user's physical and mental condition.

[1234] "Advice generation" refers to the process of creating appropriate advice based on the user's situation and data.

[1235] "Health history" refers to records of a user's past health condition.

[1236] A "database" refers to a system for effectively storing, managing, and retrieving information.

[1237] An "emotion engine" refers to software that analyzes a user's voice data to determine their emotional state.

[1238] This invention is a system that combines user authentication, speech recognition, a generative AI model, and an emotion recognition engine to improve the mental and physical care of the elderly. Specifically, the user communicates by voice using a terminal, the voice is analyzed by a server, and appropriate responses and advice are generated. The specific method for implementing this system is described below.

[1239] User authentication and profile information management

[1240] Users log in to the device using fingerprint authentication, facial recognition, or by entering a username and password. This allows the system to recognize who the individual is and provide personalized services.

[1241] The device receives the authentication information entered by the user and sends it to the server. Specifically, a standard facial recognition camera can be used for facial authentication, and a fingerprint scanner can be used for fingerprint authentication. The login information is encrypted and securely transmitted to the server.

[1242] The server compares the transmitted authentication information with the user's profile information stored in the database. If authentication is successful, the server sends the user's profile information to the device. This allows the device to provide services tailored to that user.

[1243] Speech recognition and conversation content generation

[1244] Users can converse with the device using their voice. For example, they can say, "Good morning, it's a nice day today."

[1245] The device receives voice input and converts the voice data into text using speech recognition software such as the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[1246] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate appropriate conversation content based on the received text data. The generated text data is then sent to the terminal.

[1247] The device converts the generated conversation into speech using speech synthesis software such as Amazon Polly and plays it back to the user. A concrete example of such a response might be, "Good morning. It's sunny today. Shall we go for a walk?"

[1248] emotion recognition

[1249] Users can talk to the device about their physical condition and emotions. For example, they might say, "I'm feeling a little down today."

[1250] The device records the input voice data and analyzes the emotional state using emotion recognition software such as the Microsoft Azure Emotion API. The results are then sent to the server.

[1251] The server uses a generative AI model based on the analysis results to generate conversation content and produce appropriate responses that match the user's emotional state. For example, a possible response might be, "Why don't you listen to your favorite music to change your mood?"

[1252] The device converts the generated text data into speech and plays it back to the user.

[1253] Caregiving assistance

[1254] Users can talk about specific health issues they have. For example, they might say, "My lower back has been a little sore since yesterday."

[1255] The device records the audio data and converts it into text data using the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[1256] The server retrieves the user's past health history from the database based on the received text data. Then, it uses a generative AI model to generate appropriate advice, such as a message like, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1257] The device converts the generated text data into speech using speech synthesis software and plays it back to the user.

[1258] Specific example

[1259] Example of a prompt:

[1260] Everyday conversation: "Good morning, it's a beautiful day today."

[1261] Emotion recognition: "I'm feeling a little down today."

[1262] Caregiving advice: "My lower back has been a little sore since yesterday."

[1263] Thus, the present invention integrates a series of processes, including user authentication, voice recognition, emotion recognition, conversation generation, and health management, to realize more personalized responses for the elderly. This makes it possible to provide efficient and effective mental and physical care for the elderly.

[1264] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1265] Program processing flow

[1266] User authentication and profile information management

[1267] Step 1:

[1268] The user authenticates using their fingerprint or facial recognition, or enters a username and password.

[1269] Input: Fingerprint data, facial image, username and password

[1270] Output: Authentication information

[1271] Specific action: The user places their finger on the fingerprint sensor of their smartphone.

[1272] Step 2:

[1273] The terminal receives the entered authentication information, encrypts it, and sends it to the server.

[1274] Input: Authentication information

[1275] Output: Encrypted authentication information

[1276] Specific operation: The device encrypts the received fingerprint data using AES (Advanced Encryption Standard).

[1277] Step 3:

[1278] The server compares the received authentication information with the registered information in the database.

[1279] Input: Encrypted credentials

[1280] Output: Authentication results and profile information

[1281] Specific operation: The server performs user authentication by comparing the user's profile information with that in the database. If authentication is successful, the server then retrieves the profile information.

[1282] Step 4:

[1283] If the server successfully authenticates, it will send the profile information to the device.

[1284] Input: Profile Information

[1285] Output: Profile information sent to the device

[1286] Specific operation: The server encrypts the profile information it has acquired and sends it to the terminal.

[1287] Everyday conversation

[1288] Step 1:

[1289] The user speaks to the device, saying, "Good morning, it's a nice day today."

[1290] Input: Audio data

[1291] Output: Audio data recorded by the device's microphone

[1292] Specific action: The user speaks into the smart speaker.

[1293] Step 2:

[1294] The device receives voice input and uses the Google Cloud Speech-to-Text API to convert the voice data into text.

[1295] Input: Audio data

[1296] Output: Text data

[1297] Specific operation: The device sends voice data to the API and retrieves the text data "It's a nice day today."

[1298] Step 3:

[1299] The terminal sends the converted text data to the server.

[1300] Input: Text data

[1301] Output: Text data sent to the server

[1302] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[1303] Step 4:

[1304] The server inputs the received text data into a generating AI model (e.g., GPT-3) to generate appropriate conversation content.

[1305] Input: Text data

[1306] Output: Generated conversation content

[1307] Specific operation: The server inputs the text "It's a nice day today" into GPT-3 and generates the response "Good morning. It's sunny today. Shall we go for a walk?".

[1308] Step 5:

[1309] The server sends the generated conversation content to the terminal.

[1310] Input: Generated conversation content

[1311] Output: Text data sent to the terminal

[1312] Specific operation: The server sends the generated text data to the terminal.

[1313] Step 6:

[1314] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[1315] Input: Text data of the generated conversation

[1316] Output: Audio data

[1317] Specific action: The device sends the text message "Good morning. It's sunny today. Shall we go for a walk?" to Amazon Polly and plays the audio data.

[1318] emotion recognition

[1319] Step 1:

[1320] The user speaks to the device about their physical condition and emotions.

[1321] Input: Audio data

[1322] Output: Audio data recorded by the device's microphone

[1323] Specific action: The user says, "I'm feeling a little down today."

[1324] Step 2:

[1325] The device records voice data and analyzes the emotional state using the Microsoft Azure Emotion API.

[1326] Input: Audio data

[1327] Output: Emotion analysis results

[1328] Specific operation: The device sends voice data to the API to obtain the emotional state (e.g., "feeling down").

[1329] Step 3:

[1330] The terminal sends the analysis results to the server.

[1331] Input: Sentiment analysis results

[1332] Output: Sentiment analysis results sent to the server

[1333] Specific operation: The device sends the emotion analysis results to the server using the HTTPS protocol.

[1334] Step 4:

[1335] The server generates conversation content using an AI model based on the emotion analysis results it receives.

[1336] Input: Sentiment analysis results

[1337] Output: Generated conversation content

[1338] Specific operation: The server inputs the emotional state of "feeling down" into GPT-3 and generates the response, "Why don't you try listening to your favorite music to cheer yourself up?"

[1339] Step 5:

[1340] The server sends the generated conversation content to the terminal.

[1341] Input: Generated conversation content

[1342] Output: Text data sent to the terminal

[1343] Specific operation: The server sends the generated text data to the terminal.

[1344] Step 6:

[1345] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[1346] Input: Text data of the generated conversation

[1347] Output: Audio data

[1348] Specific action: The device sends a text message to Amazon Polly saying, "Why not listen to some music you like to change your mood?" and plays the audio data.

[1349] Caregiving assistance

[1350] Step 1:

[1351] The user speaks to the device about their health problems.

[1352] Input: Audio data

[1353] Output: Audio data recorded by the device's microphone

[1354] Specific action: The user says, "My lower back has been a little sore since yesterday."

[1355] Step 2:

[1356] The device records the voice data and converts it into text data using the Google Cloud Speech-to-Text API.

[1357] Input: Audio data

[1358] Output: Text data

[1359] Specific operation: The device sends voice data to the API and retrieves text data saying, "My lower back has been a little sore since yesterday."

[1360] Step 3:

[1361] The terminal sends the converted text data to the server.

[1362] Input: Text data

[1363] Output: Text data sent to the server

[1364] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[1365] Step 4:

[1366] The server retrieves the user's past health history from the database based on text data.

[1367] Input: Text data

[1368] Output: Health history data

[1369] Specific operation: The server queries the database using the text "My lower back has been a little sore since yesterday" to retrieve past health history.

[1370] Step 5:

[1371] The server uses the generated AI model to produce appropriate advice.

[1372] Input: Health history data

[1373] Output: Generated advice

[1374] Specific operation: The server inputs health history data into GPT-3 and generates advice such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1375] Step 6:

[1376] The server sends the generated advice to the terminal.

[1377] Input: Generated advice

[1378] Output: Text data sent to the terminal

[1379] Specific operation: The server sends the generated text data to the terminal.

[1380] Step 7:

[1381] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[1382] Input: Text data of the generated advice

[1383] Output: Audio data

[1384] Specific action: The device sends a text message to Amazon Polly saying, "Please keep your lower back warm. Use a bath towel to cool it down if necessary," and plays an audio message.

[1385] In this way, by processing data based on specific input data at each step, a system is realized that provides personalized conversations and care advice to the elderly.

[1386] (Application Example 2)

[1387] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1388] In systems designed to support the safe and secure lives of the elderly, it is crucial not only to provide conversation and care assistance, but also to understand the user's emotional state in real time and provide appropriate support and advice tailored to their feelings. Furthermore, when using autonomous vehicles, it is especially important to alleviate the anxiety and tension felt by the elderly and support safe travel. Conventional technologies have difficulty adequately reflecting the user's emotions and state, making the enhancement of mental care a challenge.

[1389] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state, means for adjusting the conversation content based on the analysis results, and means for providing safety information for the autonomous vehicle. As a result, the user's emotional state can be grasped in real time, and personalized conversations and support can be provided accordingly, thereby increasing the elderly's sense of security and reducing anxiety and tension during travel. Furthermore, this makes it possible to provide an environment in which the elderly can travel with peace of mind even when using an autonomous vehicle.

[1390] "Means of user authentication" refers to methods used to verify the identity of a user when they access a system.

[1391] "Means of managing user profile information" refers to a method of centrally managing users' personal information and settings.

[1392] "Methods for converting voice input to text" refer to technologies that convert a user's voice data into text data.

[1393] "Means for sending converted text data to a server" refers to a method for converting voice input into text data and then sending that data to a server.

[1394] "Methods for generating conversation content based on received text data" refers to technologies that create appropriate dialogue content based on received text data.

[1395] "Means of converting generated conversation content into speech" refers to technology that converts the created dialogue content back into speech data.

[1396] "Means of playback as audio" refers to methods for making the generated audio data listenable to the user.

[1397] "Means for analyzing a user's emotional state" refers to technologies for analyzing a user's emotions and psychological state in real time.

[1398] "Means of adjusting conversation content based on analysis results" refers to technologies that change the content of a dialogue based on analyzed emotional data.

[1399] "Means for managing a user's health status" refers to methods for collecting and managing a user's health data.

[1400] "Means for collecting data on users' physical condition and problems" refers to technologies for collecting information on users' physical condition and problems.

[1401] "Means of generating appropriate advice based on collected data" refers to methods of creating appropriate advice based on acquired data.

[1402] "Means for providing safety information on autonomous vehicles" refers to technologies that provide users with safety information related to autonomous vehicles.

[1403] "Means for storing user profile information and health status in a database" refers to methods for storing users' personal information and health information in a database.

[1404] "A means of referring to past health history based on stored data" refers to a technology that allows users to check their past health status based on data stored in a database.

[1405] "Means of providing real-time support based on the user's emotional state" refers to technologies that provide immediate support and advice based on emotional data analyzed in real time.

[1406] This invention is a system that analyzes a user's emotional state in real time and provides conversation content and safety information for autonomous vehicles accordingly. The system consists of a terminal used by the user and a server that processes data. The specific implementation method is described below.

[1407] System Configuration

[1408] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, and functions as an interface for smartphones and autonomous vehicles. The server uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more appropriate responses.

[1409] Basic operation

[1410] User authentication and profile information management

[1411] Users log in to the system using their device via voice or touch controls. The device accepts fingerprint or facial recognition, or username and password input, and sends the authentication information to the server. The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information. If authentication is successful, the server sends the user's profile information to the device.

[1412] Everyday conversation

[1413] When a user speaks to the device, saying "Good morning, it's a nice day today," the device receives the voice input and records it as audio data. It then converts the audio data to text and sends it to the server. The server inputs the received text data into an AI model and generates appropriate conversation content. The generated conversation content is sent to the device as text data, which then converts it back into audio data and plays it back.

[1414] Emotion recognition and regulation

[1415] The device, upon receiving the user's everyday conversations and statements about their physical condition, analyzes the audio data using an emotion engine to determine their emotional state. It then sends this emotional state information to the server. Based on the received emotional information, the server appropriately adjusts the conversation content and support information to provide a personalized response.

[1416] Caregiving assistance and health management

[1417] For example, if a user says, "My lower back has been a little sore since yesterday," the device converts the audio to text and sends it to the server. The server refers to the user's health history and emotional state and generates appropriate advice. The generated advice is sent to the device as a message such as, "Keep your lower back warm. Use a bath towel to cool it if necessary," and the device converts this back to audio and plays it.

[1418] Applications in autonomous vehicles

[1419] If a user feels "a little uneasy" while using an autonomous vehicle, the terminal converts the voice input into text and analyzes it with an emotion engine. Based on the analysis results, the server provides real-time support such as, "It's okay, the car is driving safely. Please relax. Would you like some tea?" Also, if the user says, "I hear a strange noise," the system checks the vehicle's status and provides reassuring feedback in real time.

[1420] Hardware and software to be used

[1421] Hardware: Smartphone, microphone, speaker, autonomous vehicle interface

[1422] Software: Python, speech recognition library (speech_recognition), speech synthesis library (gTTS), natural language processing library (transformers)

[1423] Specific examples and prompt statements

[1424] Example 1: Anxiety

[1425] If the user says, "I'm a little worried," the system will generate a response such as, "We've determined your emotion to be negative, but it's okay. The car is driving safely, so please relax."

[1426] Specific example 2: Problem

[1427] If you say, "I hear a strange noise," and your emotion is negative, the response you'll get is, "We're checking the car's condition. There doesn't seem to be any problem, so please don't worry."

[1428] Example of a prompt:

[1429] "Generate a reassuring response regarding vehicle safety for elderly individuals who are expressing anxiety."

[1430] "Please generate relaxing responses to the negative emotions expressed by older adults."

[1431] By constructing a concrete system in this way, it is possible to enhance the mental care of the elderly and provide an environment in which they can use autonomous vehicles with peace of mind.

[1432] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1433] Step 1:

[1434] User Login

[1435] Input: The terminal receives the user's voice, fingerprint, facial recognition, or username and password.

[1436] Operation:

[1437] The terminal converts the received authentication information into text data and sends it to the server.

[1438] Output: Authentication information is sent to the server as text data.

[1439] Step 2:

[1440] Verification of authentication information and retrieval of profile information

[1441] Input: The server authenticates the user based on the authentication information received from the terminal.

[1442] Operation:

[1443] The server compares the user information with the user information in the database, and if authentication is successful, retrieves the user's profile information.

[1444] Output: Authentication results and profile information are sent to the device.

[1445] Step 3:

[1446] Processing everyday conversations

[1447] Input: The user speaks to the device (e.g., "Good morning, it's a nice day today").

[1448] Operation:

[1449] The device records voice input as data and converts the voice data into text data.

[1450] Send the converted text data to the server.

[1451] Output: The converted text data is sent to the server.

[1452] Step 4:

[1453] Generating conversation content

[1454] Input: The server inputs the received text data into the natural language processing AI.

[1455] Operation:

[1456] The server performs natural language processing to generate appropriate conversation content.

[1457] The conversation content is generated as text data and sent to the device.

[1458] Output: The generated conversation content is sent as text data to the terminal.

[1459] Step 5:

[1460] Speech conversion and playback of conversation content

[1461] Input: The terminal receives text data of the conversation received from the server.

[1462] Operation:

[1463] The device converts the received text data into audio data.

[1464] Play audio data and let the user listen.

[1465] Output: The generated audio data is played.

[1466] Step 6:

[1467] emotion recognition

[1468] Input: The user talks about everyday topics or their health.

[1469] Operation:

[1470] The device records voice input and feeds the voice data into the emotion engine.

[1471] The emotion engine analyzes the audio data and determines the emotional state.

[1472] Send emotional state information to the server.

[1473] Output: The analyzed emotional state information is sent to the server.

[1474] Step 7:

[1475] Emotion-based response generation

[1476] Input: The server generates a response based on the user's emotional state information and the content of the previous conversation.

[1477] Operation:

[1478] The server adjusts the conversation based on sentiment data and generates customized responses.

[1479] The response is sent to the terminal as text data.

[1480] Output: Text data containing the customized response is sent to the terminal.

[1481] Step 8:

[1482] Caregiving advice

[1483] Input: The user talks about their physical condition or any problems they are experiencing (e.g., "My back has been hurting since yesterday").

[1484] Operation:

[1485] The device converts the audio to text and sends it to the server.

[1486] The server references the user's health history and emotional state to generate appropriate advice.

[1487] The advice is sent to the device, converted to audio, and then played back.

[1488] Output: The generated audio data is played back to the user.

[1489] Step 9:

[1490] Support for autonomous vehicles

[1491] Input: When the user feels anxious while traveling (e.g., "I feel a little anxious").

[1492] Operation:

[1493] The device converts voice input into text and analyzes it using an emotion engine.

[1494] The server generates a reassuring response based on emotional state information.

[1495] The response is converted into audio data and played back to the user.

[1496] Output: The generated audio data is played back inside the autonomous vehicle to reassure the user.

[1497] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1498] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1499] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1500] [Third Embodiment]

[1501] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1502] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1503] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1504] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1505] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1506] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1507] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1508] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1509] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1510] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1511] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1512] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1513] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. The aim of this system is to alleviate the user's feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support.

[1514] System Configuration

[1515] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using an AI model.

[1516] Basic operation

[1517] User authentication and profile information management

[1518] User:

[1519] Users log in to the system by using voice commands or touch controls on the device.

[1520] Terminal:

[1521] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[1522] Send authentication information to the server.

[1523] server:

[1524] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[1525] If authentication is successful, the user's profile information will be sent to the device.

[1526] Everyday conversation

[1527] User:

[1528] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[1529] Terminal:

[1530] The device converts voice input into text and sends the text data to the server.

[1531] server:

[1532] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[1533] The generated conversation content is sent to the terminal as text data.

[1534] Terminal:

[1535] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[1536] Caregiving assistance

[1537] User:

[1538] The user says, "My lower back has been a little sore since yesterday."

[1539] Terminal:

[1540] The device converts voice input into text and sends the text data to the server.

[1541] server:

[1542] The server compares the received text data with the database and refers to the user's past health history.

[1543] It generates appropriate advice and sends text data to the device saying, "Keep your lower back warm. Use a bath towel to cool it down if necessary."

[1544] Terminal:

[1545] The device converts the received text data into speech and plays it back to the user.

[1546] Specific example

[1547] The following will explain this with specific examples.

[1548] Example 1: Everyday conversation

[1549] The user launches the smartphone app and performs fingerprint authentication.

[1550] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[1551] The device receives this, converts the audio to text, and sends it to the server.

[1552] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[1553] The device converts this into audio and plays it back.

[1554] Example 2: Caregiving advice

[1555] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[1556] The device converts the audio to text and sends it to the server.

[1557] The server references the user's health history and generates appropriate advice.

[1558] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1559] The device converts the audio to sound and plays it back.

[1560] As described above, this system utilizes generative AI to efficiently perform daily conversations and provide care assistance for the elderly. This system is expected to alleviate feelings of loneliness among the elderly, maintain cognitive function, and improve the quality of care.

[1561] The following describes the processing flow.

[1562] User authentication and profile information management

[1563] User

[1564] Step 1:

[1565] Launch the smartphone app.

[1566] terminal

[1567] Step 2:

[1568] It displays a screen for fingerprint authentication, facial recognition, or entry of a username and password.

[1569] User

[1570] Step 3:

[1571] You can use fingerprint authentication, facial recognition, or enter your username and password.

[1572] terminal

[1573] Step 4:

[1574] Send authentication information to the server.

[1575] server

[1576] Step 5:

[1577] The received authentication information is compared with the database to authenticate the user.

[1578] Step 6:

[1579] If authentication is successful, the user's profile information is retrieved from the database and sent to the device.

[1580] terminal

[1581] Step 7:

[1582] The received profile information is cached to prepare for the next process.

[1583] Everyday conversation

[1584] User

[1585] Step 1:

[1586] I say to the device, "Good morning, it's a beautiful day today."

[1587] terminal

[1588] Step 2:

[1589] It receives voice input and records it as audio data.

[1590] Step 3:

[1591] Converting audio data into text data (speech recognition processing).

[1592] Step 4:

[1593] Send the converted text data to the server.

[1594] server

[1595] Step 5:

[1596] The received text data is input into an AI model to generate appropriate conversation content (natural language processing).

[1597] Step 6:

[1598] The generated conversation content is sent to the terminal as text data.

[1599] terminal

[1600] Step 7:

[1601] Converts received text data into audio data (speech synthesis).

[1602] Step 8:

[1603] The generated audio data is played for the user.

[1604] Caregiving assistance

[1605] User

[1606] Step 1:

[1607] I spoke to the device, saying, "My lower back has been a little sore since yesterday."

[1608] terminal

[1609] Step 2:

[1610] It receives voice input and records it as audio data.

[1611] Step 3:

[1612] Converting audio data into text data (speech recognition processing).

[1613] Step 4:

[1614] Send the converted text data to the server.

[1615] server

[1616] Step 5:

[1617] The received text data is compared against a database to retrieve past health history.

[1618] Step 6:

[1619] Based on the acquired health history, appropriate advice is generated (health analysis and advice generation).

[1620] Step 7:

[1621] The generated advice is sent to the device as text data.

[1622] terminal

[1623] Step 8:

[1624] Converts received text data into audio data (speech synthesis).

[1625] Step 9:

[1626] The generated advice will be played back as audio.

[1627] The above is a detailed explanation of the program's processing flow.

[1628] (Example 1)

[1629] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1630] In modern society, there is a need to alleviate loneliness among the elderly, maintain cognitive function, and provide appropriate care support. However, many elderly people have limited opportunities for daily conversation and timely care, putting them at high risk of deteriorating mental and physical health. To solve this problem, a more effective and efficient system is needed.

[1631] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1632] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content using a generative AI model based on the received text data, means for converting the generated conversation content to speech, and means for playing it back as speech. This makes it possible for elderly people to receive everyday conversations and care advice without feeling lonely. Furthermore, by combining means for managing the user's health status, means for collecting data on the user's physical condition and problems, means for generating appropriate advice using a generative AI model based on the collected data, means for converting the generated advice to speech, and means for playing it back as speech, it becomes possible to provide appropriate responses based on the user's past health history.

[1633] "User authentication" is the process of verifying that the user accessing the system is indeed that specific user.

[1634] "Profile information" refers to data that includes personal information and attributes about the user.

[1635] "Means for converting voice input to text" refers to devices or software that use speech recognition technology to convert a user's speech into text format.

[1636] "Means of sending converted text data to the server" refers to the process of sending text data generated on the terminal to the server via the network.

[1637] A "generative AI model" is an algorithm or software that uses natural language processing and machine learning to generate appropriate responses from input text data.

[1638] "Means for generating conversation content" refers to the process of generating appropriate text for interacting with the user using a generation AI model based on received text data.

[1639] "Means for converting generated conversation content into speech" refers to devices or software that convert the text of generated conversation content into speech format using speech synthesis technology.

[1640] "Means of playback as audio" refers to the process of playing audio data to the user through audio equipment such as speakers.

[1641] "Means for managing health status" refers to devices and software that record and monitor a user's health information and periodically evaluate that status.

[1642] "Means of collecting data on health and problems" refers to the process of collecting information on health and problems through user feedback and questions.

[1643] "Means for generating appropriate advice" refers to the process of generating appropriate advice and suggestions for users using a generative AI model based on collected data and past history.

[1644] "A means of referencing past health history based on saved data" refers to the process of searching for a user's health information stored in a database and referencing their past health status and history.

[1645] Modes for carrying out the invention

[1646] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. This system aims to alleviate feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support by implementing user authentication, daily conversation, and care assistance functions.

[1647] System Configuration

[1648] The system consists of user terminals and servers that process data.

[1649] A terminal is a device that performs voice input and output, and includes smartphones, tablets, and smart speakers. A server is a computer device that performs natural language processing using AI models, and examples include cloud servers and dedicated servers.

[1650] Hardware and software usage

[1651] The system uses the following specific hardware and software.

[1652] Speech recognition software for converting voice input to text (e.g., Google Cloud Speech-to-Text, IBM Watson Speech to Text)

[1653] Generative AI models for generating conversation content based on text data (e.g., OpenAI GPT-3)

[1654] Text-to-speech software for converting text data into speech (e.g., Amazon Polly, Google Cloud Text-to-Speech)

[1655] Authentication technologies for user authentication (e.g., fingerprint recognition, facial recognition, username and password)

[1656] User authentication and profile information management

[1657] The user attempts to log in by using voice or touch controls on the device.

[1658] The device accepts fingerprint authentication, facial recognition, and username and password input, and sends the authentication information to the server.

[1659] The server compares the received authentication information with the database to determine whether authentication was successful. If authentication is successful, the user's profile information is sent to the device.

[1660] Everyday conversation

[1661] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[1662] The device converts voice input into text and sends that text data to the server.

[1663] The server uses an AI model to generate appropriate conversation content based on the received text data. The generated conversation content is then sent to the terminal as text data.

[1664] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[1665] Specific examples of prompt statements are as follows:

[1666] User: Good morning, it's a beautiful day today.

[1667] AI Model: Good morning. It's sunny today. Shall we go for a walk?

[1668] Caregiving assistance

[1669] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[1670] The device converts voice input into text and sends that text data to the server.

[1671] The server uses the received text data to compare and refer to the user's past health history in a database. Then, using a generative AI model, it generates appropriate advice and sends text data to the terminal such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1672] The device converts the received text data into speech and plays it back to the user.

[1673] Specific examples of prompt statements are as follows:

[1674] User: My lower back has been a little sore since yesterday.

[1675] AI Model: Keep your lower back warm. Use a bath towel to cool it down if necessary.

[1676] The above describes the embodiments for carrying out the present invention. This system provides assistance with daily conversation and care for the elderly, contributing to the alleviation of feelings of loneliness, the maintenance and improvement of cognitive function, and the improvement of the quality of care.

[1677] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1678] User Authentication

[1679] Processing flow

[1680] Step 1:

[1681] The user attempts to log in by using voice or touch controls on the device.

[1682] Input: User voice or touch operation

[1683] Output: Login attempt data

[1684] Step 2:

[1685] The device accepts fingerprint authentication, facial recognition, and username and password input.

[1686] Input: User's fingerprint data, facial data, username and password

[1687] Output: Authentication information data

[1688] Step 3:

[1689] The terminal sends the entered authentication information to the server.

[1690] Input: Authentication information data

[1691] Output: Authentication request to the server

[1692] Step 4:

[1693] The server compares the received authentication information with the database to determine whether the authentication was successful.

[1694] Input: Authentication information data

[1695] Output: Authentication success / failure result

[1696] Step 5:

[1697] If authentication is successful, the server sends the user's profile information to the device.

[1698] Input: Authentication success / failure result

[1699] Output: Profile information data

[1700] Everyday conversation

[1701] Processing flow

[1702] Step 1:

[1703] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[1704] Input: User voice input

[1705] Output: Audio data

[1706] Step 2:

[1707] The device converts voice input into text and sends that text data to the server.

[1708] Input: Audio data

[1709] Output: Text data, request to send to the server

[1710] Step 3:

[1711] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[1712] Input: Text data

[1713] Output: Generated conversation text

[1714] Step 4:

[1715] The server sends the generated conversation content to the terminal as text data.

[1716] Input: Generated conversation text

[1717] Output: Conversation text data, request to send to terminal

[1718] Step 5:

[1719] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[1720] Input: Conversation text data

[1721] Output: Audio data

[1722] Caregiving assistance

[1723] Processing flow

[1724] Step 1:

[1725] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[1726] Input: User voice input

[1727] Output: Audio data

[1728] Step 2:

[1729] The device converts voice input into text and sends that text data to the server.

[1730] Input: Audio data

[1731] Output: Text data, request to send to the server

[1732] Step 3:

[1733] The server uses the received text data to compare and reference the user's past health history from the database.

[1734] Input: Text data

[1735] Output: Health history data

[1736] Step 4:

[1737] The server uses a generative AI model to generate appropriate advice and sends the text data to the terminal.

[1738] Input: Health history data

[1739] Output: Generated advice text data, request to send to terminal

[1740] Step 5:

[1741] The device converts the received text data into speech and plays back, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1742] Input: Advice text data

[1743] Output: Audio data

[1744] (Application Example 1)

[1745] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1746] Elderly people often find food delivery services difficult to use due to their complex operation. Furthermore, they struggle to select nutritionally balanced meals, hindering proper dietary management. This highlights the growing need for food delivery applications that improve the quality of life for the elderly.

[1747] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1748] In this invention, the server includes means for generating food delivery order content using a generative AI model, means for transmitting the generated order content to a food delivery system, and means for providing the generated nutritional advice. This makes it possible for elderly people to easily order food delivery by voice and also receive appropriate nutritional advice.

[1749] User authentication is the process of identifying users who are accessing a system and verifying that they are authorized users.

[1750] "Profile information" refers to personal information and settings information about a user, and is data necessary for the system to provide services tailored to the user.

[1751] "Voice input" is the process by which a user speaks to the system through a microphone or similar device, and the system receives voice data in response.

[1752] "Converting to text" is the process of converting audio data into text data.

[1753] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate conversation content or order details based on user input.

[1754] "Food delivery order details" refers to the details of the food and drinks that a user specifies when using a food delivery service.

[1755] A "food delivery system" refers to the entire service and its operating system for delivering food to customers who order it.

[1756] "Nutritional advice" refers to expert advice designed to help users choose healthy meals.

[1757] "Converting to speech" is the process of converting text data into audio data, making it playable for the user.

[1758] "Playback as audio" refers to the process of playing the converted audio data to the user through a speaker or other device.

[1759] "Past health history" refers to past data and records regarding the user's health status, which are used to generate current advice.

[1760] This invention provides a system that simplifies the operation of food delivery services for elderly people and simultaneously provides nutritional advice. This system includes a series of processes for voice input, conversation generation using a generative AI model, order content generation, and provision of nutritional advice.

[1761] System Configuration

[1762] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using a generative AI model.

[1763] User authentication and profile information management

[1764] User:

[1765] The user launches the food delivery app and logs into the system using fingerprint authentication or similar methods.

[1766] Terminal:

[1767] The device accepts fingerprint authentication, facial recognition, or username and password input for user authentication.

[1768] Send authentication information to the server.

[1769] server:

[1770] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[1771] If authentication is successful, the user's profile information will be sent to the device.

[1772] Food delivery order generation

[1773] User:

[1774] The user speaks to the device and says, "I want to eat taco rice today."

[1775] Terminal:

[1776] The device converts voice input into text and sends the text data to the server.

[1777] server:

[1778] The server receives the text data and sends it as a prompt to the AI ​​model, which then generates an appropriate food delivery order.

[1779] The generated order details are sent to the food delivery system.

[1780] Terminal:

[1781] The terminal converts the order confirmation received from the server into audio and plays it back to the user.

[1782] Nutritional advice

[1783] User:

[1784] The user says to the device, "Lately, I've been worried about my nutritional balance."

[1785] Terminal:

[1786] The device converts voice input into text and sends the text data to the server.

[1787] server:

[1788] The server references the user's past health history and uses a generative AI model to generate appropriate nutritional advice.

[1789] The system generates and sends a message to the device saying, "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[1790] Terminal:

[1791] The device converts the received nutritional advice into audio and plays it back to the user.

[1792] Specific example

[1793] Example of a food delivery order:

[1794] Example usage: The user says, "I want to eat taco rice today."

[1795] AI: "So, you'd like taco rice. Would you like it spicy or mild?"

[1796] The user replied, "Please make it mild."

[1797] AI: "Thank you. I'll order the taco rice (mild) then."

[1798] Examples of nutritional advice:

[1799] Example of use: The user says, "Lately, I've been worried about my nutritional balance."

[1800] AI: "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[1801] Example of a prompt

[1802] "I want to eat taco rice today."

[1803] "Lately, I've been worried about nutritional balance."

[1804] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1805] Step 1:

[1806] User: The user launches the food delivery app and logs into the system using fingerprint authentication or by entering a password. During this process, the user's fingerprint information and password are entered into the system.

[1807] Step 2:

[1808] Terminal: The terminal obtains user authentication information through on-screen input forms or fingerprint sensors and sends it to the server. In this case, the entered authentication information is the data sent to the server.

[1809] Step 3:

[1810] Server: The server verifies the received authentication information and compares it with the profile information in the database. If the comparison is successful, it sends the user's profile information back to the terminal. It verifies that the authentication information is correct and outputs the user's profile data.

[1811] Step 4:

[1812] User: After successful authentication, the user speaks into the device saying, "I want to eat taco rice today."

[1813] Step 5:

[1814] Terminal: The terminal performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process. Speech recognition software (e.g., Google Speech Recognition API) is used to convert the speech input into text.

[1815] Step 6:

[1816] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[1817] Step 7:

[1818] Server: The server receives text data and sends it to the generative AI model as a prompt. This prompt becomes the input to the AI ​​model. The generative AI model (e.g., OpenAI GPT-3) is used to generate appropriate food delivery order details.

[1819] Step 8:

[1820] Server: Retrieves order details generated by the generation AI model and sends them to the food delivery system. The generated order details become the output data.

[1821] Step 9:

[1822] Terminal: Converts the order confirmation received from the server into audio. This converted audio data becomes the input data for the next process.

[1823] Step 10:

[1824] Terminal: The terminal plays back the converted order details and provides feedback to the user. It uses a voice output device to play back the text-to-speech conversion.

[1825] Step 11:

[1826] User: The user speaks to the device, saying, "Lately, I've been worried about my nutritional balance."

[1827] Step 12:

[1828] Terminal: Performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process.

[1829] Step 13:

[1830] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[1831] Step 14:

[1832] Server: The server receives text data, references the user's health history, generates prompts about nutritional balance, and sends them to an AI model. The user's health history data and prompts serve as input data.

[1833] Step 15:

[1834] Server: Retrieves nutritional advice generated by the generation AI model and sends it to the terminal as text data. The generated nutritional advice becomes the output data.

[1835] Step 16:

[1836] Terminal: Converts text data of nutritional advice received from the server into speech. This converted speech data becomes input data for the next process.

[1837] Step 17:

[1838] Device: The device plays nutritional advice converted into audio and provides feedback to the user. It uses an audio output device to play the nutritional advice.

[1839] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1840] This invention is a system that combines a generative AI, which is used to provide conversational support and care assistance for the elderly, with an emotion engine that recognizes the user's emotions. The aim of this system is to provide more effective mental care for the elderly, by offering conversations and care advice tailored to their emotional state.

[1841] System Configuration

[1842] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, while the server is a computer that uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more personalized responses.

[1843] Basic operation

[1844] User authentication and profile information management

[1845] User

[1846] Users log in to the system by using voice commands or touch controls on the device.

[1847] terminal

[1848] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[1849] Send authentication information to the server.

[1850] server

[1851] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[1852] If authentication is successful, the user's profile information will be sent to the device.

[1853] Everyday conversation

[1854] User

[1855] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[1856] terminal

[1857] The device receives voice input and records it as voice data.

[1858] The audio data is converted into text data (speech recognition processing) and sent to the server.

[1859] server

[1860] The server inputs the received text data into an AI model and generates appropriate conversation content (natural language processing).

[1861] The generated conversation content is sent to the terminal as text data.

[1862] terminal

[1863] The device converts the received text data into audio data (speech synthesis) and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[1864] emotion recognition

[1865] User

[1866] You can talk to the device about everyday topics, your health, or any problems you're facing.

[1867] terminal

[1868] It receives voice input and records it as audio data.

[1869] The system analyzes the user's emotional state from voice data using an emotion engine and sends the data to the server.

[1870] server

[1871] The server uses an emotion engine to analyze the received audio data and determine the user's emotional state.

[1872] Based on the determined emotional state, the system generates conversation content and advice.

[1873] Caregiving assistance

[1874] User

[1875] "My lower back has been a little sore since yesterday," he said.

[1876] terminal

[1877] It receives voice input and records it as audio data.

[1878] The audio data is converted into text data (speech recognition processing) and sent to the server.

[1879] server

[1880] Text data is compared against a database to retrieve past health history.

[1881] Based on health history and emotional state, the system generates appropriate advice and sends a message to the device saying, "Please keep your lower back warm. If necessary, cool it with a bath towel."

[1882] terminal

[1883] Text data is converted into audio data (speech synthesis), and then played back to the user.

[1884] Specific example

[1885] Example 1: Everyday conversation

[1886] The user launches the smartphone app and performs fingerprint authentication.

[1887] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[1888] The device receives this, converts the audio to text, and sends it to the server.

[1889] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[1890] The device converts this into audio and plays it back.

[1891] Example 2: Caregiving advice

[1892] The user says, "My lower back has been a little sore since yesterday."

[1893] The device converts the audio to text and sends it to the server.

[1894] The server references the user's health history and generates appropriate advice.

[1895] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1896] The device converts the audio to sound and plays it back.

[1897] Example 3: Emotion recognition

[1898] The user says, "I'm feeling a little down today."

[1899] The device converts speech to text and then uses an emotion engine to analyze the emotional state.

[1900] Based on the analysis results, the server generates a conversation message such as, "Why don't you listen to your favorite music to change your mood?" and sends it to the terminal.

[1901] The device converts this into audio and plays it back.

[1902] In this way, this system combines generative AI and an emotion engine to achieve more personalized responses for the elderly. As a result, it becomes possible to provide efficient and effective mental and physical care for the elderly.

[1903] The following describes the processing flow.

[1904] User authentication and profile information management

[1905] User

[1906] Step 1: Launch the smartphone app.

[1907] terminal

[1908] Step 2: Display a screen for fingerprint authentication, facial recognition, or username and password entry.

[1909] User

[1910] Step 3: Perform fingerprint authentication or facial recognition, or enter your username and password.

[1911] terminal

[1912] Step 4: Send authentication information to the server.

[1913] server

[1914] Step 5: The received authentication information is compared with the database to authenticate the user.

[1915] Step 6: If authentication is successful, retrieve the user's profile information from the database and send it to the device.

[1916] terminal

[1917] Step 7: Cache the received profile information and prepare for the next step.

[1918] Everyday conversation

[1919] User

[1920] Step 1: Speak to your device and say, "Good morning, it's a beautiful day today."

[1921] terminal

[1922] Step 2: Receive voice input and record it as audio data.

[1923] Step 3: Convert the audio data to text (speech recognition process).

[1924] Step 4: Send the converted text data to the server.

[1925] server

[1926] Step 5: Input the received text data into the AI ​​model and generate appropriate conversation content (natural language processing).

[1927] Step 6: Send the generated conversation content as text data to the device.

[1928] terminal

[1929] Step 7: Convert the received text data into audio data (speech synthesis).

[1930] Step 8: Play the generated audio data to the user.

[1931] emotion recognition

[1932] User

[1933] Step 1: Speak into the device about everyday topics, your health, or any problems you're experiencing.

[1934] terminal

[1935] Step 2: Receive voice input and record it as audio data.

[1936] Step 3: Analyze the audio data with the emotion engine to determine the emotional state.

[1937] Step 4: Send the analyzed emotional state as text data to the server.

[1938] server

[1939] Step 5: Generate conversation content and advice based on the received text data and emotional state.

[1940] Step 6: Send the generated conversation content and advice to the device.

[1941] Caregiving assistance

[1942] User

[1943] Step 1: Start by saying, "My lower back has been a little sore since yesterday."

[1944] terminal

[1945] Step 2: Receive voice input and record it as audio data.

[1946] Step 3: Convert the audio data to text (speech recognition process).

[1947] Step 4: Send the converted text data to the server.

[1948] server

[1949] Step 5: Match the text data to the database to retrieve past health history.

[1950] Step 6: Generate appropriate advice based on health history and emotional state.

[1951] Step 7: Send the generated advice as text data to the device.

[1952] terminal

[1953] Step 8: Convert the received text data into audio data (speech synthesis).

[1954] Step 9: Play the generated advice as audio.

[1955] The above is a detailed explanation of the processing flow of the system that combines the emotion engine.

[1956] (Example 2)

[1957] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1958] There is a need for systems that provide mental and physical care to the elderly in a more effective and personalized way. In particular, it is necessary to accurately understand the emotional state of the elderly when they talk about their daily lives or health, and to provide appropriate responses and advice. However, current systems have shortcomings in emotional recognition, making it difficult to respond to each individual's situation.

[1959] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1960] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content based on the received text data, means for converting the generated conversation content to voice, means for analyzing the user's voice data and determining their emotional state, means for generating conversation content and advice based on the determined emotional state, and means for playing it back as voice. This makes it possible to provide personalized conversations and care advice tailored to the emotional state of elderly people.

[1961] "User authentication" refers to the process of verifying a user's identity in order to access a system.

[1962] "Profile information" refers to data that includes personal information and specific characteristics about a user.

[1963] "Voice input" refers to audio data used to capture a user's speech into the system.

[1964] "Text conversion" refers to the process of converting audio data into text data.

[1965] A "server" refers to a computer device that processes and stores data.

[1966] "Conversation content generation" refers to the process of creating appropriate responses or messages based on input data.

[1967] "Speech conversion" refers to the process of converting text data into speech data.

[1968] "Audio playback" refers to the act of making the generated audio data playable for the user.

[1969] "Emotional state" refers to the emotional state of a user as analyzed from their speech patterns and content.

[1970] "Health status" refers to data that indicates the user's physical and mental condition.

[1971] "Advice generation" refers to the process of creating appropriate advice based on the user's situation and data.

[1972] "Health history" refers to records of a user's past health condition.

[1973] A "database" refers to a system for effectively storing, managing, and retrieving information.

[1974] An "emotion engine" refers to software that analyzes a user's voice data to determine their emotional state.

[1975] This invention is a system that combines user authentication, speech recognition, a generative AI model, and an emotion recognition engine to improve the mental and physical care of the elderly. Specifically, the user communicates by voice using a terminal, the voice is analyzed by a server, and appropriate responses and advice are generated. The specific method for implementing this system is described below.

[1976] User authentication and profile information management

[1977] Users log in to the device using fingerprint authentication, facial recognition, or by entering a username and password. This allows the system to recognize who the individual is and provide personalized services.

[1978] The device receives the authentication information entered by the user and sends it to the server. Specifically, a standard facial recognition camera can be used for facial authentication, and a fingerprint scanner can be used for fingerprint authentication. The login information is encrypted and securely transmitted to the server.

[1979] The server compares the transmitted authentication information with the user's profile information stored in the database. If authentication is successful, the server sends the user's profile information to the device. This allows the device to provide services tailored to that user.

[1980] Speech recognition and conversation content generation

[1981] Users can converse with the device using their voice. For example, they can say, "Good morning, it's a nice day today."

[1982] The device receives voice input and converts the voice data into text using speech recognition software such as the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[1983] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate appropriate conversation content based on the received text data. The generated text data is then sent to the terminal.

[1984] The device converts the generated conversation into speech using speech synthesis software such as Amazon Polly and plays it back to the user. A concrete example of such a response might be, "Good morning. It's sunny today. Shall we go for a walk?"

[1985] emotion recognition

[1986] Users can talk to the device about their physical condition and emotions. For example, they might say, "I'm feeling a little down today."

[1987] The device records the input voice data and analyzes the emotional state using emotion recognition software such as the Microsoft Azure Emotion API. The results are then sent to the server.

[1988] The server uses a generative AI model based on the analysis results to generate conversation content and produce appropriate responses that match the user's emotional state. For example, a possible response might be, "Why don't you listen to your favorite music to change your mood?"

[1989] The device converts the generated text data into speech and plays it back to the user.

[1990] Caregiving assistance

[1991] Users can talk about specific health issues they have. For example, they might say, "My lower back has been a little sore since yesterday."

[1992] The device records the audio data and converts it into text data using the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[1993] The server retrieves the user's past health history from the database based on the received text data. Then, it uses a generative AI model to generate appropriate advice, such as a message like, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[1994] The device converts the generated text data into speech using speech synthesis software and plays it back to the user.

[1995] Specific example

[1996] Example of a prompt:

[1997] Everyday conversation: "Good morning, it's a beautiful day today."

[1998] Emotion recognition: "I'm feeling a little down today."

[1999] Caregiving advice: "My lower back has been a little sore since yesterday."

[2000] Thus, the present invention integrates a series of processes, including user authentication, voice recognition, emotion recognition, conversation generation, and health management, to realize more personalized responses for the elderly. This makes it possible to provide efficient and effective mental and physical care for the elderly.

[2001] The flow of the specific processing in Example 2 will be explained using Figure 13.

[2002] Program processing flow

[2003] User authentication and profile information management

[2004] Step 1:

[2005] The user authenticates using their fingerprint or facial recognition, or enters a username and password.

[2006] Input: Fingerprint data, facial image, username and password

[2007] Output: Authentication information

[2008] Specific action: The user places their finger on the fingerprint sensor of their smartphone.

[2009] Step 2:

[2010] The terminal receives the entered authentication information, encrypts it, and sends it to the server.

[2011] Input: Authentication information

[2012] Output: Encrypted authentication information

[2013] Specific operation: The device encrypts the received fingerprint data using AES (Advanced Encryption Standard).

[2014] Step 3:

[2015] The server compares the received authentication information with the registered information in the database.

[2016] Input: Encrypted credentials

[2017] Output: Authentication results and profile information

[2018] Specific operation: The server performs user authentication by comparing the user's profile information with that in the database. If authentication is successful, the server then retrieves the profile information.

[2019] Step 4:

[2020] If the server successfully authenticates, it will send the profile information to the device.

[2021] Input: Profile Information

[2022] Output: Profile information sent to the device

[2023] Specific operation: The server encrypts the profile information it has acquired and sends it to the terminal.

[2024] Everyday conversation

[2025] Step 1:

[2026] The user speaks to the device, saying, "Good morning, it's a nice day today."

[2027] Input: Audio data

[2028] Output: Audio data recorded by the device's microphone

[2029] Specific action: The user speaks into the smart speaker.

[2030] Step 2:

[2031] The device receives voice input and uses the Google Cloud Speech-to-Text API to convert the voice data into text.

[2032] Input: Audio data

[2033] Output: Text data

[2034] Specific operation: The device sends voice data to the API and retrieves the text data "It's a nice day today."

[2035] Step 3:

[2036] The terminal sends the converted text data to the server.

[2037] Input: Text data

[2038] Output: Text data sent to the server

[2039] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[2040] Step 4:

[2041] The server inputs the received text data into a generating AI model (e.g., GPT-3) to generate appropriate conversation content.

[2042] Input: Text data

[2043] Output: Generated conversation content

[2044] Specific operation: The server inputs the text "It's a nice day today" into GPT-3 and generates the response "Good morning. It's sunny today. Shall we go for a walk?".

[2045] Step 5:

[2046] The server sends the generated conversation content to the terminal.

[2047] Input: Generated conversation content

[2048] Output: Text data sent to the terminal

[2049] Specific operation: The server sends the generated text data to the terminal.

[2050] Step 6:

[2051] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[2052] Input: Text data of the generated conversation

[2053] Output: Audio data

[2054] Specific action: The device sends the text message "Good morning. It's sunny today. Shall we go for a walk?" to Amazon Polly and plays the audio data.

[2055] emotion recognition

[2056] Step 1:

[2057] The user speaks to the device about their physical condition and emotions.

[2058] Input: Audio data

[2059] Output: Audio data recorded by the device's microphone

[2060] Specific action: The user says, "I'm feeling a little down today."

[2061] Step 2:

[2062] The device records voice data and analyzes the emotional state using the Microsoft Azure Emotion API.

[2063] Input: Audio data

[2064] Output: Emotion analysis results

[2065] Specific operation: The device sends voice data to the API to obtain the emotional state (e.g., "feeling down").

[2066] Step 3:

[2067] The terminal sends the analysis results to the server.

[2068] Input: Sentiment analysis results

[2069] Output: Sentiment analysis results sent to the server

[2070] Specific operation: The device sends the emotion analysis results to the server using the HTTPS protocol.

[2071] Step 4:

[2072] The server generates conversation content using an AI model based on the emotion analysis results it receives.

[2073] Input: Sentiment analysis results

[2074] Output: Generated conversation content

[2075] Specific operation: The server inputs the emotional state of "feeling down" into GPT-3 and generates the response, "Why don't you try listening to your favorite music to cheer yourself up?"

[2076] Step 5:

[2077] The server sends the generated conversation content to the terminal.

[2078] Input: Generated conversation content

[2079] Output: Text data sent to the terminal

[2080] Specific operation: The server sends the generated text data to the terminal.

[2081] Step 6:

[2082] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[2083] Input: Text data of the generated conversation

[2084] Output: Audio data

[2085] Specific action: The device sends a text message to Amazon Polly saying, "Why not listen to some music you like to change your mood?" and plays the audio data.

[2086] Caregiving assistance

[2087] Step 1:

[2088] The user speaks to the device about their health problems.

[2089] Input: Audio data

[2090] Output: Audio data recorded by the device's microphone

[2091] Specific action: The user says, "My lower back has been a little sore since yesterday."

[2092] Step 2:

[2093] The device records the voice data and converts it into text data using the Google Cloud Speech-to-Text API.

[2094] Input: Audio data

[2095] Output: Text data

[2096] Specific operation: The device sends voice data to the API and retrieves text data saying, "My lower back has been a little sore since yesterday."

[2097] Step 3:

[2098] The terminal sends the converted text data to the server.

[2099] Input: Text data

[2100] Output: Text data sent to the server

[2101] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[2102] Step 4:

[2103] The server retrieves the user's past health history from the database based on text data.

[2104] Input: Text data

[2105] Output: Health history data

[2106] Specific operation: The server queries the database using the text "My lower back has been a little sore since yesterday" to retrieve past health history.

[2107] Step 5:

[2108] The server uses the generated AI model to produce appropriate advice.

[2109] Input: Health history data

[2110] Output: Generated advice

[2111] Specific operation: The server inputs health history data into GPT-3 and generates advice such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2112] Step 6:

[2113] The server sends the generated advice to the terminal.

[2114] Input: Generated advice

[2115] Output: Text data sent to the terminal

[2116] Specific operation: The server sends the generated text data to the terminal.

[2117] Step 7:

[2118] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[2119] Input: Text data of the generated advice

[2120] Output: Audio data

[2121] Specific action: The device sends a text message to Amazon Polly saying, "Please keep your lower back warm. Use a bath towel to cool it down if necessary," and plays an audio message.

[2122] In this way, by processing data based on specific input data at each step, a system is realized that provides personalized conversations and care advice to the elderly.

[2123] (Application Example 2)

[2124] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[2125] In systems designed to support the safe and secure lives of the elderly, it is crucial not only to provide conversation and care assistance, but also to understand the user's emotional state in real time and provide appropriate support and advice tailored to their feelings. Furthermore, when using autonomous vehicles, it is especially important to alleviate the anxiety and tension felt by the elderly and support safe travel. Conventional technologies have difficulty adequately reflecting the user's emotions and state, making the enhancement of mental care a challenge.

[2126] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state, means for adjusting the conversation content based on the analysis results, and means for providing safety information for the autonomous vehicle. As a result, the user's emotional state can be grasped in real time, and personalized conversations and support can be provided accordingly, thereby increasing the elderly's sense of security and reducing anxiety and tension during travel. Furthermore, this makes it possible to provide an environment in which the elderly can travel with peace of mind even when using an autonomous vehicle.

[2127] "Means of user authentication" refers to methods used to verify the identity of a user when they access a system.

[2128] "Means of managing user profile information" refers to a method of centrally managing users' personal information and settings.

[2129] "Methods for converting voice input to text" refer to technologies that convert a user's voice data into text data.

[2130] "Means for sending converted text data to a server" refers to a method for converting voice input into text data and then sending that data to a server.

[2131] "Methods for generating conversation content based on received text data" refers to technologies that create appropriate dialogue content based on received text data.

[2132] "Means of converting generated conversation content into speech" refers to technology that converts the created dialogue content back into speech data.

[2133] "Means of playback as audio" refers to methods for making the generated audio data listenable to the user.

[2134] "Means for analyzing a user's emotional state" refers to technologies for analyzing a user's emotions and psychological state in real time.

[2135] "Means of adjusting conversation content based on analysis results" refers to technologies that change the content of a dialogue based on analyzed emotional data.

[2136] "Means for managing a user's health status" refers to methods for collecting and managing a user's health data.

[2137] "Means for collecting data on users' physical condition and problems" refers to technologies for collecting information on users' physical condition and problems.

[2138] "Means of generating appropriate advice based on collected data" refers to methods of creating appropriate advice based on acquired data.

[2139] "Means for providing safety information on autonomous vehicles" refers to technologies that provide users with safety information related to autonomous vehicles.

[2140] "Means for storing user profile information and health status in a database" refers to methods for storing users' personal information and health information in a database.

[2141] "A means of referring to past health history based on stored data" refers to a technology that allows users to check their past health status based on data stored in a database.

[2142] "Means of providing real-time support based on the user's emotional state" refers to technologies that provide immediate support and advice based on emotional data analyzed in real time.

[2143] This invention is a system that analyzes a user's emotional state in real time and provides conversation content and safety information for autonomous vehicles accordingly. The system consists of a terminal used by the user and a server that processes data. The specific implementation method is described below.

[2144] System Configuration

[2145] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, and functions as an interface for smartphones and autonomous vehicles. The server uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more appropriate responses.

[2146] Basic operation

[2147] User authentication and profile information management

[2148] Users log in to the system using their device via voice or touch controls. The device accepts fingerprint or facial recognition, or username and password input, and sends the authentication information to the server. The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information. If authentication is successful, the server sends the user's profile information to the device.

[2149] Everyday conversation

[2150] When a user speaks to the device, saying "Good morning, it's a nice day today," the device receives the voice input and records it as audio data. It then converts the audio data to text and sends it to the server. The server inputs the received text data into an AI model and generates appropriate conversation content. The generated conversation content is sent to the device as text data, which then converts it back into audio data and plays it back.

[2151] Emotional recognition and regulation

[2152] The device, upon receiving the user's everyday conversations and statements about their physical condition, analyzes the audio data using an emotion engine to determine their emotional state. It then sends this emotional state information to the server. Based on the received emotional information, the server appropriately adjusts the conversation content and support information to provide a personalized response.

[2153] Caregiving assistance and health management

[2154] For example, if a user says, "My lower back has been a little sore since yesterday," the device converts the audio to text and sends it to the server. The server refers to the user's health history and emotional state and generates appropriate advice. The generated advice is sent to the device as a message such as, "Keep your lower back warm. Use a bath towel to cool it if necessary," and the device converts this back to audio and plays it.

[2155] Applications in autonomous vehicles

[2156] If a user feels "a little uneasy" while using an autonomous vehicle, the terminal converts the voice input into text and analyzes it with an emotion engine. Based on the analysis results, the server provides real-time support such as, "It's okay, the car is driving safely. Please relax. Would you like some tea?" Also, if the user says, "I hear a strange noise," the system checks the vehicle's status and provides reassuring feedback in real time.

[2157] Hardware and software to be used

[2158] Hardware: Smartphone, microphone, speaker, autonomous vehicle interface

[2159] Software: Python, speech recognition library (speech_recognition), speech synthesis library (gTTS), natural language processing library (transformers)

[2160] Specific examples and prompt statements

[2161] Example 1: Anxiety

[2162] If the user says, "I'm a little worried," the system will generate a response such as, "We've determined your emotion to be negative, but it's okay. The car is driving safely, so please relax."

[2163] Specific example 2: Problem

[2164] If you say, "I hear a strange noise," and your emotion is negative, the response you'll get is, "We're checking the car's condition. There doesn't seem to be any problem, so please don't worry."

[2165] Example of a prompt:

[2166] "Generate a reassuring response regarding vehicle safety for elderly individuals who are expressing anxiety."

[2167] "Please generate relaxing responses to the negative emotions expressed by older adults."

[2168] By constructing a concrete system in this way, it is possible to enhance the mental care of the elderly and provide an environment in which they can use autonomous vehicles with peace of mind.

[2169] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2170] Step 1:

[2171] User Login

[2172] Input: The terminal receives the user's voice, fingerprint, facial recognition, or username and password.

[2173] Operation:

[2174] The terminal converts the received authentication information into text data and sends it to the server.

[2175] Output: Authentication information is sent to the server as text data.

[2176] Step 2:

[2177] Verification of authentication information and retrieval of profile information

[2178] Input: The server authenticates the user based on the authentication information received from the terminal.

[2179] Operation:

[2180] The server compares the user information with the user information in the database, and if authentication is successful, retrieves the user's profile information.

[2181] Output: Authentication results and profile information are sent to the device.

[2182] Step 3:

[2183] Processing everyday conversations

[2184] Input: The user speaks to the device (e.g., "Good morning, it's a nice day today").

[2185] Operation:

[2186] The device records voice input as data and converts the voice data into text data.

[2187] Send the converted text data to the server.

[2188] Output: The converted text data is sent to the server.

[2189] Step 4:

[2190] Generating conversation content

[2191] Input: The server inputs the received text data into the natural language processing AI.

[2192] Operation:

[2193] The server performs natural language processing to generate appropriate conversation content.

[2194] The conversation content is generated as text data and sent to the device.

[2195] Output: The generated conversation content is sent as text data to the terminal.

[2196] Step 5:

[2197] Speech conversion and playback of conversation content

[2198] Input: The terminal receives text data of the conversation received from the server.

[2199] Operation:

[2200] The device converts the received text data into audio data.

[2201] Play audio data and let the user listen.

[2202] Output: The generated audio data is played.

[2203] Step 6:

[2204] emotion recognition

[2205] Input: The user talks about everyday topics or their health.

[2206] Operation:

[2207] The device records voice input and feeds the voice data into the emotion engine.

[2208] The emotion engine analyzes the audio data and determines the emotional state.

[2209] Send emotional state information to the server.

[2210] Output: The analyzed emotional state information is sent to the server.

[2211] Step 7:

[2212] Emotion-based response generation

[2213] Input: The server generates a response based on the user's emotional state information and the content of the previous conversation.

[2214] Operation:

[2215] The server adjusts the conversation based on sentiment data and generates customized responses.

[2216] The response is sent to the terminal as text data.

[2217] Output: Text data containing the customized response is sent to the terminal.

[2218] Step 8:

[2219] Caregiving advice

[2220] Input: The user talks about their physical condition or any problems they are experiencing (e.g., "My back has been hurting since yesterday").

[2221] Operation:

[2222] The device converts the audio to text and sends it to the server.

[2223] The server references the user's health history and emotional state to generate appropriate advice.

[2224] The advice is sent to the device, converted to audio, and then played back.

[2225] Output: The generated audio data is played back to the user.

[2226] Step 9:

[2227] Support for autonomous vehicles

[2228] Input: When the user feels anxious while traveling (e.g., "I feel a little anxious").

[2229] Operation:

[2230] The device converts voice input into text and analyzes it using an emotion engine.

[2231] The server generates a reassuring response based on emotional state information.

[2232] The response is converted into audio data and played back to the user.

[2233] Output: The generated audio data is played back inside the autonomous vehicle to reassure the user.

[2234] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2235] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[2237] [Fourth Embodiment]

[2238] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[2239] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2241] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[2242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[2243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[2244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[2245] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[2246] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[2247] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2248] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2249] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[2250] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2251] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. The aim of this system is to alleviate the user's feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support.

[2252] System Configuration

[2253] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using an AI model.

[2254] Basic operation

[2255] User authentication and profile information management

[2256] User:

[2257] Users log in to the system by using voice commands or touch controls on the device.

[2258] Terminal:

[2259] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[2260] Send authentication information to the server.

[2261] server:

[2262] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[2263] If authentication is successful, the user's profile information will be sent to the device.

[2264] Everyday conversation

[2265] User:

[2266] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[2267] Terminal:

[2268] The device converts voice input into text and sends the text data to the server.

[2269] server:

[2270] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[2271] The generated conversation content is sent to the terminal as text data.

[2272] Terminal:

[2273] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[2274] Caregiving assistance

[2275] User:

[2276] The user says, "My lower back has been a little sore since yesterday."

[2277] Terminal:

[2278] The device converts voice input into text and sends the text data to the server.

[2279] server:

[2280] The server compares the received text data with the database and refers to the user's past health history.

[2281] It generates appropriate advice and sends text data to the device saying, "Keep your lower back warm. Use a bath towel to cool it down if necessary."

[2282] Terminal:

[2283] The device converts the received text data into speech and plays it back to the user.

[2284] Specific example

[2285] The following will explain this with specific examples.

[2286] Example 1: Everyday conversation

[2287] The user launches the smartphone app and performs fingerprint authentication.

[2288] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[2289] The device receives this, converts the audio to text, and sends it to the server.

[2290] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[2291] The device converts this into audio and plays it back.

[2292] Example 2: Caregiving advice

[2293] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[2294] The device converts the audio to text and sends it to the server.

[2295] The server references the user's health history and generates appropriate advice.

[2296] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2297] The device converts the audio to sound and plays it back.

[2298] As described above, this system utilizes generative AI to efficiently perform daily conversations and provide care assistance for the elderly. This system is expected to alleviate feelings of loneliness among the elderly, maintain cognitive function, and improve the quality of care.

[2299] The following describes the processing flow.

[2300] User authentication and profile information management

[2301] User

[2302] Step 1:

[2303] Launch the smartphone app.

[2304] terminal

[2305] Step 2:

[2306] It displays a screen for fingerprint authentication, facial recognition, or entry of a username and password.

[2307] User

[2308] Step 3:

[2309] You can use fingerprint authentication, facial recognition, or enter your username and password.

[2310] terminal

[2311] Step 4:

[2312] Send authentication information to the server.

[2313] server

[2314] Step 5:

[2315] The received authentication information is compared with the database to authenticate the user.

[2316] Step 6:

[2317] If authentication is successful, the user's profile information is retrieved from the database and sent to the device.

[2318] terminal

[2319] Step 7:

[2320] The received profile information is cached to prepare for the next process.

[2321] Everyday conversation

[2322] User

[2323] Step 1:

[2324] I say to the device, "Good morning, it's a beautiful day today."

[2325] terminal

[2326] Step 2:

[2327] It receives voice input and records it as audio data.

[2328] Step 3:

[2329] Converting audio data into text data (speech recognition processing).

[2330] Step 4:

[2331] Send the converted text data to the server.

[2332] server

[2333] Step 5:

[2334] The received text data is input into an AI model to generate appropriate conversation content (natural language processing).

[2335] Step 6:

[2336] The generated conversation content is sent to the terminal as text data.

[2337] terminal

[2338] Step 7:

[2339] Converts received text data into audio data (speech synthesis).

[2340] Step 8:

[2341] The generated audio data is played for the user.

[2342] Caregiving assistance

[2343] User

[2344] Step 1:

[2345] I spoke to the device, saying, "My lower back has been a little sore since yesterday."

[2346] terminal

[2347] Step 2:

[2348] It receives voice input and records it as audio data.

[2349] Step 3:

[2350] Converting audio data into text data (speech recognition processing).

[2351] Step 4:

[2352] Send the converted text data to the server.

[2353] server

[2354] Step 5:

[2355] The received text data is compared against a database to retrieve past health history.

[2356] Step 6:

[2357] Based on the acquired health history, appropriate advice is generated (health analysis and advice generation).

[2358] Step 7:

[2359] The generated advice is sent to the device as text data.

[2360] terminal

[2361] Step 8:

[2362] Converts received text data into audio data (speech synthesis).

[2363] Step 9:

[2364] The generated advice will be played back as audio.

[2365] The above is a detailed explanation of the program's processing flow.

[2366] (Example 1)

[2367] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2368] In modern society, there is a need to alleviate loneliness among the elderly, maintain cognitive function, and provide appropriate care support. However, many elderly people have limited opportunities for daily conversation and timely care, putting them at high risk of deteriorating mental and physical health. To solve this problem, a more effective and efficient system is needed.

[2369] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[2370] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content using a generative AI model based on the received text data, means for converting the generated conversation content to speech, and means for playing it back as speech. This makes it possible for elderly people to receive everyday conversations and care advice without feeling lonely. Furthermore, by combining means for managing the user's health status, means for collecting data on the user's physical condition and problems, means for generating appropriate advice using a generative AI model based on the collected data, means for converting the generated advice to speech, and means for playing it back as speech, it becomes possible to provide appropriate responses based on the user's past health history.

[2371] "User authentication" is the process of verifying that the user accessing the system is indeed that specific user.

[2372] "Profile information" refers to data that includes personal information and attributes about the user.

[2373] "Means for converting voice input to text" refers to devices or software that use speech recognition technology to convert a user's speech into text format.

[2374] "Means of sending converted text data to the server" refers to the process of sending text data generated on the terminal to the server via the network.

[2375] A "generative AI model" is an algorithm or software that uses natural language processing and machine learning to generate appropriate responses from input text data.

[2376] "Means for generating conversation content" refers to the process of generating appropriate text for interacting with the user using a generation AI model based on received text data.

[2377] "Means for converting generated conversation content into speech" refers to devices or software that convert the text of generated conversation content into speech format using speech synthesis technology.

[2378] "Means of playback as audio" refers to the process of playing audio data to the user through audio equipment such as speakers.

[2379] "Means for managing health status" refers to devices and software that record and monitor a user's health information and periodically evaluate that status.

[2380] "Means of collecting data on health and problems" refers to the process of collecting information on health and problems through user feedback and questions.

[2381] "Means for generating appropriate advice" refers to the process of generating appropriate advice and suggestions for users using a generative AI model based on collected data and past history.

[2382] "A means of referencing past health history based on saved data" refers to the process of searching for a user's health information stored in a database and referencing their past health status and history.

[2383] Modes for carrying out the invention

[2384] This invention is a system that uses generative AI to provide conversational support and care assistance for the elderly. This system aims to alleviate feelings of loneliness, maintain and improve cognitive function, and provide appropriate care support by implementing user authentication, daily conversation, and care assistance functions.

[2385] System Configuration

[2386] The system consists of user terminals and servers that process data.

[2387] A terminal is a device that performs voice input and output, and includes smartphones, tablets, and smart speakers. A server is a computer device that performs natural language processing using AI models, and examples include cloud servers and dedicated servers.

[2388] Hardware and software usage

[2389] The system uses the following specific hardware and software.

[2390] Speech recognition software for converting voice input to text (e.g., Google Cloud Speech-to-Text, IBM Watson Speech to Text)

[2391] Generative AI models for generating conversation content based on text data (e.g., OpenAI GPT-3)

[2392] Text-to-speech software for converting text data into speech (e.g., Amazon Polly, Google Cloud Text-to-Speech)

[2393] Authentication technologies for user authentication (e.g., fingerprint recognition, facial recognition, username and password)

[2394] User authentication and profile information management

[2395] The user attempts to log in by using voice or touch controls on the device.

[2396] The device accepts fingerprint authentication, facial recognition, and username and password input, and sends the authentication information to the server.

[2397] The server compares the received authentication information with the database to determine whether authentication was successful. If authentication is successful, the user's profile information is sent to the device.

[2398] Everyday conversation

[2399] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[2400] The device converts voice input into text and sends that text data to the server.

[2401] The server uses an AI model to generate appropriate conversation content based on the received text data. The generated conversation content is then sent to the terminal as text data.

[2402] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[2403] Specific examples of prompt statements are as follows:

[2404] User: Good morning, it's a beautiful day today.

[2405] AI Model: Good morning. It's sunny today. Shall we go for a walk?

[2406] Caregiving assistance

[2407] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[2408] The device converts voice input into text and sends that text data to the server.

[2409] The server uses the received text data to compare and refer to the user's past health history in a database. Then, using a generative AI model, it generates appropriate advice and sends text data to the terminal such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2410] The device converts the received text data into speech and plays it back to the user.

[2411] Specific examples of prompt statements are as follows:

[2412] User: My lower back has been a little sore since yesterday.

[2413] AI Model: Keep your lower back warm. Use a bath towel to cool it down if necessary.

[2414] The above describes the embodiments for carrying out the present invention. This system provides assistance with daily conversation and care for the elderly, contributing to the alleviation of feelings of loneliness, the maintenance and improvement of cognitive function, and the improvement of the quality of care.

[2415] The flow of the specific processing in Example 1 will be explained using Figure 11.

[2416] User Authentication

[2417] Processing flow

[2418] Step 1:

[2419] The user attempts to log in by using voice or touch controls on the device.

[2420] Input: User voice or touch operation

[2421] Output: Login attempt data

[2422] Step 2:

[2423] The device accepts fingerprint authentication, facial recognition, and username and password input.

[2424] Input: User's fingerprint data, facial data, username and password

[2425] Output: Authentication information data

[2426] Step 3:

[2427] The terminal sends the entered authentication information to the server.

[2428] Input: Authentication information data

[2429] Output: Authentication request to the server

[2430] Step 4:

[2431] The server compares the received authentication information with the database to determine whether the authentication was successful.

[2432] Input: Authentication information data

[2433] Output: Authentication success / failure result

[2434] Step 5:

[2435] If authentication is successful, the server sends the user's profile information to the device.

[2436] Input: Authentication success / failure result

[2437] Output: Profile information data

[2438] Everyday conversation

[2439] Processing flow

[2440] Step 1:

[2441] The user speaks a specific trigger word or phrase to the device (e.g., "Good morning, it's a nice day today").

[2442] Input: User voice input

[2443] Output: Audio data

[2444] Step 2:

[2445] The device converts voice input into text and sends that text data to the server.

[2446] Input: Audio data

[2447] Output: Text data, request to send to the server

[2448] Step 3:

[2449] Based on the received text data, the server uses an AI model to generate appropriate conversation content.

[2450] Input: Text data

[2451] Output: Generated conversation text

[2452] Step 4:

[2453] The server sends the generated conversation content to the terminal as text data.

[2454] Input: Generated conversation text

[2455] Output: Conversation text data, request to send to terminal

[2456] Step 5:

[2457] The device converts the received text data into speech and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[2458] Input: Conversation text data

[2459] Output: Audio data

[2460] Caregiving assistance

[2461] Processing flow

[2462] Step 1:

[2463] The user speaks to the device, saying, "My lower back has been a little sore since yesterday."

[2464] Input: User voice input

[2465] Output: Audio data

[2466] Step 2:

[2467] The device converts voice input into text and sends that text data to the server.

[2468] Input: Audio data

[2469] Output: Text data, request to send to the server

[2470] Step 3:

[2471] The server uses the received text data to compare and reference the user's past health history from the database.

[2472] Input: Text data

[2473] Output: Health history data

[2474] Step 4:

[2475] The server uses a generative AI model to generate appropriate advice and sends the text data to the terminal.

[2476] Input: Health history data

[2477] Output: Generated advice text data, request to send to terminal

[2478] Step 5:

[2479] The device converts the received text data into speech and plays back, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2480] Input: Advice text data

[2481] Output: Audio data

[2482] (Application Example 1)

[2483] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2484] Elderly people often find food delivery services difficult to use due to their complex operation. Furthermore, they struggle to select nutritionally balanced meals, hindering proper dietary management. This highlights the growing need for food delivery applications that improve the quality of life for the elderly.

[2485] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[2486] In this invention, the server includes means for generating food delivery order content using a generative AI model, means for transmitting the generated order content to a food delivery system, and means for providing the generated nutritional advice. This makes it possible for elderly people to easily order food delivery by voice and also receive appropriate nutritional advice.

[2487] User authentication is the process of identifying users who are accessing a system and verifying that they are authorized users.

[2488] "Profile information" refers to personal information and settings information about a user, and is data necessary for the system to provide services tailored to the user.

[2489] "Voice input" is the process by which a user speaks to the system through a microphone or similar device, and the system receives voice data in response.

[2490] "Converting to text" is the process of converting audio data into text data.

[2491] A "generative AI model" is an algorithm or program that uses artificial intelligence to generate conversation content or order details based on user input.

[2492] "Food delivery order details" refers to the details of the food and drinks that a user specifies when using a food delivery service.

[2493] A "food delivery system" refers to the entire service and its operating system for delivering food to customers who order it.

[2494] "Nutritional advice" refers to expert advice designed to help users choose healthy meals.

[2495] "Converting to speech" is the process of converting text data into audio data, making it playable for the user.

[2496] "Playback as audio" refers to the process of playing the converted audio data to the user through a speaker or other device.

[2497] "Past health history" refers to past data and records regarding the user's health status, which are used to generate current advice.

[2498] This invention provides a system that simplifies the operation of food delivery services for elderly people and simultaneously provides nutritional advice. This system includes a series of processes for voice input, conversation generation using a generative AI model, order content generation, and provision of nutritional advice.

[2499] System Configuration

[2500] The system consists of a user terminal and a server that processes data. The terminal is a device that performs voice input and output, while the server is a computer that performs natural language processing using a generative AI model.

[2501] User authentication and profile information management

[2502] User:

[2503] The user launches the food delivery app and logs into the system using fingerprint authentication or similar methods.

[2504] Terminal:

[2505] The device accepts fingerprint authentication, facial recognition, or username and password input for user authentication.

[2506] Send authentication information to the server.

[2507] server:

[2508] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[2509] If authentication is successful, the user's profile information will be sent to the device.

[2510] Food delivery order generation

[2511] User:

[2512] The user speaks to the device and says, "I want to eat taco rice today."

[2513] Terminal:

[2514] The device converts voice input into text and sends the text data to the server.

[2515] server:

[2516] The server receives the text data and sends it as a prompt to the AI ​​model, which then generates an appropriate food delivery order.

[2517] The generated order details are sent to the food delivery system.

[2518] Terminal:

[2519] The terminal converts the order confirmation received from the server into audio and plays it back to the user.

[2520] Nutritional advice

[2521] User:

[2522] The user says to the device, "Lately, I've been worried about my nutritional balance."

[2523] Terminal:

[2524] The device converts voice input into text and sends the text data to the server.

[2525] server:

[2526] The server references the user's past health history and uses a generative AI model to generate appropriate nutritional advice.

[2527] The system generates and sends a message to the device saying, "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[2528] Terminal:

[2529] The device converts the received nutritional advice into audio and plays it back to the user.

[2530] Specific example

[2531] Example of a food delivery order:

[2532] Example usage: The user says, "I want to eat taco rice today."

[2533] AI: "So, you'd like taco rice. Would you like it spicy or mild?"

[2534] The user replied, "Please make it mild."

[2535] AI: "Thank you. I'll order the taco rice (mild) then."

[2536] Examples of nutritional advice:

[2537] Example of use: The user says, "Lately, I've been worried about my nutritional balance."

[2538] AI: "Taco rice is nutritionally balanced. If you want to get extra vitamin C, add a salad."

[2539] Example of a prompt

[2540] "I want to eat taco rice today."

[2541] "Lately, I've been worried about nutritional balance."

[2542] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[2543] Step 1:

[2544] User: The user launches the food delivery app and logs into the system using fingerprint authentication or by entering a password. During this process, the user's fingerprint information and password are entered into the system.

[2545] Step 2:

[2546] Terminal: The terminal obtains user authentication information through on-screen input forms or fingerprint sensors and sends it to the server. In this case, the entered authentication information is the data sent to the server.

[2547] Step 3:

[2548] Server: The server verifies the received authentication information and compares it with the profile information in the database. If the comparison is successful, it sends the user's profile information back to the terminal. It verifies that the authentication information is correct and outputs the user's profile data.

[2549] Step 4:

[2550] User: After successful authentication, the user speaks into the device saying, "I want to eat taco rice today."

[2551] Step 5:

[2552] Terminal: The terminal performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process. Speech recognition software (e.g., Google Speech Recognition API) is used to convert the speech input into text.

[2553] Step 6:

[2554] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[2555] Step 7:

[2556] Server: The server receives text data and sends it to the generative AI model as a prompt. This prompt becomes the input to the AI ​​model. The generative AI model (e.g., OpenAI GPT-3) is used to generate appropriate food delivery order details.

[2557] Step 8:

[2558] Server: Retrieves order details generated by the generation AI model and sends them to the food delivery system. The generated order details become the output data.

[2559] Step 9:

[2560] Terminal: Converts the order confirmation received from the server into audio. This converted audio data becomes the input data for the next process.

[2561] Step 10:

[2562] Terminal: The terminal plays back the converted order details and provides feedback to the user. It uses a voice output device to play back the text-to-speech conversion.

[2563] Step 11:

[2564] User: The user speaks to the device, saying, "Lately, I've been worried about my nutritional balance."

[2565] Step 12:

[2566] Terminal: Performs speech recognition and converts the speech data into text data. This text data becomes the input for the next process.

[2567] Step 13:

[2568] Terminal: Sends the converted text data to the server. The sent text data becomes the input data for the next process.

[2569] Step 14:

[2570] Server: The server receives text data, references the user's health history, generates prompts about nutritional balance, and sends them to an AI model. The user's health history data and prompts serve as input data.

[2571] Step 15:

[2572] Server: Retrieves nutritional advice generated by the generation AI model and sends it to the terminal as text data. The generated nutritional advice becomes the output data.

[2573] Step 16:

[2574] Terminal: Converts text data of nutritional advice received from the server into speech. This converted speech data becomes input data for the next process.

[2575] Step 17:

[2576] Device: The device plays nutritional advice converted into audio and provides feedback to the user. It uses an audio output device to play the nutritional advice.

[2577] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[2578] This invention is a system that combines a generative AI, which is used to provide conversational support and care assistance for the elderly, with an emotion engine that recognizes the user's emotions. The aim of this system is to provide more effective mental care for the elderly, by offering conversations and care advice tailored to their emotional state.

[2579] System Configuration

[2580] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, while the server is a computer that uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more personalized responses.

[2581] Basic operation

[2582] User authentication and profile information management

[2583] User

[2584] Users log in to the system by using voice commands or touch controls on the device.

[2585] terminal

[2586] The device accepts fingerprint authentication, facial recognition, or the input of a username and password for user authentication.

[2587] Send authentication information to the server.

[2588] server

[2589] The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information.

[2590] If authentication is successful, the user's profile information will be sent to the device.

[2591] Everyday conversation

[2592] User

[2593] The user speaks to the device, saying, "Good morning, it's a beautiful day today."

[2594] terminal

[2595] The device receives voice input and records it as voice data.

[2596] The audio data is converted into text data (speech recognition processing) and sent to the server.

[2597] server

[2598] The server inputs the received text data into an AI model and generates appropriate conversation content (natural language processing).

[2599] The generated conversation content is sent to the terminal as text data.

[2600] terminal

[2601] The device converts the received text data into audio data (speech synthesis) and plays back, "Good morning. It's sunny today. Shall we go for a walk?"

[2602] emotion recognition

[2603] User

[2604] You can talk to the device about everyday topics, your health, or any problems you're facing.

[2605] terminal

[2606] It receives voice input and records it as audio data.

[2607] The system analyzes the user's emotional state from voice data using an emotion engine and sends the data to the server.

[2608] Server

[2609] The server analyzes the received voice data using an emotion engine and determines the user's emotional state.

[2610] Based on the determined emotional state, conversation content and advice are generated.

[2611] Nursing assistance

[2612] User

[2613] Say, "I've had a bit of a pain in my lower back since yesterday."

[2614] Terminal

[2615] Receives voice input and records it as voice data.

[2616] Converts the voice data into text data (voice recognition processing) and sends it to the server.

[2617] Server

[2618] Matches the text data with the database to obtain the past health history.

[2619] Generates appropriate advice based on the health history and emotional state, and sends a message such as "Please avoid getting your lower back cold. If necessary, cool it with a bath towel." to the terminal.

[2620] Terminal

[2621] Converts the text data into voice data (voice synthesis processing) and plays it back to the user.

[2622] Specific example

[2623] Example 1: Daily conversation

[2624] The user launches the smartphone app and performs fingerprint authentication.

[2625] After successful authentication, the user says, "Good morning, it's a beautiful day today."

[2626] The device receives this, converts the audio to text, and sends it to the server.

[2627] The server generates a response saying, "Good morning. It's sunny today. Shall we go for a walk?" and sends it to the terminal.

[2628] The device converts this into audio and plays it back.

[2629] Example 2: Caregiving advice

[2630] The user says, "My lower back has been a little sore since yesterday."

[2631] The device converts the audio to text and sends it to the server.

[2632] The server references the user's health history and generates appropriate advice.

[2633] A message is sent to the device saying, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2634] The device converts the audio to sound and plays it back.

[2635] Example 3: Emotion recognition

[2636] The user says, "I'm feeling a little down today."

[2637] The device converts speech to text and then uses an emotion engine to analyze the emotional state.

[2638] Based on the analysis results, the server generates a conversation message such as, "Why don't you listen to your favorite music to change your mood?" and sends it to the terminal.

[2639] The device converts this into audio and plays it back.

[2640] In this way, this system combines generative AI and an emotion engine to achieve more personalized responses for the elderly. As a result, it becomes possible to provide efficient and effective mental and physical care for the elderly.

[2641] The following describes the processing flow.

[2642] User authentication and profile information management

[2643] User

[2644] Step 1: Launch the smartphone app.

[2645] terminal

[2646] Step 2: Display a screen for fingerprint authentication, facial recognition, or username and password entry.

[2647] User

[2648] Step 3: Perform fingerprint authentication or facial recognition, or enter your username and password.

[2649] terminal

[2650] Step 4: Send authentication information to the server.

[2651] server

[2652] Step 5: The received authentication information is compared with the database to authenticate the user.

[2653] Step 6: If authentication is successful, retrieve the user's profile information from the database and send it to the device.

[2654] terminal

[2655] Step 7: Cache the received profile information and prepare for the next step.

[2656] Everyday conversation

[2657] User

[2658] Step 1: Speak to your device and say, "Good morning, it's a beautiful day today."

[2659] terminal

[2660] Step 2: Receive voice input and record it as audio data.

[2661] Step 3: Convert the audio data to text (speech recognition process).

[2662] Step 4: Send the converted text data to the server.

[2663] server

[2664] Step 5: Input the received text data into the AI ​​model and generate appropriate conversation content (natural language processing).

[2665] Step 6: Send the generated conversation content as text data to the device.

[2666] terminal

[2667] Step 7: Convert the received text data into audio data (speech synthesis).

[2668] Step 8: Play the generated audio data to the user.

[2669] emotion recognition

[2670] User

[2671] Step 1: Speak into the device about everyday topics, your health, or any problems you're experiencing.

[2672] terminal

[2673] Step 2: Receive voice input and record it as audio data.

[2674] Step 3: Analyze the audio data with the emotion engine to determine the emotional state.

[2675] Step 4: Send the analyzed emotional state as text data to the server.

[2676] server

[2677] Step 5: Generate conversation content and advice based on the received text data and emotional state.

[2678] Step 6: Send the generated conversation content and advice to the device.

[2679] Caregiving assistance

[2680] User

[2681] Step 1: Start by saying, "My lower back has been a little sore since yesterday."

[2682] terminal

[2683] Step 2: Receive voice input and record it as audio data.

[2684] Step 3: Convert the audio data to text (speech recognition process).

[2685] Step 4: Send the converted text data to the server.

[2686] server

[2687] Step 5: Match the text data to the database to retrieve past health history.

[2688] Step 6: Generate appropriate advice based on health history and emotional state.

[2689] Step 7: Send the generated advice as text data to the device.

[2690] terminal

[2691] Step 8: Convert the received text data into audio data (speech synthesis).

[2692] Step 9: Play the generated advice as audio.

[2693] The above is a detailed explanation of the processing flow of the system that combines the emotion engine.

[2694] (Example 2)

[2695] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2696] There is a need for systems that provide mental and physical care to the elderly in a more effective and personalized way. In particular, it is necessary to accurately understand the emotional state of the elderly when they talk about their daily lives or health, and to provide appropriate responses and advice. However, current systems have shortcomings in emotional recognition, making it difficult to respond to each individual's situation.

[2697] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[2698] In this invention, the server includes means for user authentication, means for managing user profile information, means for converting voice input to text, means for transmitting the converted text data to the server, means for generating conversation content based on the received text data, means for converting the generated conversation content to voice, means for analyzing the user's voice data and determining their emotional state, means for generating conversation content and advice based on the determined emotional state, and means for playing it back as voice. This makes it possible to provide personalized conversations and care advice tailored to the emotional state of elderly people.

[2699] "User authentication" refers to the process of verifying a user's identity in order to access a system.

[2700] "Profile information" refers to data that includes personal information and specific characteristics about a user.

[2701] "Voice input" refers to audio data used to capture a user's speech into the system.

[2702] "Text conversion" refers to the process of converting audio data into text data.

[2703] A "server" refers to a computer device that processes and stores data.

[2704] "Conversation content generation" refers to the process of creating appropriate responses or messages based on input data.

[2705] "Speech conversion" refers to the process of converting text data into speech data.

[2706] "Audio playback" refers to the act of making the generated audio data playable for the user.

[2707] "Emotional state" refers to the emotional state of a user as analyzed from their speech patterns and content.

[2708] "Health status" refers to data that indicates the user's physical and mental condition.

[2709] "Advice generation" refers to the process of creating appropriate advice based on the user's situation and data.

[2710] "Health history" refers to records of a user's past health condition.

[2711] A "database" refers to a system for effectively storing, managing, and retrieving information.

[2712] An "emotion engine" refers to software that analyzes a user's voice data to determine their emotional state.

[2713] This invention is a system that combines user authentication, speech recognition, a generative AI model, and an emotion recognition engine to improve the mental and physical care of the elderly. Specifically, the user communicates by voice using a terminal, the voice is analyzed by a server, and appropriate responses and advice are generated. The specific method for implementing this system is described below.

[2714] User authentication and profile information management

[2715] Users log in to the device using fingerprint authentication, facial recognition, or by entering a username and password. This allows the system to recognize who the individual is and provide personalized services.

[2716] The device receives the authentication information entered by the user and sends it to the server. Specifically, a standard facial recognition camera can be used for facial authentication, and a fingerprint scanner can be used for fingerprint authentication. The login information is encrypted and securely transmitted to the server.

[2717] The server compares the transmitted authentication information with the user's profile information stored in the database. If authentication is successful, the server sends the user's profile information to the device. This allows the device to provide services tailored to that user.

[2718] Speech recognition and conversation content generation

[2719] Users can converse with the device using their voice. For example, they can say, "Good morning, it's a nice day today."

[2720] The device receives voice input and converts the voice data into text using speech recognition software such as the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[2721] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate appropriate conversation content based on the received text data. The generated text data is then sent to the terminal.

[2722] The device converts the generated conversation into speech using speech synthesis software such as Amazon Polly and plays it back to the user. A concrete example of such a response might be, "Good morning. It's sunny today. Shall we go for a walk?"

[2723] emotion recognition

[2724] Users can talk to the device about their physical condition and emotions. For example, they might say, "I'm feeling a little down today."

[2725] The device records the input voice data and analyzes the emotional state using emotion recognition software such as the Microsoft Azure Emotion API. The results are then sent to the server.

[2726] The server uses a generative AI model based on the analysis results to generate conversation content and produce appropriate responses that match the user's emotional state. For example, a possible response might be, "Why don't you listen to your favorite music to change your mood?"

[2727] The device converts the generated text data into speech and plays it back to the user.

[2728] Caregiving assistance

[2729] Users can talk about specific health issues they have. For example, they might say, "My lower back has been a little sore since yesterday."

[2730] The device records the audio data and converts it into text data using the Google Cloud Speech-to-Text API. The converted data is then sent to the server.

[2731] The server retrieves the user's past health history from the database based on the received text data. Then, it uses a generative AI model to generate appropriate advice, such as a message like, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2732] The device converts the generated text data into speech using speech synthesis software and plays it back to the user.

[2733] Specific example

[2734] Example of a prompt:

[2735] Everyday conversation: "Good morning, it's a beautiful day today."

[2736] Emotion recognition: "I'm feeling a little down today."

[2737] Caregiving advice: "My lower back has been a little sore since yesterday."

[2738] Thus, the present invention integrates a series of processes, including user authentication, voice recognition, emotion recognition, conversation generation, and health management, to realize more personalized responses for the elderly. This makes it possible to provide efficient and effective mental and physical care for the elderly.

[2739] The flow of the specific processing in Example 2 will be explained using Figure 13.

[2740] Program processing flow

[2741] User authentication and profile information management

[2742] Step 1:

[2743] The user authenticates using their fingerprint or facial recognition, or enters a username and password.

[2744] Input: Fingerprint data, facial image, username and password

[2745] Output: Authentication information

[2746] Specific action: The user places their finger on the fingerprint sensor of their smartphone.

[2747] Step 2:

[2748] The terminal receives the entered authentication information, encrypts it, and sends it to the server.

[2749] Input: Authentication information

[2750] Output: Encrypted authentication information

[2751] Specific operation: The device encrypts the received fingerprint data using AES (Advanced Encryption Standard).

[2752] Step 3:

[2753] The server compares the received authentication information with the registered information in the database.

[2754] Input: Encrypted credentials

[2755] Output: Authentication results and profile information

[2756] Specific operation: The server performs user authentication by comparing the user's profile information with that in the database. If authentication is successful, the server then retrieves the profile information.

[2757] Step 4:

[2758] If the server successfully authenticates, it will send the profile information to the device.

[2759] Input: Profile Information

[2760] Output: Profile information sent to the device

[2761] Specific operation: The server encrypts the profile information it has acquired and sends it to the terminal.

[2762] Everyday conversation

[2763] Step 1:

[2764] The user speaks to the device, saying, "Good morning, it's a nice day today."

[2765] Input: Audio data

[2766] Output: Audio data recorded by the device's microphone

[2767] Specific action: The user speaks into the smart speaker.

[2768] Step 2:

[2769] The device receives voice input and uses the Google Cloud Speech-to-Text API to convert the voice data into text.

[2770] Input: Audio data

[2771] Output: Text data

[2772] Specific operation: The device sends voice data to the API and retrieves the text data "It's a nice day today."

[2773] Step 3:

[2774] The terminal sends the converted text data to the server.

[2775] Input: Text data

[2776] Output: Text data sent to the server

[2777] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[2778] Step 4:

[2779] The server inputs the received text data into a generating AI model (e.g., GPT-3) to generate appropriate conversation content.

[2780] Input: Text data

[2781] Output: Generated conversation content

[2782] Specific operation: The server inputs the text "It's a nice day today" into GPT-3 and generates the response "Good morning. It's sunny today. Shall we go for a walk?".

[2783] Step 5:

[2784] The server sends the generated conversation content to the terminal.

[2785] Input: Generated conversation content

[2786] Output: Text data sent to the terminal

[2787] Specific operation: The server sends the generated text data to the terminal.

[2788] Step 6:

[2789] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[2790] Input: Text data of the generated conversation

[2791] Output: Audio data

[2792] Specific action: The device sends the text message "Good morning. It's sunny today. Shall we go for a walk?" to Amazon Polly and plays the audio data.

[2793] emotion recognition

[2794] Step 1:

[2795] The user speaks to the device about their physical condition and emotions.

[2796] Input: Audio data

[2797] Output: Audio data recorded by the device's microphone

[2798] Specific action: The user says, "I'm feeling a little down today."

[2799] Step 2:

[2800] The device records voice data and analyzes the emotional state using the Microsoft Azure Emotion API.

[2801] Input: Audio data

[2802] Output: Emotion analysis results

[2803] Specific operation: The device sends voice data to the API to obtain the emotional state (e.g., "feeling down").

[2804] Step 3:

[2805] The terminal sends the analysis results to the server.

[2806] Input: Sentiment analysis results

[2807] Output: Sentiment analysis results sent to the server

[2808] Specific operation: The device sends the emotion analysis results to the server using the HTTPS protocol.

[2809] Step 4:

[2810] The server generates conversation content using an AI model based on the emotion analysis results it receives.

[2811] Input: Sentiment analysis results

[2812] Output: Generated conversation content

[2813] Specific operation: The server inputs the emotional state of "feeling down" into GPT-3 and generates the response, "Why don't you try listening to your favorite music to cheer yourself up?"

[2814] Step 5:

[2815] The server sends the generated conversation content to the terminal.

[2816] Input: Generated conversation content

[2817] Output: Text data sent to the terminal

[2818] Specific operation: The server sends the generated text data to the terminal.

[2819] Step 6:

[2820] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[2821] Input: Text data of the generated conversation

[2822] Output: Audio data

[2823] Specific action: The device sends a text message to Amazon Polly saying, "Why not listen to some music you like to change your mood?" and plays the audio data.

[2824] Caregiving assistance

[2825] Step 1:

[2826] The user speaks to the device about their health problems.

[2827] Input: Audio data

[2828] Output: Audio data recorded by the device's microphone

[2829] Specific action: The user says, "My lower back has been a little sore since yesterday."

[2830] Step 2:

[2831] The device records the voice data and converts it into text data using the Google Cloud Speech-to-Text API.

[2832] Input: Audio data

[2833] Output: Text data

[2834] Specific operation: The device sends voice data to the API and retrieves text data saying, "My lower back has been a little sore since yesterday."

[2835] Step 3:

[2836] The terminal sends the converted text data to the server.

[2837] Input: Text data

[2838] Output: Text data sent to the server

[2839] Specific action: The terminal sends text data to the server using the HTTPS protocol.

[2840] Step 4:

[2841] The server retrieves the user's past health history from the database based on text data.

[2842] Input: Text data

[2843] Output: Health history data

[2844] Specific operation: The server queries the database using the text "My lower back has been a little sore since yesterday" to retrieve past health history.

[2845] Step 5:

[2846] The server uses the generated AI model to produce appropriate advice.

[2847] Input: Health history data

[2848] Output: Generated advice

[2849] Specific operation: The server inputs health history data into GPT-3 and generates advice such as, "Please keep your lower back warm. If necessary, use a bath towel to cool it down."

[2850] Step 6:

[2851] The server sends the generated advice to the terminal.

[2852] Input: Generated advice

[2853] Output: Text data sent to the terminal

[2854] Specific operation: The server sends the generated text data to the terminal.

[2855] Step 7:

[2856] The device receives text data, which is then converted into audio data using Amazon Polly and played back to the user.

[2857] Input: Text data of the generated advice

[2858] Output: Audio data

[2859] Specific action: The device sends a text message to Amazon Polly saying, "Please keep your lower back warm. Use a bath towel to cool it down if necessary," and plays an audio message.

[2860] In this way, by processing data based on specific input data at each step, a system is realized that provides personalized conversations and care advice to the elderly.

[2861] (Application Example 2)

[2862] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2863] In systems designed to support the safe and secure lives of the elderly, it is crucial not only to provide conversation and care assistance, but also to understand the user's emotional state in real time and provide appropriate support and advice tailored to their feelings. Furthermore, when using autonomous vehicles, it is especially important to alleviate the anxiety and tension felt by the elderly and support safe travel. Conventional technologies have difficulty adequately reflecting the user's emotions and state, making the enhancement of mental care a challenge.

[2864] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing the user's emotional state, means for adjusting the conversation content based on the analysis results, and means for providing safety information for the autonomous vehicle. As a result, the user's emotional state can be grasped in real time, and personalized conversations and support can be provided accordingly, thereby increasing the elderly's sense of security and reducing anxiety and tension during travel. Furthermore, this makes it possible to provide an environment in which the elderly can travel with peace of mind even when using an autonomous vehicle.

[2865] "Means of user authentication" refers to methods used to verify the identity of a user when they access a system.

[2866] "Means of managing user profile information" refers to a method of centrally managing users' personal information and settings.

[2867] "Methods for converting voice input to text" refer to technologies that convert a user's voice data into text data.

[2868] "Means for sending converted text data to a server" refers to a method for converting voice input into text data and then sending that data to a server.

[2869] "Methods for generating conversation content based on received text data" refers to technologies that create appropriate dialogue content based on received text data.

[2870] "Means of converting generated conversation content into speech" refers to technology that converts the created dialogue content back into speech data.

[2871] "Means of playback as audio" refers to methods for making the generated audio data listenable to the user.

[2872] "Means for analyzing a user's emotional state" refers to technologies for analyzing a user's emotions and psychological state in real time.

[2873] "Means of adjusting conversation content based on analysis results" refers to technologies that change the content of a dialogue based on analyzed emotional data.

[2874] "Means for managing a user's health status" refers to methods for collecting and managing a user's health data.

[2875] "Means for collecting data on users' physical condition and problems" refers to technologies for collecting information on users' physical condition and problems.

[2876] "Means of generating appropriate advice based on collected data" refers to methods of creating appropriate advice based on acquired data.

[2877] "Means for providing safety information on autonomous vehicles" refers to technologies that provide users with safety information related to autonomous vehicles.

[2878] "Means for storing user profile information and health status in a database" refers to methods for storing users' personal information and health information in a database.

[2879] "A means of referring to past health history based on stored data" refers to a technology that allows users to check their past health status based on data stored in a database.

[2880] "Means of providing real-time support based on the user's emotional state" refers to technologies that provide immediate support and advice based on emotional data analyzed in real time.

[2881] This invention is a system that analyzes a user's emotional state in real time and provides conversation content and safety information for autonomous vehicles accordingly. The system consists of a terminal used by the user and a server that processes data. The specific implementation method is described below.

[2882] System Configuration

[2883] The system consists of a user terminal and a server that processes data. The terminal is a device that handles voice input and output, and functions as an interface for smartphones and autonomous vehicles. The server uses an AI model to perform natural language processing and emotion recognition. The emotion engine analyzes the user's voice data and determines their emotional state, enabling more appropriate responses.

[2884] Basic operation

[2885] User authentication and profile information management

[2886] Users log in to the system using their device via voice or touch controls. The device accepts fingerprint or facial recognition, or username and password input, and sends the authentication information to the server. The server authenticates the user based on the received authentication information and verifies that it matches the registered profile information. If authentication is successful, the server sends the user's profile information to the device.

[2887] Everyday conversation

[2888] When a user speaks to the device, saying "Good morning, it's a nice day today," the device receives the voice input and records it as audio data. It then converts the audio data to text and sends it to the server. The server inputs the received text data into an AI model and generates appropriate conversation content. The generated conversation content is sent to the device as text data, which then converts it back into audio data and plays it back.

[2889] Emotion recognition and regulation

[2890] The device, upon receiving the user's everyday conversations and statements about their physical condition, analyzes the audio data using an emotion engine to determine their emotional state. It then sends this emotional state information to the server. Based on the received emotional information, the server appropriately adjusts the conversation content and support information to provide a personalized response.

[2891] Caregiving assistance and health management

[2892] For example, if a user says, "My lower back has been a little sore since yesterday," the device converts the audio to text and sends it to the server. The server refers to the user's health history and emotional state and generates appropriate advice. The generated advice is sent to the device as a message such as, "Keep your lower back warm. Use a bath towel to cool it if necessary," and the device converts this back to audio and plays it.

[2893] Applications in autonomous vehicles

[2894] If a user feels "a little uneasy" while using an autonomous vehicle, the terminal converts the voice input into text and analyzes it with an emotion engine. Based on the analysis results, the server provides real-time support such as, "It's okay, the car is driving safely. Please relax. Would you like some tea?" Also, if the user says, "I hear a strange noise," the system checks the vehicle's status and provides reassuring feedback in real time.

[2895] Hardware and software to be used

[2896] Hardware: Smartphone, microphone, speaker, autonomous vehicle interface

[2897] Software: Python, speech recognition library (speech_recognition), speech synthesis library (gTTS), natural language processing library (transformers)

[2898] Specific examples and prompt statements

[2899] Example 1: Anxiety

[2900] If the user says, "I'm a little worried," the system will generate a response such as, "We've determined your emotion to be negative, but it's okay. The car is driving safely, so please relax."

[2901] Specific example 2: Problem

[2902] If you say, "I hear a strange noise," and your emotion is negative, the response you'll get is, "We're checking the car's condition. There doesn't seem to be any problem, so please don't worry."

[2903] Example of a prompt:

[2904] "Generate a reassuring response regarding vehicle safety for elderly individuals who are expressing anxiety."

[2905] "Please generate relaxing responses to the negative emotions expressed by older adults."

[2906] By constructing a concrete system in this way, it is possible to enhance the mental care of the elderly and provide an environment in which they can use autonomous vehicles with peace of mind.

[2907] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2908] Step 1:

[2909] User Login

[2910] Input: The terminal receives the user's voice, fingerprint, facial recognition, or username and password.

[2911] Operation:

[2912] The terminal converts the received authentication information into text data and sends it to the server.

[2913] Output: Authentication information is sent to the server as text data.

[2914] Step 2:

[2915] Verification of authentication information and retrieval of profile information ...

Claims

1. Means of performing user authentication, A means of managing user profile information, A means of converting voice input to text, A means of sending the converted text data to the server, A means of generating conversation content based on received text data, A means of converting the generated conversation content into speech, A means of playing it back as audio, A system that includes this.

2. Means for managing the user's health status, A means of collecting data about the user's health and problems, A means of generating appropriate advice based on collected data, A means of converting the generated advice into speech, A means of playing it back as audio, The system according to claim 1, further comprising:

3. A means for storing user profile information and health status in a database, A means of referring to past health history based on stored data, A means of generating an appropriate response based on the referenced information, The system according to claim 1, further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A