System
An AI chat system with generative capabilities and 3D avatars enhances memory recall and daily interaction for dementia patients, offering personalized exercise and dietary support to improve their quality of life and alleviate caregiver stress.
Patent Information
- Application Number
- JP2024128524
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-16
AI Technical Summary
The increasing number of dementia patients poses a significant burden on both patients and caregivers, with limited effective measures to improve their quality of life and daily communication, and current systems fail to adequately support memory recall and personalized exercise and dietary suggestions.
An AI chat system that processes user voice inputs, searches databases for generational information, and uses generative AI models to generate conversational responses and personalized exercise and dietary suggestions, supported by 3D avatars for enhanced interaction.
The system effectively evokes past memories, supports daily communication, and provides tailored exercise and dietary recommendations, improving the quality of life for dementia patients and reducing caregiver burden.
Smart Images

Figure 2026025712000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In an aging society, the increase in the number of dementia patients is becoming a major problem. The burden on patients suffering from dementia and their caregivers is immeasurable, and current countermeasures are not sufficient to address this issue. For example, while drug therapy can slow the progression of cognitive function, effective measures to improve the quality of patients' daily lives are limited. Furthermore, the stress and time burden on caregivers cannot be ignored. Therefore, new technologies are needed to elicit past memories in dementia patients and support their daily communication. [Means for solving the problem]
[0005] The present invention provides an AI chat system for eliciting past memories of dementia patients and supporting their communication. This system includes the following means.
[0006] 1. Provide a means for processing a user's voice input and receive a request when the user talks about a past memory.
[0007] 2. Provide a means to search the database for the user's generational information and related data, and obtain materials to evoke past memories (for example, items that were popular by decade, photos, movie posters, etc.).
[0008] 3. Establish a means to generate optimal conversational responses using a generative AI model, automatically generating questions and conversations to elicit the user's past memories.
[0009] 4. Provide a means to provide the generated conversational responses to the user, presenting them to the user as audio or visuals, and connecting the conversation to trace memories.
[0010] This allows dementia patients to access past memories and engage in daily communication, thereby improving their quality of life. Furthermore, by referencing the user's physical condition records and past exercise history, the system also provides comprehensive support by suggesting appropriate exercise methods and meal plans for that day. This is expected to reduce the burden on caregivers.
[0011] "Users" are dementia patients who use this system, or their caregivers.
[0012] "Voice input" is a means for processing the words a user speaks into a terminal as digital data.
[0013] A "database" is a collection of information that stores a user's personal information and materials related to past memories (photos, movie posters, etc.) and can be searched as needed.
[0014] A "generative AI model" is an algorithm that uses machine learning technology to automatically generate appropriate responses to user input.
[0015] A "conversational response" is a response to a user's question or statement generated by a generative AI model.
[0016] A "3D avatar" is a three-dimensional digital character generated based on the characteristics of a user's family or relatives and used for everyday conversations.
[0017] "Exercise methods" are suggested physical exercise and stretching methods based on the user's physical condition and past exercise history.
[0018] "Meal contents" refers to a meal menu suggested based on the user's physical condition and past eating records.
[0019] "Age information" is information about the era in which the user was born and raised, and is used to recall past memories.
[0020] The "physical condition record" is a record of information about the user's physical condition, and is used to suggest exercise methods and dietary contents.
[0021] The above are definitions of important words included in the claims. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] This invention is an AI chat system that elicits past memories in dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and also suggests exercise methods and dietary content tailored to the user's physical condition.
[0044] System Overview
[0045] 1. Voice input processing:
[0046] User: Uses the system and speaks to initiate a conversation about a past memory.
[0047] Terminal: Converts the voice into text data and sends the text data to the server.
[0048] 2. Drawer of past memories:
[0049] Server: Based on the received text data, the server searches the database for the user's age information, for example, information related to the user's childhood or school days.
[0050] Server: Using information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, utilizing materials such as popular items and events by decade, photos, and movie posters.
[0051] Server: Sends the generated conversation response to the terminal.
[0052] Terminal: The received response is conveyed to the user via voice, enabling conversations that evoke past memories.
[0053] 3. Exercise and dietary suggestions:
[0054] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[0055] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[0056] Server: Retrieves the received physical condition data and past exercise history from a database. Based on this, the server uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[0057] Server: Sends the generated proposals to the device.
[0058] Device: The received suggestions are communicated to the user via voice.
[0059] 4. Support for conversations with family and relatives:
[0060] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[0061] Terminal: Converts the voice into text data and sends the text data to the server.
[0062] Server: Retrieves information about family and relatives from a database, and uses a generative AI model to generate optimal conversation responses. A simple 3D avatar is also generated at the same time.
[0063] Server: Sends the generated conversation responses and 3D avatars to the device.
[0064] Terminal: Assists with everyday conversations by conveying received responses to the user via voice and displaying a 3D avatar.
[0065] Specific examples
[0066] Example 1: Recalling past memories
[0067] When a user says, "I want to talk about my school days,"
[0068] The terminal sends this request to the server.
[0069] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[0070] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[0071] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[0072] The user begins recounting past events to jog their memory, and the conversation continues.
[0073] Example 2: Exercise and diet suggestions
[0074] When a user says, "Which stretch should I do today?"
[0075] The terminal sends this request to the server.
[0076] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0077] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[0078] Example 3: Supporting conversations with family and relatives
[0079] When a user says, "I want to know about my grandchildren,"
[0080] The terminal sends this request to the server.
[0081] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[0082] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[0083] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[0084] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[0085] The processing flow will be explained below.
[0086] Ability to recall past memories
[0087] Step 1:
[0088] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[0089] Step 2:
[0090] The terminal converts the user's voice into text data and transmits the text data to the server.
[0091] Step 3:
[0092] The server receives the user's request and searches the database for the user's age information and related data.
[0093] Step 4:
[0094] The server uses a generative AI model to generate a conversational response such as "What did you play in elementary school?" based on age information and related data obtained from the database.
[0095] Step 5:
[0096] The server transmits the generated conversation response to the terminal.
[0097] Step 6:
[0098] The terminal converts the response received from the server into voice and conveys it to the user.
[0099] Suggestions for exercise and diet
[0100] Step 1:
[0101] The user speaks to the device, "Which stretch should I do today?"
[0102] Step 2:
[0103] The terminal converts the user's voice into text data and transmits the text data to the server.
[0104] Step 3:
[0105] The server retrieves the user's physical condition record and past exercise history from a database.
[0106] Step 4:
[0107] Based on the acquired physical condition records and exercise history, the server uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0108] Step 5:
[0109] The server sends the generated proposal to the terminal.
[0110] Step 6:
[0111] The terminal converts the proposal received from the server into voice and conveys it to the user.
[0112] Support for conversations with family and relatives
[0113] Step 1:
[0114] The user speaks to the terminal saying, "I want to know about my grandchildren."
[0115] Step 2:
[0116] The terminal converts the user's voice into text data and transmits the text data to the server.
[0117] Step 3:
[0118] The server retrieves information about the user's family and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[0119] Step 4:
[0120] Based on the acquired information, the server uses a generative AI model to generate a conversational response such as, "Do you know what your grandchild is currently learning at school?" It also generates a simple 3D avatar.
[0121] Step 5:
[0122] The server sends the generated conversational responses and 3D avatars to the device.
[0123] Step 6:
[0124] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[0125] Proposing exercise methods and meal plans tailored to the user's physical condition
[0126] Step 1:
[0127] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[0128] Step 2:
[0129] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0130] Step 3:
[0131] The server updates the database with the newly received physical condition data.
[0132] Step 4:
[0133] The server uses a generative AI model based on the received health data and previously stored data to generate suggestions such as, "We recommend that you eat something that is easy to digest today."
[0134] Step 5:
[0135] The server sends the generated proposal to the terminal.
[0136] Step 6:
[0137] The terminal converts the proposal received from the server into voice and conveys it to the user.
[0138] The above is a specific description of the processing flow.
[0139] Example 1
[0140] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0141] Today, there is a lack of appropriate systems to help dementia patients recall their past memories and support their daily communication. Furthermore, there are no suggestions for exercise or dietary content tailored to their physical condition, which results in a decline in their quality of life. Furthermore, communication with family and relatives is often difficult, leading to feelings of isolation.
[0142] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0143] In this invention, the server includes means for receiving a user's voice input and processing the voice data, means for converting the voice data into text data, means for searching a database for the user's age information and related data, means for using a generative AI model to set prompts for generating optimal conversational responses and generating response messages, means for providing the generated response messages to the user, means for referencing the user's physical condition records and past exercise history and using a generative AI model to suggest optimal exercise methods and dietary content based on the user's physical condition records and past exercise history, and means for providing the generated suggestions to the user, means for acquiring information about the user's family and relatives from a database and using a generative AI model to generate appropriate response messages based on the acquired information, and means for generating a simple 3D avatar and providing the generated 3D avatar and response message to the user. This effectively draws out the past memories of dementia patients to support conversations, enables suggestions for exercise methods and dietary content tailored to their physical condition, and enables smooth communication with family and relatives.
[0144] A "user" is an entity that uses the system to provide voice input and receive conversations and suggestions.
[0145] "Voice input" refers to voice data input by a user speaking to the system.
[0146] "Voice data" refers to a digital recording of a user's voice input.
[0147] "Text data" refers to data obtained by converting voice data into characters.
[0148] "Age information" refers to data related to a user's year of birth or a particular age group.
[0149] A "database" is a system for storing and managing related data such as age and family information.
[0150] A "generative AI model" is an artificial intelligence algorithm that generates optimal responses or suggestions based on input data.
[0151] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[0152] A "response message" is a conversational response generated by a generative AI model.
[0153] "Physical condition record" refers to a record that accumulates data related to the user's physical condition.
[0154] "Exercise history" refers to data relating to exercises that a user has performed in the past.
[0155] An "exercise method" is an exercise technique suggested based on the user's physical condition and exercise history.
[0156] "Meal content" refers to an appropriate meal plan suggested to the user.
[0157] "Family information" refers to data about the user's family and relatives.
[0158] A "3D avatar" is a three-dimensional virtual character generated on a computer.
[0159] This invention relates to an AI chat system that supports daily communication for dementia patients and improves their quality of life. This system provides comprehensive support by evoking the user's past memories, engaging in conversations based on those memories, and suggesting exercise methods and dietary content tailored to the user's physical condition.
[0160] Hardware and software used
[0161] Audio input device (microphone): Receives the user's voice and records the audio data.
[0162] Speech recognition software (e.g., Google Speech-to-Text API): converts voice data into text data.
[0163] Database system: Stores data such as the user's age, family information, health records, and exercise history.
[0164] Generative AI models (e.g., OpenAI's GPT-3): Set prompts based on received text data and generate appropriate conversational responses and suggestions.
[0165] Speech synthesis software (e.g., Amazon Polly): Converts the generated response message into audio data.
[0166] 3D avatar generation software: Generates a 3D avatar based on the user's family information.
[0167] Specific explanation of the system
[0168] 1. Voice input processing
[0169] The user speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[0170] The device receives the user's voice using a high-performance microphone and records it as voice data.
[0171] The terminal converts the recorded voice data into text data using voice recognition software and transmits the text data to the server.
[0172] 2. Database search
[0173] The server analyzes the received text data and extracts keywords such as "school days."
[0174] The server searches the database for the user's generational information, and if the user was a student in the 1960s, for example, it retrieves data on trends and events from that era.
[0175] 3. Response generation using generative AI models
[0176] The server uses a generative AI model to set prompts based on information retrieved from the database, such as "You were a student in the 1960s. Do you have any memorable experiences?"
[0177] The server sends the generated response message to the user's terminal.
[0178] 4. Providing a response message
[0179] The terminal converts the received response message into voice data using voice synthesis software.
[0180] The terminal reproduces the audio data from a speaker and conveys it to the user.
[0181] Specific examples
[0182] Example 1: Recalling past memories
[0183] When a user says, "I want to talk about my school days,"
[0184] The device records this audio and converts it into text using speech recognition software.
[0185] The server receives this text data and extracts keywords such as "school days."
[0186] The server retrieves the user's generational information from a database and researches trends and events from, for example, the 1960s.
[0187] The server uses a generative AI model to set a prompt sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?" and generates the optimal response message.
[0188] The server sends this response message to the terminal,
[0189] The device synthesizes the received message into voice and plays it back through the speaker, telling the user, "You were a student in the 1960s. Do you have any memorable experiences?"
[0190] Example 2: Exercise and diet suggestions
[0191] When a user says, "Which stretch should I do today?"
[0192] The device records this audio and converts it into text using speech recognition software.
[0193] The server receives this text data and extracts keywords such as "stretch."
[0194] The server retrieves the user's physical condition records and past exercise history from a database and uses a generative AI model to suggest the optimal exercise method.
[0195] The server sets a prompt sentence such as "Today, we recommend neck rotation exercises as a light stretch," and generates a suggestion.
[0196] The server sends this proposal to the terminal,
[0197] The device synthesizes the received suggestions into voice and plays them back through the speaker, telling the user, "Today, I recommend doing some neck rotation exercises as a light stretch."
[0198] Example 3: Supporting conversations with family and relatives
[0199] When a user says, "I want to know about my grandchildren,"
[0200] The device records this audio and converts it into text using speech recognition software.
[0201] The server receives this text data and extracts keywords such as "grandchild."
[0202] The server retrieves information about grandchildren from a database and uses a generative AI model to generate the question, "Do you know what your grandchildren are learning in school right now?"
[0203] The server also simultaneously generates a simple 3D avatar and sends it to the user's device.
[0204] The device synthesizes the received question into speech, plays it back through the speaker, and displays a 3D avatar, asking the user, "Do you know what your grandchild is learning at school right now?"
[0205] In this way, the present invention improves the quality of life of dementia patients by effectively drawing out the user's past memories and making suggestions tailored to everyday conversations and physical condition.
[0206] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0207] Step 1: Receiving voice input
[0208] User: Speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[0209] Terminal: The user's voice is received by a high-performance microphone and recorded as audio data, which becomes the input data.
[0210] Terminal: The recorded audio data is sent to the server. The output at this stage is audio data.
[0211] Step 2: Converting audio data to text data
[0212] Server: Converts the received voice data into text data using speech recognition software (e.g., Google Speech-to-Text API). This is the input data.
[0213] Server: The converted text data is used in the next processing step. The output is text data.
[0214] Step 3: Database search
[0215] Server: Analyzes the received text data and extracts keywords such as "school days." This is the input data.
[0216] Server: Searches the database for the user's generation information and retrieves data on trends and events from the corresponding era. This data processing and calculation is used to reference and search for generation information. The output is data from the corresponding era.
[0217] Step 4: Generative AI model generates a response
[0218] Server: Based on the information obtained from the database, a generative AI model (e.g., OpenAI's GPT-3) is used to set a prompt. For example, the prompt may be a text sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?"
[0219] Server: Using this prompt as input, the generative AI model generates the optimal conversational response. This response generation is performed using data calculations.
[0220] Server: Output the generated response message as text data for use in the next step.
[0221] Step 5: Send a response message
[0222] Server: Sends the generated response message to the user's terminal. This is the input data.
[0223] Terminal: Receives the response message and converts it into voice data using speech synthesis software (e.g., Amazon Polly). Voice synthesis is a type of data processing. The output is voice data.
[0224] Step 6: Play greetings
[0225] Terminal: The audio data is played back through the speaker and transmitted to the user. This is the final output.
[0226] Example flow
[0227] User: Say, "Which stretch should I do today?"
[0228] Device: Uses voice recognition software to convert the question "Which stretch should I do today?" into text data and send it to the server.
[0229] Server: Analyzes the text data and extracts keywords.
[0230] Server: Retrieves the user's physical condition records and past exercise history from the database.
[0231] Server: Using a generative AI model, it generates a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0232] Server: Sends the proposal to the device.
[0233] Device: The suggested message is generated through speech synthesis and played from the speaker. It tells the user, "Today, we recommend you do some light stretching exercises, such as neck rotations."
[0234] In this way, the processing steps from the user's voice input to the final playback of the response message are concretely carried out.
[0235] (Application example 1)
[0236] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0237] This invention relates to an AI chat system that evokes past memories in dementia patients and improves their quality of life. Currently, many nursing homes lack support for retrieving individual memories for dementia patients, resulting in insufficient mental stability and communication. It is also difficult to provide dementia patients with appropriate exercise and dietary recommendations tailored to their physical condition, and communication with their families is also difficult. To address these issues, a system is needed that effectively engages dementia patients in conversation and suggests health management strategies based on their past living conditions.
[0238] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0239] In this invention, the server includes a means for processing user voice input, a means for searching a database for user age information and related data, a means for generating optimal conversational responses using a generative AI model, a means for providing the generated conversational responses to the user by voice and eliciting past memories, a means for using the generative AI model to hold conversations based on the past living situations of dementia patients and support daily communication, and an application installed on a robot installed in a care facility, which makes it possible to increase the mental stability of dementia patients and provide effective communication support.
[0240] "Means for processing user voice input" refers to a function for receiving voice data uttered by a user, recognizing it, and converting it into text data.
[0241] "Means of searching for a user's age information and related data from a database" refers to the function of searching for and retrieving the necessary information from a database that holds information related to the user's age, interests, and past life.
[0242] "Means for generating optimal conversational responses using a generative AI model" refers to a function that uses AI technology based on text data to generate optimal responses for users.
[0243] "Means for providing the generated conversational response to the user by voice and eliciting past memories" refers to a function for providing the user with an automatically generated text response by voice output, thereby eliciting the user's past memories.
[0244] "A means of using a generative AI model to hold conversations based on the past living conditions of dementia patients and support daily communication" refers to a function that makes full use of AI technology to conduct conversations based on content related to the past experiences and lives of dementia patients, thereby promoting daily interaction.
[0245] "Applications installed on robots installed in nursing facilities" refers to application software that is installed on robot devices used in nursing facilities and that is used to interact with dementia patients.
[0246] "A means for referring to the user's health records and past exercise history, and proposing optimal exercise methods and dietary content based on the generated health data" refers to a function that refers to the user's previous health information and exercise data, and provides appropriate exercise and dietary advice based on the user's current health condition.
[0247] "Means for providing the generated suggestions to the user by voice" refers to a function for conveying the generated exercise and diet suggestions to the user by voice.
[0248] "Means for a robot installed in a care facility to demonstrate suggested content to a user" refers to a function in which a robot demonstrates appropriate exercise methods and other instructions to a user.
[0249] "Means of obtaining information about the user's family and relatives from a database and generating a simple 3D avatar based on that information" refers to a function that obtains data about the user's family and relatives from a database and generates an avatar based on that data.
[0250] "Means of using the generated 3D avatar to engage in everyday conversation with the user and provide supplementary animations" refers to the function of using the generated 3D avatar as an interface to assist in conversation with the user and display animations as a visual aid.
[0251] "Means for providing information about family members and relatives to the user by voice" refers to a function that conveys information about family members and relatives obtained from a database to the user by voice.
[0252] This invention is an AI chat system that elicits past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary recommendations tailored to the user's physical condition. This system is implemented as an application installed on a robot in a nursing home.
[0253] System Overview
[0254] 1. Voice input processing
[0255] The user uses the system to speak to initiate a conversation about a past memory.
[0256] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[0257] 2. Drawer of past memories
[0258] The server searches a database for information about the user's age based on the received text data, such as information about the user's childhood or school days.
[0259] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[0260] The server transmits the generated conversation response to the terminal.
[0261] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[0262] 3. Exercise and dietary suggestions
[0263] The user reports their physical condition to the device, for example, by saying, "I feel a little tired today."
[0264] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0265] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[0266] The server sends the generated proposal to the terminal.
[0267] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[0268] 4. Support for conversations with family and relatives
[0269] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[0270] The terminal converts the voice into text data and transmits the text data to the server.
[0271] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[0272] The server sends the generated conversational responses and 3D avatars to the terminal.
[0273] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[0274] Specific examples
[0275] Example 1: Recalling past memories
[0276] The user says, "I want to talk about my old school days."
[0277] The terminal sends this request to the server.
[0278] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[0279] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[0280] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[0281] The user begins recounting past events to jog their memory, and the conversation continues.
[0282] Example prompt sentence:
[0283] User: I want to talk about my old school days.
[0284] AI: What did you play in elementary school?
[0285] Example 2: Exercise and diet suggestions
[0286] The user says, "Which stretch should I do today?"
[0287] The terminal sends this request to the server.
[0288] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0289] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[0290] If necessary, the robot will demonstrate neck rotation.
[0291] Example prompt sentence:
[0292] User: Which stretches should I do today?
[0293] AI: Today, I recommend some gentle stretching exercises, such as neck rotations.
[0294] Example 3: Supporting conversations with family and relatives
[0295] The user says, "I want to know about my grandchildren."
[0296] The terminal sends this request to the server.
[0297] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[0298] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[0299] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[0300] Example prompt sentence:
[0301] User: I want to know about my grandchildren
[0302] AI: Do you know what your grandchildren are learning in school right now?
[0303] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[0304] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0305] Step 1:
[0306] The user uses the system to speak to initiate a conversation about a past memory.
[0307] Input: User's voice data
[0308] How it works: A microphone on a robot installed in a care home captures voice input.
[0309] Step 2:
[0310] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[0311] Input: User's voice data
[0312] What it does: It uses speech recognition software (e.g., the speech_recognition library) to convert speech to text, which is then sent over the Internet to a server.
[0313] Output: Text data sent to the server
[0314] Step 3:
[0315] The server searches the database for the user's age information based on the received text data.
[0316] Input: User's text data
[0317] What it does: The server performs a database query to retrieve data about the user's age and related past memories. For example, if the user types "I want to talk about my old school days," it searches for information related to the user's childhood and school days.
[0318] Output: User age and related data
[0319] Step 4:
[0320] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[0321] Input: User demographics and related data
[0322] Specific operation: Query a generative AI model (e.g., OpenAI GPT-3) with a prompt sentence to generate an appropriate conversational response. Reference data such as trends and events relevant to each generation are also used as reference.
[0323] Output: Generated conversation response
[0324] Step 5:
[0325] The server transmits the generated conversation response to the terminal.
[0326] Input: Generated conversation response
[0327] Specific operation: The response text generated by the server is sent to the terminal via the Internet.
[0328] Output: Conversation response sent to the terminal
[0329] Step 6:
[0330] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[0331] Input: Conversation response sent by the server
[0332] What it does: It uses speech synthesis software (e.g., the pyttsx3 library) to play the text data as speech, allowing the user to continue the conversation by listening to the voice response.
[0333] Output: The audio response provided to the user
[0334] Step 7:
[0335] The user reports their physical condition to the device, for example, saying, "I feel a little tired today."
[0336] Input: User's voice data
[0337] What happens: The device's microphone captures audio input.
[0338] Step 8:
[0339] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0340] Input: Audio data about the user's physical condition
[0341] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[0342] Output: Text data about your health condition sent to the server
[0343] Step 9:
[0344] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[0345] Input: Text data about the user's physical condition and past exercise history
[0346] Specific operation: The server retrieves the user's past exercise history from the database and uses a generative AI model to suggest appropriate exercise methods and dietary recommendations.
[0347] Output: Generated exercise and diet suggestions
[0348] Step 10:
[0349] The server sends the generated proposal to the terminal.
[0350] Input: Generated proposal
[0351] Specific operation: The proposal content is sent from the server to the device via the Internet.
[0352] Output: Suggestions sent to the device
[0353] Step 11:
[0354] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[0355] Input: Suggestion from the server
[0356] Specific behavior: Using speech synthesis software, the robot plays back the suggestions as voice, and shows the user the appropriate exercise method.
[0357] Output: Audio suggestions and demonstrations provided to the user
[0358] Step 12:
[0359] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[0360] Input: User's voice data
[0361] What happens: The device's microphone captures audio input.
[0362] Step 13:
[0363] The terminal converts the voice into text data and transmits the text data to the server.
[0364] Input: User's voice data
[0365] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[0366] Output: Text data sent to the server
[0367] Step 14:
[0368] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[0369] Input: User's text data
[0370] How it works: The server performs database queries to retrieve information about family and relatives, then uses a generative AI model to generate optimal conversational responses, while simultaneously generating a simple 3D avatar.
[0371] Output: Generated conversational responses and 3D avatars
[0372] Step 15:
[0373] The server sends the generated conversational responses and 3D avatars to the terminal.
[0374] Input: Generated conversational responses and 3D avatars
[0375] Specific operation: The server generates a response text and sends a 3D avatar to the device via the Internet.
[0376] Output: Speech responses and 3D avatar sent to device
[0377] Step 16:
[0378] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[0379] Input: Conversational responses and 3D avatars sent from the server
[0380] Specific operations: Uses speech synthesis software to play text data as speech, and displays a 3D avatar on a display device to assist in conversation with the user.
[0381] Output: Audio response and 3D avatar representation provided to the user
[0382] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0383] This invention is an AI chat system that draws out past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to the user's physical condition. In addition, by combining it with an emotion engine that recognizes the user's emotions, more personalized responses and suggestions are possible.
[0384] System Overview
[0385] 1. Voice input processing:
[0386] User: Uses the system to initiate a conversation about a past memory.
[0387] Terminal: Converts the voice into text data and sends the text data to the server.
[0388] 2. Emotion Recognition with Emotion Engine:
[0389] Server: Analyzes the received text data and recognizes the user's emotions using an emotion engine.
[0390] The emotion engine analyzes emotions from the user's voice input and text data to determine emotional states such as joy, sadness, and anger.
[0391] 3. Drawer of past memories:
[0392] Server: Retrieves the user's age information and related data from a database, for example, retrieves information related to the user's childhood or school days.
[0393] Server: Based on information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, adjusting the content of the conversation based on emotions recognized by the emotion engine.
[0394] Server: Sends the generated conversation response to the terminal.
[0395] Terminal: Converts the received response into speech and conveys it to the user.
[0396] 4. Exercise and dietary suggestions:
[0397] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[0398] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[0399] Server: Retrieves the user's physical condition records and past exercise history from a database, and uses a generative AI model to suggest optimal exercise methods and dietary recommendations. The server also adjusts the suggestions by taking into account emotions recognized by the emotion engine.
[0400] Server: Sends the generated proposals to the device.
[0401] Device: Converts the received suggestions into speech and conveys them to the user.
[0402] 5. Support for conversations with family and relatives:
[0403] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[0404] Terminal: Converts the voice into text data and sends the text data to the server.
[0405] Server: Retrieves information about family and relatives from a database and uses a generative AI model to generate optimal conversational responses. A simple 3D avatar is also generated at the same time. The conversational responses are adjusted based on emotions recognized by the emotion engine.
[0406] Server: Sends the generated conversation responses and 3D avatars to the device.
[0407] Terminal: Converts the received response into voice and displays a 3D avatar to communicate with the user.
[0408] Specific examples
[0409] Example 1: Recalling past memories
[0410] When a user says, "I want to talk about my school days,"
[0411] The terminal sends this request to the server.
[0412] The server retrieves the user's age and related data from the database. For example, if the user is having a good time, it generates questions about the pastimes and movies they enjoyed at the time.
[0413] The server uses a generative AI model based on information retrieved from the database to tailor the conversational response based on the emotions recognized by the emotion engine, generating the question, "What did you play at elementary school?"
[0414] The server sends the generated question to the terminal, which then transmits it to the user as voice.
[0415] The user begins to recount past events to jog their memory, and the conversation continues.
[0416] Example 2: Exercise and diet suggestions
[0417] When a user says, "Which stretch should I do today?"
[0418] The terminal sends this request to the server.
[0419] The server references the user's physical condition records and past exercise history, and uses a generative AI model to determine the optimal exercise method and diet. The suggestions are adjusted based on the emotions recognized by the emotion engine. For example, if the user is tired, the server might suggest, "Today, we recommend neck rotation exercises as a light stretch."
[0420] The server transmits the generated proposal to the terminal, which then conveys it to the user as voice.
[0421] Example 3: Supporting conversations with family and relatives
[0422] When a user says, "I want to know about my grandchildren,"
[0423] The terminal sends this request to the server.
[0424] The server uses a generative AI model to generate conversational responses based on information about the grandchild (such as name, school, and interests) retrieved from a database. The conversation content is adjusted based on emotions recognized by an emotion engine. A simple 3D avatar is also generated.
[0425] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[0426] In this way, the present invention provides a concrete means to improve the quality of life of dementia patients and reduce the burden on caregivers. By combining it with an emotion engine, personalized responses and suggestions that take into account the user's emotional state become possible, realizing more personal and effective support.
[0427] The processing flow will be explained below.
[0428] Ability to recall past memories
[0429] Step 1:
[0430] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[0431] Step 2:
[0432] The terminal converts the user's voice into text data and transmits the text data to the server.
[0433] Step 3:
[0434] The server analyzes the text data and uses an emotion engine to recognize the user's emotions, such as whether they are happy or sad.
[0435] Step 4:
[0436] The server retrieves the user's age information and related data from a database, for example, information related to the user's childhood or school days.
[0437] Step 5:
[0438] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database. The content of the conversation is adjusted based on emotions recognized by the emotion engine. For example, if the user seems to be having fun, the server generates a question such as, "Did you have any happy memories from that time?"
[0439] Step 6:
[0440] The server transmits the generated conversation response to the terminal.
[0441] Step 7:
[0442] The terminal converts the response received from the server into voice and conveys it to the user.
[0443] Suggestions for exercise and diet
[0444] Step 1:
[0445] The user speaks to the device, "Which stretch should I do today?"
[0446] Step 2:
[0447] The terminal converts the user's voice into text data and transmits the text data to the server.
[0448] Step 3:
[0449] The server retrieves the user's physical condition record and past exercise history from a database.
[0450] Step 4:
[0451] The server uses a generative AI model to suggest optimal exercise methods and dietary recommendations based on the user's physical condition data and past exercise history. The server also adjusts the recommendations by taking into account the user's emotional state, as recognized by the emotion engine. For example, if the user expresses feelings of fatigue, the server might suggest, "We recommend getting plenty of rest and doing some light stretching today."
[0452] Step 5:
[0453] The server sends the generated proposal to the terminal.
[0454] Step 6:
[0455] The terminal converts the proposal received from the server into voice and conveys it to the user.
[0456] Support for conversations with family and relatives
[0457] Step 1:
[0458] The user speaks to the terminal saying, "I want to know about my grandchildren."
[0459] Step 2:
[0460] The terminal converts the user's voice into text data and transmits the text data to the server.
[0461] Step 3:
[0462] The server analyzes the text data and uses an emotion engine to recognize the user's emotions, for example, capturing the joy and excitement when the user talks about their grandchildren.
[0463] Step 4:
[0464] The server retrieves information about family members and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[0465] Step 5:
[0466] The server uses a generative AI model based on the acquired information to generate optimal conversational responses based on the emotions recognized by the emotion engine. For example, if the user is excited about their grandchild, the server generates a question such as, "What is your grandchild's favorite pastime these days?"
[0467] Step 6:
[0468] The server also simultaneously generates a simple 3D avatar and sends a conversation response including this avatar to the terminal.
[0469] Step 7:
[0470] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[0471] Proposing exercise methods and meal plans tailored to the user's physical condition
[0472] Step 1:
[0473] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[0474] Step 2:
[0475] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0476] Step 3:
[0477] The server updates the database with the newly received physical condition data.
[0478] Step 4:
[0479] The server uses a generative AI model to suggest optimal exercise methods and meal plans based on the received physical condition data and previously stored data. The suggestions are adjusted taking into account the emotions recognized by the emotion engine. For example, if the user expresses feelings of fatigue, the server generates a suggestion such as, "Today's meal is recommended to help you relax."
[0480] Step 5:
[0481] The server sends the generated proposal to the terminal.
[0482] Step 6:
[0483] The terminal converts the proposal received from the server into voice and conveys it to the user.
[0484] The above is a description of the process flow broken down into specific steps.
[0485] Example 2
[0486] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0487] In modern society, the number of dementia patients is increasing, and methods to improve their quality of life are needed. In particular, it is necessary to enrich the lives of dementia patients and reduce the burden on caregivers by evoking past memories and providing personalized responses based on emotions. However, conventional technologies lack the accuracy of voice input and the provision of personalized responses, making it difficult to provide appropriate support tailored to the specific needs of dementia patients.
[0488] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for processing a user's voice input, means for converting the user's voice input into text, means for transmitting the text data to the server, means for receiving the text data at the server and recognizing emotions with an emotion engine, means for searching a database for user age information and related data, means for generating optimal conversational responses based on the results of the emotion engine using a generative AI model, means for transmitting the generated conversational responses from the server to the terminal, means for converting the conversational responses received by the terminal into speech, and means for providing the generated speech to the user. This enables personalized responses and suggestions based on the emotions and past memories of dementia patients, improving their quality of life and reducing the burden on caregivers.
[0489] The "means for processing user speech input" is a means for capturing speech from the user and sending it to the next processing step.
[0490] "Means for converting user voice input into text" refers to a speech recognition system for converting voice data into text data.
[0491] "Means for transmitting text data to a server" refers to a communication means for transmitting the converted text data to a server via a network.
[0492] "Means for recognizing emotions using an emotion engine" refers to an emotion analysis system for analyzing and recognizing a user's emotional state from text data.
[0493] "Means for searching a database for user generation information and related data" refers to a search function for obtaining a user generation information and related data from a database.
[0494] "Means for generating optimal conversational responses based on the results of an emotion engine using a generative AI model" refers to a system for generating optimal responses based on the results of emotion analysis using an AI model.
[0495] The "means for transmitting the generated conversation response from the server to the terminal" refers to a communication means for transmitting the generated conversation response from the server to the terminal.
[0496] "Means for converting the conversational response received by the terminal into speech" refers to a speech synthesis system that converts the received text-based conversational response into speech.
[0497] The "means for providing the generated audio to the user" refers to an output device such as a speaker or a headphone that allows the user to hear the generated audio.
[0498] This invention is an AI system that elicits past memories in dementia patients and improves their quality of life. The system aims to evoke the user's past memories, support daily communication, and suggest exercise methods and dietary content tailored to the user's physical condition. In addition, by combining it with an emotion engine that recognizes the user's emotions, more personalized responses and suggestions are possible.
[0499] System configuration
[0500] The system uses the following hardware and software:
[0501] A microphone to capture the user's voice input
[0502] Speech recognition software (e.g., Google Cloud Speech-to-Text API)
[0503] Sentiment analysis software (e.g., IBM Watson Tone Analyzer)
[0504] A database (e.g., MySQL or MongoDB)
[0505] Generative AI models (e.g., OpenAI GPT-3)
[0506] A speech synthesis engine (e.g., Amazon Polly)
[0507] Speakers or headphones for audio output to the user
[0508] System Operation
[0509] Voice Input Processing
[0510] When a user speaks to the system, the device's microphone captures the speech, and speech recognition software converts the speech data into text data, which is then sent to the server.
[0511] Emotion recognition by emotion engine
[0512] The server inputs the received text data into emotion analysis software to analyze the user's emotions. The emotion analysis software analyzes keywords and context within the text to recognize emotions such as joy, sadness, and anger.
[0513] Drawer of past memories
[0514] The server searches a database for the user's age information and related data. For example, it obtains information about the user's childhood and school days. Based on the obtained information and the results of emotion analysis, the server inputs a prompt sentence into the generative AI model to generate an optimal conversational response. An example of a prompt sentence is, "The user wants to talk about their school days. They appear quietly pleased." The generated conversational response is sent to the device, which then converts it into speech using a speech synthesis engine and conveys it to the user.
[0515] Suggestions for exercise and diet
[0516] When a user reports their physical condition, the device converts it into text and sends it to a server. The server retrieves the user's physical condition records and past exercise history from a database, and, taking into account the results of emotion analysis, uses a generative AI model to suggest optimal exercise methods and dietary content. For example, a suggestion might be generated such as, "Today, we recommend neck rotation exercises as a light stretch." The generated suggestion is sent to the device, which converts it into voice and conveys it to the user.
[0517] Support for conversations with family and relatives
[0518] When a user wants to know about their family or relatives, they speak their request into the device. The device converts the voice input into text and sends it to the server. The server retrieves information about family and relatives from a database and generates a conversational response using a generative AI model based on that information. A simple 3D avatar is also generated. For example, in response to a request such as "I want to know about my grandchildren," a response such as "Your grandchild is currently 8 years old and loves soccer" is generated. The generated conversational response and 3D avatar are sent to the device, which converts it into speech and displays the 3D avatar to communicate with the user.
[0519] Specific examples
[0520] Below are some specific examples of how the system can be used.
[0521] Example 1: Recalling past memories
[0522] When a user says, "I want to talk about my school days," the device sends this request to the server. The server retrieves the user's age information and related data from a database, and generates a question based on the results of emotion analysis: "What did you play in elementary school?" The device converts this question into speech and conveys it to the user, who then begins talking about past events to jog their memories.
[0523] Example 2: Exercise and diet suggestions
[0524] When the user says, "Which stretch should I do today?", the device sends this request to the server. The server retrieves physical condition records and past exercise history from a database, and, taking into account the results of emotion analysis, suggests, "Today, I recommend neck rotations as a light stretch." The device converts this into voice and conveys it to the user.
[0525] Example 3: Supporting conversations with family and relatives
[0526] When a user says, "I want to know about my grandchild," the device sends this request to the server. The server retrieves information about the family from a database and generates a 3D avatar along with the response, "Your grandchild is now 8 years old and loves soccer." The device then provides this as voice and displays the 3D avatar.
[0527] As described above, the system of the present invention can provide personalized responses and suggestions based on the user's emotional state and past memories to improve the quality of life of dementia patients.
[0528] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0529] Step 1:
[0530] A user speaks aloud. For example, the user says, "I'd like to talk about my old school days." The input is the user's voice data, and the output is the voice data itself.
[0531] Step 2:
[0532] The device converts speech to text. The device's microphone captures the user's speech data and uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the speech to text data. The input is the user's speech data, and the output is text data. Specifically, the speech recognition software performs phonemic analysis and temporarily stores the recognition results as text on the device.
[0533] Step 3:
[0534] The terminal sends text data to the server. The converted text data is sent to the server using the HTTPS protocol. The input is text data, and the output is the text data sent to the server.
[0535] Step 4:
[0536] The server receives text data. The server receives text data sent from the terminal. The input is the text data sent from the terminal, and the output is the text data saved on the server.
[0537] Step 5:
[0538] The server recognizes emotions using an emotion engine. The server inputs the received text data into emotion analysis software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data, and the output is data indicating the user's emotional state (e.g., joy, sadness, anger, etc.). Specifically, the emotion analysis engine analyzes keywords and context within the text and recognizes the emotional state numerically and categorically.
[0539] Step 6:
[0540] The server searches for the user's past data. The server retrieves the user's age information and related data from a database (e.g., MySQL or MongoDB). The input is the user's profile information, and the output is the retrieved past data. Specifically, the server issues an SQL query to retrieve the required information.
[0541] Step 7:
[0542] The server generates a response using a generative AI model. Based on the acquired past data and the results of emotion analysis, the server inputs a prompt into the generative AI model (e.g., OpenAI GPT-3) to generate an optimal conversational response. The input is past data and emotional state data, and the output is the generated conversational response. An example of a specific prompt is, "The user wants to talk about their school days. They appear quietly pleased."
[0543] Step 8:
[0544] The server sends the generated response to the terminal. The generated conversation response is sent to the terminal using the HTTPS protocol. The input is the generated conversation response, and the output is the conversation response sent to the terminal.
[0545] Step 9:
[0546] The device converts the response into speech and transmits it. The device converts the received conversational response into speech using a speech synthesis engine (e.g., Amazon Polly) and transmits it to the user. The input is a text-format conversational response, and the output is a speech-format conversational response. Specifically, the audio is played back to the user through the device's speakers or headphones.
[0547] Step 10:
[0548] A user reports how they are feeling, for example, "I feel a little tired today." The input is the user's voice data, and the output is the voice data itself.
[0549] Step 11:
[0550] The device converts the voice into text and sends it to the server. As in step 2, it uses voice recognition software and sends the converted text data to the server. The input is the user's voice data, and the output is the text data and the text data sent to the server.
[0551] Step 12:
[0552] The server proposes exercise methods and dietary recommendations. The server retrieves health records and past exercise history from a database, and generates optimal recommendations using a generative AI model that takes into account the results of emotion analysis. The inputs are health data, past exercise history, and emotional state data, and the output is the generated recommendations.
[0553] Step 13:
[0554] The device converts the proposal content into voice and conveys it. The received proposal content is converted into voice using a speech synthesis engine and conveyed to the user. The input is the proposal content in text format, and the output is the proposal content in voice format.
[0555] Step 14:
[0556] The user asks a question about their family, for example, "I want to know about my grandchildren." The input is the user's voice data, and the output is the voice data itself.
[0557] Step 15:
[0558] The device converts the voice into text and sends it to the server. As in step 2, it uses voice recognition software and sends the converted text data to the server. The input is the user's voice data, and the output is the text data and the text data sent to the server.
[0559] Step 16:
[0560] The server retrieves family information and generates conversational responses. The server retrieves information about family members and relatives from a database, generates conversational responses using a generative AI model, and also generates a simple 3D avatar. The input is family information and emotional state data, and the output is the generated conversational responses and 3D avatar.
[0561] Step 17:
[0562] The device converts the response into speech and displays a 3D avatar. The received conversational response is converted into speech using a speech synthesis engine and then displays a 3D avatar. The input is text-format conversational response and 3D avatar data, and the output is audio-format conversational response and display of a 3D avatar. The response is communicated to the user through the device's speaker and display.
[0563] (Application example 2)
[0564] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0565] Currently, there are systems that can retrieve past memories to improve the quality of life of dementia patients, but they lack personalized responses and suggestions that take into account the user's emotional state. Furthermore, there are no established methods for providing appropriate product recommendations or virtual shopping experiences. It is necessary to solve these problems and provide a more personalized experience.
[0566] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for processing a user's voice input, means for searching a database for the user's age information and related data, means for generating optimal conversational responses using a generative AI model, means for providing the generated conversational responses to the user, means for recognizing the user's emotions using an emotion engine, means for suggesting personalized products based on past purchase history and memories, means for displaying product information in real time using a virtual reality display, and means for conveying the generated suggestions to the user audio and visually. This enables personalized suggestions and responses that take the user's emotional state into consideration, making it possible to provide a more personalized and high-quality virtual shopping experience.
[0567] The "means for processing voice input" is a means for acquiring voice information from a user and converting the voice information into text data.
[0568] "Age information" is information about the user's time series, such as the user's date of birth and age, that is necessary to retrieve past memories.
[0569] A "means for searching from a database" is a means for finding and obtaining necessary data from a database that stores specific information.
[0570] A "generative AI model" is a collection of algorithms that utilizes artificial intelligence techniques to generate optimal conversational responses and suggestions based on input data.
[0571] The "means for generating a conversational response" is a means for generating an appropriate reply based on information obtained from the user.
[0572] The "means for providing" refers to a means for conveying the generated conversational response or suggestion content to the user.
[0573] An "emotion engine" is an algorithm that analyzes emotions from a user's voice and text data and recognizes emotional states such as joy, sadness, and anger.
[0574] "Purchase history" is a record of products that a user has purchased in the past, and is information that is used to make personalized product suggestions.
[0575] "Memories" are specific events or memories that a user has experienced in the past, and are the information that forms the basis for personalized responses and suggestions.
[0576] A "virtual reality display" is a display device that uses virtual reality technology to allow users to obtain information visually.
[0577] The "means for displaying product information in real time" refers to a means for visually presenting product information to the user immediately at the current time.
[0578] The "means of communicating to the user by voice and visual means" refers to means of communicating the generated information and proposal contents to the user by voice (auditory) and visual (visual).
[0579] This invention is a system that takes into account the user's emotional state, taps into past memories, and provides personalized product recommendations and virtual shopping experiences. The system combines an emotion engine and a generative AI model to achieve more personal and effective responses.
[0580] System configuration
[0581] The system of the present invention uses the following hardware and software:
[0582] Hardware: microphone, smart glasses or head-mounted display
[0583] software:
[0584] Speech recognition (Google Speech Recognition API)
[0585] Emotion Recognition (EmotionRecognizer library)
[0586] Product recommendation system (RecommendationEngine library)
[0587] Speech synthesis (pyttsx3 library)
[0588] VR display (vr_display module)
[0589] Overview of the process
[0590] 1. Voice input processing:
[0591] The user speaks to the system, making a request such as "Tell me about my recent purchases."
[0592] The device captures audio and converts it into text using the Google Speech Recognition API.
[0593] 2. Emotion Recognition with Emotion Engine:
[0594] The server analyzes the converted text data using EmotionRecognizer to recognize the user's emotions.
[0595] The perceived emotions are classified into emotional states such as joy, sadness, and anger.
[0596] 3. Recalling past memories and product suggestions:
[0597] The server retrieves the user's age information and past purchase history from a database and uses a generative AI model to generate optimal conversational responses and personalized product suggestions.
[0598] The suggestions are adjusted taking into account the emotions recognized by the emotion engine.
[0599] 4. Real-time visual and audio feedback:
[0600] Suggested product information and conversational responses are visually displayed to the user in real time using a VR display.
[0601] The device uses a speech synthesis engine (pyttsx3) to communicate the suggestions to the user aloud.
[0602] Specific examples
[0603] Example 1: Recalling past memories
[0604] The user says, "I want to talk about my old school days."
[0605] The device sends this request to a server, which looks up the user's demographic and related data and uses a generative AI model to generate a conversational response.
[0606] The server generates a question such as "What did you play at elementary school?" and the terminal conveys this to the user by voice.
[0607] Example 2: Product proposal
[0608] The user says, "Tell me about my recent purchases."
[0609] The device sends this request to the server, which retrieves the user's purchase history from a database, uses an emotion engine to recognize the user's happiness, and then suggests promotional items.
[0610] The suggestions are conveyed by a speech synthesis engine, such as "Here are some products you recently purchased. We also recommend these as new promotions," and are simultaneously displayed on the VR display.
[0611] Prompt Sentence Examples
[0612] Provide the following prompt as input to the system:
[0613] "When a user says, 'Tell me about the latte you just bought,' the emotion engine recognizes that the user is happy and suggests in a soft tone, 'That was a delicious latte. How about trying a new flavor next time as part of a special promotion?'"
[0614] In this way, the present invention uses a system that combines emotion engines to provide personalized conversational responses and product suggestions that take into account the user's emotional state, resulting in a more personalized and effective virtual shopping experience.
[0615] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0616] Step 1:
[0617] The user speaks, for example making a request such as "Tell me about my recent purchases," which is captured as voice input.
[0618] Step 2:
[0619] The device captures the user's voice. It takes in the voice data through the microphone and converts the voice into text using the Google Speech Recognition API. The input is voice data and the output is text data.
[0620] Step 3:
[0621] The server receives the converted text data. This text data is analyzed using the EmotionRecognizer library to recognize the user's emotions. The input is text data, and the output is the emotional state (joy, sadness, anger, etc.).
[0622] Step 4:
[0623] Based on the emotion recognition results, the server retrieves the user's age information and past purchase history from the database. The input is the user ID and emotional state, and the output is age information and purchase history.
[0624] Step 5:
[0625] Based on the generation information and purchase history acquired by the server, a generative AI model is used to generate optimal conversational responses and product suggestions. The responses and suggestions are adjusted taking into account the emotional state recognized by the emotion engine. The input is generational information, purchase history, and emotional state, and the output is the generated conversational responses or product suggestions.
[0626] Step 6:
[0627] The server sends the generated conversational responses and product suggestions to the device. The device receives this and converts the text data into speech using a speech synthesis engine (pyttsx3). It then visually displays the product information in real time using a VR display. The input is the generated conversational responses or product suggestions, and the output is audio and visual information.
[0628] Step 7:
[0629] The user confirms the information they have received through audio and visuals. For example, they may hear a voice message saying, "Here are the products you recently purchased. We also recommend these as new promotions," while product information is displayed on the VR display. This allows users to easily check products that are relevant to their interests.
[0630] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0631] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0632] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0633] [Second embodiment]
[0634] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0635] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0636] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0637] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0638] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0639] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0640] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0641] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0642] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0643] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0644] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0645] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0646] This invention is an AI chat system that elicits past memories in dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and also suggests exercise methods and dietary content tailored to the user's physical condition.
[0647] System Overview
[0648] 1. Voice input processing:
[0649] User: Uses the system and speaks to initiate a conversation about a past memory.
[0650] Terminal: Converts the voice into text data and sends the text data to the server.
[0651] 2. Drawer of past memories:
[0652] Server: Based on the received text data, the server searches the database for the user's age information, for example, information related to the user's childhood or school days.
[0653] Server: Using information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, utilizing materials such as popular items and events by decade, photos, and movie posters.
[0654] Server: Sends the generated conversation response to the terminal.
[0655] Terminal: The received response is conveyed to the user via voice, enabling conversations that evoke past memories.
[0656] 3. Exercise and dietary suggestions:
[0657] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[0658] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[0659] Server: Retrieves the received physical condition data and past exercise history from a database. Based on this, the server uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[0660] Server: Sends the generated proposals to the device.
[0661] Device: The received suggestions are communicated to the user via voice.
[0662] 4. Support for conversations with family and relatives:
[0663] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[0664] Terminal: Converts the voice into text data and sends the text data to the server.
[0665] Server: Retrieves information about family and relatives from a database, and uses a generative AI model to generate optimal conversation responses. A simple 3D avatar is also generated at the same time.
[0666] Server: Sends the generated conversation responses and 3D avatars to the device.
[0667] Terminal: Assists with everyday conversations by conveying received responses to the user via voice and displaying a 3D avatar.
[0668] Specific examples
[0669] Example 1: Recalling past memories
[0670] When a user says, "I want to talk about my school days,"
[0671] The terminal sends this request to the server.
[0672] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[0673] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[0674] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[0675] The user begins recounting past events to jog their memory, and the conversation continues.
[0676] Example 2: Exercise and diet suggestions
[0677] When a user says, "Which stretch should I do today?"
[0678] The terminal sends this request to the server.
[0679] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0680] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[0681] Example 3: Supporting conversations with family and relatives
[0682] When a user says, "I want to know about my grandchildren,"
[0683] The terminal sends this request to the server.
[0684] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[0685] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[0686] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[0687] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[0688] The processing flow will be explained below.
[0689] Ability to recall past memories
[0690] Step 1:
[0691] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[0692] Step 2:
[0693] The terminal converts the user's voice into text data and transmits the text data to the server.
[0694] Step 3:
[0695] The server receives the user's request and searches the database for the user's age information and related data.
[0696] Step 4:
[0697] The server uses a generative AI model to generate a conversational response such as "What did you play in elementary school?" based on age information and related data obtained from the database.
[0698] Step 5:
[0699] The server transmits the generated conversation response to the terminal.
[0700] Step 6:
[0701] The terminal converts the response received from the server into voice and conveys it to the user.
[0702] Suggestions for exercise and diet
[0703] Step 1:
[0704] The user speaks to the device, "Which stretch should I do today?"
[0705] Step 2:
[0706] The terminal converts the user's voice into text data and transmits the text data to the server.
[0707] Step 3:
[0708] The server retrieves the user's physical condition record and past exercise history from a database.
[0709] Step 4:
[0710] Based on the acquired physical condition records and exercise history, the server uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0711] Step 5:
[0712] The server sends the generated proposal to the terminal.
[0713] Step 6:
[0714] The terminal converts the proposal received from the server into voice and conveys it to the user.
[0715] Support for conversations with family and relatives
[0716] Step 1:
[0717] The user speaks to the terminal saying, "I want to know about my grandchildren."
[0718] Step 2:
[0719] The terminal converts the user's voice into text data and transmits the text data to the server.
[0720] Step 3:
[0721] The server retrieves information about the user's family and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[0722] Step 4:
[0723] Based on the acquired information, the server uses a generative AI model to generate a conversational response such as, "Do you know what your grandchild is currently learning at school?" It also generates a simple 3D avatar.
[0724] Step 5:
[0725] The server sends the generated conversational responses and 3D avatars to the device.
[0726] Step 6:
[0727] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[0728] Proposing exercise methods and meal plans tailored to the user's physical condition
[0729] Step 1:
[0730] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[0731] Step 2:
[0732] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0733] Step 3:
[0734] The server updates the database with the newly received physical condition data.
[0735] Step 4:
[0736] The server uses a generative AI model based on the received health data and previously stored data to generate suggestions such as, "We recommend that you eat something that is easy to digest today."
[0737] Step 5:
[0738] The server sends the generated proposal to the terminal.
[0739] Step 6:
[0740] The terminal converts the proposal received from the server into voice and conveys it to the user.
[0741] The above is a specific description of the processing flow.
[0742] Example 1
[0743] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0744] Today, there is a lack of appropriate systems to help dementia patients recall their past memories and support their daily communication. Furthermore, there are no suggestions for exercise or dietary content tailored to their physical condition, which results in a decline in their quality of life. Furthermore, communication with family and relatives is often difficult, leading to feelings of isolation.
[0745] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0746] In this invention, the server includes means for receiving a user's voice input and processing the voice data, means for converting the voice data into text data, means for searching a database for the user's age information and related data, means for using a generative AI model to set prompts for generating optimal conversational responses and generating response messages, means for providing the generated response messages to the user, means for referencing the user's physical condition records and past exercise history and using a generative AI model to suggest optimal exercise methods and dietary content based on the user's physical condition records and past exercise history, and means for providing the generated suggestions to the user, means for acquiring information about the user's family and relatives from a database and using a generative AI model to generate appropriate response messages based on the acquired information, and means for generating a simple 3D avatar and providing the generated 3D avatar and response message to the user. This effectively draws out the past memories of dementia patients to support conversations, enables suggestions for exercise methods and dietary content tailored to their physical condition, and enables smooth communication with family and relatives.
[0747] A "user" is an entity that uses the system to provide voice input and receive conversations and suggestions.
[0748] "Voice input" refers to voice data input by a user speaking to the system.
[0749] "Voice data" refers to a digital recording of a user's voice input.
[0750] "Text data" refers to data obtained by converting voice data into characters.
[0751] "Age information" refers to data related to a user's year of birth or a particular age group.
[0752] A "database" is a system for storing and managing related data such as age and family information.
[0753] A "generative AI model" is an artificial intelligence algorithm that generates optimal responses or suggestions based on input data.
[0754] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[0755] A "response message" is a conversational response generated by a generative AI model.
[0756] "Physical condition record" refers to a record that accumulates data related to the user's physical condition.
[0757] "Exercise history" refers to data relating to exercises that a user has performed in the past.
[0758] An "exercise method" is an exercise technique suggested based on the user's physical condition and exercise history.
[0759] "Meal content" refers to an appropriate meal plan suggested to the user.
[0760] "Family information" refers to data about the user's family and relatives.
[0761] A "3D avatar" is a three-dimensional virtual character generated on a computer.
[0762] This invention relates to an AI chat system that supports daily communication for dementia patients and improves their quality of life. This system provides comprehensive support by evoking the user's past memories, engaging in conversations based on those memories, and suggesting exercise methods and dietary content tailored to the user's physical condition.
[0763] Hardware and software used
[0764] Audio input device (microphone): Receives the user's voice and records the audio data.
[0765] Speech recognition software (e.g., Google Speech-to-Text API): converts voice data into text data.
[0766] Database system: Stores data such as the user's age, family information, health records, and exercise history.
[0767] Generative AI models (e.g., OpenAI's GPT-3): Set prompts based on received text data and generate appropriate conversational responses and suggestions.
[0768] Speech synthesis software (e.g., Amazon Polly): Converts the generated response message into audio data.
[0769] 3D avatar generation software: Generates a 3D avatar based on the user's family information.
[0770] Specific explanation of the system
[0771] 1. Voice input processing
[0772] The user speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[0773] The device receives the user's voice using a high-performance microphone and records it as voice data.
[0774] The terminal converts the recorded voice data into text data using voice recognition software and transmits the text data to the server.
[0775] 2. Database search
[0776] The server analyzes the received text data and extracts keywords such as "school days."
[0777] The server searches the database for the user's generational information, and if the user was a student in the 1960s, for example, it retrieves data on trends and events from that era.
[0778] 3. Response generation using generative AI models
[0779] The server uses a generative AI model to set prompts based on information retrieved from the database, such as "You were a student in the 1960s. Do you have any memorable experiences?"
[0780] The server sends the generated response message to the user's terminal.
[0781] 4. Providing a response message
[0782] The terminal converts the received response message into voice data using voice synthesis software.
[0783] The terminal reproduces the audio data from a speaker and conveys it to the user.
[0784] Specific examples
[0785] Example 1: Recalling past memories
[0786] When a user says, "I want to talk about my school days,"
[0787] The device records this audio and converts it into text using speech recognition software.
[0788] The server receives this text data and extracts keywords such as "school days."
[0789] The server retrieves the user's generational information from a database and researches trends and events from, for example, the 1960s.
[0790] The server uses a generative AI model to set a prompt sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?" and generates the optimal response message.
[0791] The server sends this response message to the terminal,
[0792] The device synthesizes the received message into voice and plays it back through the speaker, telling the user, "You were a student in the 1960s. Do you have any memorable experiences?"
[0793] Example 2: Exercise and diet suggestions
[0794] When a user says, "Which stretch should I do today?"
[0795] The device records this audio and converts it into text using speech recognition software.
[0796] The server receives this text data and extracts keywords such as "stretch."
[0797] The server retrieves the user's physical condition records and past exercise history from a database and uses a generative AI model to suggest the optimal exercise method.
[0798] The server sets a prompt sentence such as "Today, we recommend neck rotation exercises as a light stretch," and generates a suggestion.
[0799] The server sends this proposal to the terminal,
[0800] The device synthesizes the received suggestions into voice and plays them back through the speaker, telling the user, "Today, I recommend doing some neck rotation exercises as a light stretch."
[0801] Example 3: Supporting conversations with family and relatives
[0802] When a user says, "I want to know about my grandchildren,"
[0803] The device records this audio and converts it into text using speech recognition software.
[0804] The server receives this text data and extracts keywords such as "grandchild."
[0805] The server retrieves information about grandchildren from a database and uses a generative AI model to generate the question, "Do you know what your grandchildren are learning in school right now?"
[0806] The server also simultaneously generates a simple 3D avatar and sends it to the user's device.
[0807] The device synthesizes the received question into speech, plays it back through the speaker, and displays a 3D avatar, asking the user, "Do you know what your grandchild is learning at school right now?"
[0808] In this way, the present invention improves the quality of life of dementia patients by effectively drawing out the user's past memories and making suggestions tailored to everyday conversations and physical condition.
[0809] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0810] Step 1: Receiving voice input
[0811] User: Speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[0812] Terminal: The user's voice is received by a high-performance microphone and recorded as audio data, which becomes the input data.
[0813] Terminal: The recorded audio data is sent to the server. The output at this stage is audio data.
[0814] Step 2: Converting audio data to text data
[0815] Server: Converts the received voice data into text data using speech recognition software (e.g., Google Speech-to-Text API). This is the input data.
[0816] Server: The converted text data is used in the next processing step. The output is text data.
[0817] Step 3: Database search
[0818] Server: Analyzes the received text data and extracts keywords such as "school days." This is the input data.
[0819] Server: Searches the database for the user's generation information and retrieves data on trends and events from the corresponding era. This data processing and calculation is used to reference and search for generation information. The output is data from the corresponding era.
[0820] Step 4: Generative AI model generates a response
[0821] Server: Based on the information obtained from the database, a generative AI model (e.g., OpenAI's GPT-3) is used to set a prompt. For example, the prompt may be a text sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?"
[0822] Server: Using this prompt as input, the generative AI model generates the optimal conversational response. This response generation is performed using data calculations.
[0823] Server: Output the generated response message as text data for use in the next step.
[0824] Step 5: Send a response message
[0825] Server: Sends the generated response message to the user's terminal. This is the input data.
[0826] Terminal: Receives the response message and converts it into voice data using speech synthesis software (e.g., Amazon Polly). Voice synthesis is a type of data processing. The output is voice data.
[0827] Step 6: Play greetings
[0828] Terminal: The audio data is played back through the speaker and transmitted to the user. This is the final output.
[0829] Example flow
[0830] User: Say, "Which stretch should I do today?"
[0831] Device: Uses voice recognition software to convert the question "Which stretch should I do today?" into text data and send it to the server.
[0832] Server: Analyzes the text data and extracts keywords.
[0833] Server: Retrieves the user's physical condition records and past exercise history from the database.
[0834] Server: Using a generative AI model, it generates a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0835] Server: Sends the proposal to the device.
[0836] Device: The suggested message is generated through speech synthesis and played from the speaker. It tells the user, "Today, we recommend you do some light stretching exercises, such as neck rotations."
[0837] In this way, the processing steps from the user's voice input to the final playback of the response message are concretely carried out.
[0838] (Application example 1)
[0839] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0840] This invention relates to an AI chat system that evokes past memories in dementia patients and improves their quality of life. Currently, many nursing homes lack support for retrieving individual memories for dementia patients, resulting in insufficient mental stability and communication. It is also difficult to provide dementia patients with appropriate exercise and dietary recommendations tailored to their physical condition, and communication with their families is also difficult. To address these issues, a system is needed that effectively engages dementia patients in conversation and suggests health management strategies based on their past living conditions.
[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0842] In this invention, the server includes a means for processing user voice input, a means for searching a database for user age information and related data, a means for generating optimal conversational responses using a generative AI model, a means for providing the generated conversational responses to the user by voice and eliciting past memories, a means for using the generative AI model to hold conversations based on the past living situations of dementia patients and support daily communication, and an application installed on a robot installed in a care facility, which makes it possible to increase the mental stability of dementia patients and provide effective communication support.
[0843] "Means for processing user voice input" refers to a function for receiving voice data uttered by a user, recognizing it, and converting it into text data.
[0844] "Means of searching for a user's age information and related data from a database" refers to the function of searching for and retrieving the necessary information from a database that holds information related to the user's age, interests, and past life.
[0845] "Means for generating optimal conversational responses using a generative AI model" refers to a function that uses AI technology based on text data to generate optimal responses for users.
[0846] "Means for providing the generated conversational response to the user by voice and eliciting past memories" refers to a function for providing the user with an automatically generated text response by voice output, thereby eliciting the user's past memories.
[0847] "A means of using a generative AI model to hold conversations based on the past living conditions of dementia patients and support daily communication" refers to a function that makes full use of AI technology to conduct conversations based on content related to the past experiences and lives of dementia patients, thereby promoting daily interaction.
[0848] "Applications installed on robots installed in nursing facilities" refers to application software that is installed on robot devices used in nursing facilities and that is used to interact with dementia patients.
[0849] "A means for referring to the user's health records and past exercise history, and proposing optimal exercise methods and dietary content based on the generated health data" refers to a function that refers to the user's previous health information and exercise data, and provides appropriate exercise and dietary advice based on the user's current health condition.
[0850] "Means for providing the generated suggestions to the user by voice" refers to a function for conveying the generated exercise and diet suggestions to the user by voice.
[0851] "Means for a robot installed in a care facility to demonstrate suggested content to a user" refers to a function in which a robot demonstrates appropriate exercise methods and other instructions to a user.
[0852] "Means of obtaining information about the user's family and relatives from a database and generating a simple 3D avatar based on that information" refers to a function that obtains data about the user's family and relatives from a database and generates an avatar based on that data.
[0853] "Means of using the generated 3D avatar to engage in everyday conversation with the user and provide supplementary animations" refers to the function of using the generated 3D avatar as an interface to assist in conversation with the user and display animations as a visual aid.
[0854] "Means for providing information about family members and relatives to the user by voice" refers to a function that conveys information about family members and relatives obtained from a database to the user by voice.
[0855] This invention is an AI chat system that elicits past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary recommendations tailored to the user's physical condition. This system is implemented as an application installed on a robot in a nursing home.
[0856] System Overview
[0857] 1. Voice input processing
[0858] The user uses the system to speak to initiate a conversation about a past memory.
[0859] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[0860] 2. Drawer of past memories
[0861] The server searches a database for information about the user's age based on the received text data, such as information about the user's childhood or school days.
[0862] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[0863] The server transmits the generated conversation response to the terminal.
[0864] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[0865] 3. Exercise and dietary suggestions
[0866] The user reports their physical condition to the device, for example, by saying, "I feel a little tired today."
[0867] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0868] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[0869] The server sends the generated proposal to the terminal.
[0870] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[0871] 4. Support for conversations with family and relatives
[0872] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[0873] The terminal converts the voice into text data and transmits the text data to the server.
[0874] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[0875] The server sends the generated conversational responses and 3D avatars to the terminal.
[0876] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[0877] Specific examples
[0878] Example 1: Recalling past memories
[0879] The user says, "I want to talk about my old school days."
[0880] The terminal sends this request to the server.
[0881] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[0882] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[0883] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[0884] The user begins recounting past events to jog their memory, and the conversation continues.
[0885] Example prompt sentence:
[0886] User: I want to talk about my old school days.
[0887] AI: What did you play in elementary school?
[0888] Example 2: Exercise and diet suggestions
[0889] The user says, "Which stretch should I do today?"
[0890] The terminal sends this request to the server.
[0891] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[0892] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[0893] If necessary, the robot will demonstrate neck rotation.
[0894] Example prompt sentence:
[0895] User: Which stretches should I do today?
[0896] AI: Today, I recommend some gentle stretching exercises, such as neck rotations.
[0897] Example 3: Supporting conversations with family and relatives
[0898] The user says, "I want to know about my grandchildren."
[0899] The terminal sends this request to the server.
[0900] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[0901] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[0902] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[0903] Example prompt sentence:
[0904] User: I want to know about my grandchildren
[0905] AI: Do you know what your grandchildren are learning in school right now?
[0906] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[0907] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0908] Step 1:
[0909] The user uses the system to speak to initiate a conversation about a past memory.
[0910] Input: User's voice data
[0911] How it works: A microphone on a robot installed in a care home captures voice input.
[0912] Step 2:
[0913] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[0914] Input: User's voice data
[0915] What it does: It uses speech recognition software (e.g., the speech_recognition library) to convert speech to text, which is then sent over the Internet to a server.
[0916] Output: Text data sent to the server
[0917] Step 3:
[0918] The server searches the database for the user's age information based on the received text data.
[0919] Input: User's text data
[0920] What it does: The server performs a database query to retrieve data about the user's age and related past memories. For example, if the user types "I want to talk about my old school days," it searches for information related to the user's childhood and school days.
[0921] Output: User age and related data
[0922] Step 4:
[0923] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[0924] Input: User demographics and related data
[0925] Specific operation: Query a generative AI model (e.g., OpenAI GPT-3) with a prompt sentence to generate an appropriate conversational response. Reference data such as trends and events relevant to each generation are also used as reference.
[0926] Output: Generated conversation response
[0927] Step 5:
[0928] The server transmits the generated conversation response to the terminal.
[0929] Input: Generated conversation response
[0930] Specific operation: The response text generated by the server is sent to the terminal via the Internet.
[0931] Output: Conversation response sent to the terminal
[0932] Step 6:
[0933] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[0934] Input: Conversation response sent by the server
[0935] What it does: It uses speech synthesis software (e.g., the pyttsx3 library) to play the text data as speech, allowing the user to continue the conversation by listening to the voice response.
[0936] Output: The audio response provided to the user
[0937] Step 7:
[0938] The user reports their physical condition to the device, for example, saying, "I feel a little tired today."
[0939] Input: User's voice data
[0940] What happens: The device's microphone captures audio input.
[0941] Step 8:
[0942] The terminal converts the physical condition data into text data and transmits the text data to a server.
[0943] Input: Audio data about the user's physical condition
[0944] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[0945] Output: Text data about your health condition sent to the server
[0946] Step 9:
[0947] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[0948] Input: Text data about the user's physical condition and past exercise history
[0949] Specific operation: The server retrieves the user's past exercise history from the database and uses a generative AI model to suggest appropriate exercise methods and dietary recommendations.
[0950] Output: Generated exercise and diet suggestions
[0951] Step 10:
[0952] The server sends the generated proposal to the terminal.
[0953] Input: Generated proposal
[0954] Specific operation: The proposal content is sent from the server to the device via the Internet.
[0955] Output: Suggestions sent to the device
[0956] Step 11:
[0957] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[0958] Input: Suggestion from the server
[0959] Specific behavior: Using speech synthesis software, the robot plays back the suggestions as voice, and shows the user the appropriate exercise method.
[0960] Output: Audio suggestions and demonstrations provided to the user
[0961] Step 12:
[0962] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[0963] Input: User's voice data
[0964] What happens: The device's microphone captures audio input.
[0965] Step 13:
[0966] The terminal converts the voice into text data and transmits the text data to the server.
[0967] Input: User's voice data
[0968] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[0969] Output: Text data sent to the server
[0970] Step 14:
[0971] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[0972] Input: User's text data
[0973] How it works: The server performs database queries to retrieve information about family and relatives, then uses a generative AI model to generate optimal conversational responses, while simultaneously generating a simple 3D avatar.
[0974] Output: Generated conversational responses and 3D avatars
[0975] Step 15:
[0976] The server sends the generated conversational responses and 3D avatars to the terminal.
[0977] Input: Generated conversational responses and 3D avatars
[0978] Specific operation: The server generates a response text and sends a 3D avatar to the device via the Internet.
[0979] Output: Speech responses and 3D avatar sent to device
[0980] Step 16:
[0981] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[0982] Input: Conversational responses and 3D avatars sent from the server
[0983] Specific operations: Uses speech synthesis software to play text data as speech, and displays a 3D avatar on a display device to assist in conversation with the user.
[0984] Output: Audio response and 3D avatar representation provided to the user
[0985] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0986] This invention is an AI chat system that draws out past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to the user's physical condition. In addition, by combining it with an emotion engine that recognizes the user's emotions, more personalized responses and suggestions are possible.
[0987] System Overview
[0988] 1. Voice input processing:
[0989] User: Uses the system to initiate a conversation about a past memory.
[0990] Terminal: Converts the voice into text data and sends the text data to the server.
[0991] 2. Emotion Recognition with Emotion Engine:
[0992] Server: Analyzes the received text data and recognizes the user's emotions using an emotion engine.
[0993] The emotion engine analyzes emotions from the user's voice input and text data to determine emotional states such as joy, sadness, and anger.
[0994] 3. Drawer of past memories:
[0995] Server: Retrieves the user's age information and related data from a database, for example, retrieves information related to the user's childhood or school days.
[0996] Server: Based on information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, adjusting the content of the conversation based on emotions recognized by the emotion engine.
[0997] Server: Sends the generated conversation response to the terminal.
[0998] Terminal: Converts the received response into speech and conveys it to the user.
[0999] 4. Exercise and dietary suggestions:
[1000] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[1001] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[1002] Server: Retrieves the user's physical condition records and past exercise history from a database, and uses a generative AI model to suggest optimal exercise methods and dietary recommendations. The server also adjusts the suggestions by taking into account emotions recognized by the emotion engine.
[1003] Server: Sends the generated proposals to the device.
[1004] Device: Converts the received suggestions into speech and conveys them to the user.
[1005] 5. Support for conversations with family and relatives:
[1006] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[1007] Terminal: Converts the voice into text data and sends the text data to the server.
[1008] Server: Retrieves information about family and relatives from a database and uses a generative AI model to generate optimal conversational responses. A simple 3D avatar is also generated at the same time. The conversational responses are adjusted based on emotions recognized by the emotion engine.
[1009] Server: Sends the generated conversation responses and 3D avatars to the device.
[1010] Terminal: Converts the received response into voice and displays a 3D avatar to communicate with the user.
[1011] Specific examples
[1012] Example 1: Recalling past memories
[1013] When a user says, "I want to talk about my school days,"
[1014] The terminal sends this request to the server.
[1015] The server retrieves the user's age and related data from the database. For example, if the user is having a good time, it generates questions about the pastimes and movies they enjoyed at the time.
[1016] The server uses a generative AI model based on information retrieved from the database to tailor the conversational response based on the emotions recognized by the emotion engine, generating the question, "What did you play at elementary school?"
[1017] The server sends the generated question to the terminal, which then transmits it to the user as voice.
[1018] The user begins to recount past events to jog their memory, and the conversation continues.
[1019] Example 2: Exercise and diet suggestions
[1020] When a user says, "Which stretch should I do today?"
[1021] The terminal sends this request to the server.
[1022] The server references the user's physical condition records and past exercise history, and uses a generative AI model to determine the optimal exercise method and diet. The suggestions are adjusted based on the emotions recognized by the emotion engine. For example, if the user is tired, the server might suggest, "Today, we recommend neck rotation exercises as a light stretch."
[1023] The server transmits the generated proposal to the terminal, which then conveys it to the user as voice.
[1024] Example 3: Supporting conversations with family and relatives
[1025] When a user says, "I want to know about my grandchildren,"
[1026] The terminal sends this request to the server.
[1027] The server uses a generative AI model to generate conversational responses based on information about the grandchild (such as name, school, and interests) retrieved from a database. The conversation content is adjusted based on emotions recognized by an emotion engine. A simple 3D avatar is also generated.
[1028] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[1029] In this way, the present invention provides a concrete means to improve the quality of life of dementia patients and reduce the burden on caregivers. By combining it with an emotion engine, personalized responses and suggestions that take into account the user's emotional state become possible, realizing more personal and effective support.
[1030] The processing flow will be explained below.
[1031] Ability to recall past memories
[1032] Step 1:
[1033] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[1034] Step 2:
[1035] The terminal converts the user's voice into text data and transmits the text data to the server.
[1036] Step 3:
[1037] The server analyzes the text data and uses an emotion engine to recognize the user's emotions, such as whether they are happy or sad.
[1038] Step 4:
[1039] The server retrieves the user's age information and related data from a database, for example, information related to the user's childhood or school days.
[1040] Step 5:
[1041] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database. The content of the conversation is adjusted based on emotions recognized by the emotion engine. For example, if the user seems to be having fun, the server generates a question such as, "Did you have any happy memories from that time?"
[1042] Step 6:
[1043] The server transmits the generated conversation response to the terminal.
[1044] Step 7:
[1045] The terminal converts the response received from the server into voice and conveys it to the user.
[1046] Suggestions for exercise and diet
[1047] Step 1:
[1048] The user speaks to the device, "Which stretch should I do today?"
[1049] Step 2:
[1050] The terminal converts the user's voice into text data and transmits the text data to the server.
[1051] Step 3:
[1052] The server retrieves the user's physical condition record and past exercise history from a database.
[1053] Step 4:
[1054] The server uses a generative AI model to suggest optimal exercise methods and dietary recommendations based on the user's physical condition data and past exercise history. The server also adjusts the recommendations by taking into account the user's emotional state, as recognized by the emotion engine. For example, if the user expresses feelings of fatigue, the server might suggest, "We recommend getting plenty of rest and doing some light stretching today."
[1055] Step 5:
[1056] The server sends the generated proposal to the terminal.
[1057] Step 6:
[1058] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1059] Support for conversations with family and relatives
[1060] Step 1:
[1061] The user speaks to the terminal saying, "I want to know about my grandchildren."
[1062] Step 2:
[1063] The terminal converts the user's voice into text data and transmits the text data to the server.
[1064] Step 3:
[1065] The server analyzes the text data and uses an emotion engine to recognize the user's emotions, for example, capturing the joy and excitement when the user talks about their grandchildren.
[1066] Step 4:
[1067] The server retrieves information about family members and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[1068] Step 5:
[1069] The server uses a generative AI model based on the acquired information to generate optimal conversational responses based on the emotions recognized by the emotion engine. For example, if the user is excited about their grandchild, the server generates a question such as, "What is your grandchild's favorite pastime these days?"
[1070] Step 6:
[1071] The server also simultaneously generates a simple 3D avatar and sends a conversation response including this avatar to the terminal.
[1072] Step 7:
[1073] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[1074] Proposing exercise methods and meal plans tailored to the user's physical condition
[1075] Step 1:
[1076] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[1077] Step 2:
[1078] The terminal converts the physical condition data into text data and transmits the text data to a server.
[1079] Step 3:
[1080] The server updates the database with the newly received physical condition data.
[1081] Step 4:
[1082] The server uses a generative AI model to suggest optimal exercise methods and meal plans based on the received physical condition data and previously stored data. The suggestions are adjusted taking into account the emotions recognized by the emotion engine. For example, if the user expresses feelings of fatigue, the server generates a suggestion such as, "Today's meal is recommended to help you relax."
[1083] Step 5:
[1084] The server sends the generated proposal to the terminal.
[1085] Step 6:
[1086] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1087] The above is a description of the process flow broken down into specific steps.
[1088] Example 2
[1089] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1090] In modern society, the number of dementia patients is increasing, and methods to improve their quality of life are needed. In particular, it is necessary to enrich the lives of dementia patients and reduce the burden on caregivers by evoking past memories and providing personalized responses based on emotions. However, conventional technologies lack the accuracy of voice input and the provision of personalized responses, making it difficult to provide appropriate support tailored to the specific needs of dementia patients.
[1091] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for processing a user's voice input, means for converting the user's voice input into text, means for transmitting the text data to the server, means for receiving the text data at the server and recognizing emotions with an emotion engine, means for searching a database for user age information and related data, means for generating optimal conversational responses based on the results of the emotion engine using a generative AI model, means for transmitting the generated conversational responses from the server to the terminal, means for converting the conversational responses received by the terminal into speech, and means for providing the generated speech to the user. This enables personalized responses and suggestions based on the emotions and past memories of dementia patients, improving their quality of life and reducing the burden on caregivers.
[1092] The "means for processing user speech input" is a means for capturing speech from the user and sending it to the next processing step.
[1093] "Means for converting user voice input into text" refers to a speech recognition system for converting voice data into text data.
[1094] "Means for transmitting text data to a server" refers to a communication means for transmitting the converted text data to a server via a network.
[1095] "Means for recognizing emotions using an emotion engine" refers to an emotion analysis system for analyzing and recognizing a user's emotional state from text data.
[1096] "Means for searching a database for user generation information and related data" refers to a search function for obtaining a user generation information and related data from a database.
[1097] "Means for generating optimal conversational responses based on the results of an emotion engine using a generative AI model" refers to a system for generating optimal responses based on the results of emotion analysis using an AI model.
[1098] The "means for transmitting the generated conversation response from the server to the terminal" refers to a communication means for transmitting the generated conversation response from the server to the terminal.
[1099] "Means for converting the conversational response received by the terminal into speech" refers to a speech synthesis system that converts the received text-based conversational response into speech.
[1100] The "means for providing the generated audio to the user" refers to an output device such as a speaker or a headphone that allows the user to hear the generated audio.
[1101] This invention is an AI system that elicits past memories in dementia patients and improves their quality of life. The system aims to evoke the user's past memories, support daily communication, and suggest exercise methods and dietary content tailored to the user's physical condition. In addition, by combining it with an emotion engine that recognizes the user's emotions, more personalized responses and suggestions are possible.
[1102] System configuration
[1103] The system uses the following hardware and software:
[1104] A microphone to capture the user's voice input
[1105] Speech recognition software (e.g., Google Cloud Speech-to-Text API)
[1106] Sentiment analysis software (e.g., IBM Watson Tone Analyzer)
[1107] A database (e.g., MySQL or MongoDB)
[1108] Generative AI models (e.g., OpenAI GPT-3)
[1109] A speech synthesis engine (e.g., Amazon Polly)
[1110] Speakers or headphones for audio output to the user
[1111] System Operation
[1112] Voice Input Processing
[1113] When a user speaks to the system, the device's microphone captures the speech, and speech recognition software converts the speech data into text data, which is then sent to the server.
[1114] Emotion recognition by emotion engine
[1115] The server inputs the received text data into emotion analysis software to analyze the user's emotions. The emotion analysis software analyzes keywords and context within the text to recognize emotions such as joy, sadness, and anger.
[1116] Drawer of past memories
[1117] The server searches a database for the user's age information and related data. For example, it obtains information about the user's childhood and school days. Based on the obtained information and the results of emotion analysis, the server inputs a prompt sentence into the generative AI model to generate an optimal conversational response. An example of a prompt sentence is, "The user wants to talk about their school days. They appear quietly pleased." The generated conversational response is sent to the device, which then converts it into speech using a speech synthesis engine and conveys it to the user.
[1118] Suggestions for exercise and diet
[1119] When a user reports their physical condition, the device converts it into text and sends it to a server. The server retrieves the user's physical condition records and past exercise history from a database, and, taking into account the results of emotion analysis, uses a generative AI model to suggest optimal exercise methods and dietary content. For example, a suggestion might be generated such as, "Today, we recommend neck rotation exercises as a light stretch." The generated suggestion is sent to the device, which converts it into voice and conveys it to the user.
[1120] Support for conversations with family and relatives
[1121] When a user wants to know about their family or relatives, they speak their request into the device. The device converts the voice input into text and sends it to the server. The server retrieves information about family and relatives from a database and generates a conversational response using a generative AI model based on that information. A simple 3D avatar is also generated. For example, in response to a request such as "I want to know about my grandchildren," a response such as "Your grandchild is currently 8 years old and loves soccer" is generated. The generated conversational response and 3D avatar are sent to the device, which converts it into speech and displays the 3D avatar to communicate with the user.
[1122] Specific examples
[1123] Below are some specific examples of how the system can be used.
[1124] Example 1: Recalling past memories
[1125] When a user says, "I want to talk about my school days," the device sends this request to the server. The server retrieves the user's age information and related data from a database, and generates a question based on the results of emotion analysis: "What did you play in elementary school?" The device converts this question into speech and conveys it to the user, who then begins talking about past events to jog their memories.
[1126] Example 2: Exercise and diet suggestions
[1127] When the user says, "Which stretch should I do today?", the device sends this request to the server. The server retrieves physical condition records and past exercise history from a database, and, taking into account the results of emotion analysis, suggests, "Today, I recommend neck rotations as a light stretch." The device converts this into voice and conveys it to the user.
[1128] Example 3: Supporting conversations with family and relatives
[1129] When a user says, "I want to know about my grandchild," the device sends this request to the server. The server retrieves information about the family from a database and generates a 3D avatar along with the response, "Your grandchild is now 8 years old and loves soccer." The device then provides this as voice and displays the 3D avatar.
[1130] As described above, the system of the present invention can provide personalized responses and suggestions based on the user's emotional state and past memories to improve the quality of life of dementia patients.
[1131] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1132] Step 1:
[1133] A user speaks aloud. For example, the user says, "I'd like to talk about my old school days." The input is the user's voice data, and the output is the voice data itself.
[1134] Step 2:
[1135] The device converts speech to text. The device's microphone captures the user's speech data and uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the speech to text data. The input is the user's speech data, and the output is text data. Specifically, the speech recognition software performs phonemic analysis and temporarily stores the recognition results as text on the device.
[1136] Step 3:
[1137] The terminal sends text data to the server. The converted text data is sent to the server using the HTTPS protocol. The input is text data, and the output is the text data sent to the server.
[1138] Step 4:
[1139] The server receives text data. The server receives text data sent from the terminal. The input is the text data sent from the terminal, and the output is the text data saved on the server.
[1140] Step 5:
[1141] The server recognizes emotions using an emotion engine. The server inputs the received text data into emotion analysis software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data, and the output is data indicating the user's emotional state (e.g., joy, sadness, anger, etc.). Specifically, the emotion analysis engine analyzes keywords and context within the text and recognizes the emotional state numerically and categorically.
[1142] Step 6:
[1143] The server searches for the user's past data. The server retrieves the user's age information and related data from a database (e.g., MySQL or MongoDB). The input is the user's profile information, and the output is the retrieved past data. Specifically, the server issues an SQL query to retrieve the required information.
[1144] Step 7:
[1145] The server generates a response using a generative AI model. Based on the acquired past data and the results of emotion analysis, the server inputs a prompt into the generative AI model (e.g., OpenAI GPT-3) to generate an optimal conversational response. The input is past data and emotional state data, and the output is the generated conversational response. An example of a specific prompt is, "The user wants to talk about their school days. They appear quietly pleased."
[1146] Step 8:
[1147] The server sends the generated response to the terminal. The generated conversation response is sent to the terminal using the HTTPS protocol. The input is the generated conversation response, and the output is the conversation response sent to the terminal.
[1148] Step 9:
[1149] The device converts the response into speech and transmits it. The device converts the received conversational response into speech using a speech synthesis engine (e.g., Amazon Polly) and transmits it to the user. The input is a text-format conversational response, and the output is a speech-format conversational response. Specifically, the audio is played back to the user through the device's speakers or headphones.
[1150] Step 10:
[1151] A user reports how they are feeling, for example, "I feel a little tired today." The input is the user's voice data, and the output is the voice data itself.
[1152] Step 11:
[1153] The device converts the voice into text and sends it to the server. As in step 2, it uses voice recognition software and sends the converted text data to the server. The input is the user's voice data, and the output is the text data and the text data sent to the server.
[1154] Step 12:
[1155] The server proposes exercise methods and dietary recommendations. The server retrieves health records and past exercise history from a database, and generates optimal recommendations using a generative AI model that takes into account the results of emotion analysis. The inputs are health data, past exercise history, and emotional state data, and the output is the generated recommendations.
[1156] Step 13:
[1157] The device converts the proposal content into voice and conveys it. The received proposal content is converted into voice using a speech synthesis engine and conveyed to the user. The input is the proposal content in text format, and the output is the proposal content in voice format.
[1158] Step 14:
[1159] The user asks a question about their family, for example, "I want to know about my grandchildren." The input is the user's voice data, and the output is the voice data itself.
[1160] Step 15:
[1161] The device converts the voice into text and sends it to the server. As in step 2, it uses voice recognition software and sends the converted text data to the server. The input is the user's voice data, and the output is the text data and the text data sent to the server.
[1162] Step 16:
[1163] The server retrieves family information and generates conversational responses. The server retrieves information about family members and relatives from a database, generates conversational responses using a generative AI model, and also generates a simple 3D avatar. The input is family information and emotional state data, and the output is the generated conversational responses and 3D avatar.
[1164] Step 17:
[1165] The device converts the response into speech and displays a 3D avatar. The received conversational response is converted into speech using a speech synthesis engine and then displays a 3D avatar. The input is text-format conversational response and 3D avatar data, and the output is audio-format conversational response and display of a 3D avatar. The response is communicated to the user through the device's speaker and display.
[1166] (Application example 2)
[1167] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1168] Currently, there are systems that can retrieve past memories to improve the quality of life of dementia patients, but they lack personalized responses and suggestions that take into account the user's emotional state. Furthermore, there are no established methods for providing appropriate product recommendations or virtual shopping experiences. It is necessary to solve these problems and provide a more personalized experience.
[1169] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for processing a user's voice input, means for searching a database for the user's age information and related data, means for generating optimal conversational responses using a generative AI model, means for providing the generated conversational responses to the user, means for recognizing the user's emotions using an emotion engine, means for suggesting personalized products based on past purchase history and memories, means for displaying product information in real time using a virtual reality display, and means for conveying the generated suggestions to the user audio and visually. This enables personalized suggestions and responses that take the user's emotional state into consideration, making it possible to provide a more personalized and high-quality virtual shopping experience.
[1170] The "means for processing voice input" is a means for acquiring voice information from a user and converting the voice information into text data.
[1171] "Age information" is information about the user's time series, such as the user's date of birth and age, that is necessary to retrieve past memories.
[1172] A "means for searching from a database" is a means for finding and obtaining necessary data from a database that stores specific information.
[1173] A "generative AI model" is a collection of algorithms that utilizes artificial intelligence techniques to generate optimal conversational responses and suggestions based on input data.
[1174] The "means for generating a conversational response" is a means for generating an appropriate reply based on information obtained from the user.
[1175] The "means for providing" refers to a means for conveying the generated conversational response or suggestion content to the user.
[1176] An "emotion engine" is an algorithm that analyzes emotions from a user's voice and text data and recognizes emotional states such as joy, sadness, and anger.
[1177] "Purchase history" is a record of products that a user has purchased in the past, and is information that is used to make personalized product suggestions.
[1178] "Memories" are specific events or memories that a user has experienced in the past, and are the information that forms the basis for personalized responses and suggestions.
[1179] A "virtual reality display" is a display device that uses virtual reality technology to allow users to obtain information visually.
[1180] The "means for displaying product information in real time" refers to a means for visually presenting product information to the user immediately at the current time.
[1181] The "means of communicating to the user by voice and visual means" refers to means of communicating the generated information and proposal contents to the user by voice (auditory) and visual (visual).
[1182] This invention is a system that takes into account the user's emotional state, taps into past memories, and provides personalized product recommendations and virtual shopping experiences. The system combines an emotion engine and a generative AI model to achieve more personal and effective responses.
[1183] System configuration
[1184] The system of the present invention uses the following hardware and software:
[1185] Hardware: microphone, smart glasses or head-mounted display
[1186] software:
[1187] Speech recognition (Google Speech Recognition API)
[1188] Emotion Recognition (EmotionRecognizer library)
[1189] Product recommendation system (RecommendationEngine library)
[1190] Speech synthesis (pyttsx3 library)
[1191] VR display (vr_display module)
[1192] Overview of the process
[1193] 1. Voice input processing:
[1194] The user speaks to the system, making a request such as "Tell me about my recent purchases."
[1195] The device captures audio and converts it into text using the Google Speech Recognition API.
[1196] 2. Emotion Recognition with Emotion Engine:
[1197] The server analyzes the converted text data using EmotionRecognizer to recognize the user's emotions.
[1198] The perceived emotions are classified into emotional states such as joy, sadness, and anger.
[1199] 3. Recalling past memories and product suggestions:
[1200] The server retrieves the user's age information and past purchase history from a database and uses a generative AI model to generate optimal conversational responses and personalized product suggestions.
[1201] The suggestions are adjusted taking into account the emotions recognized by the emotion engine.
[1202] 4. Real-time visual and audio feedback:
[1203] Suggested product information and conversational responses are visually displayed to the user in real time using a VR display.
[1204] The device uses a speech synthesis engine (pyttsx3) to communicate the suggestions to the user aloud.
[1205] Specific examples
[1206] Example 1: Recalling past memories
[1207] The user says, "I want to talk about my old school days."
[1208] The device sends this request to a server, which looks up the user's demographic and related data and uses a generative AI model to generate a conversational response.
[1209] The server generates a question such as "What did you play at elementary school?" and the terminal conveys this to the user by voice.
[1210] Example 2: Product proposal
[1211] The user says, "Tell me about my recent purchases."
[1212] The device sends this request to the server, which retrieves the user's purchase history from a database, uses an emotion engine to recognize the user's happiness, and then suggests promotional items.
[1213] The suggestions are conveyed by a speech synthesis engine, such as "Here are some products you recently purchased. We also recommend these as new promotions," and are simultaneously displayed on the VR display.
[1214] Prompt Sentence Examples
[1215] Provide the following prompt as input to the system:
[1216] "When a user says, 'Tell me about the latte you just bought,' the emotion engine recognizes that the user is happy and suggests in a soft tone, 'That was a delicious latte. How about trying our new flavor next time as part of a special promotion?'"
[1217] In this way, the present invention uses a system that combines emotion engines to provide personalized conversational responses and product suggestions that take into account the user's emotional state, resulting in a more personalized and effective virtual shopping experience.
[1218] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1219] Step 1:
[1220] The user speaks, for example making a request such as "Tell me about my recent purchases," which is captured as voice input.
[1221] Step 2:
[1222] The device captures the user's voice. It takes in the voice data through the microphone and converts the voice into text using the Google Speech Recognition API. The input is voice data and the output is text data.
[1223] Step 3:
[1224] The server receives the converted text data. This text data is analyzed using the EmotionRecognizer library to recognize the user's emotions. The input is text data, and the output is the emotional state (joy, sadness, anger, etc.).
[1225] Step 4:
[1226] Based on the emotion recognition results, the server retrieves the user's age information and past purchase history from the database. The input is the user ID and emotional state, and the output is age information and purchase history.
[1227] Step 5:
[1228] Based on the generation information and purchase history acquired by the server, a generative AI model is used to generate optimal conversational responses and product suggestions. The responses and suggestions are adjusted taking into account the emotional state recognized by the emotion engine. The input is generational information, purchase history, and emotional state, and the output is the generated conversational responses or product suggestions.
[1229] Step 6:
[1230] The server sends the generated conversational responses and product suggestions to the device. The device receives this and converts the text data into speech using a speech synthesis engine (pyttsx3). It then visually displays the product information in real time using a VR display. The input is the generated conversational responses or product suggestions, and the output is audio and visual information.
[1231] Step 7:
[1232] The user confirms the information they have received through audio and visuals. For example, they may hear a voice message saying, "Here are the products you recently purchased. We also recommend these as new promotions," while product information is displayed on the VR display. This allows users to easily check products that are relevant to their interests.
[1233] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1234] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1235] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1236] [Third embodiment]
[1237] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1238] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1239] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1240] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1241] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1243] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1244] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1245] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1246] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1247] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1248] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1249] This invention is an AI chat system that elicits past memories in dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and also suggests exercise methods and dietary content tailored to the user's physical condition.
[1250] System Overview
[1251] 1. Voice input processing:
[1252] User: Uses the system and speaks to initiate a conversation about a past memory.
[1253] Terminal: Converts the voice into text data and sends the text data to the server.
[1254] 2. Drawer of past memories:
[1255] Server: Based on the received text data, the server searches the database for the user's age information, for example, information related to the user's childhood or school days.
[1256] Server: Using information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, utilizing materials such as popular items and events by decade, photos, and movie posters.
[1257] Server: Sends the generated conversation response to the terminal.
[1258] Terminal: The received response is conveyed to the user via voice, enabling conversations that evoke past memories.
[1259] 3. Exercise and dietary suggestions:
[1260] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[1261] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[1262] Server: Retrieves the received physical condition data and past exercise history from a database. Based on this, the server uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[1263] Server: Sends the generated proposals to the device.
[1264] Device: The received suggestions are communicated to the user via voice.
[1265] 4. Support for conversations with family and relatives:
[1266] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[1267] Terminal: Converts the voice into text data and sends the text data to the server.
[1268] Server: Retrieves information about family and relatives from a database, and uses a generative AI model to generate optimal conversation responses. A simple 3D avatar is also generated at the same time.
[1269] Server: Sends the generated conversation responses and 3D avatars to the device.
[1270] Terminal: Assists with everyday conversations by conveying received responses to the user via voice and displaying a 3D avatar.
[1271] Specific examples
[1272] Example 1: Recalling past memories
[1273] When a user says, "I want to talk about my school days,"
[1274] The terminal sends this request to the server.
[1275] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[1276] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[1277] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[1278] The user begins recounting past events to jog their memory, and the conversation continues.
[1279] Example 2: Exercise and dietary suggestions
[1280] When a user says, "Which stretch should I do today?"
[1281] The terminal sends this request to the server.
[1282] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[1283] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[1284] Example 3: Supporting conversations with family and relatives
[1285] When a user says, "I want to know about my grandchildren,"
[1286] The terminal sends this request to the server.
[1287] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[1288] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[1289] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[1290] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[1291] The processing flow will be explained below.
[1292] Ability to recall past memories
[1293] Step 1:
[1294] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[1295] Step 2:
[1296] The terminal converts the user's voice into text data and transmits the text data to the server.
[1297] Step 3:
[1298] The server receives the user's request and searches the database for the user's age information and related data.
[1299] Step 4:
[1300] The server uses a generative AI model to generate a conversational response such as "What did you play in elementary school?" based on age information and related data obtained from the database.
[1301] Step 5:
[1302] The server transmits the generated conversation response to the terminal.
[1303] Step 6:
[1304] The terminal converts the response received from the server into voice and conveys it to the user.
[1305] Suggestions for exercise and diet
[1306] Step 1:
[1307] The user speaks to the device, "Which stretch should I do today?"
[1308] Step 2:
[1309] The terminal converts the user's voice into text data and transmits the text data to the server.
[1310] Step 3:
[1311] The server retrieves the user's physical condition record and past exercise history from a database.
[1312] Step 4:
[1313] Based on the acquired physical condition records and exercise history, the server uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[1314] Step 5:
[1315] The server sends the generated proposal to the terminal.
[1316] Step 6:
[1317] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1318] Support for conversations with family and relatives
[1319] Step 1:
[1320] The user speaks to the terminal saying, "I want to know about my grandchildren."
[1321] Step 2:
[1322] The terminal converts the user's voice into text data and transmits the text data to the server.
[1323] Step 3:
[1324] The server retrieves information about the user's family and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[1325] Step 4:
[1326] Based on the acquired information, the server uses a generative AI model to generate a conversational response such as, "Do you know what your grandchild is currently learning at school?" It also generates a simple 3D avatar.
[1327] Step 5:
[1328] The server sends the generated conversational responses and 3D avatars to the device.
[1329] Step 6:
[1330] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[1331] Proposing exercise methods and meal plans tailored to the user's physical condition
[1332] Step 1:
[1333] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[1334] Step 2:
[1335] The terminal converts the physical condition data into text data and transmits the text data to a server.
[1336] Step 3:
[1337] The server updates the database with the newly received physical condition data.
[1338] Step 4:
[1339] The server uses a generative AI model based on the received health data and previously stored data to generate suggestions such as, "We recommend that you eat something that is easy to digest today."
[1340] Step 5:
[1341] The server sends the generated proposal to the terminal.
[1342] Step 6:
[1343] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1344] The above is a specific description of the processing flow.
[1345] Example 1
[1346] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1347] Today, there is a lack of appropriate systems to help dementia patients recall their past memories and support their daily communication. Furthermore, there are no suggestions for exercise or dietary content tailored to their physical condition, which results in a decline in their quality of life. Furthermore, communication with family and relatives is often difficult, leading to feelings of isolation.
[1348] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1349] In this invention, the server includes means for receiving a user's voice input and processing the voice data, means for converting the voice data into text data, means for searching a database for the user's age information and related data, means for using a generative AI model to set prompts for generating optimal conversational responses and generating response messages, means for providing the generated response messages to the user, means for referencing the user's physical condition records and past exercise history and using a generative AI model to suggest optimal exercise methods and dietary content based on the user's physical condition records and past exercise history, and means for providing the generated suggestions to the user, means for acquiring information about the user's family and relatives from a database and using a generative AI model to generate appropriate response messages based on the acquired information, and means for generating a simple 3D avatar and providing the generated 3D avatar and response message to the user. This effectively draws out the past memories of dementia patients to support conversations, enables suggestions for exercise methods and dietary content tailored to their physical condition, and enables smooth communication with family and relatives.
[1350] A "user" is an entity that uses the system to provide voice input and receive conversations and suggestions.
[1351] "Voice input" refers to voice data input by a user speaking to the system.
[1352] "Voice data" refers to a digital recording of a user's voice input.
[1353] "Text data" refers to data obtained by converting voice data into characters.
[1354] "Age information" refers to data related to a user's year of birth or a particular age group.
[1355] A "database" is a system for storing and managing related data such as age and family information.
[1356] A "generative AI model" is an artificial intelligence algorithm that generates optimal responses or suggestions based on input data.
[1357] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[1358] A "response message" is a conversational response generated by a generative AI model.
[1359] "Physical condition record" refers to a record that accumulates data related to the user's physical condition.
[1360] "Exercise history" refers to data relating to exercises that a user has performed in the past.
[1361] An "exercise method" is an exercise technique suggested based on the user's physical condition and exercise history.
[1362] "Meal content" refers to an appropriate meal plan suggested to the user.
[1363] "Family information" refers to data about the user's family and relatives.
[1364] A "3D avatar" is a three-dimensional virtual character generated on a computer.
[1365] This invention relates to an AI chat system that supports daily communication for dementia patients and improves their quality of life. This system provides comprehensive support by evoking the user's past memories, engaging in conversations based on those memories, and suggesting exercise methods and dietary content tailored to the user's physical condition.
[1366] Hardware and software used
[1367] Audio input device (microphone): Receives the user's voice and records the audio data.
[1368] Speech recognition software (e.g., Google Speech-to-Text API): converts voice data into text data.
[1369] Database system: Stores data such as the user's age, family information, health records, and exercise history.
[1370] Generative AI models (e.g., OpenAI's GPT-3): Set prompts based on received text data and generate appropriate conversational responses and suggestions.
[1371] Speech synthesis software (e.g., Amazon Polly): Converts the generated response message into audio data.
[1372] 3D avatar generation software: Generates a 3D avatar based on the user's family information.
[1373] Specific explanation of the system
[1374] 1. Voice input processing
[1375] The user speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[1376] The device receives the user's voice using a high-performance microphone and records it as voice data.
[1377] The terminal converts the recorded voice data into text data using voice recognition software and transmits the text data to the server.
[1378] 2. Database search
[1379] The server analyzes the received text data and extracts keywords such as "school days."
[1380] The server searches the database for the user's generational information, and if the user was a student in the 1960s, for example, it retrieves data on trends and events from that era.
[1381] 3. Response generation using generative AI models
[1382] The server uses a generative AI model to set prompts based on information retrieved from the database, such as "You were a student in the 1960s. Do you have any memorable experiences?"
[1383] The server sends the generated response message to the user's terminal.
[1384] 4. Providing a response message
[1385] The terminal converts the received response message into voice data using voice synthesis software.
[1386] The terminal reproduces the audio data from a speaker and conveys it to the user.
[1387] Specific examples
[1388] Example 1: Recalling past memories
[1389] When a user says, "I want to talk about my school days,"
[1390] The device records this audio and converts it into text using speech recognition software.
[1391] The server receives this text data and extracts keywords such as "school days."
[1392] The server retrieves the user's generational information from a database and researches trends and events from, for example, the 1960s.
[1393] The server uses a generative AI model to set a prompt sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?" and generates the optimal response message.
[1394] The server sends this response message to the terminal,
[1395] The device synthesizes the received message into voice and plays it back through the speaker, telling the user, "You were a student in the 1960s. Do you have any memorable experiences?"
[1396] Example 2: Exercise and dietary suggestions
[1397] When a user says, "Which stretch should I do today?"
[1398] The device records this audio and converts it into text using speech recognition software.
[1399] The server receives this text data and extracts keywords such as "stretch."
[1400] The server retrieves the user's physical condition records and past exercise history from a database and uses a generative AI model to suggest the optimal exercise method.
[1401] The server sets a prompt sentence such as "Today, we recommend neck rotation exercises as a light stretch," and generates a suggestion.
[1402] The server sends this proposal to the terminal,
[1403] The device synthesizes the received suggestions into voice and plays them back through the speaker, telling the user, "Today, I recommend doing some neck rotation exercises as a light stretch."
[1404] Example 3: Supporting conversations with family and relatives
[1405] When a user says, "I want to know about my grandchildren,"
[1406] The device records this audio and converts it into text using speech recognition software.
[1407] The server receives this text data and extracts keywords such as "grandchild."
[1408] The server retrieves information about grandchildren from a database and uses a generative AI model to generate the question, "Do you know what your grandchildren are learning in school right now?"
[1409] The server also simultaneously generates a simple 3D avatar and sends it to the user's device.
[1410] The device synthesizes the received question into speech, plays it back through the speaker, and displays a 3D avatar, asking the user, "Do you know what your grandchild is learning at school right now?"
[1411] In this way, the present invention improves the quality of life of dementia patients by effectively drawing out the user's past memories and making suggestions tailored to everyday conversations and physical condition.
[1412] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1413] Step 1: Receiving voice input
[1414] User: Speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[1415] Terminal: The user's voice is received by a high-performance microphone and recorded as audio data, which becomes the input data.
[1416] Terminal: The recorded audio data is sent to the server. The output at this stage is audio data.
[1417] Step 2: Converting audio data to text data
[1418] Server: Converts the received voice data into text data using speech recognition software (e.g., Google Speech-to-Text API). This is the input data.
[1419] Server: The converted text data is used in the next processing step. The output is text data.
[1420] Step 3: Database search
[1421] Server: Analyzes the received text data and extracts keywords such as "school days." This is the input data.
[1422] Server: Searches the database for the user's generation information and retrieves data on trends and events from the corresponding era. This data processing and calculation is used to reference and search for generation information. The output is data from the corresponding era.
[1423] Step 4: Generative AI model generates a response
[1424] Server: Based on the information obtained from the database, a generative AI model (e.g., OpenAI's GPT-3) is used to set a prompt. For example, the prompt may be a text sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?"
[1425] Server: Using this prompt as input, the generative AI model generates the optimal conversational response. This response generation is performed using data calculations.
[1426] Server: Output the generated response message as text data for use in the next step.
[1427] Step 5: Send a response message
[1428] Server: Sends the generated response message to the user's terminal. This is the input data.
[1429] Terminal: Receives the response message and converts it into voice data using speech synthesis software (e.g., Amazon Polly). Voice synthesis is a type of data processing. The output is voice data.
[1430] Step 6: Play greetings
[1431] Terminal: The audio data is played back through the speaker and transmitted to the user. This is the final output.
[1432] Example flow
[1433] User: Say, "Which stretch should I do today?"
[1434] Device: Uses voice recognition software to convert the question "Which stretch should I do today?" into text data and send it to the server.
[1435] Server: Analyzes the text data and extracts keywords.
[1436] Server: Retrieves the user's physical condition records and past exercise history from the database.
[1437] Server: Using a generative AI model, it generates a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[1438] Server: Sends the proposal to the device.
[1439] Device: The suggested message is generated through speech synthesis and played from the speaker. It tells the user, "Today, we recommend you do some light stretching exercises, such as neck rotations."
[1440] In this way, the processing steps from the user's voice input to the final playback of the response message are concretely carried out.
[1441] (Application example 1)
[1442] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1443] This invention relates to an AI chat system that evokes past memories in dementia patients and improves their quality of life. Currently, many nursing homes lack support for retrieving individual memories for dementia patients, resulting in insufficient mental stability and communication. It is also difficult to provide dementia patients with appropriate exercise and dietary recommendations tailored to their physical condition, and communication with their families is also difficult. To address these issues, a system is needed that effectively engages dementia patients in conversation and suggests health management strategies based on their past living conditions.
[1444] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1445] In this invention, the server includes a means for processing user voice input, a means for searching a database for user age information and related data, a means for generating optimal conversational responses using a generative AI model, a means for providing the generated conversational responses to the user by voice and eliciting past memories, a means for using the generative AI model to hold conversations based on the past living situations of dementia patients and support daily communication, and an application installed on a robot installed in a care facility, which makes it possible to increase the mental stability of dementia patients and provide effective communication support.
[1446] "Means for processing user voice input" refers to a function for receiving voice data uttered by a user, recognizing it, and converting it into text data.
[1447] "Means of searching for a user's age information and related data from a database" refers to the function of searching for and retrieving the necessary information from a database that holds information related to the user's age, interests, and past life.
[1448] "Means for generating optimal conversational responses using a generative AI model" refers to a function that uses AI technology based on text data to generate optimal responses for users.
[1449] "Means for providing the generated conversational response to the user by voice and eliciting past memories" refers to a function for providing the user with an automatically generated text response by voice output, thereby eliciting the user's past memories.
[1450] "A means of using a generative AI model to hold conversations based on the past living conditions of dementia patients and support daily communication" refers to a function that makes full use of AI technology to conduct conversations based on content related to the past experiences and lives of dementia patients, thereby promoting daily interaction.
[1451] "Applications installed on robots installed in nursing facilities" refers to application software that is installed on robot devices used in nursing facilities and that is used to interact with dementia patients.
[1452] "A means for referring to the user's health records and past exercise history, and proposing optimal exercise methods and dietary content based on the generated health data" refers to a function that refers to the user's previous health information and exercise data, and provides appropriate exercise and dietary advice based on the user's current health condition.
[1453] "Means for providing the generated suggestions to the user by voice" refers to a function for conveying the generated exercise and diet suggestions to the user by voice.
[1454] "Means for a robot installed in a care facility to demonstrate suggested content to a user" refers to a function in which a robot demonstrates appropriate exercise methods and other instructions to a user.
[1455] "Means of obtaining information about the user's family and relatives from a database and generating a simple 3D avatar based on that information" refers to a function that obtains data about the user's family and relatives from a database and generates an avatar based on that data.
[1456] "Means of using the generated 3D avatar to engage in everyday conversation with the user and provide supplementary animations" refers to the function of using the generated 3D avatar as an interface to assist in conversation with the user and display animations as a visual aid.
[1457] "Means for providing information about family members and relatives to the user by voice" refers to a function that conveys information about family members and relatives obtained from a database to the user by voice.
[1458] This invention is an AI chat system that elicits past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary recommendations tailored to the user's physical condition. This system is implemented as an application installed on a robot in a nursing home.
[1459] System Overview
[1460] 1. Voice input processing
[1461] The user uses the system to speak to initiate a conversation about a past memory.
[1462] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[1463] 2. Drawer of past memories
[1464] The server searches a database for information about the user's age based on the received text data, such as information about the user's childhood or school days.
[1465] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[1466] The server transmits the generated conversation response to the terminal.
[1467] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[1468] 3. Exercise and dietary suggestions
[1469] The user reports their physical condition to the device, for example, by saying, "I feel a little tired today."
[1470] The terminal converts the physical condition data into text data and transmits the text data to a server.
[1471] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[1472] The server sends the generated proposal to the terminal.
[1473] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[1474] 4. Support for conversations with family and relatives
[1475] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[1476] The terminal converts the voice into text data and transmits the text data to the server.
[1477] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[1478] The server sends the generated conversational responses and 3D avatars to the terminal.
[1479] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[1480] Specific examples
[1481] Example 1: Recalling past memories
[1482] The user says, "I want to talk about my old school days."
[1483] The terminal sends this request to the server.
[1484] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[1485] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[1486] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[1487] The user begins recounting past events to jog their memory, and the conversation continues.
[1488] Example prompt sentence:
[1489] User: I want to talk about my old school days.
[1490] AI: What did you play in elementary school?
[1491] Example 2: Exercise and diet suggestions
[1492] The user says, "Which stretch should I do today?"
[1493] The terminal sends this request to the server.
[1494] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[1495] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[1496] If necessary, the robot will demonstrate neck rotation.
[1497] Example prompt sentence:
[1498] User: Which stretches should I do today?
[1499] AI: Today, I recommend some gentle stretching exercises, such as neck rotations.
[1500] Example 3: Supporting conversations with family and relatives
[1501] The user says, "I want to know about my grandchildren."
[1502] The terminal sends this request to the server.
[1503] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[1504] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[1505] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[1506] Example prompt sentence:
[1507] User: I want to know about my grandchildren
[1508] AI: Do you know what your grandchildren are learning in school right now?
[1509] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[1510] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1511] Step 1:
[1512] The user uses the system to speak to initiate a conversation about a past memory.
[1513] Input: User's voice data
[1514] How it works: A microphone on a robot installed in a care home captures voice input.
[1515] Step 2:
[1516] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[1517] Input: User's voice data
[1518] What it does: It uses speech recognition software (e.g., the speech_recognition library) to convert speech to text, which is then sent over the Internet to a server.
[1519] Output: Text data sent to the server
[1520] Step 3:
[1521] The server searches the database for the user's age information based on the received text data.
[1522] Input: User's text data
[1523] What it does: The server performs a database query to retrieve data about the user's age and related past memories. For example, if the user types "I want to talk about my old school days," it searches for information related to the user's childhood and school days.
[1524] Output: User age and related data
[1525] Step 4:
[1526] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[1527] Input: User demographics and related data
[1528] Specific operation: Query a generative AI model (e.g., OpenAI GPT-3) with a prompt sentence to generate an appropriate conversational response. Reference data such as trends and events relevant to each generation are also used as reference.
[1529] Output: Generated conversation response
[1530] Step 5:
[1531] The server transmits the generated conversation response to the terminal.
[1532] Input: Generated conversation response
[1533] Specific operation: The response text generated by the server is sent to the terminal via the Internet.
[1534] Output: Conversation response sent to the terminal
[1535] Step 6:
[1536] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[1537] Input: Conversation response sent by the server
[1538] What it does: It uses speech synthesis software (e.g., the pyttsx3 library) to play the text data as speech, allowing the user to continue the conversation by listening to the voice response.
[1539] Output: The audio response provided to the user
[1540] Step 7:
[1541] The user reports their physical condition to the device, for example, saying, "I feel a little tired today."
[1542] Input: User's voice data
[1543] What happens: The device's microphone captures audio input.
[1544] Step 8:
[1545] The terminal converts the physical condition data into text data and transmits the text data to a server.
[1546] Input: Audio data about the user's physical condition
[1547] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[1548] Output: Text data about your health condition sent to the server
[1549] Step 9:
[1550] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[1551] Input: Text data about the user's physical condition and past exercise history
[1552] Specific operation: The server retrieves the user's past exercise history from the database and uses a generative AI model to suggest appropriate exercise methods and dietary recommendations.
[1553] Output: Generated exercise and diet suggestions
[1554] Step 10:
[1555] The server sends the generated proposal to the terminal.
[1556] Input: Generated proposal
[1557] Specific operation: The proposal content is sent from the server to the device via the Internet.
[1558] Output: Suggestions sent to the device
[1559] Step 11:
[1560] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[1561] Input: Suggestion from the server
[1562] Specific behavior: Using speech synthesis software, the robot plays back the suggestions as voice, and shows the user the appropriate exercise method.
[1563] Output: Audio suggestions and demonstrations provided to the user
[1564] Step 12:
[1565] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[1566] Input: User's voice data
[1567] What happens: The device's microphone captures audio input.
[1568] Step 13:
[1569] The terminal converts the voice into text data and transmits the text data to the server.
[1570] Input: User's voice data
[1571] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[1572] Output: Text data sent to the server
[1573] Step 14:
[1574] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[1575] Input: User's text data
[1576] How it works: The server performs database queries to retrieve information about family and relatives, then uses a generative AI model to generate optimal conversational responses, while simultaneously generating a simple 3D avatar.
[1577] Output: Generated conversational responses and 3D avatars
[1578] Step 15:
[1579] The server sends the generated conversational responses and 3D avatars to the terminal.
[1580] Input: Generated conversational responses and 3D avatars
[1581] Specific operation: The server generates a response text and sends a 3D avatar to the device via the Internet.
[1582] Output: Speech responses and 3D avatar sent to device
[1583] Step 16:
[1584] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[1585] Input: Conversational responses and 3D avatars sent from the server
[1586] Specific operations: Uses speech synthesis software to play text data as speech, and displays a 3D avatar on a display device to assist in conversation with the user.
[1587] Output: Audio response and 3D avatar representation provided to the user
[1588] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1589] This invention is an AI chat system that draws out past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to the user's physical condition. In addition, by combining it with an emotion engine that recognizes the user's emotions, more personalized responses and suggestions are possible.
[1590] System Overview
[1591] 1. Voice input processing:
[1592] User: Uses the system to initiate a conversation about a past memory.
[1593] Terminal: Converts the voice into text data and sends the text data to the server.
[1594] 2. Emotion Recognition with Emotion Engine:
[1595] Server: Analyzes the received text data and recognizes the user's emotions using an emotion engine.
[1596] The emotion engine analyzes emotions from the user's voice input and text data to determine emotional states such as joy, sadness, and anger.
[1597] 3. Drawer of past memories:
[1598] Server: Retrieves the user's age information and related data from a database, for example, retrieves information related to the user's childhood or school days.
[1599] Server: Based on information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, adjusting the content of the conversation based on emotions recognized by the emotion engine.
[1600] Server: Sends the generated conversation response to the terminal.
[1601] Terminal: Converts the received response into speech and conveys it to the user.
[1602] 4. Exercise and dietary suggestions:
[1603] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[1604] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[1605] Server: Retrieves the user's physical condition records and past exercise history from a database, and uses a generative AI model to suggest optimal exercise methods and dietary recommendations. The server also adjusts the suggestions by taking into account emotions recognized by the emotion engine.
[1606] Server: Sends the generated proposals to the device.
[1607] Device: Converts the received suggestions into speech and conveys them to the user.
[1608] 5. Support for conversations with family and relatives:
[1609] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[1610] Terminal: Converts the voice into text data and sends the text data to the server.
[1611] Server: Retrieves information about family and relatives from a database and uses a generative AI model to generate optimal conversational responses. A simple 3D avatar is also generated at the same time. The conversational responses are adjusted based on emotions recognized by the emotion engine.
[1612] Server: Sends the generated conversation responses and 3D avatars to the device.
[1613] Terminal: Converts the received response into voice and displays a 3D avatar to communicate with the user.
[1614] Specific examples
[1615] Example 1: Recalling past memories
[1616] When a user says, "I want to talk about my school days,"
[1617] The terminal sends this request to the server.
[1618] The server retrieves the user's age and related data from the database. For example, if the user is having a good time, it generates questions about the pastimes and movies they enjoyed at the time.
[1619] The server uses a generative AI model based on information retrieved from the database to tailor the conversational response based on the emotions recognized by the emotion engine, generating the question, "What did you play at elementary school?"
[1620] The server sends the generated question to the terminal, which then transmits it to the user as voice.
[1621] The user begins to recount past events to jog their memory, and the conversation continues.
[1622] Example 2: Exercise and dietary suggestions
[1623] When a user says, "Which stretch should I do today?"
[1624] The terminal sends this request to the server.
[1625] The server references the user's physical condition records and past exercise history, and uses a generative AI model to determine the optimal exercise method and diet. The suggestions are adjusted based on the emotions recognized by the emotion engine. For example, if the user is tired, the server might suggest, "Today, we recommend neck rotation exercises as a light stretch."
[1626] The server transmits the generated proposal to the terminal, which then conveys it to the user as voice.
[1627] Example 3: Supporting conversations with family and relatives
[1628] When a user says, "I want to know about my grandchildren,"
[1629] The terminal sends this request to the server.
[1630] The server uses a generative AI model to generate conversational responses based on information about the grandchild (such as name, school, and interests) retrieved from a database. The conversation content is adjusted based on emotions recognized by an emotion engine. A simple 3D avatar is also generated.
[1631] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[1632] In this way, the present invention provides a concrete means to improve the quality of life of dementia patients and reduce the burden on caregivers. By combining it with an emotion engine, personalized responses and suggestions that take into account the user's emotional state become possible, realizing more personal and effective support.
[1633] The processing flow will be explained below.
[1634] Ability to recall past memories
[1635] Step 1:
[1636] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[1637] Step 2:
[1638] The terminal converts the user's voice into text data and transmits the text data to the server.
[1639] Step 3:
[1640] The server analyzes the text data and uses an emotion engine to recognize the user's emotions, such as whether they are happy or sad.
[1641] Step 4:
[1642] The server retrieves the user's age information and related data from a database, for example, information related to the user's childhood or school days.
[1643] Step 5:
[1644] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database. The content of the conversation is adjusted based on emotions recognized by the emotion engine. For example, if the user seems to be having fun, the server generates a question such as, "Did you have any happy memories from that time?"
[1645] Step 6:
[1646] The server transmits the generated conversation response to the terminal.
[1647] Step 7:
[1648] The terminal converts the response received from the server into voice and conveys it to the user.
[1649] Suggestions for exercise and diet
[1650] Step 1:
[1651] The user speaks to the device, "Which stretch should I do today?"
[1652] Step 2:
[1653] The terminal converts the user's voice into text data and transmits the text data to the server.
[1654] Step 3:
[1655] The server retrieves the user's physical condition record and past exercise history from a database.
[1656] Step 4:
[1657] The server uses a generative AI model to suggest optimal exercise methods and dietary recommendations based on the user's physical condition data and past exercise history. The server also adjusts the recommendations by taking into account the user's emotional state, as recognized by the emotion engine. For example, if the user expresses feelings of fatigue, the server might suggest, "We recommend getting plenty of rest and doing some light stretching today."
[1658] Step 5:
[1659] The server sends the generated proposal to the terminal.
[1660] Step 6:
[1661] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1662] Support for conversations with family and relatives
[1663] Step 1:
[1664] The user speaks to the terminal saying, "I want to know about my grandchildren."
[1665] Step 2:
[1666] The terminal converts the user's voice into text data and transmits the text data to the server.
[1667] Step 3:
[1668] The server analyzes the text data and uses an emotion engine to recognize the user's emotions, for example, capturing the joy and excitement when the user talks about their grandchildren.
[1669] Step 4:
[1670] The server retrieves information about family members and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[1671] Step 5:
[1672] The server uses a generative AI model based on the acquired information to generate optimal conversational responses based on the emotions recognized by the emotion engine. For example, if the user is excited about their grandchild, the server generates a question such as, "What is your grandchild's favorite pastime these days?"
[1673] Step 6:
[1674] The server also simultaneously generates a simple 3D avatar and sends a conversation response including this avatar to the terminal.
[1675] Step 7:
[1676] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[1677] Proposing exercise methods and meal plans tailored to the user's physical condition
[1678] Step 1:
[1679] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[1680] Step 2:
[1681] The terminal converts the physical condition data into text data and transmits the text data to a server.
[1682] Step 3:
[1683] The server updates the database with the newly received physical condition data.
[1684] Step 4:
[1685] The server uses a generative AI model to suggest optimal exercise methods and meal plans based on the received physical condition data and previously stored data. The suggestions are adjusted taking into account the emotions recognized by the emotion engine. For example, if the user expresses feelings of fatigue, the server generates a suggestion such as, "Today's meal is recommended to help you relax."
[1686] Step 5:
[1687] The server sends the generated proposal to the terminal.
[1688] Step 6:
[1689] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1690] The above is a description of the process flow broken down into specific steps.
[1691] Example 2
[1692] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1693] In modern society, the number of dementia patients is increasing, and methods to improve their quality of life are needed. In particular, it is necessary to enrich the lives of dementia patients and reduce the burden on caregivers by evoking past memories and providing personalized responses based on emotions. However, conventional technologies lack the accuracy of voice input and the provision of personalized responses, making it difficult to provide appropriate support tailored to the specific needs of dementia patients.
[1694] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for processing a user's voice input, means for converting the user's voice input into text, means for transmitting the text data to the server, means for receiving the text data at the server and recognizing emotions with an emotion engine, means for searching a database for user age information and related data, means for generating optimal conversational responses based on the results of the emotion engine using a generative AI model, means for transmitting the generated conversational responses from the server to the terminal, means for converting the conversational responses received by the terminal into speech, and means for providing the generated speech to the user. This enables personalized responses and suggestions based on the emotions and past memories of dementia patients, improving their quality of life and reducing the burden on caregivers.
[1695] The "means for processing user speech input" is a means for capturing speech from the user and sending it to the next processing step.
[1696] "Means for converting user voice input into text" refers to a speech recognition system for converting voice data into text data.
[1697] "Means for transmitting text data to a server" refers to a communication means for transmitting the converted text data to a server via a network.
[1698] "Means for recognizing emotions using an emotion engine" refers to an emotion analysis system for analyzing and recognizing a user's emotional state from text data.
[1699] "Means for searching a database for user generation information and related data" refers to a search function for obtaining a user generation information and related data from a database.
[1700] "Means for generating optimal conversational responses based on the results of an emotion engine using a generative AI model" refers to a system for generating optimal responses based on the results of emotion analysis using an AI model.
[1701] The "means for transmitting the generated conversation response from the server to the terminal" refers to a communication means for transmitting the generated conversation response from the server to the terminal.
[1702] "Means for converting the conversational response received by the terminal into speech" refers to a speech synthesis system that converts the received text-based conversational response into speech.
[1703] The "means for providing the generated audio to the user" refers to an output device such as a speaker or a headphone that allows the user to hear the generated audio.
[1704] This invention is an AI system that elicits past memories in dementia patients and improves their quality of life. The system aims to evoke the user's past memories, support daily communication, and suggest exercise methods and dietary content tailored to the user's physical condition. In addition, by combining it with an emotion engine that recognizes the user's emotions, more personalized responses and suggestions are possible.
[1705] System configuration
[1706] The system uses the following hardware and software:
[1707] A microphone to capture the user's voice input
[1708] Speech recognition software (e.g., Google Cloud Speech-to-Text API)
[1709] Sentiment analysis software (e.g., IBM Watson Tone Analyzer)
[1710] A database (e.g., MySQL or MongoDB)
[1711] Generative AI models (e.g., OpenAI GPT-3)
[1712] A speech synthesis engine (e.g., Amazon Polly)
[1713] Speakers or headphones for audio output to the user
[1714] System Operation
[1715] Voice Input Processing
[1716] When a user speaks to the system, the device's microphone captures the speech, and speech recognition software converts the speech data into text data, which is then sent to the server.
[1717] Emotion recognition by emotion engine
[1718] The server inputs the received text data into emotion analysis software to analyze the user's emotions. The emotion analysis software analyzes keywords and context within the text to recognize emotions such as joy, sadness, and anger.
[1719] Drawer of past memories
[1720] The server searches a database for the user's age information and related data. For example, it obtains information about the user's childhood and school days. Based on the obtained information and the results of emotion analysis, the server inputs a prompt sentence into the generative AI model to generate an optimal conversational response. An example of a prompt sentence is, "The user wants to talk about their school days. They appear quietly pleased." The generated conversational response is sent to the device, which then converts it into speech using a speech synthesis engine and conveys it to the user.
[1721] Suggestions for exercise and diet
[1722] When a user reports their physical condition, the device converts it into text and sends it to a server. The server retrieves the user's physical condition records and past exercise history from a database, and, taking into account the results of emotion analysis, uses a generative AI model to suggest optimal exercise methods and dietary content. For example, a suggestion might be generated such as, "Today, we recommend neck rotation exercises as a light stretch." The generated suggestion is sent to the device, which converts it into voice and conveys it to the user.
[1723] Support for conversations with family and relatives
[1724] When a user wants to know about their family or relatives, they speak their request into the device. The device converts the voice input into text and sends it to the server. The server retrieves information about family and relatives from a database and generates a conversational response using a generative AI model based on that information. A simple 3D avatar is also generated. For example, in response to a request such as "I want to know about my grandchildren," a response such as "Your grandchild is currently 8 years old and loves soccer" is generated. The generated conversational response and 3D avatar are sent to the device, which converts it into speech and displays the 3D avatar to communicate with the user.
[1725] Specific examples
[1726] Below are some specific examples of how the system can be used.
[1727] Example 1: Recalling past memories
[1728] When a user says, "I want to talk about my school days," the device sends this request to the server. The server retrieves the user's age information and related data from a database, and generates a question based on the results of emotion analysis: "What did you play in elementary school?" The device converts this question into speech and conveys it to the user, who then begins talking about past events to jog their memories.
[1729] Example 2: Exercise and diet suggestions
[1730] When the user says, "Which stretch should I do today?", the device sends this request to the server. The server retrieves physical condition records and past exercise history from a database, and, taking into account the results of emotion analysis, suggests, "Today, I recommend neck rotations as a light stretch." The device converts this into voice and conveys it to the user.
[1731] Example 3: Supporting conversations with family and relatives
[1732] When a user says, "I want to know about my grandchild," the device sends this request to the server. The server retrieves information about the family from a database and generates a 3D avatar along with the response, "Your grandchild is now 8 years old and loves soccer." The device then provides this as voice and displays the 3D avatar.
[1733] As described above, the system of the present invention can provide personalized responses and suggestions based on the user's emotional state and past memories to improve the quality of life of dementia patients.
[1734] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1735] Step 1:
[1736] A user speaks aloud. For example, the user says, "I'd like to talk about my old school days." The input is the user's voice data, and the output is the voice data itself.
[1737] Step 2:
[1738] The device converts speech to text. The device's microphone captures the user's speech data and uses speech recognition software (e.g., Google Cloud Speech-to-Text API) to convert the speech to text data. The input is the user's speech data, and the output is text data. Specifically, the speech recognition software performs phonemic analysis and temporarily stores the recognition results as text on the device.
[1739] Step 3:
[1740] The terminal sends text data to the server. The converted text data is sent to the server using the HTTPS protocol. The input is text data, and the output is the text data sent to the server.
[1741] Step 4:
[1742] The server receives text data. The server receives text data sent from the terminal. The input is the text data sent from the terminal, and the output is the text data saved on the server.
[1743] Step 5:
[1744] The server recognizes emotions using an emotion engine. The server inputs the received text data into emotion analysis software (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The input is text data, and the output is data indicating the user's emotional state (e.g., joy, sadness, anger, etc.). Specifically, the emotion analysis engine analyzes keywords and context within the text and recognizes the emotional state numerically and categorically.
[1745] Step 6:
[1746] The server searches for the user's past data. The server retrieves the user's age information and related data from a database (e.g., MySQL or MongoDB). The input is the user's profile information, and the output is the retrieved past data. Specifically, the server issues an SQL query to retrieve the required information.
[1747] Step 7:
[1748] The server generates a response using a generative AI model. Based on the acquired past data and the results of emotion analysis, the server inputs a prompt into the generative AI model (e.g., OpenAI GPT-3) to generate an optimal conversational response. The input is past data and emotional state data, and the output is the generated conversational response. An example of a specific prompt is, "The user wants to talk about their school days. They appear quietly pleased."
[1749] Step 8:
[1750] The server sends the generated response to the terminal. The generated conversation response is sent to the terminal using the HTTPS protocol. The input is the generated conversation response, and the output is the conversation response sent to the terminal.
[1751] Step 9:
[1752] The device converts the response into speech and transmits it. The device converts the received conversational response into speech using a speech synthesis engine (e.g., Amazon Polly) and transmits it to the user. The input is a text-format conversational response, and the output is a speech-format conversational response. Specifically, the audio is played back to the user through the device's speakers or headphones.
[1753] Step 10:
[1754] A user reports how they are feeling, for example, "I feel a little tired today." The input is the user's voice data, and the output is the voice data itself.
[1755] Step 11:
[1756] The device converts the voice into text and sends it to the server. As in step 2, it uses voice recognition software and sends the converted text data to the server. The input is the user's voice data, and the output is the text data and the text data sent to the server.
[1757] Step 12:
[1758] The server proposes exercise methods and dietary recommendations. The server retrieves health records and past exercise history from a database, and generates optimal recommendations using a generative AI model that takes into account the results of emotion analysis. The inputs are health data, past exercise history, and emotional state data, and the output is the generated recommendations.
[1759] Step 13:
[1760] The device converts the proposal content into voice and conveys it. The received proposal content is converted into voice using a speech synthesis engine and conveyed to the user. The input is the proposal content in text format, and the output is the proposal content in voice format.
[1761] Step 14:
[1762] The user asks a question about their family, for example, "I want to know about my grandchildren." The input is the user's voice data, and the output is the voice data itself.
[1763] Step 15:
[1764] The device converts the voice into text and sends it to the server. As in step 2, it uses voice recognition software and sends the converted text data to the server. The input is the user's voice data, and the output is the text data and the text data sent to the server.
[1765] Step 16:
[1766] The server retrieves family information and generates conversational responses. The server retrieves information about family members and relatives from a database, generates conversational responses using a generative AI model, and also generates a simple 3D avatar. The input is family information and emotional state data, and the output is the generated conversational responses and 3D avatar.
[1767] Step 17:
[1768] The device converts the response into speech and displays a 3D avatar. The received conversational response is converted into speech using a speech synthesis engine and then displays a 3D avatar. The input is text-format conversational response and 3D avatar data, and the output is audio-format conversational response and display of a 3D avatar. The response is communicated to the user through the device's speaker and display.
[1769] (Application example 2)
[1770] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1771] Currently, there are systems that can retrieve past memories to improve the quality of life of dementia patients, but they lack personalized responses and suggestions that take into account the user's emotional state. Furthermore, there are no established methods for providing appropriate product recommendations or virtual shopping experiences. It is necessary to solve these problems and provide a more personalized experience.
[1772] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for processing a user's voice input, means for searching a database for the user's age information and related data, means for generating optimal conversational responses using a generative AI model, means for providing the generated conversational responses to the user, means for recognizing the user's emotions using an emotion engine, means for suggesting personalized products based on past purchase history and memories, means for displaying product information in real time using a virtual reality display, and means for conveying the generated suggestions to the user audio and visually. This enables personalized suggestions and responses that take the user's emotional state into consideration, making it possible to provide a more personalized and high-quality virtual shopping experience.
[1773] The "means for processing voice input" is a means for acquiring voice information from a user and converting the voice information into text data.
[1774] "Age information" is information about the user's time series, such as the user's date of birth and age, that is necessary to retrieve past memories.
[1775] A "means for searching from a database" is a means for finding and obtaining necessary data from a database that stores specific information.
[1776] A "generative AI model" is a collection of algorithms that utilizes artificial intelligence techniques to generate optimal conversational responses and suggestions based on input data.
[1777] The "means for generating a conversational response" is a means for generating an appropriate reply based on information obtained from the user.
[1778] The "means for providing" refers to a means for conveying the generated conversational response or suggestion content to the user.
[1779] An "emotion engine" is an algorithm that analyzes emotions from a user's voice and text data and recognizes emotional states such as joy, sadness, and anger.
[1780] "Purchase history" is a record of products that a user has purchased in the past, and is information that is used to make personalized product suggestions.
[1781] "Memories" are specific events or memories that a user has experienced in the past, and are the information that forms the basis for personalized responses and suggestions.
[1782] A "virtual reality display" is a display device that uses virtual reality technology to allow users to obtain information visually.
[1783] The "means for displaying product information in real time" refers to a means for visually presenting product information to the user immediately at the current time.
[1784] The "means of communicating to the user by voice and visual means" refers to means of communicating the generated information and proposal contents to the user by voice (auditory) and visual (visual).
[1785] This invention is a system that takes into account the user's emotional state, taps into past memories, and provides personalized product recommendations and virtual shopping experiences. The system combines an emotion engine and a generative AI model to achieve more personal and effective responses.
[1786] System configuration
[1787] The system of the present invention uses the following hardware and software:
[1788] Hardware: microphone, smart glasses or head-mounted display
[1789] software:
[1790] Speech recognition (Google Speech Recognition API)
[1791] Emotion Recognition (EmotionRecognizer library)
[1792] Product recommendation system (RecommendationEngine library)
[1793] Speech synthesis (pyttsx3 library)
[1794] VR display (vr_display module)
[1795] Overview of the process
[1796] 1. Voice input processing:
[1797] The user speaks to the system, making a request such as "Tell me about my recent purchases."
[1798] The device captures audio and converts it into text using the Google Speech Recognition API.
[1799] 2. Emotion Recognition with Emotion Engine:
[1800] The server analyzes the converted text data using EmotionRecognizer to recognize the user's emotions.
[1801] The perceived emotions are classified into emotional states such as joy, sadness, and anger.
[1802] 3. Recalling past memories and product suggestions:
[1803] The server retrieves the user's age information and past purchase history from a database and uses a generative AI model to generate optimal conversational responses and personalized product suggestions.
[1804] The suggestions are adjusted taking into account the emotions recognized by the emotion engine.
[1805] 4. Real-time visual and audio feedback:
[1806] Suggested product information and conversational responses are visually displayed to the user in real time using a VR display.
[1807] The device uses a speech synthesis engine (pyttsx3) to communicate the suggestions to the user aloud.
[1808] Specific examples
[1809] Example 1: Recalling past memories
[1810] The user says, "I want to talk about my old school days."
[1811] The device sends this request to a server, which looks up the user's demographic and related data and uses a generative AI model to generate a conversational response.
[1812] The server generates a question such as "What did you play at elementary school?" and the terminal conveys this to the user by voice.
[1813] Example 2: Product proposal
[1814] The user says, "Tell me about my recent purchases."
[1815] The device sends this request to the server, which retrieves the user's purchase history from a database, uses an emotion engine to recognize the user's happiness, and then suggests promotional items.
[1816] The suggestions are conveyed by a speech synthesis engine, such as "Here are some products you recently purchased. We also recommend these as new promotions," and are simultaneously displayed on the VR display.
[1817] Prompt Sentence Examples
[1818] Provide the following prompt as input to the system:
[1819] "When a user says, 'Tell me about the latte you just bought,' the emotion engine recognizes that the user is happy and suggests in a soft tone, 'That was a delicious latte. How about trying a new flavor next time as part of a special promotion?'"
[1820] In this way, the present invention uses a system that combines emotion engines to provide personalized conversational responses and product suggestions that take into account the user's emotional state, resulting in a more personalized and effective virtual shopping experience.
[1821] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1822] Step 1:
[1823] The user speaks, for example making a request such as "Tell me about my recent purchases," which is captured as voice input.
[1824] Step 2:
[1825] The device captures the user's voice. It takes in the voice data through the microphone and converts the voice into text using the Google Speech Recognition API. The input is voice data and the output is text data.
[1826] Step 3:
[1827] The server receives the converted text data. This text data is analyzed using the EmotionRecognizer library to recognize the user's emotions. The input is text data, and the output is the emotional state (joy, sadness, anger, etc.).
[1828] Step 4:
[1829] Based on the emotion recognition results, the server retrieves the user's age information and past purchase history from the database. The input is the user ID and emotional state, and the output is age information and purchase history.
[1830] Step 5:
[1831] Based on the generation information and purchase history acquired by the server, a generative AI model is used to generate optimal conversational responses and product suggestions. The responses and suggestions are adjusted taking into account the emotional state recognized by the emotion engine. The input is generational information, purchase history, and emotional state, and the output is the generated conversational responses or product suggestions.
[1832] Step 6:
[1833] The server sends the generated conversational responses and product suggestions to the device. The device receives this and converts the text data into speech using a speech synthesis engine (pyttsx3). It then visually displays the product information in real time using a VR display. The input is the generated conversational responses or product suggestions, and the output is audio and visual information.
[1834] Step 7:
[1835] The user confirms the information they have received through audio and visuals. For example, they may hear a voice message saying, "Here are the products you recently purchased. We also recommend these as new promotions," while product information is displayed on the VR display. This allows users to easily check products that are relevant to their interests.
[1836] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1837] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1838] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1839] [Fourth embodiment]
[1840] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1841] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1842] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1843] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1844] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1845] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1846] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1847] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1848] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1849] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1850] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1851] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1852] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1853] This invention is an AI chat system that elicits past memories in dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and also suggests exercise methods and dietary content tailored to the user's physical condition.
[1854] System Overview
[1855] 1. Voice input processing:
[1856] User: Uses the system and speaks to initiate a conversation about a past memory.
[1857] Terminal: Converts the voice into text data and sends the text data to the server.
[1858] 2. Drawer of past memories:
[1859] Server: Based on the received text data, the server searches the database for the user's age information, for example, information related to the user's childhood or school days.
[1860] Server: Using information retrieved from the database, the server uses a generative AI model to generate optimal conversational responses, utilizing materials such as popular items and events by decade, photos, and movie posters.
[1861] Server: Sends the generated conversation response to the terminal.
[1862] Terminal: The received response is conveyed to the user via voice, enabling conversations that evoke past memories.
[1863] 3. Exercise and dietary suggestions:
[1864] User: Reports their physical condition to the device. For example, they say, "I'm feeling a little tired today."
[1865] Terminal: Converts the physical condition data into text data and sends the text data to the server.
[1866] Server: Retrieves the received physical condition data and past exercise history from a database. Based on this, the server uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[1867] Server: Sends the generated proposals to the device.
[1868] Device: The received suggestions are communicated to the user via voice.
[1869] 4. Support for conversations with family and relatives:
[1870] User: If you want to know information about your family or relatives, you can say to your device, for example, "I want to know about my grandchildren."
[1871] Terminal: Converts the voice into text data and sends the text data to the server.
[1872] Server: Retrieves information about family and relatives from a database, and uses a generative AI model to generate optimal conversation responses. A simple 3D avatar is also generated at the same time.
[1873] Server: Sends the generated conversation responses and 3D avatars to the device.
[1874] Terminal: Assists with everyday conversations by conveying received responses to the user via voice and displaying a 3D avatar.
[1875] Specific examples
[1876] Example 1: Recalling past memories
[1877] When a user says, "I want to talk about my school days,"
[1878] The terminal sends this request to the server.
[1879] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[1880] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[1881] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[1882] The user begins recounting past events to jog their memory, and the conversation continues.
[1883] Example 2: Exercise and dietary suggestions
[1884] When a user says, "Which stretch should I do today?"
[1885] The terminal sends this request to the server.
[1886] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[1887] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[1888] Example 3: Supporting conversations with family and relatives
[1889] When a user says, "I want to know about my grandchildren,"
[1890] The terminal sends this request to the server.
[1891] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[1892] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[1893] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[1894] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[1895] The processing flow will be explained below.
[1896] Ability to recall past memories
[1897] Step 1:
[1898] The user speaks to the terminal saying, "I'd like to talk about my old school days."
[1899] Step 2:
[1900] The terminal converts the user's voice into text data and transmits the text data to the server.
[1901] Step 3:
[1902] The server receives the user's request and searches the database for the user's age information and related data.
[1903] Step 4:
[1904] The server uses a generative AI model to generate a conversational response such as "What did you play in elementary school?" based on age information and related data obtained from the database.
[1905] Step 5:
[1906] The server transmits the generated conversation response to the terminal.
[1907] Step 6:
[1908] The terminal converts the response received from the server into voice and conveys it to the user.
[1909] Suggestions for exercise and diet
[1910] Step 1:
[1911] The user speaks to the device, "Which stretch should I do today?"
[1912] Step 2:
[1913] The terminal converts the user's voice into text data and transmits the text data to the server.
[1914] Step 3:
[1915] The server retrieves the user's physical condition record and past exercise history from a database.
[1916] Step 4:
[1917] Based on the acquired physical condition records and exercise history, the server uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[1918] Step 5:
[1919] The server sends the generated proposal to the terminal.
[1920] Step 6:
[1921] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1922] Support for conversations with family and relatives
[1923] Step 1:
[1924] The user speaks to the terminal saying, "I want to know about my grandchildren."
[1925] Step 2:
[1926] The terminal converts the user's voice into text data and transmits the text data to the server.
[1927] Step 3:
[1928] The server retrieves information about the user's family and relatives from a database, such as the names, ages, and recent activities of grandchildren.
[1929] Step 4:
[1930] Based on the acquired information, the server uses a generative AI model to generate a conversational response such as, "Do you know what your grandchild is currently learning at school?" It also generates a simple 3D avatar.
[1931] Step 5:
[1932] The server sends the generated conversational responses and 3D avatars to the device.
[1933] Step 6:
[1934] The device converts the response received from the server into voice and displays a 3D avatar to communicate with the user.
[1935] Proposing exercise methods and meal plans tailored to the user's physical condition
[1936] Step 1:
[1937] The user reports their physical condition for the day to the terminal, for example, by inputting "I feel a little tired today."
[1938] Step 2:
[1939] The terminal converts the physical condition data into text data and transmits the text data to a server.
[1940] Step 3:
[1941] The server updates the database with the newly received physical condition data.
[1942] Step 4:
[1943] The server uses a generative AI model based on the received health data and previously stored data to generate suggestions such as, "We recommend that you eat something that is easy to digest today."
[1944] Step 5:
[1945] The server sends the generated proposal to the terminal.
[1946] Step 6:
[1947] The terminal converts the proposal received from the server into voice and conveys it to the user.
[1948] The above is a specific description of the processing flow.
[1949] Example 1
[1950] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1951] Today, there is a lack of appropriate systems to help dementia patients recall their past memories and support their daily communication. Furthermore, there are no suggestions for exercise or dietary content tailored to their physical condition, which results in a decline in their quality of life. Furthermore, communication with family and relatives is often difficult, leading to feelings of isolation.
[1952] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1953] In this invention, the server includes means for receiving a user's voice input and processing the voice data, means for converting the voice data into text data, means for searching a database for the user's age information and related data, means for using a generative AI model to set prompts for generating optimal conversational responses and generating response messages, means for providing the generated response messages to the user, means for referencing the user's physical condition records and past exercise history and using a generative AI model to suggest optimal exercise methods and dietary content based on the user's physical condition records and past exercise history, and means for providing the generated suggestions to the user, means for acquiring information about the user's family and relatives from a database and using a generative AI model to generate appropriate response messages based on the acquired information, and means for generating a simple 3D avatar and providing the generated 3D avatar and response message to the user. This effectively draws out the past memories of dementia patients to support conversations, enables suggestions for exercise methods and dietary content tailored to their physical condition, and enables smooth communication with family and relatives.
[1954] A "user" is an entity that uses the system to provide voice input and receive conversations and suggestions.
[1955] "Voice input" refers to voice data input by a user speaking to the system.
[1956] "Voice data" refers to a digital recording of a user's voice input.
[1957] "Text data" refers to data obtained by converting voice data into characters.
[1958] "Age information" refers to data related to a user's year of birth or a particular age group.
[1959] A "database" is a system for storing and managing related data such as age and family information.
[1960] A "generative AI model" is an artificial intelligence algorithm that generates optimal responses or suggestions based on input data.
[1961] A "prompt sentence" is an input sentence that gives instructions to a generative AI model.
[1962] A "response message" is a conversational response generated by a generative AI model.
[1963] "Physical condition record" refers to a record that accumulates data related to the user's physical condition.
[1964] "Exercise history" refers to data relating to exercises that a user has performed in the past.
[1965] An "exercise method" is an exercise technique suggested based on the user's physical condition and exercise history.
[1966] "Meal content" refers to an appropriate meal plan suggested to the user.
[1967] "Family information" refers to data about the user's family and relatives.
[1968] A "3D avatar" is a three-dimensional virtual character generated on a computer.
[1969] This invention relates to an AI chat system that supports daily communication for dementia patients and improves their quality of life. This system provides comprehensive support by evoking the user's past memories, engaging in conversations based on those memories, and suggesting exercise methods and dietary content tailored to the user's physical condition.
[1970] Hardware and software used
[1971] Audio input device (microphone): Receives the user's voice and records the audio data.
[1972] Speech recognition software (e.g., Google Speech-to-Text API): converts voice data into text data.
[1973] Database system: Stores data such as the user's age, family information, health records, and exercise history.
[1974] Generative AI models (e.g., OpenAI's GPT-3): Set prompts based on received text data and generate appropriate conversational responses and suggestions.
[1975] Speech synthesis software (e.g., Amazon Polly): Converts the generated response message into audio data.
[1976] 3D avatar generation software: Generates a 3D avatar based on the user's family information.
[1977] Specific explanation of the system
[1978] 1. Voice input processing
[1979] The user speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[1980] The device receives the user's voice using a high-performance microphone and records it as voice data.
[1981] The terminal converts the recorded voice data into text data using voice recognition software and transmits the text data to the server.
[1982] 2. Database search
[1983] The server analyzes the received text data and extracts keywords such as "school days."
[1984] The server searches the database for the user's generational information, and if the user was a student in the 1960s, for example, it retrieves data on trends and events from that era.
[1985] 3. Response generation using generative AI models
[1986] The server uses a generative AI model to set prompts based on information retrieved from the database, such as "You were a student in the 1960s. Do you have any memorable experiences?"
[1987] The server sends the generated response message to the user's terminal.
[1988] 4. Providing a response message
[1989] The terminal converts the received response message into voice data using voice synthesis software.
[1990] The terminal reproduces the audio data from a speaker and conveys it to the user.
[1991] Specific examples
[1992] Example 1: Recalling past memories
[1993] When a user says, "I want to talk about my school days,"
[1994] The device records this audio and converts it into text using speech recognition software.
[1995] The server receives this text data and extracts keywords such as "school days."
[1996] The server retrieves the user's generational information from a database and researches trends and events from, for example, the 1960s.
[1997] The server uses a generative AI model to set a prompt sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?" and generates the optimal response message.
[1998] The server sends this response message to the terminal,
[1999] The device synthesizes the received message into voice and plays it back through the speaker, telling the user, "You were a student in the 1960s. Do you have any memorable experiences?"
[2000] Example 2: Exercise and diet suggestions
[2001] When a user says, "Which stretch should I do today?"
[2002] The device records this audio and converts it into text using speech recognition software.
[2003] The server receives this text data and extracts keywords such as "stretch."
[2004] The server retrieves the user's physical condition records and past exercise history from a database and uses a generative AI model to suggest the optimal exercise method.
[2005] The server sets a prompt sentence such as "Today, we recommend neck rotation exercises as a light stretch," and generates a suggestion.
[2006] The server sends this proposal to the terminal,
[2007] The device synthesizes the received suggestions into voice and plays them back through the speaker, telling the user, "Today, I recommend doing some neck rotation exercises as a light stretch."
[2008] Example 3: Supporting conversations with family and relatives
[2009] When a user says, "I want to know about my grandchildren,"
[2010] The device records this audio and converts it into text using speech recognition software.
[2011] The server receives this text data and extracts keywords such as "grandchild."
[2012] The server retrieves information about grandchildren from a database and uses a generative AI model to generate the question, "Do you know what your grandchildren are learning in school right now?"
[2013] The server also simultaneously generates a simple 3D avatar and sends it to the user's device.
[2014] The device synthesizes the received question into speech, plays it back through the speaker, and displays a 3D avatar, asking the user, "Do you know what your grandchild is learning at school right now?"
[2015] In this way, the present invention improves the quality of life of dementia patients by effectively drawing out the user's past memories and making suggestions tailored to everyday conversations and physical condition.
[2016] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2017] Step 1: Receiving voice input
[2018] User: Speaks to the system about a topic related to a past memory, for example, "I'd like to talk about my old school days."
[2019] Terminal: The user's voice is received by a high-performance microphone and recorded as audio data, which becomes the input data.
[2020] Terminal: The recorded audio data is sent to the server. The output at this stage is audio data.
[2021] Step 2: Converting audio data to text data
[2022] Server: Converts the received voice data into text data using speech recognition software (e.g., Google Speech-to-Text API). This is the input data.
[2023] Server: The converted text data is used in the next processing step. The output is text data.
[2024] Step 3: Database search
[2025] Server: Analyzes the received text data and extracts keywords such as "school days." This is the input data.
[2026] Server: Searches the database for the user's generation information and retrieves data on trends and events from the corresponding era. This data processing and calculation is used to reference and search for generation information. The output is data from the corresponding era.
[2027] Step 4: Generative AI model generates a response
[2028] Server: Based on the information obtained from the database, a generative AI model (e.g., OpenAI's GPT-3) is used to set a prompt. For example, the prompt may be a text sentence such as, "You were a student in the 1960s. Do you have any memorable experiences?"
[2029] Server: Using this prompt as input, the generative AI model generates the optimal conversational response. This response generation is performed using data calculations.
[2030] Server: Output the generated response message as text data for use in the next step.
[2031] Step 5: Send a response message
[2032] Server: Sends the generated response message to the user's terminal. This is the input data.
[2033] Terminal: Receives the response message and converts it into voice data using speech synthesis software (e.g., Amazon Polly). Voice synthesis is a type of data processing. The output is voice data.
[2034] Step 6: Play greetings
[2035] Terminal: The audio data is played back through the speaker and transmitted to the user. This is the final output.
[2036] Example flow
[2037] User: Say, "Which stretch should I do today?"
[2038] Device: Uses voice recognition software to convert the question "Which stretch should I do today?" into text data and send it to the server.
[2039] Server: Analyzes the text data and extracts keywords.
[2040] Server: Retrieves the user's physical condition records and past exercise history from the database.
[2041] Server: Using a generative AI model, it generates a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[2042] Server: Sends the proposal to the device.
[2043] Device: The suggested message is generated through speech synthesis and played from the speaker. It tells the user, "Today, we recommend you do some light stretching exercises, such as neck rotations."
[2044] In this way, the processing steps from the user's voice input to the final playback of the response message are concretely carried out.
[2045] (Application example 1)
[2046] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2047] This invention relates to an AI chat system that evokes past memories in dementia patients and improves their quality of life. Currently, many nursing homes lack support for retrieving individual memories for dementia patients, resulting in insufficient mental stability and communication. It is also difficult to provide dementia patients with appropriate exercise and dietary recommendations tailored to their physical condition, and communication with their families is also difficult. To address these issues, a system is needed that effectively engages dementia patients in conversation and suggests health management strategies based on their past living conditions.
[2048] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2049] In this invention, the server includes a means for processing user voice input, a means for searching a database for user age information and related data, a means for generating optimal conversational responses using a generative AI model, a means for providing the generated conversational responses to the user by voice and eliciting past memories, a means for using the generative AI model to hold conversations based on the past living situations of dementia patients and support daily communication, and an application installed on a robot installed in a care facility, which makes it possible to increase the mental stability of dementia patients and provide effective communication support.
[2050] "Means for processing user voice input" refers to a function for receiving voice data uttered by a user, recognizing it, and converting it into text data.
[2051] "Means of searching for a user's age information and related data from a database" refers to the function of searching for and retrieving the necessary information from a database that holds information related to the user's age, interests, and past life.
[2052] "Means for generating optimal conversational responses using a generative AI model" refers to a function that uses AI technology based on text data to generate optimal responses for users.
[2053] "Means for providing the generated conversational response to the user by voice and eliciting past memories" refers to a function for providing the user with an automatically generated text response by voice output, thereby eliciting the user's past memories.
[2054] "A means of using a generative AI model to hold conversations based on the past living conditions of dementia patients and support daily communication" refers to a function that makes full use of AI technology to conduct conversations based on content related to the past experiences and lives of dementia patients, thereby promoting daily interaction.
[2055] "Applications installed on robots installed in nursing facilities" refers to application software that is installed on robot devices used in nursing facilities and that is used to interact with dementia patients.
[2056] "A means for referring to the user's health records and past exercise history, and proposing optimal exercise methods and dietary content based on the generated health data" refers to a function that refers to the user's previous health information and exercise data, and provides appropriate exercise and dietary advice based on the user's current health condition.
[2057] "Means for providing the generated suggestions to the user by voice" refers to a function for conveying the generated exercise and diet suggestions to the user by voice.
[2058] "Means for a robot installed in a care facility to demonstrate suggested content to a user" refers to a function in which a robot demonstrates appropriate exercise methods and other instructions to a user.
[2059] "Means of obtaining information about the user's family and relatives from a database and generating a simple 3D avatar based on that information" refers to a function that obtains data about the user's family and relatives from a database and generates an avatar based on that data.
[2060] "Means of using the generated 3D avatar to engage in everyday conversation with the user and provide supplementary animations" refers to the function of using the generated 3D avatar as an interface to assist in conversation with the user and display animations as a visual aid.
[2061] "Means for providing information about family members and relatives to the user by voice" refers to a function that conveys information about family members and relatives obtained from a database to the user by voice.
[2062] This invention is an AI chat system that elicits past memories of dementia patients and improves their quality of life. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary recommendations tailored to the user's physical condition. This system is implemented as an application installed on a robot in a nursing home.
[2063] System Overview
[2064] 1. Voice input processing
[2065] The user uses the system to speak to initiate a conversation about a past memory.
[2066] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[2067] 2. Drawer of past memories
[2068] The server searches a database for information about the user's age based on the received text data, such as information about the user's childhood or school days.
[2069] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[2070] The server transmits the generated conversation response to the terminal.
[2071] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[2072] 3. Exercise and dietary suggestions
[2073] The user reports their physical condition to the device, for example, by saying, "I feel a little tired today."
[2074] The terminal converts the physical condition data into text data and transmits the text data to a server.
[2075] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[2076] The server sends the generated proposal to the terminal.
[2077] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[2078] 4. Support for conversations with family and relatives
[2079] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[2080] The terminal converts the voice into text data and transmits the text data to the server.
[2081] The server retrieves information about family and relatives from a database, and then uses a generative AI model to generate optimal conversation responses. It also simultaneously generates a simple 3D avatar.
[2082] The server sends the generated conversational responses and 3D avatars to the terminal.
[2083] The device communicates the received responses to the user aloud and displays a 3D avatar to assist in everyday conversations.
[2084] Specific examples
[2085] Example 1: Recalling past memories
[2086] The user says, "I want to talk about my old school days."
[2087] The terminal sends this request to the server.
[2088] The server obtains the user's generational information from a database and extracts information about items and events that were popular during that era.
[2089] For example, if it is a period when a unique game or a particular movie was popular, the system can use that information to generate a question such as, "What games did you play in elementary school?"
[2090] The server sends the generated question to the terminal, which then conveys it to the user as voice.
[2091] The user begins recounting past events to jog their memory, and the conversation continues.
[2092] Example prompt sentence:
[2093] User: I want to talk about my old school days.
[2094] AI: What did you play in elementary school?
[2095] Example 2: Exercise and diet suggestions
[2096] The user says, "Which stretch should I do today?"
[2097] The terminal sends this request to the server.
[2098] The server references the user's physical condition records and past exercise history and uses a generative AI model to generate a suggestion such as, "Today, we recommend neck rotation exercises as a light stretch."
[2099] The server transmits the generated suggestions to the terminal, which then conveys them to the user as voice.
[2100] If necessary, the robot will demonstrate neck rotation.
[2101] Example prompt sentence:
[2102] User: Which stretches should I do today?
[2103] AI: Today, I recommend some gentle stretching exercises, such as neck rotations.
[2104] Example 3: Supporting conversations with family and relatives
[2105] The user says, "I want to know about my grandchildren."
[2106] The terminal sends this request to the server.
[2107] The server retrieves information about the grandchildren (such as their names, schools, interests, etc.) from a database.
[2108] Based on this, the server uses a generative AI model to generate a question such as, "Do you know what your grandchild is learning at school right now?" and simultaneously generates a simple 3D avatar.
[2109] The server sends the generated question and 3D avatar to the terminal, which then conveys the question to the user as voice and displays the 3D avatar.
[2110] Example prompt sentence:
[2111] User: I want to know about my grandchildren
[2112] AI: Do you know what your grandchildren are learning in school right now?
[2113] In this way, the present invention provides a concrete means for improving the quality of life of dementia patients and reducing the burden on caregivers. This system evokes the user's past memories, supports daily communication, and even suggests exercise methods and dietary content tailored to their physical condition, thereby providing comprehensive support.
[2114] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2115] Step 1:
[2116] The user uses the system to speak to initiate a conversation about a past memory.
[2117] Input: User's voice data
[2118] How it works: A microphone on a robot installed in a care home captures voice input.
[2119] Step 2:
[2120] The terminal (a robot installed in a nursing facility) converts the voice into text data and sends the text data to a server.
[2121] Input: User's voice data
[2122] What it does: It uses speech recognition software (e.g., the speech_recognition library) to convert speech to text, which is then sent over the Internet to a server.
[2123] Output: Text data sent to the server
[2124] Step 3:
[2125] The server searches the database for the user's age information based on the received text data.
[2126] Input: User's text data
[2127] What it does: The server performs a database query to retrieve data about the user's age and related past memories. For example, if the user types "I want to talk about my old school days," it searches for information related to the user's childhood and school days.
[2128] Output: User age and related data
[2129] Step 4:
[2130] The server uses a generative AI model to generate optimal conversational responses based on information retrieved from the database, utilizing materials such as popular items and events by decade, photos, and movie posters.
[2131] Input: User demographics and related data
[2132] Specific operation: Query a generative AI model (e.g., OpenAI GPT-3) with a prompt sentence to generate an appropriate conversational response. Reference data such as trends and events relevant to each generation are also used as reference.
[2133] Output: Generated conversation response
[2134] Step 5:
[2135] The server transmits the generated conversation response to the terminal.
[2136] Input: Generated conversation response
[2137] Specific operation: The response text generated by the server is sent to the terminal via the Internet.
[2138] Output: Conversation response sent to the terminal
[2139] Step 6:
[2140] The device then conveys the received response to the user by voice, enabling a conversation that evokes past memories.
[2141] Input: Conversation response sent by the server
[2142] What it does: It uses speech synthesis software (e.g., the pyttsx3 library) to play the text data as speech, allowing the user to continue the conversation by listening to the voice response.
[2143] Output: The audio response provided to the user
[2144] Step 7:
[2145] The user reports their physical condition to the device, for example, saying, "I feel a little tired today."
[2146] Input: User's voice data
[2147] What happens: The device's microphone captures audio input.
[2148] Step 8:
[2149] The terminal converts the physical condition data into text data and transmits the text data to a server.
[2150] Input: Audio data about the user's physical condition
[2151] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[2152] Output: Text data about your health condition sent to the server
[2153] Step 9:
[2154] The server retrieves the received physical condition data and past exercise history from a database, and based on that, uses a generative AI model to suggest optimal exercise methods and dietary recommendations.
[2155] Input: Text data about the user's physical condition and past exercise history
[2156] Specific operation: The server retrieves the user's past exercise history from the database and uses a generative AI model to suggest appropriate exercise methods and dietary recommendations.
[2157] Output: Generated exercise and diet suggestions
[2158] Step 10:
[2159] The server sends the generated proposal to the terminal.
[2160] Input: Generated proposal
[2161] Specific operation: The proposal content is sent from the server to the device via the Internet.
[2162] Output: Suggestions sent to the device
[2163] Step 11:
[2164] The terminal will communicate the received suggestions to the user via voice, and if necessary, the robot will demonstrate how to exercise.
[2165] Input: Suggestion from the server
[2166] Specific behavior: Using speech synthesis software, the robot plays back the suggestions as voice, and shows the user the appropriate exercise method.
[2167] Output: Audio suggestions and demonstrations provided to the user
[2168] Step 12:
[2169] When a user wants to know information about family members or relatives, the user speaks to the terminal, for example, saying, "I want to know about my grandchildren."
[2170] Input: User's voice data
[2171] What happens: The device's microphone captures audio input.
[2172] Step 13:
[2173] The terminal converts the voice into text data and transmits the text data to the server.
[2174] Input: User's voice data
[2175] What it does: It uses speech recognition software to convert speech into text, which is then sent over the internet to a server.
[2176] Output: Text data sent to the server
[2177] Step 14:
[2178] The server retrieves information about family and relative...
Claims
1. means for processing a user's voice input to retrieve past memories; A means of searching the database for user age information and related data; a means for generating optimal conversational responses using a generative AI model; The system includes means for providing the generated conversational response to a user.
2. A means for referencing the user's physical condition record and past exercise history; A means to suggest optimal exercise methods and dietary content based on the user's physical condition data, 2. The system of claim 1, further comprising means for providing the generated suggestions to the user.
3. means for retrieving information about the user's family and relatives from a database; A means to generate a simple 3D avatar based on the acquired information, The system according to claim 1, further comprising means for providing the generated 3D avatar to the user and for engaging in everyday conversation with the user.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A