System

A system using language processing, data storage, and image recognition technologies monitors and supports the daily lives of elderly people with dementia, addressing the lack of effective care by providing real-time advice and reducing caregiver burden.

JP2026034296APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137417
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

Smart Images

  • Figure 2026034296000001_ABST
    Figure 2026034296000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a language processing means for collecting information on the basis of interaction with a user, a data holding means for holding the collected information in a database, an image recognition means for analyzing a life habit by using an image recognition technique, and an assist means for generating appropriate advice on the basis of an analysis result and providing the user with the advice.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, with the increasing number of elderly people with dementia, there is a demand for technology to help them live safe and independent daily lives. Families and caregivers need a way to understand the living conditions of elderly people even from a distance and provide appropriate care, but conventional technology may not provide sufficient support. In particular, there is a lack of effective measures for memory impairment and lifestyle disorders, which are symptoms specific to dementia. Therefore, there is a demand for a system that monitors the lives of elderly people and provides the necessary support. [Means for solving the problem]

[0005] To solve the above problems, the present invention proposes the following means. The system of the present invention includes a language processing means for collecting information based on dialogue with a user, a data storage means for storing the collected information in a database, an image recognition means for analyzing lifestyle habits using image recognition technology, and an assistance means for generating appropriate advice based on the analysis results and providing it to the user. This system can comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. Furthermore, by combining the means for monitoring the living environment using a camera device and acquiring image data, more detailed analysis and judgment of lifestyle habits become possible, supporting the safe and healthy living of elderly people.

[0006] "Users" are those who use the system, and in this case primarily refer to elderly people with dementia.

[0007] "Dialogue" is the act of communication between a user and a system through language.

[0008] A "language processing means" is a device or software that uses natural language processing techniques to analyze interactions with a user and to collect and generate information.

[0009] "Data storage means" refers to a database system and related technologies for storing collected information temporarily or long-term.

[0010] "Image recognition means" refers to a device or program that analyzes image data obtained from a camera or sensor and recognizes the objects and actions contained therein.

[0011] "Lifestyle habits" refers to the user's daily actions and habits (e.g., diet, clothing, sleep, etc.), which are the subject of monitoring and analysis.

[0012] The "assisting means" is a device or program for providing appropriate advice and warnings to the user based on the analysis results.

[0013] "Camera device" refers to a video camera or still camera for capturing video data to monitor the living environment.

[0014] "Image data" refers to visual data (still images and video) captured by a camera device. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system supports the lives of elderly people by collecting information through dialogue with the user and analyzing the user's lifestyle habits using image recognition technology.

[0037] System configuration

[0038] The system consists of the following main components:

[0039] 1. Language processing means: Collects necessary information through dialogue with the user.

[0040] 2. Data retention measures: Collected information is stored in a database.

[0041] 3. Image recognition means: Analyzes image data acquired using a camera device.

[0042] 4. Assistance: Providing appropriate advice based on the collected information and analysis results.

[0043] Implementation of language processing measures

[0044] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[0045] Implementing data retention measures

[0046] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[0047] Implementation of image recognition measures

[0048] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[0049] Implementation of assistance measures

[0050] The server generates appropriate advice based on the collected information and image recognition results and provides it to the user. For example, it generates advice such as, "Don't forget to eat, start eating now." This advice is notified to the user via their device. Similar information is also notified to family members as needed.

[0051] Example

[0052] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that no food is found. It then asks the user through language processing means, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now." This advice is notified to the user through the device, and the information is also sent to family members if necessary.

[0053] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[0054] The processing flow will be explained below.

[0055] Step 1:

[0056] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[0057] Step 2:

[0058] The device sends the captured image data to the server, where it is sent in real time over the network.

[0059] Step 3:

[0060] The server pre-processes the received image data, including image resizing, noise reduction, and color correction.

[0061] Step 4:

[0062] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[0063] Step 5:

[0064] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[0065] Step 6:

[0066] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[0067] Step 7:

[0068] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[0069] Step 8:

[0070] The user responds to the initial message through the terminal, either by voice or text.

[0071] Step 9:

[0072] The device captures the user's response and converts it into text data using voice recognition technology.

[0073] Step 10:

[0074] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[0075] Step 11:

[0076] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[0077] Step 12:

[0078] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[0079] Step 13:

[0080] The server makes the final decision and generates appropriate advice, e.g., "Don't forget to eat, let's eat now."

[0081] Step 14:

[0082] The server sends the generated advice to the device, which then notifies the user via voice or text.

[0083] Step 15:

[0084] The server will notify the family of the same information as necessary. For example, it may notify the family by email or SMS that "The user has not eaten a meal today."

[0085] Through the above steps, it is possible to comprehensively support the user's daily life and provide a safe living environment.

[0086] Example 1

[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0088] This invention relates to a system that comprehensively supports the lives of elderly people with dementia and reduces the burden on their families and caregivers. Conventional systems have difficulty accurately understanding a user's lifestyle and health status, and have been unable to provide appropriate advice or notifications. Therefore, more effective information collection and analysis is needed to maintain a user's health and improve their quality of life.

[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0090] In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a storage means, an image recognition means for analyzing image data acquired using a camera device, and an assistance means for generating appropriate notifications or advice based on the analysis results and the collected information and providing them to the user. This makes it possible to appropriately grasp the user's behavior and health condition in real time and quickly provide necessary advice or notifications.

[0091] "User" refers to the individual or family member who uses the system, and is often targeted at elderly people with dementia.

[0092] "Language processing means" refers to technology that collects information from dialogue with the user and converts it into text data.

[0093] "Data retention means" refers to technology that stores collected information in a storage means such as a database, making it available for later reference.

[0094] A "camera device" is hardware for acquiring image data, and is used to monitor the user's living environment.

[0095] "Image recognition means" refers to technology that analyzes image data acquired by a camera device and determines the user's behavior and situation.

[0096] "Assistance means" refers to technology that generates and provides appropriate notifications and advice to users based on collected information and analysis results.

[0097] "Analysis results" refers to the analysis data obtained by the language processing means and the image recognition means.

[0098] "Notification" means a message or alert provided by the System to a User or their Family Members.

[0099] "Advice" refers to the guidance the system provides to the user to encourage improvements in their lifestyle and health.

[0100] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and provide an environment in which they can live with peace of mind. This system consists of four main components: language processing means, data storage means, image recognition means, and assistance means.

[0101] Language Processing Methods

[0102] The server uses natural language processing technology (e.g., a generative AI model such as OpenAI's GPT-4) to analyze the user's dialogue and collect necessary information. The dialogue with the user is conducted through the device, which recognizes the user's voice and transmits it as text data to the server. For example, if the device asks the user, "Hello, how are you?" and the user responds, "I'm a little tired today," that information is collected by the server.

[0103] Data Retention Methods

[0104] The collected information is stored in a database (for example, an RDBMS such as MySQL (registered trademark) or PostgreSQL) by the server. This records the user's most recent actions and status, and makes them available for later reference. For example, data such as "The user responded that they were tired at 10:00 AM on October 13, 2023" is stored.

[0105] Image Recognition Method

[0106] The camera device installed in the device periodically captures the user's living environment and sends the image data to a server. The server then uses image recognition technology to analyze this data. For example, if the camera analyzes an image of the user's room and no food is found, it can determine that the user has not had breakfast.

[0107] Assistance Method

[0108] The server generates appropriate notifications or advice based on the collected information and the results of image recognition analysis, and provides them to the user. For example, it generates a notification saying, "Don't forget to eat, let's eat now." This notification is sent to the user via their device, and similar notifications are also sent to family members if necessary.

[0109] Specific examples

[0110] For example, one morning, the device captures the state of the user's room with a camera and sends the image data to the server. The server analyzes this image data and confirms that no food is found. The device then asks the user, "How was your breakfast today?" through language processing means. If the user replies, "I forgot," this information is saved in the database. The server then generates advice such as, "Don't forget to eat, start eating now," and notifies this to the user via the device. Information that "the user may not have had breakfast" is also sent to the user's family.

[0111] Example prompt sentence:

[0112] Ask the user, "Hello, how are you?". Before the user asks, "How was your breakfast today?", save the user's most recent meal history in a database. Then, analyze the camera footage and generate an advice if the user hasn't eaten, saying, "Don't forget to eat, eat now."

[0113] This system provides comprehensive support for users' daily lives, reducing the burden on their families and caregivers, and provides appropriate advice based on the collected information, creating an environment in which users can live with peace of mind.

[0114] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0115] Step 1: Initiate a conversation with the user

[0116] Input: The initial message the terminal is set to (e.g. "Hello, how are you?")

[0117] Action: The device asks the user a question using voice.

[0118] Output: The user responds verbally (e.g., "I'm a little tired today").

[0119] Step 2: Collect and analyze user responses

[0120] Input: User's voice response

[0121] How it works: The device uses voice recognition technology to convert the user's response into text data, which it then sends to the server.

[0122] Output: Text data (e.g., "I'm a little tired today")

[0123] Step 3: Information analysis using natural language processing

[0124] Input: Text data (e.g., "I'm a little tired today")

[0125] How it works: The server uses a generative AI model (e.g., OpenAI's GPT-4) to analyze text data and evaluate the user's state.

[0126] Output: Analysis result (e.g. "The user is tired")

[0127] Step 4: Storing information in a database

[0128] Input: Analysis result (e.g. "The user is tired")

[0129] How it works: The server stores the analysis results in a database.

[0130] Output: Stored data (e.g., "User responded that he was tired on October 13, 2023 at 10:00 AM")

[0131] Step 5: Capture camera images

[0132] Input: Timer setting for camera device (e.g. capture every hour)

[0133] Operation: The camera device installed on the device captures an image and sends the image data to the server.

[0134] Output: Image data (e.g., a photo of the user's room)

[0135] Step 6: Data analysis using image recognition

[0136] Input: Image data (e.g., a photo of the user's room)

[0137] How it works: The server uses image recognition technology to analyze the image data and determine the user's situation (e.g., if there is no food in sight, it determines that the user is not eating).

[0138] Output: Image analysis results (e.g., "The user has not eaten")

[0139] Step 7: Generate and notify advice

[0140] Input: User information analysis results and image analysis results

[0141] Operation: The server generates appropriate advice based on the collected information and image recognition results. The generated advice (e.g., "Don't forget to eat, let's eat now") is notified to the user via the device. If necessary, the information is also sent to family members.

[0142] Output: Advice notification (e.g., notification from the device to the user saying "Don't forget to eat, let's eat now" and sharing the information with family members)

[0143] (Application example 1)

[0144] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0145] The present invention aims to provide an environment in which elderly people with dementia can use self-driving vehicles with peace of mind. Currently, there is a lack of means to ensure the safe and comfortable travel of elderly people with dementia in self-driving vehicles, which creates a risk of elderly people behaving incorrectly. Specifically, this poses problems such as forgetting to eat and not taking care of their health. In response to this, the present invention solves these issues by monitoring the condition of elderly people in real time and providing appropriate warnings and instructions.

[0146] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0147] In this invention, the server includes language processing means for collecting information based on dialogue with the user, data storage means for storing the collected information in a database, image recognition means for analyzing lifestyle habits using image recognition technology, assistance means for generating appropriate advice based on the analysis results and providing it to the user, means for monitoring the condition of the elderly person using a camera device installed in the vehicle and acquiring image data, and means for analyzing the acquired image and audio data and providing instructions and warnings tailored to the elderly person's condition. This enables an environment in which elderly people with dementia can use self-driving vehicles safely and comfortably.

[0148] The "language processing means" is a means for collecting necessary information based on a dialogue with the user.

[0149] "Data storage means" refers to a means for storing collected information in a database.

[0150] The "image recognition means" is a means for analyzing a user's lifestyle habits using image recognition technology.

[0151] The "assisting means" is a means for generating appropriate advice based on the analysis results and providing it to the user.

[0152] The "camera device" is a device that is installed in a vehicle to monitor the condition of the elderly person and acquire image data.

[0153] The "image and audio data analysis means" is a means for analyzing the acquired image and audio data and providing instructions and warnings that are tailored to the elderly person's situation.

[0154] The system of this invention is designed to provide an environment in which elderly people with dementia can safely use autonomous vehicles. The system consists of the following main components:

[0155] 1. Language Processing Methods

[0156] The server collects necessary information through dialogue with the user. Specifically, it responds to questions and greetings from the user through voice input, and analyzes the response using natural language processing technology. For example, if an elderly person responds "I forgot" to the question "How was your breakfast today?", this information is processed and recorded as text data. The software used includes the SpeechRecognition library and the gTTS library.

[0157] 2. Data Retention Methods

[0158] The server stores the collected information in a database. This database is used to record the user's daily behavior and status history. For example, information such as "I did not eat breakfast at 10:00 AM on October 12, 2023" is recorded. This allows the system to provide appropriate advice while referring to past history.

[0159] 3. Image Recognition Methods

[0160] A camera device installed inside the vehicle monitors the elderly person's condition and periodically captures image data. The captured image data is sent to a server and analyzed using image recognition technology. Specifically, the OpenCV library is used to process the images and perform object detection and behavior recognition. For example, if no food is found, it is determined that the elderly person has not eaten.

[0161] 4. Assistance methods

[0162] The server generates appropriate advice based on the collected information and the results of image recognition analysis, and provides it to the user. The advice is communicated to the user via voice. For example, advice such as "Don't forget to eat, start eating now" is generated and communicated to the user through a speaker.

[0163] 5. Image and audio data analysis methods

[0164] The acquired image and audio data is processed in real time. This method is used to provide instructions and warnings tailored to the user's situation. The image and audio data captured by the camera are used for analysis, and the server comprehensively analyzes these to provide optimal advice.

[0165] Specific example explanation

[0166] For example, one morning, the server uses a camera inside the vehicle to capture the state of an elderly person's room. Analysis of the image data confirms that food is missing. The server then asks aloud, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and advises, "Don't forget to eat. Let's eat now." This advice is communicated to the user through the speaker, and the information is also sent to family members if necessary.

[0167] Prompt Sentence Examples

[0168] Here are some examples of specific prompts:

[0169] Elderly Co-Driver Assistance prompts:

[0170] 1. Capture an image with the camera and send it to an image analysis API to determine if food is in the image.

[0171] 2. A voice prompt asks the elderly person, "How was your breakfast today?"

[0172] 3. If the user answers "I forgot," the system provides the advice "Don't forget to eat, eat now."

[0173] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[0174] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0175] Step 1:

[0176] The server uses a camera device to monitor the condition of the elderly person in the vehicle and captures image data.

[0177] Input: Real-time video from a camera device.

[0178] Data processing and calculation: Images are captured using the OpenCV library and saved as still images.

[0179] Output: The captured image data.

[0180] Step 2:

[0181] The server sends the captured image data to an image recognition API to determine whether food is in the image.

[0182] Input: The captured image data.

[0183] Data processing and calculation: Send image data to the image recognition API and obtain the analysis results.

[0184] Output: Analysis result (e.g., determination that food is not visible).

[0185] Step 3:

[0186] Based on the analysis results, the server asks the user aloud, "How was your breakfast today?"

[0187] Input: Analysis results (e.g. food not shown).

[0188] Data processing and calculation: The question is converted into an audio file using the gTTS library and played through the speaker.

[0189] Output: A spoken question to the user.

[0190] Step 4:

[0191] The user responds to the voice questions and the server receives the response as voice input.

[0192] Input: The user's spoken response.

[0193] Data processing and calculation: Convert speech to text using the SpeechRecognition library.

[0194] Output: User response as text data (e.g., "I forgot").

[0195] Step 5:

[0196] The server records the user's response in a database.

[0197] Input: User response as text data.

[0198] Data processing and calculation: Connect to the database and record the response content in an appropriate format.

[0199] Output: The recorded database entries.

[0200] Step 6:

[0201] Based on the user's response, the server generates advice such as "Don't forget to eat, start eating now," and notifies the user by voice.

[0202] Input: User response as text data.

[0203] Data processing and calculation: The advice sentence is converted into an audio file using the gTTS library and played through the speaker.

[0204] Output: Audio advice to the user.

[0205] Step 7:

[0206] The server notifies the user's family of the user's condition and advice as necessary.

[0207] Input: User responses and advice as text data.

[0208] Data processing and calculation: Send text messages and emails to family members through the notification system.

[0209] Output: Notification message to family members.

[0210] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0211] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and an emotion engine to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[0212] System configuration

[0213] The system consists of the following main components:

[0214] 1. Language processing means: Collects necessary information through dialogue with the user.

[0215] 2. Data retention measures: Collected information is stored in a database.

[0216] 3. Image recognition means: Analyzes image data acquired using a camera device.

[0217] 4. Emotion Engine: Analyzes the emotional state from user interactions and image data.

[0218] 5. Assistance: Providing appropriate advice based on the collected information and analysis results.

[0219] Implementation of language processing measures

[0220] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[0221] Implementing data retention measures

[0222] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[0223] Implementation of image recognition measures

[0224] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[0225] Emotion Engine Implementation

[0226] The server runs an emotion engine based on the user's dialogue, voice tone, and even image data sent from the device. This engine analyzes the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.). The analysis results can be used to more specifically determine the most appropriate advice.

[0227] Implementation of assistance measures

[0228] The server generates appropriate advice based on the collected information and the analysis results of the emotion engine, and provides it to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, it can take an approach such as "Is there anything troubling you?" This advice is notified to the user via their device. Similar information can also be notified to family members if necessary.

[0229] Example

[0230] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that there is no food in sight. It also uses an emotion engine to determine that the user's facial expression is "gloomy." The server then uses language processing to ask the user, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now. Also, let us know if you need any help." This advice is notified to the user via the device, and the information is also sent to family members if necessary.

[0231] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding an emotion engine, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[0232] The processing flow will be explained below.

[0233] Step 1:

[0234] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[0235] Step 2:

[0236] The device sends the captured image data to the server, where it is sent in real time over the network.

[0237] Step 3:

[0238] The server pre-processes the image data it receives, including image resizing, noise reduction, and color correction.

[0239] Step 4:

[0240] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[0241] Step 5:

[0242] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[0243] Step 6:

[0244] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[0245] Step 7:

[0246] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[0247] Step 8:

[0248] The user responds to the initial message through the terminal, either by voice or text.

[0249] Step 9:

[0250] The device captures the user's response and converts it into text data using voice recognition technology.

[0251] Step 10:

[0252] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[0253] Step 11:

[0254] The server uses an emotion engine to analyze the emotional state from the user's response and image data, e.g., to identify the emotional state such as "sad" or "happy."

[0255] Step 12:

[0256] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[0257] Step 13:

[0258] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[0259] Step 14:

[0260] The server makes the final decision and generates appropriate advice. For example, "Don't forget to eat, let's eat now. Also, are you feeling sad? Can you tell me something?"

[0261] Step 15:

[0262] The server sends the generated advice to the device, which then notifies the user via voice or text.

[0263] Step 16:

[0264] The server notifies the family of the same information as necessary. For example, it may notify the family by email or SMS, saying, "The user does not seem to have eaten a meal today. Also, he may be depressed."

[0265] Through the above steps, it is possible to provide comprehensive support for the user's daily life and provide appropriate responses according to their emotional state.

[0266] Example 2

[0267] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0268] The problem that this invention aims to solve is to provide appropriate support to elderly people with dementia and their families based on their individual lifestyle habits and emotional states, and to create an environment where elderly people with dementia can live with peace of mind. To achieve this, it is important to understand the user's daily behavior and emotions in real time and provide appropriate advice. However, existing systems do not adequately analyze emotions or recognize lifestyle habits, making it difficult to provide effective support to users and their families.

[0269] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a database, an image recognition means for analyzing lifestyle habits using image recognition technology, an emotion analysis means for analyzing the emotional state from the content of the user's dialogue and image data, and an assistance means for generating appropriate advice based on the analysis results and the emotional state and providing the advice to the user. This makes it possible to grasp the user's daily behavior and emotions in real time and provide appropriate support based on them.

[0270] "Language processing means" refers to devices or software that collect information based on dialogue with a user and analyze that information using natural language processing technology.

[0271] "Data storage means" refers to devices or software that store collected information in a database so that it can be referenced later.

[0272] "Image recognition means" refers to a technology or device that analyzes image data acquired by a camera device and determines the lifestyle habits of a user.

[0273] "Emotion analysis means" refers to a device or software that analyzes the emotional state of a user from the content of their dialogue or image data.

[0274] "Assistance means" refers to a device or software that generates appropriate advice based on the analysis results and emotional state and provides it to the user.

[0275] The term "camera device" refers to a device for monitoring a user's living environment and acquiring image data.

[0276] "Means for analyzing lifestyle habits" refers to technology or devices that analyze and judge a user's behavior and habits based on acquired image data.

[0277] "Emotional state" refers to the user's psychological state and feelings, as judged from the content of the user's dialogue, facial expression, tone of voice, etc.

[0278] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and emotion analysis technology to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[0279] Implementation of language processing measures

[0280] The server uses natural language processing technology to analyze the dialogue with the user and collect the necessary information. To do this, natural language processing (NLP) software such as Google® Cloud Natural Language API and IBM Watson® is used. For example, the server asks the user, "Hello, how are you?" and converts the user's response, "I'm not feeling well," into text data using speech recognition technology (e.g., Google Speech-to-Text API) and analyzes it.

[0281] Implementing data retention measures

[0282] The collected information is stored in a database by the server. MySQL or PostgreSQL is used as the database. This records the user's most recent actions and status, allowing them to be referenced later. For example, it may record "No meals at 10:00 AM on October 12, 2023."

[0283] Implementation of image recognition measures

[0284] A camera device (e.g., Raspberry Pi camera module) installed on the device periodically captures image data and sends it to the server. The server then analyzes this data using image recognition technology (e.g., OpenCV, TENSORFLOW (registered trademark)). For example, if no food is found in the video, it will determine that the person is not eating.

[0285] Implementing sentiment analysis measures

[0286] The server uses emotion analysis technology (e.g., Microsoft® Azure® Emotion API) based on the content of the user's dialogue, voice tone, and image data sent from the device to analyze the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.).

[0287] Implementation of assistance measures

[0288] The server uses a generative AI model to generate appropriate advice based on the collected information and the results of emotion analysis. The generated advice is sent to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API) and provided to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, an approach such as "Is there anything troubling you?" is taken. This advice is notified to the user via the device, and similar information is also notified to family members if necessary.

[0289] Example

[0290] Example 1: Morning dialogue scenario

[0291] At 6:00 a.m., the device speaks to the user, asking, "Good morning. How are you feeling today?" If the user responds, "I'm not feeling very good today," the device captures this voice and sends it to the server. The server converts this into text data and uses NLP technology to analyze the "bad mood." The analysis determines that the user is feeling "depressed," which is supported by emotion analysis technology. The server generates advice such as, "Let's find something to cheer you up together. Is there anything troubling you?" and notifies the user of this via the device.

[0292] Example 2: Lunch Check Scenario

[0293] At 12:00, the device captures the state of the room with its camera and sends it to the server. The server analyzes the image data and determines that "food is missing." It also uses emotion analysis technology to determine that the user's facial expression is "gloomy." The server then asks the user, "How was your lunch today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat. Let's eat now. Also, let us know if you need any help," and notifies the user through the device.

[0294] Prompt Sentence Examples

[0295] Example prompts to be input to the generative AI model:

[0296] "Please explain in detail the processing steps of the daily life support system for elderly people with dementia. Please include the specific operations and technologies for each step, the names of the hardware and software used, and examples of use."

[0297] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding emotion analysis technology, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[0298] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0299] Step 1:

[0300] A user accesses the system, and the device asks a question via the microphone, such as "Hello, how are you?" The device obtains the user's response (voice). This response is captured by the device's microphone (input data: user's voice). Specifically, the device activates its voice recognition function to capture the voice signal and sends it to the server as a digital voice file (output data: digital voice file).

[0301] Step 2:

[0302] The server uses speech recognition technology (e.g., Google Speech-to-Text API) to analyze the received audio file and convert it into text data (input data: digital audio file, output data: text data). Specifically, the server samples the audio data and uses a speech recognition model to convert it into a string of characters.

[0303] Step 3:

[0304] The server analyzes the converted text data using natural language processing technology (e.g., Google Cloud Natural Language API), thereby extracting semantic information from the user's response (input data: text data, output data: semantic information). Specifically, the server analyzes the text using techniques such as tokenization, part-of-speech tagging, and sentiment analysis to extract information about the user's state and emotions.

[0305] Step 4:

[0306] The device periodically captures image data of the environment. The camera device acquires image data of the user's living environment (input data: environmental image). Specifically, the device activates the camera at set time intervals, captures images, and sends them to the server (output data: environmental image file).

[0307] Step 5:

[0308] The server analyzes the received image data using image recognition technology (e.g., OpenCV, TensorFlow). It determines lifestyle habits from the images (input data: environmental image files, output data: analysis results). Specifically, the server runs a model to recognize the user's behavior (e.g., not eating) from the images and obtains the analysis results.

[0309] Step 6:

[0310] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) to analyze the user's emotional state from the content of the user's dialogue and image data. This identifies the user's emotional state (input data: text data, environmental image files, output data: emotional state). Specifically, the server uses a model that analyzes the user's emotions (e.g., joy, sadness, anger, etc.) from the text data and image data.

[0311] Step 7:

[0312] The server uses a generative AI model to generate appropriate advice based on the collected information (conversation content, lifestyle habits, emotional state) (input data: analysis results, emotional state, output data: advice). Specifically, the server uses a generative model (e.g., GPT-3 (registered trademark) or a similar model) to generate advice and support messages that are most suited to the user's current situation.

[0313] Step 8:

[0314] The server sends the generated advice to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API). The device notifies the user of the advice (input data: advice, output data: notification to user). Specifically, the device communicates the received advice to the user using a voice output device or a display device.

[0315] Step 9:

[0316] The server saves all conversation content, analysis results, and notification content in a database (input data: all recorded data, output data: saved data). MySQL or PostgreSQL is used as the database. Specifically, the server stores the data in the database in an appropriate format so that it can be saved and referenced later. This data can be referenced later or used for analysis.

[0317] (Application example 2)

[0318] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0319] Elderly people, especially those with dementia, require assistance in various aspects of daily life. However, this places a heavy burden on family members and caregivers, and even situations where elderly people can act independently, such as shopping at a store, can be accompanied by anxiety. Therefore, there is a need for effective systems that support the safety and independence of elderly people and reduce the burden on family members and caregivers. In particular, there is a need for technology that can analyze the user's emotional state and provide appropriate assistance based on that analysis.

[0320] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0321] In this invention, the server includes a natural language processing means for collecting information based on a dialogue with a user, an information storage means for storing the collected information in a database, an image recognition means for analyzing daily activities using image recognition technology, a support means for generating appropriate advice based on the analysis results and providing it to the user, an emotion analysis means for analyzing the user's emotional state, and a display means for supporting elderly people's shopping and daily life in physical stores. This allows elderly people to move around safely and independently in physical stores and reduces the burden on family members and caregivers.

[0322] "Natural language processing means" refers to technology for collecting information through dialogue with users and analyzing it as text data.

[0323] "Information storage means" refers to a system that stores collected information in a database and makes it possible to refer to it as needed.

[0324] "Image recognition means" refers to technology that analyzes image data acquired through devices such as cameras and recognizes specific objects or situations.

[0325] "Support means" refers to a system that provides appropriate advice and guidance to users based on collected information and analysis results.

[0326] "Emotion analysis means" refers to technology that analyzes the content of a user's dialogue and image data to identify the user's emotional state.

[0327] "Display means" refers to technology that provides users with visual or auditory information to assist elderly people in their shopping and daily life in physical stores.

[0328] The system of the present invention is designed to support elderly people in their shopping and daily life in brick-and-mortar stores. The detailed configuration of the system and how it is implemented will be described below.

[0329] System Configuration

[0330] Hardware

[0331] Smart glasses: A device worn by the user that contains a camera and microphone.

[0332] Server: A central system for data processing and storage.

[0333] software

[0334] Natural language processing means: Manages dialogue with users using services such as Amazon AWS (registered trademark)'s Lex service or Google's Dialogflow.

[0335] Information storage method: Information is stored using a database management system such as SQLite3 or MySQL.

[0336] Image recognition method: Uses image recognition technologies such as OpenCV and YOLOv3.

[0337] Sentiment analysis method: Sentiment analysis is performed using deep learning libraries such as Transformers and Pytorch.

[0338] Display means: Uses the display function of the smart glasses to provide information visually to the user.

[0339] Program processing explanation

[0340] The server first collects input information from the camera and microphone while the user is wearing the smart glasses. A natural language processing means then uses the conversation with the user to analyze the information as text data. For example, if the user says, "I'm not feeling well today," the content is converted into text data.

[0341] The collected information is stored in a database by an information storage means, and this information is available for future reference and used to track the user's progress.

[0342] The camera on the device periodically captures image data, which is then sent to a server. The server then uses image recognition to analyze the image data and identify, for example, the status of a product shelf. For example, when a specific product is in view, information about that product is notified to the user.

[0343] The emotional state of the user is analyzed from the content of the user's dialogue and image data through the emotion analysis means. Based on the analysis results, advice and guidance appropriate to the user's current emotional state is generated and provided to the user through the display means.

[0344] Specific examples

[0345] 1. A user puts on smart glasses and enters a physical store.

[0346] 2. Immediately after entering the store, the natural language processing means asks the user, "How are you feeling today?" The user replies, "I'm a little tired."

[0347] 3. The server stores this information in a database and uses emotion analysis means to recognize the emotion "tired."

[0348] 4. The device's camera scans the shelves, and image recognition detects immune-boosting health foods.

[0349] 5. The user is notified through the display means, "This is a health food that boosts your immunity. Why not give it a try?"

[0350] Prompt Sentence Examples

[0351] 1. "Hi, how are you?"

[0352] 2. "How was your meal today?"

[0353] 3. "Did you find what you were looking for?"

[0354] In this way, this system covers all the technologies necessary to provide an environment where elderly people can shop independently and safely. By combining comprehensive information processing by the server with interaction through smart glasses, it is possible to improve the quality of life of users.

[0355] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0356] Processing Steps

[0357] Step 1:

[0358] The device collects user input information through the smart glasses' camera and microphone. The input data includes the user's voice and the video captured by the camera. This input data is sent to the server as image data for the camera video and as audio data for the user's voice.

[0359] Step 2:

[0360] After the server receives the voice data, it uses natural language processing means to convert the voice data into text data. For example, it analyzes the voice of the question "Hello, how are you?" and generates text data. The input is voice data and the output is text data.

[0361] Step 3:

[0362] The server continues the dialogue with the user based on the converted text data. The converted text data is analyzed using natural language processing means to generate appropriate questions and answers. The generated dialogue content is sent to the terminal as voice. The input is the converted text data, and the output is the generated dialogue content.

[0363] Step 4:

[0364] The device receives the generated dialogue content sent from the server and notifies the user by voice, for example, asking the user through the smart glasses, "How are you feeling today?" The user's voice response is collected again and sent to the server.

[0365] Step 5:

[0366] The server receives the image data and analyzes the everyday environment using image recognition means. Specifically, it identifies objects from camera footage, for example, identifying products on a shelf. The input is image data, and the output is data on the recognized objects.

[0367] Step 6:

[0368] The server generates appropriate advice and guidance based on the analysis results. It generates information to be provided to the user for the recognized object and creates guidance in text format. The input is the recognized object data, and the output is the text data of the advice and guidance.

[0369] Step 7:

[0370] The server uses emotion analysis means to analyze the user's emotional state from the user's voice and image data. For example, if the user replies, "I'm not feeling well today," the server analyzes the text data and voice tone to identify the user's emotional state. The input is the voice and image data, and the output is the analyzed emotional data.

[0371] Step 8:

[0372] The server generates more appropriate dialogue content based on the generated advice and guidance and the analyzed emotional data. For example, if the server determines that the user is tired, it generates a dialogue recommending "health foods that boost the immune system." The input is the text data of the advice and guidance and emotional data, and the output is the optimized dialogue content.

[0373] Step 9:

[0374] The device receives the generated dialogue and notifies the user through the smart glasses. It provides the user with visual and auditory information to assist with shopping. For example, it displays a message such as, "This is a health food that boosts your immunity. Why not give it a try?"

[0375] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0376] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0377] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0378] [Second embodiment]

[0379] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0380] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0381] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0382] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0383] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0384] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0385] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0386] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0387] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0388] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0389] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0390] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0391] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system supports the lives of elderly people by collecting information through dialogue with the user and analyzing the user's lifestyle habits using image recognition technology.

[0392] System configuration

[0393] The system consists of the following main components:

[0394] 1. Language processing means: Collects necessary information through dialogue with the user.

[0395] 2. Data retention measures: Collected information is stored in a database.

[0396] 3. Image recognition means: Analyzes image data acquired using a camera device.

[0397] 4. Assistance: Providing appropriate advice based on the collected information and analysis results.

[0398] Implementation of language processing measures

[0399] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[0400] Implementing data retention measures

[0401] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[0402] Implementation of image recognition measures

[0403] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[0404] Implementation of assistance measures

[0405] The server generates appropriate advice based on the collected information and image recognition results and provides it to the user. For example, it generates advice such as, "Don't forget to eat, start eating now." This advice is notified to the user via their device. Similar information is also notified to family members as needed.

[0406] Example

[0407] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that no food is found. It then asks the user through language processing means, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now." This advice is notified to the user through the device, and the information is also sent to family members if necessary.

[0408] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[0409] The processing flow will be explained below.

[0410] Step 1:

[0411] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[0412] Step 2:

[0413] The device sends the captured image data to the server, where it is sent in real time over the network.

[0414] Step 3:

[0415] The server pre-processes the received image data, including image resizing, noise reduction, and color correction.

[0416] Step 4:

[0417] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[0418] Step 5:

[0419] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[0420] Step 6:

[0421] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[0422] Step 7:

[0423] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[0424] Step 8:

[0425] The user responds to the initial message through the terminal, either by voice or text.

[0426] Step 9:

[0427] The device captures the user's response and converts it into text data using voice recognition technology.

[0428] Step 10:

[0429] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[0430] Step 11:

[0431] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[0432] Step 12:

[0433] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[0434] Step 13:

[0435] The server makes the final decision and generates appropriate advice, e.g., "Don't forget to eat, let's eat now."

[0436] Step 14:

[0437] The server sends the generated advice to the device, which then notifies the user via voice or text.

[0438] Step 15:

[0439] The server will notify the family of the same information as necessary. For example, it may notify the family by email or SMS that "The user has not eaten a meal today."

[0440] Through the above steps, it is possible to comprehensively support the user's daily life and provide a safe living environment.

[0441] Example 1

[0442] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0443] This invention relates to a system that comprehensively supports the lives of elderly people with dementia and reduces the burden on their families and caregivers. Conventional systems have difficulty accurately understanding a user's lifestyle and health status, and have been unable to provide appropriate advice or notifications. Therefore, more effective information collection and analysis is needed to maintain a user's health and improve their quality of life.

[0444] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0445] In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a storage means, an image recognition means for analyzing image data acquired using a camera device, and an assistance means for generating appropriate notifications or advice based on the analysis results and the collected information and providing them to the user. This makes it possible to appropriately grasp the user's behavior and health condition in real time and quickly provide necessary advice or notifications.

[0446] "User" refers to the individual or family member who uses the system, and is often targeted at elderly people with dementia.

[0447] "Language processing means" refers to technology that collects information from dialogue with the user and converts it into text data.

[0448] "Data retention means" refers to technology that stores collected information in a storage means such as a database, making it available for later reference.

[0449] A "camera device" is hardware for acquiring image data, and is used to monitor the user's living environment.

[0450] "Image recognition means" refers to technology that analyzes image data acquired by a camera device and determines the user's behavior and situation.

[0451] "Assistance means" refers to technology that generates and provides appropriate notifications and advice to users based on collected information and analysis results.

[0452] "Analysis results" refers to the analysis data obtained by the language processing means and the image recognition means.

[0453] "Notification" means a message or alert provided by the System to a User or their Family Members.

[0454] "Advice" refers to the guidance the system provides to the user to encourage improvements in their lifestyle and health.

[0455] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and provide an environment in which they can live with peace of mind. This system consists of four main components: language processing means, data storage means, image recognition means, and assistance means.

[0456] Language Processing Methods

[0457] The server uses natural language processing technology (for example, a generative AI model such as OpenAI's GPT-4) to analyze the user's dialogue and collect necessary information. The dialogue with the user is conducted through the device, which recognizes the user's voice and sends it to the server as text data. For example, if the device asks the user, "Hello, how are you?" and the user responds, "I'm a little tired today," that information is collected by the server.

[0458] Data Retention Methods

[0459] The collected information is stored by the server in a database (for example, an RDBMS such as MySQL or PostgreSQL). This records the user's most recent actions and status, and can be referenced later. For example, data such as "The user responded that they were tired at 10:00 AM on October 13, 2023" is stored.

[0460] Image Recognition Method

[0461] The camera device installed in the device periodically captures the user's living environment and sends the image data to a server. The server then uses image recognition technology to analyze this data. For example, if the camera analyzes an image of the user's room and no food is found, it can determine that the user has not had breakfast.

[0462] Assistance Method

[0463] The server generates appropriate notifications or advice based on the collected information and the results of image recognition analysis, and provides them to the user. For example, it generates a notification saying, "Don't forget to eat, let's eat now." This notification is sent to the user via their device, and similar notifications are also sent to family members if necessary.

[0464] Specific examples

[0465] For example, one morning, the device captures the state of the user's room with a camera and sends the image data to the server. The server analyzes this image data and confirms that no food is found. The device then asks the user, "How was your breakfast today?" through language processing means. If the user replies, "I forgot," this information is saved in the database. The server then generates advice such as, "Don't forget to eat, start eating now," and notifies this to the user via the device. Information that "the user may not have had breakfast" is also sent to the user's family.

[0466] Example prompt sentence:

[0467] Ask the user, "Hello, how are you?". Before the user asks, "How was your breakfast today?", save the user's most recent meal history in a database. Then, analyze the camera footage and generate an advice if the user hasn't eaten, saying, "Don't forget to eat, eat now."

[0468] This system provides comprehensive support for users' daily lives, reducing the burden on their families and caregivers, and provides appropriate advice based on the collected information, creating an environment in which users can live with peace of mind.

[0469] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0470] Step 1: Initiate a conversation with the user

[0471] Input: The initial message the terminal is set to (e.g. "Hello, how are you?")

[0472] Action: The device asks the user a question using voice.

[0473] Output: The user responds verbally (e.g., "I'm a little tired today").

[0474] Step 2: Collect and analyze user responses

[0475] Input: User's voice response

[0476] How it works: The device uses voice recognition technology to convert the user's response into text data, which it then sends to the server.

[0477] Output: Text data (e.g., "I'm a little tired today")

[0478] Step 3: Information analysis using natural language processing

[0479] Input: Text data (e.g., "I'm a little tired today")

[0480] How it works: The server uses a generative AI model (e.g., OpenAI's GPT-4) to analyze text data and evaluate the user's state.

[0481] Output: Analysis result (e.g. "The user is tired")

[0482] Step 4: Storing information in a database

[0483] Input: Analysis result (e.g. "The user is tired")

[0484] How it works: The server stores the analysis results in a database.

[0485] Output: Stored data (e.g., "User responded that he was tired on October 13, 2023 at 10:00 AM")

[0486] Step 5: Capture camera images

[0487] Input: Timer setting for camera device (e.g. capture every hour)

[0488] Operation: The camera device installed on the device captures an image and sends the image data to the server.

[0489] Output: Image data (e.g., a photo of the user's room)

[0490] Step 6: Data analysis using image recognition

[0491] Input: Image data (e.g., a photo of the user's room)

[0492] How it works: The server uses image recognition technology to analyze the image data and determine the user's situation (e.g., if there is no food in sight, it determines that the user is not eating).

[0493] Output: Image analysis results (e.g., "The user has not eaten")

[0494] Step 7: Generate and notify advice

[0495] Input: User information analysis results and image analysis results

[0496] Operation: The server generates appropriate advice based on the collected information and image recognition results. The generated advice (e.g., "Don't forget to eat, let's eat now") is notified to the user via the device. If necessary, the information is also sent to family members.

[0497] Output: Advice notification (e.g., notification from the device to the user saying "Don't forget to eat, let's eat now" and sharing the information with family members)

[0498] (Application example 1)

[0499] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0500] The present invention aims to provide an environment in which elderly people with dementia can use self-driving vehicles with peace of mind. Currently, there is a lack of means to ensure the safe and comfortable travel of elderly people with dementia in self-driving vehicles, which creates a risk of elderly people behaving incorrectly. Specifically, this poses problems such as forgetting to eat and not taking care of their health. In response to this, the present invention solves these issues by monitoring the condition of elderly people in real time and providing appropriate warnings and instructions.

[0501] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0502] In this invention, the server includes language processing means for collecting information based on dialogue with the user, data storage means for storing the collected information in a database, image recognition means for analyzing lifestyle habits using image recognition technology, assistance means for generating appropriate advice based on the analysis results and providing it to the user, means for monitoring the condition of the elderly person using a camera device installed in the vehicle and acquiring image data, and means for analyzing the acquired image and audio data and providing instructions and warnings tailored to the elderly person's condition. This enables an environment in which elderly people with dementia can use self-driving vehicles safely and comfortably.

[0503] The "language processing means" is a means for collecting necessary information based on a dialogue with the user.

[0504] "Data storage means" refers to a means for storing collected information in a database.

[0505] The "image recognition means" is a means for analyzing a user's lifestyle habits using image recognition technology.

[0506] The "assisting means" is a means for generating appropriate advice based on the analysis results and providing it to the user.

[0507] The "camera device" is a device that is installed in a vehicle to monitor the condition of the elderly person and acquire image data.

[0508] The "image and audio data analysis means" is a means for analyzing the acquired image and audio data and providing instructions and warnings that are tailored to the elderly person's situation.

[0509] The system of this invention is designed to provide an environment in which elderly people with dementia can safely use autonomous vehicles. The system consists of the following main components:

[0510] 1. Language Processing Methods

[0511] The server collects necessary information through dialogue with the user. Specifically, it responds to questions and greetings from the user through voice input, and analyzes the response using natural language processing technology. For example, if an elderly person responds "I forgot" to the question "How was your breakfast today?", this information is processed and recorded as text data. The software used includes the SpeechRecognition library and the gTTS library.

[0512] 2. Data Retention Methods

[0513] The server stores the collected information in a database. This database is used to record the user's daily behavior and status history. For example, information such as "I did not eat breakfast at 10:00 AM on October 12, 2023" is recorded. This allows the system to provide appropriate advice while referring to past history.

[0514] 3. Image Recognition Methods

[0515] A camera device installed inside the vehicle monitors the elderly person's condition and periodically captures image data. The captured image data is sent to a server and analyzed using image recognition technology. Specifically, the OpenCV library is used to process the images and perform object detection and behavior recognition. For example, if no food is found, it is determined that the elderly person has not eaten.

[0516] 4. Assistance methods

[0517] The server generates appropriate advice based on the collected information and the results of image recognition analysis, and provides it to the user. The advice is communicated to the user via voice. For example, advice such as "Don't forget to eat, start eating now" is generated and communicated to the user through a speaker.

[0518] 5. Image and audio data analysis methods

[0519] The acquired image and audio data is processed in real time. This method is used to provide instructions and warnings tailored to the user's situation. The image and audio data captured by the camera are used for analysis, and the server comprehensively analyzes these to provide optimal advice.

[0520] Specific example explanation

[0521] For example, one morning, the server uses a camera inside the vehicle to capture the state of an elderly person's room. Analysis of the image data confirms that food is missing. The server then asks aloud, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and advises, "Don't forget to eat. Let's eat now." This advice is communicated to the user through the speaker, and the information is also sent to family members if necessary.

[0522] Prompt Sentence Examples

[0523] Here are some examples of specific prompts:

[0524] Elderly Co-Driver Assistance prompts:

[0525] 1. Capture an image with the camera and send it to an image analysis API to determine if food is in the image.

[0526] 2. A voice prompt asks the elderly person, "How was your breakfast today?"

[0527] 3. If the user answers "I forgot," the system provides the advice "Don't forget to eat, eat now."

[0528] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[0529] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0530] Step 1:

[0531] The server uses a camera device to monitor the condition of the elderly person in the vehicle and captures image data.

[0532] Input: Real-time video from a camera device.

[0533] Data processing and calculation: Images are captured using the OpenCV library and saved as still images.

[0534] Output: The captured image data.

[0535] Step 2:

[0536] The server sends the captured image data to an image recognition API to determine whether food is in the image.

[0537] Input: The captured image data.

[0538] Data processing and calculation: Send image data to the image recognition API and obtain the analysis results.

[0539] Output: Analysis result (e.g., determination that food is not visible).

[0540] Step 3:

[0541] Based on the analysis results, the server asks the user aloud, "How was your breakfast today?"

[0542] Input: Analysis results (e.g. food not shown).

[0543] Data processing and calculation: The question is converted into an audio file using the gTTS library and played through the speaker.

[0544] Output: A spoken question to the user.

[0545] Step 4:

[0546] The user responds to the voice questions and the server receives the response as voice input.

[0547] Input: The user's spoken response.

[0548] Data processing and calculation: Convert speech to text using the SpeechRecognition library.

[0549] Output: User response as text data (e.g., "I forgot").

[0550] Step 5:

[0551] The server records the user's response in a database.

[0552] Input: User response as text data.

[0553] Data processing and calculation: Connect to the database and record the response content in an appropriate format.

[0554] Output: The recorded database entries.

[0555] Step 6:

[0556] Based on the user's response, the server generates advice such as "Don't forget to eat, start eating now," and notifies the user by voice.

[0557] Input: User response as text data.

[0558] Data processing and calculation: The advice sentence is converted into an audio file using the gTTS library and played through the speaker.

[0559] Output: Audio advice to the user.

[0560] Step 7:

[0561] The server notifies the user's family of the user's condition and advice as necessary.

[0562] Input: User responses and advice as text data.

[0563] Data processing and calculation: Send text messages and emails to family members through the notification system.

[0564] Output: Notification message to family members.

[0565] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0566] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and an emotion engine to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[0567] System configuration

[0568] The system consists of the following main components:

[0569] 1. Language processing means: Collects necessary information through dialogue with the user.

[0570] 2. Data retention measures: Collected information is stored in a database.

[0571] 3. Image recognition means: Analyzes image data acquired using a camera device.

[0572] 4. Emotion Engine: Analyzes the emotional state from user interactions and image data.

[0573] 5. Assistance: Providing appropriate advice based on the collected information and analysis results.

[0574] Implementation of language processing measures

[0575] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[0576] Implementing data retention measures

[0577] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[0578] Implementation of image recognition measures

[0579] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[0580] Emotion Engine Implementation

[0581] The server runs an emotion engine based on the user's dialogue, voice tone, and even image data sent from the device. This engine analyzes the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.). The analysis results can be used to more specifically determine the most appropriate advice.

[0582] Implementation of assistance measures

[0583] The server generates appropriate advice based on the collected information and the analysis results of the emotion engine, and provides it to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, it can take an approach such as "Is there anything troubling you?" This advice is notified to the user via their device. Similar information can also be notified to family members if necessary.

[0584] Example

[0585] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that there is no food in sight. It also uses an emotion engine to determine that the user's facial expression is "gloomy." The server then uses language processing to ask the user, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now. Also, let us know if you need any help." This advice is notified to the user via the device, and the information is also sent to family members if necessary.

[0586] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding an emotion engine, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[0587] The processing flow will be explained below.

[0588] Step 1:

[0589] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[0590] Step 2:

[0591] The device sends the captured image data to the server, where it is sent in real time over the network.

[0592] Step 3:

[0593] The server pre-processes the image data it receives, including image resizing, noise reduction, and color correction.

[0594] Step 4:

[0595] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[0596] Step 5:

[0597] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[0598] Step 6:

[0599] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[0600] Step 7:

[0601] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[0602] Step 8:

[0603] The user responds to the initial message through the terminal, either by voice or text.

[0604] Step 9:

[0605] The device captures the user's response and converts it into text data using voice recognition technology.

[0606] Step 10:

[0607] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[0608] Step 11:

[0609] The server uses an emotion engine to analyze the emotional state from the user's response and image data, e.g., to identify the emotional state such as "sad" or "happy."

[0610] Step 12:

[0611] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[0612] Step 13:

[0613] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[0614] Step 14:

[0615] The server makes the final decision and generates appropriate advice. For example, "Don't forget to eat, let's eat now. Also, are you feeling sad? Can you tell me something?"

[0616] Step 15:

[0617] The server sends the generated advice to the device, which then notifies the user via voice or text.

[0618] Step 16:

[0619] The server notifies the family of the same information as necessary. For example, it may notify the family by email or SMS, saying, "The user does not seem to have eaten a meal today. Also, he may be depressed."

[0620] Through the above steps, it is possible to provide comprehensive support for the user's daily life and provide appropriate responses according to their emotional state.

[0621] Example 2

[0622] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0623] The problem that this invention aims to solve is to provide appropriate support to elderly people with dementia and their families based on their individual lifestyle habits and emotional states, and to create an environment where elderly people with dementia can live with peace of mind. To achieve this, it is important to understand the user's daily behavior and emotions in real time and provide appropriate advice. However, existing systems do not adequately analyze emotions or recognize lifestyle habits, making it difficult to provide effective support to users and their families.

[0624] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a database, an image recognition means for analyzing lifestyle habits using image recognition technology, an emotion analysis means for analyzing the emotional state from the content of the user's dialogue and image data, and an assistance means for generating appropriate advice based on the analysis results and the emotional state and providing the advice to the user. This makes it possible to grasp the user's daily behavior and emotions in real time and provide appropriate support based on them.

[0625] "Language processing means" refers to devices or software that collect information based on dialogue with a user and analyze that information using natural language processing technology.

[0626] "Data storage means" refers to devices or software that store collected information in a database so that it can be referenced later.

[0627] "Image recognition means" refers to a technology or device that analyzes image data acquired by a camera device and determines the lifestyle habits of a user.

[0628] "Emotion analysis means" refers to a device or software that analyzes the emotional state of a user from the content of their dialogue or image data.

[0629] "Assistance means" refers to a device or software that generates appropriate advice based on the analysis results and emotional state and provides it to the user.

[0630] The term "camera device" refers to a device for monitoring a user's living environment and acquiring image data.

[0631] "Means for analyzing lifestyle habits" refers to technology or devices that analyze and judge a user's behavior and habits based on acquired image data.

[0632] "Emotional state" refers to the user's psychological state and feelings, as judged from the content of the user's dialogue, facial expression, tone of voice, etc.

[0633] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and emotion analysis technology to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[0634] Implementation of language processing measures

[0635] The server uses natural language processing technology to analyze the conversation with the user and collect the necessary information. To do this, it uses software such as Google Cloud Natural Language API and IBM Watson as natural language processing technology (NLP). For example, the server asks the user, "Hello, how are you?" and the user's response, "I'm not feeling well," is converted into text data using speech recognition technology (e.g., Google Speech-to-Text API) and analyzed.

[0636] Implementing data retention measures

[0637] The collected information is stored in a database by the server. MySQL or PostgreSQL is used as the database. This records the user's most recent actions and status, allowing them to be referenced later. For example, it may record "No meals at 10:00 AM on October 12, 2023."

[0638] Implementation of image recognition measures

[0639] A camera device (e.g., a Raspberry Pi camera module) installed on the device periodically captures image data and sends it to a server. The server then analyzes the data using image recognition technology (e.g., OpenCV, TensorFlow). For example, if no food is found in the video, it determines that the person has not eaten.

[0640] Implementing sentiment analysis measures

[0641] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) based on the user's dialogue content, voice tone, and image data sent from the device to analyze the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.).

[0642] Implementation of assistance measures

[0643] The server uses a generative AI model to generate appropriate advice based on the collected information and the results of emotion analysis. The generated advice is sent to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API) and provided to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, an approach such as "Is there anything troubling you?" is taken. This advice is notified to the user via the device, and similar information is also notified to family members if necessary.

[0644] Example

[0645] Example 1: Morning dialogue scenario

[0646] At 6:00 a.m., the device speaks to the user, asking, "Good morning. How are you feeling today?" If the user responds, "I'm not feeling very good today," the device captures this voice and sends it to the server. The server converts this into text data and uses NLP technology to analyze the "bad mood." The analysis determines that the user is feeling "depressed," which is supported by emotion analysis technology. The server generates advice such as, "Let's find something to cheer you up together. Is there anything troubling you?" and notifies the user of this via the device.

[0647] Example 2: Lunch Check Scenario

[0648] At 12:00, the device captures the state of the room with its camera and sends it to the server. The server analyzes the image data and determines that "food is missing." It also uses emotion analysis technology to determine that the user's facial expression is "gloomy." The server then asks the user, "How was your lunch today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat. Let's eat now. Also, let us know if you need any help," and notifies the user through the device.

[0649] Prompt Sentence Examples

[0650] Example prompts to be input to the generative AI model:

[0651] "Please explain in detail the processing steps of the daily life support system for elderly people with dementia. Please include the specific operations and technologies for each step, the names of the hardware and software used, and examples of use."

[0652] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding emotion analysis technology, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[0653] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0654] Step 1:

[0655] A user accesses the system, and the device asks a question via the microphone, such as "Hello, how are you?" The device obtains the user's response (voice). This response is captured by the device's microphone (input data: user's voice). Specifically, the device activates its voice recognition function to capture the voice signal and sends it to the server as a digital voice file (output data: digital voice file).

[0656] Step 2:

[0657] The server uses speech recognition technology (e.g., Google Speech-to-Text API) to analyze the received audio file and convert it into text data (input data: digital audio file, output data: text data). Specifically, the server samples the audio data and uses a speech recognition model to convert it into a string of characters.

[0658] Step 3:

[0659] The server analyzes the converted text data using natural language processing technology (e.g., Google Cloud Natural Language API), thereby extracting semantic information from the user's response (input data: text data, output data: semantic information). Specifically, the server analyzes the text using techniques such as tokenization, part-of-speech tagging, and sentiment analysis to extract information about the user's state and emotions.

[0660] Step 4:

[0661] The device periodically captures image data of the environment. The camera device acquires image data of the user's living environment (input data: environmental image). Specifically, the device activates the camera at set time intervals, captures images, and sends them to the server (output data: environmental image file).

[0662] Step 5:

[0663] The server analyzes the received image data using image recognition technology (e.g., OpenCV, TensorFlow). It determines lifestyle habits from the images (input data: environmental image files, output data: analysis results). Specifically, the server runs a model to recognize the user's behavior (e.g., not eating) from the images and obtains the analysis results.

[0664] Step 6:

[0665] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) to analyze the user's emotional state from the content of the user's dialogue and image data. This identifies the user's emotional state (input data: text data, environmental image files, output data: emotional state). Specifically, the server uses a model that analyzes the user's emotions (e.g., joy, sadness, anger, etc.) from the text data and image data.

[0666] Step 7:

[0667] The server uses a generative AI model to generate appropriate advice based on the collected information (conversation content, lifestyle habits, emotional state) (input data: analysis results, emotional state, output data: advice). Specifically, the server uses a generative model (e.g., GPT-3 or a similar model) to generate advice or support messages that are most appropriate for the user's current situation.

[0668] Step 8:

[0669] The server sends the generated advice to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API). The device notifies the user of the advice (input data: advice, output data: notification to user). Specifically, the device communicates the received advice to the user using a voice output device or a display device.

[0670] Step 9:

[0671] The server saves all conversation content, analysis results, and notification content in a database (input data: all recorded data, output data: saved data). MySQL or PostgreSQL is used as the database. Specifically, the server stores the data in the database in an appropriate format so that it can be saved and referenced later. This data can be referenced later or used for analysis.

[0672] (Application example 2)

[0673] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0674] Elderly people, especially those with dementia, require assistance in various aspects of daily life. However, this places a heavy burden on family members and caregivers, and even situations where elderly people can act independently, such as shopping at a store, can be accompanied by anxiety. Therefore, there is a need for effective systems that support the safety and independence of elderly people and reduce the burden on family members and caregivers. In particular, there is a need for technology that can analyze the user's emotional state and provide appropriate assistance based on that analysis.

[0675] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0676] In this invention, the server includes a natural language processing means for collecting information based on a dialogue with a user, an information storage means for storing the collected information in a database, an image recognition means for analyzing daily activities using image recognition technology, a support means for generating appropriate advice based on the analysis results and providing it to the user, an emotion analysis means for analyzing the user's emotional state, and a display means for supporting elderly people's shopping and daily life in physical stores. This allows elderly people to move around safely and independently in physical stores and reduces the burden on family members and caregivers.

[0677] "Natural language processing means" refers to technology for collecting information through dialogue with users and analyzing it as text data.

[0678] "Information storage means" refers to a system that stores collected information in a database and makes it possible to refer to it as needed.

[0679] "Image recognition means" refers to technology that analyzes image data acquired through devices such as cameras and recognizes specific objects or situations.

[0680] "Support means" refers to a system that provides appropriate advice and guidance to users based on collected information and analysis results.

[0681] "Emotion analysis means" refers to technology that analyzes the content of a user's dialogue and image data to identify the user's emotional state.

[0682] "Display means" refers to technology that provides users with visual or auditory information to assist elderly people in their shopping and daily life in physical stores.

[0683] The system of the present invention is designed to support elderly people in their shopping and daily life in brick-and-mortar stores. The detailed configuration of the system and how it is implemented will be described below.

[0684] System Configuration

[0685] Hardware

[0686] Smart glasses: A device worn by the user that contains a camera and microphone.

[0687] Server: A central system for data processing and storage.

[0688] software

[0689] Natural language processing means: Manage dialogue with users using services such as Amazon AWS's Lex service or Google's Dialogflow.

[0690] Information storage method: Information is stored using a database management system such as SQLite3 or MySQL.

[0691] Image recognition method: Uses image recognition technologies such as OpenCV and YOLOv3.

[0692] Sentiment analysis method: Sentiment analysis is performed using deep learning libraries such as Transformers and Pytorch.

[0693] Display means: Uses the display function of the smart glasses to provide information visually to the user.

[0694] Program processing explanation

[0695] The server first collects input information from the camera and microphone while the user is wearing the smart glasses. A natural language processing means then uses the conversation with the user to analyze the information as text data. For example, if the user says, "I'm not feeling well today," the content is converted into text data.

[0696] The collected information is stored in a database by an information storage means, and this information is available for future reference and used to track the user's progress.

[0697] The camera on the device periodically captures image data, which is then sent to a server. The server then uses image recognition to analyze the image data and identify, for example, the status of a product shelf. For example, when a specific product is in view, information about that product is notified to the user.

[0698] The emotional state of the user is analyzed from the content of the user's dialogue and image data through the emotion analysis means. Based on the analysis results, advice and guidance appropriate to the user's current emotional state is generated and provided to the user through the display means.

[0699] Specific examples

[0700] 1. A user puts on smart glasses and enters a physical store.

[0701] 2. Immediately after entering the store, the natural language processing means asks the user, "How are you feeling today?" The user replies, "I'm a little tired."

[0702] 3. The server stores this information in a database and uses emotion analysis means to recognize the emotion "tired."

[0703] 4. The device's camera scans the shelves, and image recognition detects immune-boosting health foods.

[0704] 5. The user is notified through the display means, "This is a health food that boosts your immunity. Why not give it a try?"

[0705] Prompt Sentence Examples

[0706] 1. "Hi, how are you?"

[0707] 2. "How was your meal today?"

[0708] 3. "Did you find what you were looking for?"

[0709] In this way, this system covers all the technologies necessary to provide an environment where elderly people can shop independently and safely. By combining comprehensive information processing by the server with interaction through smart glasses, it is possible to improve the quality of life of users.

[0710] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0711] Processing Steps

[0712] Step 1:

[0713] The device collects user input information through the smart glasses' camera and microphone. The input data includes the user's voice and the video captured by the camera. This input data is sent to the server as image data for the camera video and as audio data for the user's voice.

[0714] Step 2:

[0715] After the server receives the voice data, it uses natural language processing means to convert the voice data into text data. For example, it analyzes the voice of the question "Hello, how are you?" and generates text data. The input is voice data and the output is text data.

[0716] Step 3:

[0717] The server continues the dialogue with the user based on the converted text data. The converted text data is analyzed using natural language processing means to generate appropriate questions and answers. The generated dialogue content is sent to the terminal as voice. The input is the converted text data, and the output is the generated dialogue content.

[0718] Step 4:

[0719] The device receives the generated dialogue content sent from the server and notifies the user by voice, for example, asking the user through the smart glasses, "How are you feeling today?" The user's voice response is collected again and sent to the server.

[0720] Step 5:

[0721] The server receives the image data and analyzes the everyday environment using image recognition means. Specifically, it identifies objects from camera footage, for example, identifying products on a shelf. The input is image data, and the output is data on the recognized objects.

[0722] Step 6:

[0723] The server generates appropriate advice and guidance based on the analysis results. It generates information to be provided to the user for the recognized object and creates guidance in text format. The input is the recognized object data, and the output is the text data of the advice and guidance.

[0724] Step 7:

[0725] The server uses emotion analysis means to analyze the user's emotional state from the user's voice and image data. For example, if the user replies, "I'm not feeling well today," the server analyzes the text data and voice tone to identify the user's emotional state. The input is the voice and image data, and the output is the analyzed emotional data.

[0726] Step 8:

[0727] The server generates more appropriate dialogue content based on the generated advice and guidance and the analyzed emotional data. For example, if the server determines that the user is tired, it generates a dialogue recommending "health foods that boost the immune system." The input is the text data of the advice and guidance and emotional data, and the output is the optimized dialogue content.

[0728] Step 9:

[0729] The device receives the generated dialogue and notifies the user through the smart glasses. It provides the user with visual and auditory information to assist with shopping. For example, it displays a message such as, "This is a health food that boosts your immunity. Why not give it a try?"

[0730] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0731] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0732] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0733] [Third embodiment]

[0734] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0735] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0736] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0737] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0738] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0739] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0740] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0741] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0742] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0743] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0744] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0745] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0746] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system supports the lives of elderly people by collecting information through dialogue with the user and analyzing the user's lifestyle habits using image recognition technology.

[0747] System configuration

[0748] The system consists of the following main components:

[0749] 1. Language processing means: Collects necessary information through dialogue with the user.

[0750] 2. Data retention measures: Collected information is stored in a database.

[0751] 3. Image recognition means: Analyzes image data acquired using a camera device.

[0752] 4. Assistance: Providing appropriate advice based on the collected information and analysis results.

[0753] Implementation of language processing measures

[0754] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[0755] Implementing data retention measures

[0756] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[0757] Implementation of image recognition measures

[0758] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[0759] Implementation of assistance measures

[0760] The server generates appropriate advice based on the collected information and image recognition results and provides it to the user. For example, it generates advice such as, "Don't forget to eat, start eating now." This advice is notified to the user via their device. Similar information is also notified to family members as needed.

[0761] Example

[0762] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that no food is found. It then asks the user through language processing means, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now." This advice is notified to the user through the device, and the information is also sent to family members if necessary.

[0763] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[0764] The processing flow will be explained below.

[0765] Step 1:

[0766] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[0767] Step 2:

[0768] The device sends the captured image data to the server, where it is sent in real time over the network.

[0769] Step 3:

[0770] The server pre-processes the received image data, including image resizing, noise reduction, and color correction.

[0771] Step 4:

[0772] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[0773] Step 5:

[0774] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[0775] Step 6:

[0776] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[0777] Step 7:

[0778] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[0779] Step 8:

[0780] The user responds to the initial message through the terminal, either by voice or text.

[0781] Step 9:

[0782] The device captures the user's response and converts it into text data using voice recognition technology.

[0783] Step 10:

[0784] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[0785] Step 11:

[0786] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[0787] Step 12:

[0788] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[0789] Step 13:

[0790] The server makes the final decision and generates appropriate advice, e.g., "Don't forget to eat, let's eat now."

[0791] Step 14:

[0792] The server sends the generated advice to the device, which then notifies the user via voice or text.

[0793] Step 15:

[0794] The server will notify the family of the same information as necessary. For example, it may notify the family by email or SMS that "The user has not eaten a meal today."

[0795] Through the above steps, it is possible to comprehensively support the user's daily life and provide a safe living environment.

[0796] Example 1

[0797] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0798] This invention relates to a system that comprehensively supports the lives of elderly people with dementia and reduces the burden on their families and caregivers. Conventional systems have difficulty accurately understanding a user's lifestyle and health status, and have been unable to provide appropriate advice or notifications. Therefore, more effective information collection and analysis is needed to maintain a user's health and improve their quality of life.

[0799] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0800] In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a storage means, an image recognition means for analyzing image data acquired using a camera device, and an assistance means for generating appropriate notifications or advice based on the analysis results and the collected information and providing them to the user. This makes it possible to appropriately grasp the user's behavior and health condition in real time and quickly provide necessary advice or notifications.

[0801] "User" refers to the individual or family member who uses the system, and is often targeted at elderly people with dementia.

[0802] "Language processing means" refers to technology that collects information from dialogue with the user and converts it into text data.

[0803] "Data retention means" refers to technology that stores collected information in a storage means such as a database, making it available for later reference.

[0804] A "camera device" is hardware for acquiring image data, and is used to monitor the user's living environment.

[0805] "Image recognition means" refers to technology that analyzes image data acquired by a camera device and determines the user's behavior and situation.

[0806] "Assistance means" refers to technology that generates and provides appropriate notifications and advice to users based on collected information and analysis results.

[0807] "Analysis results" refers to the analysis data obtained by the language processing means and the image recognition means.

[0808] "Notification" means a message or alert provided by the System to a User or their Family Members.

[0809] "Advice" refers to the guidance the system provides to the user to encourage improvements in their lifestyle and health.

[0810] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and provide an environment in which they can live with peace of mind. This system consists of four main components: language processing means, data storage means, image recognition means, and assistance means.

[0811] Language Processing Methods

[0812] The server uses natural language processing technology (for example, a generative AI model such as OpenAI's GPT-4) to analyze the user's dialogue and collect necessary information. The dialogue with the user is conducted through the device, which recognizes the user's voice and sends it to the server as text data. For example, if the device asks the user, "Hello, how are you?" and the user responds, "I'm a little tired today," that information is collected by the server.

[0813] Data Retention Methods

[0814] The collected information is stored by the server in a database (for example, an RDBMS such as MySQL or PostgreSQL). This records the user's most recent actions and status, and can be referenced later. For example, data such as "The user responded that they were tired at 10:00 AM on October 13, 2023" is stored.

[0815] Image Recognition Method

[0816] The camera device installed in the device periodically captures the user's living environment and sends the image data to a server. The server then uses image recognition technology to analyze this data. For example, if the camera analyzes an image of the user's room and no food is found, it can determine that the user has not had breakfast.

[0817] Assistance Method

[0818] The server generates appropriate notifications or advice based on the collected information and the results of image recognition analysis, and provides them to the user. For example, it generates a notification saying, "Don't forget to eat, let's eat now." This notification is sent to the user via their device, and similar notifications are also sent to family members if necessary.

[0819] Specific examples

[0820] For example, one morning, the device captures the state of the user's room with a camera and sends the image data to the server. The server analyzes this image data and confirms that no food is found. The device then asks the user, "How was your breakfast today?" through language processing means. If the user replies, "I forgot," this information is saved in the database. The server then generates advice such as, "Don't forget to eat, start eating now," and notifies this to the user via the device. Information that "the user may not have had breakfast" is also sent to the user's family.

[0821] Example prompt sentence:

[0822] Ask the user, "Hello, how are you?". Before the user asks, "How was your breakfast today?", save the user's most recent meal history in a database. Then, analyze the camera footage and generate an advice if the user hasn't eaten, saying, "Don't forget to eat, eat now."

[0823] This system provides comprehensive support for users' daily lives, reducing the burden on their families and caregivers, and provides appropriate advice based on the collected information, creating an environment in which users can live with peace of mind.

[0824] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0825] Step 1: Initiate a conversation with the user

[0826] Input: The initial message the terminal is set to (e.g. "Hello, how are you?")

[0827] Action: The device asks the user a question using voice.

[0828] Output: The user responds verbally (e.g., "I'm a little tired today").

[0829] Step 2: Collect and analyze user responses

[0830] Input: User's voice response

[0831] How it works: The device uses voice recognition technology to convert the user's response into text data, which it then sends to the server.

[0832] Output: Text data (e.g., "I'm a little tired today")

[0833] Step 3: Information analysis using natural language processing

[0834] Input: Text data (e.g., "I'm a little tired today")

[0835] How it works: The server uses a generative AI model (e.g., OpenAI's GPT-4) to analyze text data and evaluate the user's state.

[0836] Output: Analysis result (e.g. "The user is tired")

[0837] Step 4: Storing information in a database

[0838] Input: Analysis result (e.g. "The user is tired")

[0839] How it works: The server stores the analysis results in a database.

[0840] Output: Stored data (e.g., "User responded that he was tired on October 13, 2023 at 10:00 AM")

[0841] Step 5: Capture camera images

[0842] Input: Timer setting for camera device (e.g. capture every hour)

[0843] Operation: The camera device installed on the device captures an image and sends the image data to the server.

[0844] Output: Image data (e.g., a photo of the user's room)

[0845] Step 6: Data analysis using image recognition

[0846] Input: Image data (e.g., a photo of the user's room)

[0847] How it works: The server uses image recognition technology to analyze the image data and determine the user's situation (e.g., if there is no food in sight, it determines that the user is not eating).

[0848] Output: Image analysis results (e.g., "The user has not eaten")

[0849] Step 7: Generate and notify advice

[0850] Input: User information analysis results and image analysis results

[0851] Operation: The server generates appropriate advice based on the collected information and image recognition results. The generated advice (e.g., "Don't forget to eat, let's eat now") is notified to the user via the device. If necessary, the information is also sent to family members.

[0852] Output: Advice notification (e.g., notification from the device to the user saying "Don't forget to eat, let's eat now" and sharing the information with family members)

[0853] (Application example 1)

[0854] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0855] The present invention aims to provide an environment in which elderly people with dementia can use self-driving vehicles with peace of mind. Currently, there is a lack of means to ensure the safe and comfortable travel of elderly people with dementia in self-driving vehicles, which creates a risk of elderly people behaving incorrectly. Specifically, this poses problems such as forgetting to eat and not taking care of their health. In response to this, the present invention solves these issues by monitoring the condition of elderly people in real time and providing appropriate warnings and instructions.

[0856] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0857] In this invention, the server includes language processing means for collecting information based on dialogue with the user, data storage means for storing the collected information in a database, image recognition means for analyzing lifestyle habits using image recognition technology, assistance means for generating appropriate advice based on the analysis results and providing it to the user, means for monitoring the condition of the elderly person using a camera device installed in the vehicle and acquiring image data, and means for analyzing the acquired image and audio data and providing instructions and warnings tailored to the elderly person's condition. This enables an environment in which elderly people with dementia can use self-driving vehicles safely and comfortably.

[0858] The "language processing means" is a means for collecting necessary information based on a dialogue with the user.

[0859] "Data storage means" refers to a means for storing collected information in a database.

[0860] The "image recognition means" is a means for analyzing a user's lifestyle habits using image recognition technology.

[0861] The "assisting means" is a means for generating appropriate advice based on the analysis results and providing it to the user.

[0862] The "camera device" is a device that is installed in a vehicle to monitor the condition of the elderly person and acquire image data.

[0863] The "image and audio data analysis means" is a means for analyzing the acquired image and audio data and providing instructions and warnings that are tailored to the elderly person's situation.

[0864] The system of this invention is designed to provide an environment in which elderly people with dementia can safely use autonomous vehicles. The system consists of the following main components:

[0865] 1. Language Processing Methods

[0866] The server collects necessary information through dialogue with the user. Specifically, it responds to questions and greetings from the user through voice input, and analyzes the response using natural language processing technology. For example, if an elderly person responds "I forgot" to the question "How was your breakfast today?", this information is processed and recorded as text data. The software used includes the SpeechRecognition library and the gTTS library.

[0867] 2. Data Retention Methods

[0868] The server stores the collected information in a database. This database is used to record the user's daily behavior and status history. For example, information such as "I did not eat breakfast at 10:00 AM on October 12, 2023" is recorded. This allows the system to provide appropriate advice while referring to past history.

[0869] 3. Image Recognition Methods

[0870] A camera device installed inside the vehicle monitors the elderly person's condition and periodically captures image data. The captured image data is sent to a server and analyzed using image recognition technology. Specifically, the OpenCV library is used to process the images and perform object detection and behavior recognition. For example, if no food is found, it is determined that the elderly person has not eaten.

[0871] 4. Assistance methods

[0872] The server generates appropriate advice based on the collected information and the results of image recognition analysis, and provides it to the user. The advice is communicated to the user via voice. For example, advice such as "Don't forget to eat, start eating now" is generated and communicated to the user through a speaker.

[0873] 5. Image and audio data analysis methods

[0874] The acquired image and audio data is processed in real time. This method is used to provide instructions and warnings tailored to the user's situation. The image and audio data captured by the camera are used for analysis, and the server comprehensively analyzes these to provide optimal advice.

[0875] Specific example explanation

[0876] For example, one morning, the server uses a camera inside the vehicle to capture the state of an elderly person's room. Analysis of the image data confirms that food is missing. The server then asks aloud, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and advises, "Don't forget to eat. Let's eat now." This advice is communicated to the user through the speaker, and the information is also sent to family members if necessary.

[0877] Prompt Sentence Examples

[0878] Here are some examples of specific prompts:

[0879] Elderly Co-Driver Assistance prompts:

[0880] 1. Capture an image with the camera and send it to an image analysis API to determine if food is in the image.

[0881] 2. A voice prompt asks the elderly person, "How was your breakfast today?"

[0882] 3. If the user answers "I forgot," the system provides the advice "Don't forget to eat, eat now."

[0883] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[0884] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0885] Step 1:

[0886] The server uses a camera device to monitor the condition of the elderly person in the vehicle and captures image data.

[0887] Input: Real-time video from a camera device.

[0888] Data processing and calculation: Images are captured using the OpenCV library and saved as still images.

[0889] Output: The captured image data.

[0890] Step 2:

[0891] The server sends the captured image data to an image recognition API to determine whether food is in the image.

[0892] Input: The captured image data.

[0893] Data processing and calculation: Send image data to the image recognition API and obtain the analysis results.

[0894] Output: Analysis result (e.g., determination that food is not visible).

[0895] Step 3:

[0896] Based on the analysis results, the server asks the user aloud, "How was your breakfast today?"

[0897] Input: Analysis results (e.g. food not shown).

[0898] Data processing and calculation: The question is converted into an audio file using the gTTS library and played through the speaker.

[0899] Output: A spoken question to the user.

[0900] Step 4:

[0901] The user responds to the voice questions and the server receives the response as voice input.

[0902] Input: The user's spoken response.

[0903] Data processing and calculation: Convert speech to text using the SpeechRecognition library.

[0904] Output: User response as text data (e.g., "I forgot").

[0905] Step 5:

[0906] The server records the user's response in a database.

[0907] Input: User response as text data.

[0908] Data processing and calculation: Connect to the database and record the response content in an appropriate format.

[0909] Output: The recorded database entries.

[0910] Step 6:

[0911] Based on the user's response, the server generates advice such as "Don't forget to eat, start eating now," and notifies the user by voice.

[0912] Input: User response as text data.

[0913] Data processing and calculation: The advice sentence is converted into an audio file using the gTTS library and played through the speaker.

[0914] Output: Audio advice to the user.

[0915] Step 7:

[0916] The server notifies the user's family of the user's condition and advice as necessary.

[0917] Input: User responses and advice as text data.

[0918] Data processing and calculation: Send text messages and emails to family members through the notification system.

[0919] Output: Notification message to family members.

[0920] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0921] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and an emotion engine to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[0922] System configuration

[0923] The system consists of the following main components:

[0924] 1. Language processing means: Collects necessary information through dialogue with the user.

[0925] 2. Data retention measures: Collected information is stored in a database.

[0926] 3. Image recognition means: Analyzes image data acquired using a camera device.

[0927] 4. Emotion Engine: Analyzes the emotional state from user interactions and image data.

[0928] 5. Assistance: Providing appropriate advice based on the collected information and analysis results.

[0929] Implementation of language processing measures

[0930] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[0931] Implementing data retention measures

[0932] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[0933] Implementation of image recognition measures

[0934] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[0935] Emotion Engine Implementation

[0936] The server runs an emotion engine based on the user's dialogue, voice tone, and even image data sent from the device. This engine analyzes the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.). The analysis results can be used to more specifically determine the most appropriate advice.

[0937] Implementation of assistance measures

[0938] The server generates appropriate advice based on the collected information and the analysis results of the emotion engine, and provides it to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, it can take an approach such as "Is there anything troubling you?" This advice is notified to the user via their device. Similar information can also be notified to family members if necessary.

[0939] Example

[0940] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that there is no food in sight. It also uses an emotion engine to determine that the user's facial expression is "gloomy." The server then uses language processing to ask the user, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now. Also, let us know if you need any help." This advice is notified to the user via the device, and the information is also sent to family members if necessary.

[0941] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding an emotion engine, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[0942] The processing flow will be explained below.

[0943] Step 1:

[0944] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[0945] Step 2:

[0946] The device sends the captured image data to the server, where it is sent in real time over the network.

[0947] Step 3:

[0948] The server pre-processes the image data it receives, including image resizing, noise reduction, and color correction.

[0949] Step 4:

[0950] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[0951] Step 5:

[0952] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[0953] Step 6:

[0954] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[0955] Step 7:

[0956] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[0957] Step 8:

[0958] The user responds to the initial message through the terminal, either by voice or text.

[0959] Step 9:

[0960] The device captures the user's response and converts it into text data using voice recognition technology.

[0961] Step 10:

[0962] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[0963] Step 11:

[0964] The server uses an emotion engine to analyze the emotional state from the user's response and image data, e.g., to identify the emotional state such as "sad" or "happy."

[0965] Step 12:

[0966] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[0967] Step 13:

[0968] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[0969] Step 14:

[0970] The server makes the final decision and generates appropriate advice. For example, "Don't forget to eat, let's eat now. Also, are you feeling sad? Can you tell me something?"

[0971] Step 15:

[0972] The server sends the generated advice to the device, which then notifies the user via voice or text.

[0973] Step 16:

[0974] The server notifies the family of the same information as necessary. For example, it may notify the family by email or SMS, saying, "The user does not seem to have eaten a meal today. Also, he may be depressed."

[0975] Through the above steps, it is possible to provide comprehensive support for the user's daily life and provide appropriate responses according to their emotional state.

[0976] Example 2

[0977] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0978] The problem that this invention aims to solve is to provide appropriate support to elderly people with dementia and their families based on their individual lifestyle habits and emotional states, and to create an environment where elderly people with dementia can live with peace of mind. To achieve this, it is important to understand the user's daily behavior and emotions in real time and provide appropriate advice. However, existing systems do not adequately analyze emotions or recognize lifestyle habits, making it difficult to provide effective support to users and their families.

[0979] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a database, an image recognition means for analyzing lifestyle habits using image recognition technology, an emotion analysis means for analyzing the emotional state from the content of the user's dialogue and image data, and an assistance means for generating appropriate advice based on the analysis results and the emotional state and providing the advice to the user. This makes it possible to grasp the user's daily behavior and emotions in real time and provide appropriate support based on them.

[0980] "Language processing means" refers to devices or software that collect information based on dialogue with a user and analyze that information using natural language processing technology.

[0981] "Data storage means" refers to devices or software that store collected information in a database so that it can be referenced later.

[0982] "Image recognition means" refers to a technology or device that analyzes image data acquired by a camera device and determines the lifestyle habits of a user.

[0983] "Emotion analysis means" refers to a device or software that analyzes the emotional state of a user from the content of their dialogue or image data.

[0984] "Assistance means" refers to a device or software that generates appropriate advice based on the analysis results and emotional state and provides it to the user.

[0985] The term "camera device" refers to a device for monitoring a user's living environment and acquiring image data.

[0986] "Means for analyzing lifestyle habits" refers to technology or devices that analyze and judge a user's behavior and habits based on acquired image data.

[0987] "Emotional state" refers to the user's psychological state and feelings, as judged from the content of the user's dialogue, facial expression, tone of voice, etc.

[0988] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and emotion analysis technology to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[0989] Implementation of language processing measures

[0990] The server uses natural language processing technology to analyze the conversation with the user and collect the necessary information. To do this, it uses software such as Google Cloud Natural Language API and IBM Watson as natural language processing technology (NLP). For example, the server asks the user, "Hello, how are you?" and the user's response, "I'm not feeling well," is converted into text data using speech recognition technology (e.g., Google Speech-to-Text API) and analyzed.

[0991] Implementing data retention measures

[0992] The collected information is stored in a database by the server. MySQL or PostgreSQL is used as the database. This records the user's most recent actions and status, allowing them to be referenced later. For example, it may record "No meals at 10:00 AM on October 12, 2023."

[0993] Implementation of image recognition measures

[0994] A camera device (e.g., a Raspberry Pi camera module) installed on the device periodically captures image data and sends it to a server. The server then analyzes the data using image recognition technology (e.g., OpenCV, TensorFlow). For example, if no food is found in the video, it determines that the person has not eaten.

[0995] Implementing sentiment analysis measures

[0996] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) based on the user's dialogue content, voice tone, and image data sent from the device to analyze the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.).

[0997] Implementation of assistance measures

[0998] The server uses a generative AI model to generate appropriate advice based on the collected information and the results of emotion analysis. The generated advice is sent to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API) and provided to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, an approach such as "Is there anything troubling you?" is taken. This advice is notified to the user via the device, and similar information is also notified to family members if necessary.

[0999] Example

[1000] Example 1: Morning dialogue scenario

[1001] At 6:00 a.m., the device speaks to the user, asking, "Good morning. How are you feeling today?" If the user responds, "I'm not feeling very good today," the device captures this voice and sends it to the server. The server converts this into text data and uses NLP technology to analyze the "bad mood." The analysis determines that the user is feeling "depressed," which is supported by emotion analysis technology. The server generates advice such as, "Let's find something to cheer you up together. Is there anything troubling you?" and notifies the user of this via the device.

[1002] Example 2: Lunch Check Scenario

[1003] At 12:00, the device captures the state of the room with its camera and sends it to the server. The server analyzes the image data and determines that "food is missing." It also uses emotion analysis technology to determine that the user's facial expression is "gloomy." The server then asks the user, "How was your lunch today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat. Let's eat now. Also, let us know if you need any help," and notifies the user through the device.

[1004] Prompt Sentence Examples

[1005] Example prompts to be input to the generative AI model:

[1006] "Please explain in detail the processing steps of the daily life support system for elderly people with dementia. Please include the specific operations and technologies for each step, the names of the hardware and software used, and examples of use."

[1007] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding emotion analysis technology, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[1008] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1009] Step 1:

[1010] A user accesses the system, and the device asks a question via the microphone, such as "Hello, how are you?" The device obtains the user's response (voice). This response is captured by the device's microphone (input data: user's voice). Specifically, the device activates its voice recognition function to capture the voice signal and sends it to the server as a digital voice file (output data: digital voice file).

[1011] Step 2:

[1012] The server uses speech recognition technology (e.g., Google Speech-to-Text API) to analyze the received audio file and convert it into text data (input data: digital audio file, output data: text data). Specifically, the server samples the audio data and uses a speech recognition model to convert it into a string of characters.

[1013] Step 3:

[1014] The server analyzes the converted text data using natural language processing technology (e.g., Google Cloud Natural Language API), thereby extracting semantic information from the user's response (input data: text data, output data: semantic information). Specifically, the server analyzes the text using techniques such as tokenization, part-of-speech tagging, and sentiment analysis to extract information about the user's state and emotions.

[1015] Step 4:

[1016] The device periodically captures image data of the environment. The camera device acquires image data of the user's living environment (input data: environmental image). Specifically, the device activates the camera at set time intervals, captures images, and sends them to the server (output data: environmental image file).

[1017] Step 5:

[1018] The server analyzes the received image data using image recognition technology (e.g., OpenCV, TensorFlow). It determines lifestyle habits from the images (input data: environmental image files, output data: analysis results). Specifically, the server runs a model to recognize the user's behavior (e.g., not eating) from the images and obtains the analysis results.

[1019] Step 6:

[1020] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) to analyze the user's emotional state from the content of the user's dialogue and image data. This identifies the user's emotional state (input data: text data, environmental image files, output data: emotional state). Specifically, the server uses a model that analyzes the user's emotions (e.g., joy, sadness, anger, etc.) from the text data and image data.

[1021] Step 7:

[1022] The server uses a generative AI model to generate appropriate advice based on the collected information (conversation content, lifestyle habits, emotional state) (input data: analysis results, emotional state, output data: advice). Specifically, the server uses a generative model (e.g., GPT-3 or a similar model) to generate advice or support messages that are most appropriate for the user's current situation.

[1023] Step 8:

[1024] The server sends the generated advice to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API). The device notifies the user of the advice (input data: advice, output data: notification to user). Specifically, the device communicates the received advice to the user using a voice output device or a display device.

[1025] Step 9:

[1026] The server saves all conversation content, analysis results, and notification content in a database (input data: all recorded data, output data: saved data). MySQL or PostgreSQL is used as the database. Specifically, the server stores the data in the database in an appropriate format so that it can be saved and referenced later. This data can be referenced later or used for analysis.

[1027] (Application example 2)

[1028] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1029] Elderly people, especially those with dementia, require assistance in various aspects of daily life. However, this places a heavy burden on family members and caregivers, and even situations where elderly people can act independently, such as shopping at a store, can be accompanied by anxiety. Therefore, there is a need for effective systems that support the safety and independence of elderly people and reduce the burden on family members and caregivers. In particular, there is a need for technology that can analyze the user's emotional state and provide appropriate assistance based on that analysis.

[1030] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1031] In this invention, the server includes a natural language processing means for collecting information based on a dialogue with a user, an information storage means for storing the collected information in a database, an image recognition means for analyzing daily activities using image recognition technology, a support means for generating appropriate advice based on the analysis results and providing it to the user, an emotion analysis means for analyzing the user's emotional state, and a display means for supporting elderly people's shopping and daily life in physical stores. This allows elderly people to move around safely and independently in physical stores and reduces the burden on family members and caregivers.

[1032] "Natural language processing means" refers to technology for collecting information through dialogue with users and analyzing it as text data.

[1033] "Information storage means" refers to a system that stores collected information in a database and makes it possible to refer to it as needed.

[1034] "Image recognition means" refers to technology that analyzes image data acquired through devices such as cameras and recognizes specific objects or situations.

[1035] "Support means" refers to a system that provides appropriate advice and guidance to users based on collected information and analysis results.

[1036] "Emotion analysis means" refers to technology that analyzes the content of a user's dialogue and image data to identify the user's emotional state.

[1037] "Display means" refers to technology that provides users with visual or auditory information to assist elderly people in their shopping and daily life in physical stores.

[1038] The system of the present invention is designed to support elderly people in their shopping and daily life in brick-and-mortar stores. The detailed configuration of the system and how it is implemented will be described below.

[1039] System Configuration

[1040] Hardware

[1041] Smart glasses: A device worn by the user that contains a camera and microphone.

[1042] Server: A central system for data processing and storage.

[1043] software

[1044] Natural language processing means: Manage dialogue with users using services such as Amazon AWS's Lex service or Google's Dialogflow.

[1045] Information storage method: Information is stored using a database management system such as SQLite3 or MySQL.

[1046] Image recognition method: Uses image recognition technologies such as OpenCV and YOLOv3.

[1047] Sentiment analysis method: Sentiment analysis is performed using deep learning libraries such as Transformers and Pytorch.

[1048] Display means: Uses the display function of the smart glasses to provide information visually to the user.

[1049] Program processing explanation

[1050] The server first collects input information from the camera and microphone while the user is wearing the smart glasses. A natural language processing means then uses the conversation with the user to analyze the information as text data. For example, if the user says, "I'm not feeling well today," the content is converted into text data.

[1051] The collected information is stored in a database by an information storage means, and this information is available for future reference and used to track the user's progress.

[1052] The camera on the device periodically captures image data, which is then sent to a server. The server then uses image recognition to analyze the image data and identify, for example, the status of a product shelf. For example, when a specific product is in view, information about that product is notified to the user.

[1053] The emotional state of the user is analyzed from the content of the user's dialogue and image data through the emotion analysis means. Based on the analysis results, advice and guidance appropriate to the user's current emotional state is generated and provided to the user through the display means.

[1054] Specific examples

[1055] 1. A user puts on smart glasses and enters a physical store.

[1056] 2. Immediately after entering the store, the natural language processing means asks the user, "How are you feeling today?" The user replies, "I'm a little tired."

[1057] 3. The server stores this information in a database and uses emotion analysis means to recognize the emotion "tired."

[1058] 4. The device's camera scans the shelves, and image recognition detects immune-boosting health foods.

[1059] 5. The user is notified through the display means, "This is a health food that boosts your immunity. Why not give it a try?"

[1060] Prompt Sentence Examples

[1061] 1. "Hi, how are you?"

[1062] 2. "How was your meal today?"

[1063] 3. "Did you find what you were looking for?"

[1064] In this way, this system covers all the technologies necessary to provide an environment where elderly people can shop independently and safely. By combining comprehensive information processing by the server with interaction through smart glasses, it is possible to improve the quality of life of users.

[1065] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1066] Processing Steps

[1067] Step 1:

[1068] The device collects user input information through the smart glasses' camera and microphone. The input data includes the user's voice and the video captured by the camera. This input data is sent to the server as image data for the camera video and as audio data for the user's voice.

[1069] Step 2:

[1070] After the server receives the voice data, it uses natural language processing means to convert the voice data into text data. For example, it analyzes the voice of the question "Hello, how are you?" and generates text data. The input is voice data and the output is text data.

[1071] Step 3:

[1072] The server continues the dialogue with the user based on the converted text data. The converted text data is analyzed using natural language processing means to generate appropriate questions and answers. The generated dialogue content is sent to the terminal as voice. The input is the converted text data, and the output is the generated dialogue content.

[1073] Step 4:

[1074] The device receives the generated dialogue content sent from the server and notifies the user by voice, for example, asking the user through the smart glasses, "How are you feeling today?" The user's voice response is collected again and sent to the server.

[1075] Step 5:

[1076] The server receives the image data and analyzes the everyday environment using image recognition means. Specifically, it identifies objects from camera footage, for example, identifying products on a shelf. The input is image data, and the output is data on the recognized objects.

[1077] Step 6:

[1078] The server generates appropriate advice and guidance based on the analysis results. It generates information to be provided to the user for the recognized object and creates guidance in text format. The input is the recognized object data, and the output is the text data of the advice and guidance.

[1079] Step 7:

[1080] The server uses emotion analysis means to analyze the user's emotional state from the user's voice and image data. For example, if the user replies, "I'm not feeling well today," the server analyzes the text data and voice tone to identify the user's emotional state. The input is the voice and image data, and the output is the analyzed emotional data.

[1081] Step 8:

[1082] The server generates more appropriate dialogue content based on the generated advice and guidance and the analyzed emotional data. For example, if the server determines that the user is tired, it generates a dialogue recommending "health foods that boost the immune system." The input is the text data of the advice and guidance and emotional data, and the output is the optimized dialogue content.

[1083] Step 9:

[1084] The device receives the generated dialogue and notifies the user through the smart glasses. It provides the user with visual and auditory information to assist with shopping. For example, it displays a message such as, "This is a health food that boosts your immunity. Why not give it a try?"

[1085] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1086] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1087] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1088] [Fourth embodiment]

[1089] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1090] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1091] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1092] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1093] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1094] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1095] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1096] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1097] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1098] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1099] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1100] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1101] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1102] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system supports the lives of elderly people by collecting information through dialogue with the user and analyzing the user's lifestyle habits using image recognition technology.

[1103] System configuration

[1104] The system consists of the following main components:

[1105] 1. Language processing means: Collects necessary information through dialogue with the user.

[1106] 2. Data retention measures: Collected information is stored in a database.

[1107] 3. Image recognition means: Analyzes image data acquired using a camera device.

[1108] 4. Assistance: Providing appropriate advice based on the collected information and analysis results.

[1109] Implementation of language processing measures

[1110] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[1111] Implementing data retention measures

[1112] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[1113] Implementation of image recognition measures

[1114] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[1115] Implementation of assistance measures

[1116] The server generates appropriate advice based on the collected information and image recognition results and provides it to the user. For example, it generates advice such as, "Don't forget to eat, start eating now." This advice is notified to the user via their device. Similar information is also notified to family members as needed.

[1117] Example

[1118] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that no food is found. It then asks the user through language processing means, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now." This advice is notified to the user through the device, and the information is also sent to family members if necessary.

[1119] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[1120] The processing flow will be explained below.

[1121] Step 1:

[1122] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[1123] Step 2:

[1124] The device sends the captured image data to the server, where it is sent in real time over the network.

[1125] Step 3:

[1126] The server pre-processes the received image data, including image resizing, noise reduction, and color correction.

[1127] Step 4:

[1128] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[1129] Step 5:

[1130] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[1131] Step 6:

[1132] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[1133] Step 7:

[1134] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[1135] Step 8:

[1136] The user responds to the initial message through the terminal, either by voice or text.

[1137] Step 9:

[1138] The device captures the user's response and converts it into text data using voice recognition technology.

[1139] Step 10:

[1140] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[1141] Step 11:

[1142] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[1143] Step 12:

[1144] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[1145] Step 13:

[1146] The server makes the final decision and generates appropriate advice, e.g., "Don't forget to eat, let's eat now."

[1147] Step 14:

[1148] The server sends the generated advice to the device, which then notifies the user via voice or text.

[1149] Step 15:

[1150] The server will notify the family of the same information as necessary. For example, it may notify the family by email or SMS that "The user has not eaten a meal today."

[1151] Through the above steps, it is possible to comprehensively support the user's daily life and provide a safe living environment.

[1152] Example 1

[1153] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1154] This invention relates to a system that comprehensively supports the lives of elderly people with dementia and reduces the burden on their families and caregivers. Conventional systems have difficulty accurately understanding a user's lifestyle and health status, and have been unable to provide appropriate advice or notifications. Therefore, more effective information collection and analysis is needed to maintain a user's health and improve their quality of life.

[1155] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1156] In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a storage means, an image recognition means for analyzing image data acquired using a camera device, and an assistance means for generating appropriate notifications or advice based on the analysis results and the collected information and providing them to the user. This makes it possible to appropriately grasp the user's behavior and health condition in real time and quickly provide necessary advice or notifications.

[1157] "User" refers to the individual or family member who uses the system, and is often targeted at elderly people with dementia.

[1158] "Language processing means" refers to technology that collects information from dialogue with the user and converts it into text data.

[1159] "Data retention means" refers to technology that stores collected information in a storage means such as a database, making it available for later reference.

[1160] A "camera device" is hardware for acquiring image data, and is used to monitor the user's living environment.

[1161] "Image recognition means" refers to technology that analyzes image data acquired by a camera device and determines the user's behavior and situation.

[1162] "Assistance means" refers to technology that generates and provides appropriate notifications and advice to users based on collected information and analysis results.

[1163] "Analysis results" refers to the analysis data obtained by the language processing means and the image recognition means.

[1164] "Notification" means a message or alert provided by the System to a User or their Family Members.

[1165] "Advice" refers to the guidance the system provides to the user to encourage improvements in their lifestyle and health.

[1166] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and provide an environment in which they can live with peace of mind. This system consists of four main components: language processing means, data storage means, image recognition means, and assistance means.

[1167] Language Processing Methods

[1168] The server uses natural language processing technology (for example, a generative AI model such as OpenAI's GPT-4) to analyze the user's dialogue and collect necessary information. The dialogue with the user is conducted through the device, which recognizes the user's voice and sends it to the server as text data. For example, if the device asks the user, "Hello, how are you?" and the user responds, "I'm a little tired today," that information is collected by the server.

[1169] Data Retention Methods

[1170] The collected information is stored by the server in a database (for example, an RDBMS such as MySQL or PostgreSQL). This records the user's most recent actions and status, and can be referenced later. For example, data such as "The user responded that they were tired at 10:00 AM on October 13, 2023" is stored.

[1171] Image Recognition Method

[1172] The camera device installed in the device periodically captures the user's living environment and sends the image data to a server. The server then uses image recognition technology to analyze this data. For example, if the camera analyzes an image of the user's room and no food is found, it can determine that the user has not had breakfast.

[1173] Assistance Method

[1174] The server generates appropriate notifications or advice based on the collected information and the results of image recognition analysis, and provides them to the user. For example, it generates a notification saying, "Don't forget to eat, let's eat now." This notification is sent to the user via their device, and similar notifications are also sent to family members if necessary.

[1175] Specific examples

[1176] For example, one morning, the device captures the state of the user's room with a camera and sends the image data to the server. The server analyzes this image data and confirms that no food is found. The device then asks the user, "How was your breakfast today?" through language processing means. If the user replies, "I forgot," this information is saved in the database. The server then generates advice such as, "Don't forget to eat, start eating now," and notifies this to the user via the device. Information that "the user may not have had breakfast" is also sent to the user's family.

[1177] Example prompt sentence:

[1178] Ask the user, "Hello, how are you?". Before the user asks, "How was your breakfast today?", save the user's most recent meal history in a database. Then, analyze the camera footage and generate an advice if the user hasn't eaten, saying, "Don't forget to eat, eat now."

[1179] This system provides comprehensive support for users' daily lives, reducing the burden on their families and caregivers, and provides appropriate advice based on the collected information, creating an environment in which users can live with peace of mind.

[1180] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1181] Step 1: Initiate a conversation with the user

[1182] Input: The initial message the terminal is set to (e.g. "Hello, how are you?")

[1183] Action: The device asks the user a question using voice.

[1184] Output: The user responds verbally (e.g., "I'm a little tired today").

[1185] Step 2: Collect and analyze user responses

[1186] Input: User's voice response

[1187] How it works: The device uses voice recognition technology to convert the user's response into text data, which it then sends to the server.

[1188] Output: Text data (e.g., "I'm a little tired today")

[1189] Step 3: Information analysis using natural language processing

[1190] Input: Text data (e.g., "I'm a little tired today")

[1191] How it works: The server uses a generative AI model (e.g., OpenAI's GPT-4) to analyze text data and evaluate the user's state.

[1192] Output: Analysis result (e.g. "The user is tired")

[1193] Step 4: Storing information in a database

[1194] Input: Analysis result (e.g. "The user is tired")

[1195] How it works: The server stores the analysis results in a database.

[1196] Output: Stored data (e.g., "User responded that he was tired on October 13, 2023 at 10:00 AM")

[1197] Step 5: Capture camera images

[1198] Input: Timer setting for camera device (e.g. capture every hour)

[1199] Operation: The camera device installed on the device captures an image and sends the image data to the server.

[1200] Output: Image data (e.g., a photo of the user's room)

[1201] Step 6: Data analysis using image recognition

[1202] Input: Image data (e.g., a photo of the user's room)

[1203] How it works: The server uses image recognition technology to analyze the image data and determine the user's situation (e.g., if there is no food in sight, it determines that the user is not eating).

[1204] Output: Image analysis results (e.g., "The user has not eaten")

[1205] Step 7: Generate and notify advice

[1206] Input: User information analysis results and image analysis results

[1207] Operation: The server generates appropriate advice based on the collected information and image recognition results. The generated advice (e.g., "Don't forget to eat, let's eat now") is notified to the user via the device. If necessary, the information is also sent to family members.

[1208] Output: Advice notification (e.g., notification from the device to the user saying "Don't forget to eat, let's eat now" and sharing the information with family members)

[1209] (Application example 1)

[1210] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1211] The present invention aims to provide an environment in which elderly people with dementia can use self-driving vehicles with peace of mind. Currently, there is a lack of means to ensure the safe and comfortable travel of elderly people with dementia in self-driving vehicles, which creates a risk of elderly people behaving incorrectly. Specifically, this poses problems such as forgetting to eat and not taking care of their health. In response to this, the present invention solves these issues by monitoring the condition of elderly people in real time and providing appropriate warnings and instructions.

[1212] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1213] In this invention, the server includes language processing means for collecting information based on dialogue with the user, data storage means for storing the collected information in a database, image recognition means for analyzing lifestyle habits using image recognition technology, assistance means for generating appropriate advice based on the analysis results and providing it to the user, means for monitoring the condition of the elderly person using a camera device installed in the vehicle and acquiring image data, and means for analyzing the acquired image and audio data and providing instructions and warnings tailored to the elderly person's condition. This enables an environment in which elderly people with dementia can use self-driving vehicles safely and comfortably.

[1214] The "language processing means" is a means for collecting necessary information based on a dialogue with the user.

[1215] "Data storage means" refers to a means for storing collected information in a database.

[1216] The "image recognition means" is a means for analyzing a user's lifestyle habits using image recognition technology.

[1217] The "assisting means" is a means for generating appropriate advice based on the analysis results and providing it to the user.

[1218] The "camera device" is a device that is installed in a vehicle to monitor the condition of the elderly person and acquire image data.

[1219] The "image and audio data analysis means" is a means for analyzing the acquired image and audio data and providing instructions and warnings that are tailored to the elderly person's situation.

[1220] The system of this invention is designed to provide an environment in which elderly people with dementia can safely use autonomous vehicles. The system consists of the following main components:

[1221] 1. Language Processing Methods

[1222] The server collects necessary information through dialogue with the user. Specifically, it responds to questions and greetings from the user through voice input, and analyzes the response using natural language processing technology. For example, if an elderly person responds "I forgot" to the question "How was your breakfast today?", this information is processed and recorded as text data. The software used includes the SpeechRecognition library and the gTTS library.

[1223] 2. Data Retention Methods

[1224] The server stores the collected information in a database. This database is used to record the user's daily behavior and status history. For example, information such as "I did not eat breakfast at 10:00 AM on October 12, 2023" is recorded. This allows the system to provide appropriate advice while referring to past history.

[1225] 3. Image Recognition Methods

[1226] A camera device installed inside the vehicle monitors the elderly person's condition and periodically captures image data. The captured image data is sent to a server and analyzed using image recognition technology. Specifically, the OpenCV library is used to process the images and perform object detection and behavior recognition. For example, if no food is found, it is determined that the elderly person has not eaten.

[1227] 4. Assistance methods

[1228] The server generates appropriate advice based on the collected information and the results of image recognition analysis, and provides it to the user. The advice is communicated to the user via voice. For example, advice such as "Don't forget to eat, start eating now" is generated and communicated to the user through a speaker.

[1229] 5. Image and audio data analysis methods

[1230] The acquired image and audio data is processed in real time. This method is used to provide instructions and warnings tailored to the user's situation. The image and audio data captured by the camera are used for analysis, and the server comprehensively analyzes these to provide optimal advice.

[1231] Specific example explanation

[1232] For example, one morning, the server uses a camera inside the vehicle to capture the state of an elderly person's room. Analysis of the image data confirms that food is missing. The server then asks aloud, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and advises, "Don't forget to eat. Let's eat now." This advice is communicated to the user through the speaker, and the information is also sent to family members if necessary.

[1233] Prompt Sentence Examples

[1234] Here are some examples of specific prompts:

[1235] Elderly Co-Driver Assistance prompts:

[1236] 1. Capture an image with the camera and send it to an image analysis API to determine if food is in the image.

[1237] 2. A voice prompt asks the elderly person, "How was your breakfast today?"

[1238] 3. If the user answers "I forgot," the system provides the advice "Don't forget to eat, eat now."

[1239] In this way, the system of the present invention aims to provide comprehensive support for the lives of elderly people with dementia and reduce the burden on their families and caregivers.

[1240] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1241] Step 1:

[1242] The server uses a camera device to monitor the condition of the elderly person in the vehicle and captures image data.

[1243] Input: Real-time video from a camera device.

[1244] Data processing and calculation: Images are captured using the OpenCV library and saved as still images.

[1245] Output: The captured image data.

[1246] Step 2:

[1247] The server sends the captured image data to an image recognition API to determine whether food is in the image.

[1248] Input: The captured image data.

[1249] Data processing and calculation: Send image data to the image recognition API and obtain the analysis results.

[1250] Output: Analysis result (e.g., determination that food is not visible).

[1251] Step 3:

[1252] Based on the analysis results, the server asks the user aloud, "How was your breakfast today?"

[1253] Input: Analysis results (e.g. food not shown).

[1254] Data processing and calculation: The question is converted into an audio file using the gTTS library and played through the speaker.

[1255] Output: A spoken question to the user.

[1256] Step 4:

[1257] The user responds to the voice questions and the server receives the response as voice input.

[1258] Input: The user's spoken response.

[1259] Data processing and calculation: Convert speech to text using the SpeechRecognition library.

[1260] Output: User response as text data (e.g., "I forgot").

[1261] Step 5:

[1262] The server records the user's response in a database.

[1263] Input: User response as text data.

[1264] Data processing and calculation: Connect to the database and record the response content in an appropriate format.

[1265] Output: The recorded database entries.

[1266] Step 6:

[1267] Based on the user's response, the server generates advice such as "Don't forget to eat, start eating now," and notifies the user by voice.

[1268] Input: User response as text data.

[1269] Data processing and calculation: The advice sentence is converted into an audio file using the gTTS library and played through the speaker.

[1270] Output: Audio advice to the user.

[1271] Step 7:

[1272] The server notifies the user's family of the user's condition and advice as necessary.

[1273] Input: User responses and advice as text data.

[1274] Data processing and calculation: Send text messages and emails to family members through the notification system.

[1275] Output: Notification message to family members.

[1276] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1277] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and an emotion engine to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[1278] System configuration

[1279] The system consists of the following main components:

[1280] 1. Language processing means: Collects necessary information through dialogue with the user.

[1281] 2. Data retention measures: Collected information is stored in a database.

[1282] 3. Image recognition means: Analyzes image data acquired using a camera device.

[1283] 4. Emotion Engine: Analyzes the emotional state from user interactions and image data.

[1284] 5. Assistance: Providing appropriate advice based on the collected information and analysis results.

[1285] Implementation of language processing measures

[1286] The server uses natural language processing technology to analyze the user's conversation and collect the necessary information. For example, the conversation might start with a simple greeting like "Hello, how are you?" and then ask questions like "How was your meal today?" The user's response is analyzed as text data using speech recognition technology.

[1287] Implementing data retention measures

[1288] The collected information is stored in a database by the server. This allows the user's most recent actions and status to be recorded and referenced later. For example, a record such as "No meals at 10:00 AM on October 12, 2023" may be recorded.

[1289] Implementation of image recognition measures

[1290] The camera device on the device periodically captures image data and sends it to a server. The server then uses image recognition technology to analyze the data. For example, if no food is visible in the video, it will determine that the person is not eating.

[1291] Emotion Engine Implementation

[1292] The server runs an emotion engine based on the user's dialogue, voice tone, and even image data sent from the device. This engine analyzes the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.). The analysis results can be used to more specifically determine the most appropriate advice.

[1293] Implementation of assistance measures

[1294] The server generates appropriate advice based on the collected information and the analysis results of the emotion engine, and provides it to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, it can take an approach such as "Is there anything troubling you?" This advice is notified to the user via their device. Similar information can also be notified to family members if necessary.

[1295] Example

[1296] For example, one morning, the device captures the state of the user's room with a camera and sends it to the server. The server analyzes the image data and confirms that there is no food in sight. It also uses an emotion engine to determine that the user's facial expression is "gloomy." The server then uses language processing to ask the user, "How was your breakfast today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat, let's eat now. Also, let us know if you need any help." This advice is notified to the user via the device, and the information is also sent to family members if necessary.

[1297] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding an emotion engine, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[1298] The processing flow will be explained below.

[1299] Step 1:

[1300] The device monitors the elderly person's living environment through a camera. Captures are set to occur periodically. For example, one image is captured every hour.

[1301] Step 2:

[1302] The device sends the captured image data to the server, where it is sent in real time over the network.

[1303] Step 3:

[1304] The server pre-processes the image data it receives, including image resizing, noise reduction, and color correction.

[1305] Step 4:

[1306] The server analyzes the pre-processed image data using image recognition models to detect objects (e.g., food, drink, clothing, etc.) in the image.

[1307] Step 5:

[1308] The server uses data from the detected object to determine a specific state (e.g., whether it is eating or wearing appropriate clothing).

[1309] Step 6:

[1310] The server stores the result of the judgment in a database. This records the user's most recent actions and status. For example, it records "I did not eat at 10:00 AM on October 12, 2023."

[1311] Step 7:

[1312] The server generates an initial message and sends it to the device. Example: "Hello, how are you?"

[1313] Step 8:

[1314] The user responds to the initial message through the terminal, either by voice or text.

[1315] Step 9:

[1316] The device captures the user's response and converts it into text data using voice recognition technology.

[1317] Step 10:

[1318] The server analyzes the voice recognition results (text data) and understands the user's intent. Example: Confirming whether the user has eaten a meal.

[1319] Step 11:

[1320] The server uses an emotion engine to analyze the emotional state from the user's response and image data, e.g., to identify the emotional state such as "sad" or "happy."

[1321] Step 12:

[1322] The server generates additional questions as needed and sends them to the device. Example: "How was your meal today?"

[1323] Step 13:

[1324] The user responds to additional questions, and the terminal again uses voice recognition technology to convert the responses into text data, which is then sent to the server.

[1325] Step 14:

[1326] The server makes the final decision and generates appropriate advice. For example, "Don't forget to eat, let's eat now. Also, are you feeling sad? Can you tell me something?"

[1327] Step 15:

[1328] The server sends the generated advice to the device, which then notifies the user via voice or text.

[1329] Step 16:

[1330] The server notifies the family of the same information as necessary. For example, it may notify the family by email or SMS, saying, "The user does not seem to have eaten a meal today. Also, he may be depressed."

[1331] Through the above steps, it is possible to provide comprehensive support for the user's daily life and provide appropriate responses according to their emotional state.

[1332] Example 2

[1333] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1334] The problem that this invention aims to solve is to provide appropriate support to elderly people with dementia and their families based on their individual lifestyle habits and emotional states, and to create an environment where elderly people with dementia can live with peace of mind. To achieve this, it is important to understand the user's daily behavior and emotions in real time and provide appropriate advice. However, existing systems do not adequately analyze emotions or recognize lifestyle habits, making it difficult to provide effective support to users and their families.

[1335] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a language processing means for collecting information based on a dialogue with the user, a data storage means for storing the collected information in a database, an image recognition means for analyzing lifestyle habits using image recognition technology, an emotion analysis means for analyzing the emotional state from the content of the user's dialogue and image data, and an assistance means for generating appropriate advice based on the analysis results and the emotional state and providing the advice to the user. This makes it possible to grasp the user's daily behavior and emotions in real time and provide appropriate support based on them.

[1336] "Language processing means" refers to devices or software that collect information based on dialogue with a user and analyze that information using natural language processing technology.

[1337] "Data storage means" refers to devices or software that store collected information in a database so that it can be referenced later.

[1338] "Image recognition means" refers to a technology or device that analyzes image data acquired by a camera device and determines the lifestyle habits of a user.

[1339] "Emotion analysis means" refers to a device or software that analyzes the emotional state of a user from the content of their dialogue or image data.

[1340] "Assistance means" refers to a device or software that generates appropriate advice based on the analysis results and emotional state and provides it to the user.

[1341] The term "camera device" refers to a device for monitoring a user's living environment and acquiring image data.

[1342] "Means for analyzing lifestyle habits" refers to technology or devices that analyze and judge a user's behavior and habits based on acquired image data.

[1343] "Emotional state" refers to the user's psychological state and feelings, as judged from the content of the user's dialogue, facial expression, tone of voice, etc.

[1344] The system of this invention aims to support the daily lives of elderly people with dementia and their families, and to provide an environment in which they can live with peace of mind. This system collects information through dialogue with the user, and furthermore, utilizes image recognition technology and emotion analysis technology to analyze the user's lifestyle habits and emotional state, thereby providing comprehensive support for the elderly's lives.

[1345] Implementation of language processing measures

[1346] The server uses natural language processing technology to analyze the conversation with the user and collect the necessary information. To do this, it uses software such as Google Cloud Natural Language API and IBM Watson as natural language processing technology (NLP). For example, the server asks the user, "Hello, how are you?" and the user's response, "I'm not feeling well," is converted into text data using speech recognition technology (e.g., Google Speech-to-Text API) and analyzed.

[1347] Implementing data retention measures

[1348] The collected information is stored in a database by the server. MySQL or PostgreSQL is used as the database. This records the user's most recent actions and status, allowing them to be referenced later. For example, it may record "No meals at 10:00 AM on October 12, 2023."

[1349] Implementation of image recognition measures

[1350] A camera device (e.g., a Raspberry Pi camera module) installed on the device periodically captures image data and sends it to a server. The server then analyzes the data using image recognition technology (e.g., OpenCV, TensorFlow). For example, if no food is found in the video, it determines that the person has not eaten.

[1351] Implementing sentiment analysis measures

[1352] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) based on the user's dialogue content, voice tone, and image data sent from the device to analyze the user's current emotional state (e.g., joy, sadness, anger, surprise, etc.).

[1353] Implementation of assistance measures

[1354] The server uses a generative AI model to generate appropriate advice based on the collected information and the results of emotion analysis. The generated advice is sent to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API) and provided to the user. For example, if the emotional state is determined to be "sadness," in addition to the usual advice, an approach such as "Is there anything troubling you?" is taken. This advice is notified to the user via the device, and similar information is also notified to family members if necessary.

[1355] Example

[1356] Example 1: Morning dialogue scenario

[1357] At 6:00 a.m., the device speaks to the user, asking, "Good morning. How are you feeling today?" If the user responds, "I'm not feeling very good today," the device captures this voice and sends it to the server. The server converts this into text data and uses NLP technology to analyze the "bad mood." The analysis determines that the user is feeling "depressed," which is supported by emotion analysis technology. The server generates advice such as, "Let's find something to cheer you up together. Is there anything troubling you?" and notifies the user of this via the device.

[1358] Example 2: Lunch Check Scenario

[1359] At 12:00, the device captures the state of the room with its camera and sends it to the server. The server analyzes the image data and determines that "food is missing." It also uses emotion analysis technology to determine that the user's facial expression is "gloomy." The server then asks the user, "How was your lunch today?" If the user replies, "I forgot," the server stores this information in a database and generates advice such as, "Don't forget to eat. Let's eat now. Also, let us know if you need any help," and notifies the user through the device.

[1360] Prompt Sentence Examples

[1361] Example prompts to be input to the generative AI model:

[1362] "Please explain in detail the processing steps of the daily life support system for elderly people with dementia. Please include the specific operations and technologies for each step, the names of the hardware and software used, and examples of use."

[1363] In this way, the system of this invention aims to comprehensively support the lives of elderly people with dementia and reduce the burden on their families and caregivers. By adding emotion analysis technology, it becomes possible to respond to the user's emotional state, providing higher quality lifestyle support.

[1364] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1365] Step 1:

[1366] A user accesses the system, and the device asks a question via the microphone, such as "Hello, how are you?" The device obtains the user's response (voice). This response is captured by the device's microphone (input data: user's voice). Specifically, the device activates its voice recognition function to capture the voice signal and sends it to the server as a digital voice file (output data: digital voice file).

[1367] Step 2:

[1368] The server uses speech recognition technology (e.g., Google Speech-to-Text API) to analyze the received audio file and convert it into text data (input data: digital audio file, output data: text data). Specifically, the server samples the audio data and uses a speech recognition model to convert it into a string of characters.

[1369] Step 3:

[1370] The server analyzes the converted text data using natural language processing technology (e.g., Google Cloud Natural Language API), thereby extracting semantic information from the user's response (input data: text data, output data: semantic information). Specifically, the server analyzes the text using techniques such as tokenization, part-of-speech tagging, and sentiment analysis to extract information about the user's state and emotions.

[1371] Step 4:

[1372] The device periodically captures image data of the environment. The camera device acquires image data of the user's living environment (input data: environmental image). Specifically, the device activates the camera at set time intervals, captures images, and sends them to the server (output data: environmental image file).

[1373] Step 5:

[1374] The server analyzes the received image data using image recognition technology (e.g., OpenCV, TensorFlow). It determines lifestyle habits from the images (input data: environmental image files, output data: analysis results). Specifically, the server runs a model to recognize the user's behavior (e.g., not eating) from the images and obtains the analysis results.

[1375] Step 6:

[1376] The server uses emotion analysis technology (e.g., Microsoft Azure's Emotion API) to analyze the user's emotional state from the content of the user's dialogue and image data. This identifies the user's emotional state (input data: text data, environmental image files, output data: emotional state). Specifically, the server uses a model that analyzes the user's emotions (e.g., joy, sadness, anger, etc.) from the text data and image data.

[1377] Step 7:

[1378] The server uses a generative AI model to generate appropriate advice based on the collected information (conversation content, lifestyle habits, emotional state) (input data: analysis results, emotional state, output data: advice). Specifically, the server uses a generative model (e.g., GPT-3 or a similar model) to generate advice or support messages that are most appropriate for the user's current situation.

[1379] Step 8:

[1380] The server sends the generated advice to the device in text format or speech synthesis format (e.g., Google Text-to-Speech API). The device notifies the user of the advice (input data: advice, output data: notification to user). Specifically, the device communicates the received advice to the user using a voice output device or a display device.

[1381] Step 9:

[1382] The server saves all conversation content, analysis results, and notification content in a database (input data: all recorded data, output data: saved data). MySQL or PostgreSQL is used as the database. Specifically, the server stores the data in the database in an appropriate format so that it can be saved and referenced later. This data can be referenced later or used for analysis.

[1383] (Application example 2)

[1384] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1385] Elderly people, especially those with dementia, require assistance in various aspects of daily life. However, this places a heavy burden on family members and caregivers, and even situations where elderly people can act independently, such as shopping at a store, can be accompanied by anxiety. Therefore, there is a need for effective systems that support the safety and independence of elderly people and reduce the burden on family members and caregivers. In particular, there is a need for technology that can analyze the user's emotional state and provide appropriate assistance based on that analysis.

[1386] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1387] In this invention, the server includes a natural language processing means for collecting information based on a dialogue with a user, an information storage means for storing the collected information in a database, an image recognition means for analyzing daily activities using image recognition technology, a support means for generating appropriate advice based on the analysis results and providing it to the user, an emotion analysis means for analyzing the user's emotional state, and a display means for supporting elderly people's shopping and daily life in physical stores. This allows elderly people to move around safely and independently in physical stores and reduces the burden on family members and caregivers.

[1388] "Natural language processing means" refers to technology for collecting information through dialogue with users and analyzing it as text data.

[1389] "Information storage means" refers to a system that stores collected information in a database and makes it possible to refer to it as needed.

[1390] "Image recognition means" refers to technology that analyzes image data acquired through devices such as cameras and recognizes specific objects or situations.

[1391] "Support means" refers to a system that provides appropriate advice and guidance to users based on collected information and analysis results.

[1392] "Emotion analysis means" refers to technology that analyzes the content of a user's dialogue and image data to identify the user's emotional state.

[1393] "Display means" refers to technology that provides users with visual or auditory information to assist elderly people in their shopping and daily life in physical stores.

[1394] The system of the present invention is designed to support elderly people in their shopping and daily life in brick-and-mortar stores. The detailed configuration of the system and how it is implemented will be described below.

[1395] System Configuration

[1396] Hardware

[1397] Smart glasses: A device worn by the user that contains a camera and microphone.

[1398] Server: A central system for data processing and storage.

[1399] software

[1400] Natural language processing means: Manage dialogue with users using services such as Amazon AWS's Lex service or Google's Dialogflow.

[1401] Information storage method: Information is stored using a database management system such as SQLite3 or MySQL.

[1402] Image recognition method: Uses image recognition technologies such as OpenCV and YOLOv3.

[1403] Sentiment analysis method: Sentiment analysis is performed using deep learning libraries such as Transformers and Pytorch.

[1404] Display means: Uses the display function of the smart glasses to provide information visually to the user.

[1405] Program processing explanation

[1406] The server first collects input information from the camera and microphone while the user is wearing the smart glasses. A natural language processing means then uses the conversation with the user to analyze the information as text data. For example, if the user says, "I'm not feeling well today," the content is converted into text data.

[1407] The collected information is stored in a database by an information storage means, and this information is available for future reference and used to track the user's progress.

[1408] The camera on the device periodically captures image data, which is then sent to a server. The server then uses image recognition to analyze the image data and identify, for example, the status of a product shelf. For example, when a specific product is in view, information about that product is notified to the user.

[1409] The emotional state of the user is analyzed from the content of the user's dialogue and image data through the emotion analysis means. Based on the analysis results, advice and guidance appropriate to the user's current emotional state is generated and provided to the user through the display means.

[1410] Specific examples

[1411] 1. A user puts on smart glasses and enters a physical store.

[1412] 2. Immediately after entering the store, the natural language processing means asks the user, "How are you feeling today?" The user replies, "I'm a little tired."

[1413] 3. The server stores this information in a database and uses emotion analysis means to recognize the emotion "tired."

[1414] 4. The device's camera scans the shelves, and image recognition detects immune-boosting health foods.

[1415] 5. The user is notified through the display means, "This is a health food that boosts your immunity. Why not give it a try?"

[1416] Prompt Sentence Examples

[1417] 1. "Hi, how are you?"

[1418] 2. "How was your meal today?"

[1419] 3. "Did you find what you were looking for?"

[1420] In this way, this system covers all the technologies necessary to provide an environment where elderly people can shop independently and safely. By combining comprehensive information processing by the server with interaction through smart glasses, it is possible to improve the quality of life of users.

[1421] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1422] Processing Steps

[1423] Step 1:

[1424] The device collects user input information through the smart glasses' camera and microphone. The input data includes the user's voice and the video captured by the camera. This input data is sent to the server as image data for the camera video and as audio data for the user's voice.

[1425] Step 2:

[1426] After the server receives the voice data, it uses natural language processing means to convert the voice data into text data. For example, it analyzes the voice of the question "Hello, how are you?" and generates text data. The input is voice data and the output is text data.

[1427] Step 3:

[1428] The server continues the dialogue with the user based on the converted text data. The converted text data is analyzed using natural language processing means to generate appropriate questions and answers. The generated dialogue content is sent to the terminal as voice. The input is the converted text data, and the output is the generated dialogue content.

[1429] Step 4:

[1430] The device receives the generated dialogue content sent from the server and notifies the user by voice, for example, asking the user through the smart glasses, "How are you feeling today?" The user's voice response is collected again and sent to the server.

[1431] Step 5:

[1432] The server receives the image data and analyzes the everyday environment using image recognition means. Specifically, it identifies objects from camera footage, for example, identifying products on a shelf. The input is image data, and the output is data on the recognized objects.

[1433] Step 6:

[1434] The server generates appropriate advice and guidance based on the analysis results. It generates information to be provided to the user for the recognized object and creates guidance in text format. The input is the recognized object data, and the output is the text data of the advice and guidance.

[1435] Step 7:

[1436] The server uses emotion analysis means to analyze the user's emotional state from the user's voice and image data. For example, if the user replies, "I'm not feeling well today," the server analyzes the text data and voice tone to identify the user's emotional state. The input is the voice and image data, and the output is the analyzed emotional data.

[1437] Step 8:

[1438] The server generates more appropriate dialogue content based on the generated advice and guidance and the analyzed emotional data. For example, if the server determines that the user is tired, it generates a dialogue recommending "health foods that boost the immune system." The input is the text data of the advice and guidance and emotional data, and the output is the optimized dialogue content.

[1439] Step 9:

[1440] The device receives the generated dialogue and notifies the user through the smart glasses. It provides the user with visual and auditory information to assist with shopping. For example, it displays a message such as, "This is a health food that boosts your immunity. Why not give it a try?"

[1441] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1443] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1444] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1445] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1446] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1447] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1448] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1449] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1450] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1451] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1452] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1453] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1454] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1455] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1456] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1457] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1458] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1459] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1460] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1461] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1462] The following is further disclosed regarding the above embodiment.

[1463] (Claim 1)

[1464] a language processing means for collecting information based on a dialogue with a user;

[1465] a data storage means for storing the collected information in a database;

[1466] Image recognition means for analyzing lifestyle habits using image recognition technology;

[1467] an assisting means for generating appropriate advice based on the analysis results and providing the advice to the user;

[1468] A system including:

[1469] (Claim 2)

[1470] a means for monitoring a living environment using a camera device and acquiring image data;

[1471] 10. The system of claim 1.

[1472] (Claim 3)

[1473] a means for processing the acquired image data in real time to determine the user's lifestyle habits;

[1474] 10. The system of claim 1.

[1475] "Example 1"

[1476] (Claim 1)

[1477] a language processing means for collecting information based on a dialogue with a user;

[1478] a data storage means for storing the collected information in a storage means;

[1479] image recognition means for analyzing image data acquired using a camera device;

[1480] an assisting means for generating appropriate notifications or advice based on the analysis results and the collected information and providing the same to the user;

[1481] A system including:

[1482] (Claim 2)

[1483] means for monitoring the environment using a camera device and acquiring image data;

[1484] 10. The system of claim 1.

[1485] (Claim 3)

[1486] a means for processing the acquired image data in real time to determine the user's behavior or situation;

[1487] 10. The system of claim 1.

[1488] "Application Example 1"

[1489] (Claim 1)

[1490] a language processing means for collecting information based on a dialogue with a user;

[1491] a data storage means for storing the collected information in a database;

[1492] Image recognition means for analyzing lifestyle habits using image recognition technology;

[1493] an assisting means for generating appropriate advice based on the analysis results and providing the advice to the user;

[1494] a means for monitoring the condition of the elderly person using a camera device installed in the vehicle and acquiring image data;

[1495] A means for analyzing the acquired image and audio data and providing instructions and warnings tailored to the elderly person's situation;

[1496] A system including:

[1497] (Claim 2)

[1498] A means for monitoring a living environment using an in-vehicle camera device and acquiring image data;

[1499] 10. The system of claim 1.

[1500] (Claim 3)

[1501] a means for processing the acquired image data in real time to determine the user's lifestyle habits;

[1502] a means for analyzing the elderly person's vocal responses and recording the information in a database as needed;

[1503] 10. The system of claim 1.

[1504] "Example 2: Combining Emotion Engines"

[1505] (Claim 1)

[1506] a language processing means for collecting information based on a dialogue with a user;

[1507] a data storage means for storing the collected information in a database;

[1508] Image recognition means for analyzing lifestyle habits using image recognition technology;

[1509] emotion analysis means for analyzing an emotional state of a user from the content of the user's dialogue and image data;

[1510] an assisting means for generating appropriate advice based on the analysis results and the emotional state and providing the advice to the user;

[1511] A system including:

[1512] (Claim 2)

[1513] a means for monitoring a living environment using a camera device and acquiring image data;

[1514] 10. The system of claim 1.

[1515] (Claim 3)

[1516] means for processing the acquired image data in real time to determine the user's lifestyle habits and emotional state;

[1517] 10. The system of claim 1.

[1518] "Application example 2 when combining emotion engines"

[1519] (Claim 1)

[1520] natural language processing means for collecting information based on a dialogue with a user;

[1521] an information storage means for storing the collected information in a database;

[1522] Image recognition means for analyzing daily activities using image recognition technology;

[1523] a support means for generating appropriate advice based on the analysis results and providing the advice to the user;

[1524] emotion analysis means for analyzing the emotional state of a user;

[1525] Display means to support elderly people in their shopping and daily life in physical stores,

[1526] A system including:

[1527] (Claim 2)

[1528] means for monitoring the daily environment using a camera device and acquiring image data;

[1529] 10. The system of claim 1.

[1530] (Claim 3)

[1531] means for processing the acquired image data in real time to determine the user's daily activities;

[1532] 10. The system of claim 1. [Explanation of symbols]

[1533] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a language processing means for collecting information based on a dialogue with a user; a data storage means for storing the collected information in a database; Image recognition means for analyzing lifestyle habits using image recognition technology; an assisting means for generating appropriate advice based on the analysis results and providing the advice to the user; A system including:

2. a means for monitoring a living environment using a camera device and acquiring image data; The system of claim 1 .

3. a means for processing the acquired image data in real time to determine the user's lifestyle habits; The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A