System
A system with a terminal and server analyzes audio and video data to support elderly individuals' health and daily activities, addressing loneliness and health issues by providing feedback and reminders.
Patent Information
- Application Number
- JP2024123821
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Elderly people living alone face challenges such as loneliness, dementia, depression, poor diet, forgetfulness in taking medication, and difficulty in managing their daily schedules, which are often not effectively monitored by family members living far away.
A system comprising a terminal with a microphone, camera, speaker, location information acquisition means, and sensor device, connected to a server that analyzes audio and video data to provide feedback on diet, medication intake, and schedule management, generating alerts and reminders as needed.
The system effectively monitors and supports the health and daily activities of elderly individuals, reducing health risks and improving their quality of life by providing precise feedback and reminders.
Smart Images

Figure 2026022304000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Elderly people living alone are at risk of many problems, including loneliness, dementia, depression, lack of diet and exercise, forgetting to take medication, and going missing. It can be particularly difficult for family members who live far away to regularly monitor the health status of their elderly loved ones and provide appropriate support. To address this issue, multifaceted support is needed for the elderly throughout their lives. [Means for solving the problem]
[0005] The present invention provides a system including a terminal equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device, a server that receives and analyzes audio and video data transmitted from the terminal, and a means for notifying a user based on the analysis results from the server. The system realizes the following specific means:
[0006] 1. A means for the device to take a photo of the meal contents, the server to analyze the image and calculate the calories and nutrients, and the device to provide the user with feedback based on the analysis results.
[0007] 2. A means for the terminal to take a picture of the medication intake status, the server to analyze the video to confirm the medication intake, and if any abnormalities are found, generate an alert and the terminal notifies the user.
[0008] In this way, it is possible to remotely check the lifestyle and health status of elderly people and provide appropriate feedback and advice.
[0009] A "microphone" is a device for collecting sound and converting sound waves into electrical signals.
[0010] A "camera" is a device that converts light into electrical signals and records them as still or moving images.
[0011] A "speaker" is a device that converts electrical signals into sound waves and outputs sound.
[0012] A "location information acquisition means" is a device that detects a physical location using technology such as GPS and acquires that information.
[0013] A "sensor device" is a device that detects a physical quantity and converts it into an electrical signal. Examples include a heart rate sensor and a pedometer.
[0014] A "terminal" is a device that integrates the above-mentioned multiple devices and interacts with a user.
[0015] A "server" is a computer system that receives data sent from a terminal, analyzes it, and returns the results.
[0016] A "user" is an individual who utilizes the system to receive input and feedback about aspects of their life.
[0017] "Analysis results" are information obtained as a result of processing performed by the server based on the data received.
[0018] "Notification" is the act of a device providing information to a user through audio or visual means.
[0019] "Meal content" refers to the types and amounts of food a user consumes during a meal.
[0020] A "calorie" is a unit that indicates the amount of energy contained in food.
[0021] "Nutrients" are the components of food that the body needs, such as vitamins, minerals, and proteins.
[0022] "Medicine intake status" is information about which medicines a user takes and how they take them.
[0023] An "alert" is a notification that notifies the user of an abnormality or a situation that requires attention. [Brief explanation of the drawings]
[0024] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0025] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0026] First, the terms used in the following description will be explained.
[0027] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0028] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0029] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0030] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0031] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0032] [First embodiment]
[0033] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0034] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0035] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0036] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0037] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0038] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0039] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0040] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0041] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0043] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0045] This invention is a system that includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device to support the daily lives of elderly people, and a server that receives and analyzes data from these devices. Specific embodiments of the system program and its processing are described below.
[0046] Program processing
[0047] Dietary management
[0048] When a user wants to eat, they say "I'm going to start eating" to the camera. The device then activates the camera and captures the video of the meal. The captured video is then sent from the device to the server.
[0049] The server analyzes the received video to recognize the meal contents and retrieves the calories and nutrients obtained from the database. The analysis results are then sent from the server to the device, which then verbally informs the user of the next action to be taken and an evaluation of the meal.
[0050] For example, suppose a user is eating salad and chicken for lunch. The device captures the video, which is then received and analyzed by the server. As a result of the analysis, the calories and nutrients of the salad and chicken are calculated, and it is determined that they are lacking in vitamin C. Based on this information, the device provides feedback to the user, such as "Add some fruit containing vitamin C to your next meal."
[0051] Medication Management
[0052] When it's time for the user to take their medicine, they say "I'll take my medicine" to the device. In response, the device activates its camera, captures the video, and sends it to the server. The server analyzes the video data to confirm the type of medicine and the intake status. It checks whether the medicine has been taken properly and generates an alert if there is a problem.
[0053] For example, if a user is about to take their morning medicine, the device will record the action and send it to the server, which will then use the video to check whether the medicine was taken properly. If the user forgets to take their morning medicine, the device will notify the user, asking, "Did you forget to take your morning medicine?"
[0054] Schedule management
[0055] The user says to the device, "I have a dentist appointment tomorrow at 3 p.m." The device converts this speech into text and sends it to the server, where it analyzes the text data to identify the appointment date and time and records it on the calendar.
[0056] Thirty minutes before the scheduled time, the server generates a reminder and sends it to the device, which then issues a voice reminder saying, "It's almost time for your 3:00 PM dentist appointment."
[0057] Specific examples
[0058] 1. Dietary Management:
[0059] A user eats oatmeal and a banana for breakfast.
[0060] The device captures the video and sends it to the server.
[0061] The server analyzed the calories and nutrients in the oatmeal and banana and determined that they were lacking in dietary fiber.
[0062] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[0063] 2. Medication Management:
[0064] The user takes regular medication before going to bed.
[0065] The device takes a photo of the situation and sends it to a server to confirm the intake status.
[0066] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[0067] No problem: Under normal circumstances, the device will notify you that it was successful.
[0068] 3. Schedule Management:
[0069] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[0070] The device sends this information to the server and records it in the calendar.
[0071] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[0072] In this way, the entire system can provide multifaceted support for users' daily lives and ensure the safety and health of the elderly.
[0073] The processing flow will be explained below.
[0074] Dietary management
[0075] Step 1:
[0076] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[0077] Step 2:
[0078] The device recognizes the voice and activates the camera to capture footage of the meal.
[0079] Step 3:
[0080] The device sends the captured video to the server.
[0081] Step 4:
[0082] The server analyzes the received video data.
[0083] The server uses computer vision technology to identify the food.
[0084] The server retrieves calorie and nutrient information about the food from a database.
[0085] Step 5:
[0086] The server evaluates the diet based on the analysis results, identifies missing nutrients, and generates recommended menus.
[0087] Step 6:
[0088] The server sends the analysis results and suggestions to the terminal.
[0089] Step 7:
[0090] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[0091] Medication Management
[0092] Step 1:
[0093] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[0094] Step 2:
[0095] The device will recognize the voice and activate the camera.
[0096] Step 3:
[0097] The device captures a video of the medicine and sends the data to a server.
[0098] Step 4:
[0099] The server receives the video data and analyzes it.
[0100] The server identifies the type of drug and the intake situation from the video.
[0101] Step 5:
[0102] The server checks whether proper intake has been achieved and generates an alert if there is a shortage or abnormality.
[0103] Step 6:
[0104] The server sends the alert information to the terminal.
[0105] Step 7:
[0106] The device will notify the user by voice, "Did you forget your morning medicine?"
[0107] Schedule management
[0108] Step 1:
[0109] The user says to the device, "I have a dentist appointment tomorrow at 3:00 PM."
[0110] Step 2:
[0111] The device converts the speech into text and sends the data to the server.
[0112] Step 3:
[0113] The server receives the text data and analyzes it.
[0114] The server uses natural language processing technology to identify the scheduled date, time, and content.
[0115] Step 4:
[0116] The server records the event on the calendar based on the analysis results.
[0117] Step 5:
[0118] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[0119] Step 6:
[0120] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[0121] summary
[0122] Through these detailed processing steps, the "Overly Nosy Dog" system provides multifaceted support for the lives of the elderly, allowing users to live with peace of mind.
[0123] Example 1
[0124] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0125] There is a need for systems that support the lives of elderly people and efficiently manage their health and daily schedules. However, existing systems often lack the precision of analyzing voice input and camera footage, and provide inaccurate feedback to users. Furthermore, important tasks such as medication intake and dietary management are often not automated. This can make it difficult for elderly people to properly take medication and manage their nutrition, potentially increasing health risks. Furthermore, schedule management must be done manually, which can lead to forgetfulness.
[0126] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0127] In this invention, the server includes a terminal equipped with a microphone, a camera, a speaker, location information acquisition means, and a sensor device, an information processing device that receives and analyzes audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the information processing device, means for the terminal to photograph meal contents, the information processing device to analyze the images to calculate calories and nutrients, and the terminal to provide feedback based on the analysis results, means for the terminal to photograph medication intake status, the information processing device to analyze the images to check medication intake, and if there is an abnormality, generate an alert and notify the user, means for the terminal to convert audio to text, the information processing device to analyze the text data to identify an appointment and record it on a calendar, and means for the information processing device to generate a reminder a certain time before the appointment time and notify the user, thereby improving the quality of life of elderly people, reducing health risks, and enabling efficient management of daily life.
[0128] A "microphone" is a device that converts sound into an electrical signal.
[0129] A "camera" is a device that converts light into an electrical signal and captures images.
[0130] A "speaker" is a device that reproduces electrical signals as sound.
[0131] A "location information acquisition means" is a device that has the function of determining the current location of a device using GPS or other location information technology.
[0132] A "sensor device" is a device that detects changes in the physical environment and outputs them as an electrical signal.
[0133] A "terminal" is a device that transmits and receives information between a user and a system, and is a multifunction device that includes a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[0134] An "information processing device" is a computer device that receives and analyzes audio and video data transmitted from a terminal.
[0135] The "analysis results" are data obtained from the audio and video data analyzed by the information processing device, and are information that is fed back to the user.
[0136] "Feedback" refers to advice, prompts for action, and notifications provided to users based on the analysis results.
[0137] A "calorie" is a unit that indicates the amount of energy contained in food.
[0138] "Nutrients" are components contained in food and are substances necessary for the growth and maintenance of health of the body.
[0139] "Medicine intake status" refers to whether the user is taking prescribed medication appropriately.
[0140] "Abnormal" refers to a state that is not an expected normal state, and includes, for example, failure to take medication or not following schedules.
[0141] An "alert" is a notification that notifies the user when an abnormality or a condition requiring attention occurs.
[0142] "Speech-to-text" is the process of converting spoken words into text using speech recognition technology.
[0143] A "calendar" is a schedule management tool that records date information for managing appointments and important events.
[0144] "Reminder" is a function that notifies and reminds the user of pre-set schedules and tasks.
[0145] This invention is a system designed to support the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and an information processing device that receives and analyzes audio and video data transmitted from these devices. This system notifies the user based on the results of analyzing various types of data.
[0146] Hardware and Software Configuration
[0147] 1. Device:
[0148] Microphone: Used to collect audio data.
[0149] Camera: Used to collect video data. Example: Raspberry Pi camera module.
[0150] Speaker: Used to provide audio notifications to the user.
[0151] Location information acquisition means: Includes GPS modules, etc.
[0152] Sensor devices: Used to collect information about the user's surrounding environment.
[0153] 2. Information processing equipment:
[0154] Server: A central control device that receives and analyzes audio and video data sent from terminals.
[0155] Database: A database such as MySQL is used to store meal and medication information.
[0156] Image processing software: Analyzes received video data using OpenCV, TensorFlow, etc.
[0157] Speech recognition and speech synthesis software, such as Google Speech-to-Text API and Google Text-to-Speech API, to analyze and generate voice data.
[0158] Program processing
[0159] Dietary management
[0160] When a user starts eating, they say "I'm going to start eating" to the device. The device detects this voice command, activates the camera to capture video of the meal, and sends it to the server. The server analyzes the received video and identifies the meal contents. Based on the identified meal contents, it retrieves calorie and nutrient information from the database and sends the analysis results to the device. The device then notifies the user of the results by voice, giving advice such as "Add a fruit containing vitamin C to your next meal."
[0161] Medication Management
[0162] When the user wants to take their medicine, they say, "I'm taking my medicine." The device activates the camera and sends a video of the medicine being taken to the server. The server analyzes the video data and checks the type of medicine and the state of ingestion. If there is a problem, an alert is generated and the device notifies the user, "Did you forget to take your morning medicine?"
[0163] Schedule management
[0164] When a user says, "I have a dentist appointment tomorrow at 3 p.m.", the device converts the speech to text and sends it to the server. The server analyzes the text data to identify the appointment date and time and records it on the calendar. 30 minutes before the scheduled time, the server generates a reminder and the device notifies the user, "I have a doctor's appointment at 1 p.m."
[0165] Specific examples
[0166] 1. Dietary Management:
[0167] The user says they eat oatmeal and a banana for breakfast.
[0168] The device captures the video and sends it to the server.
[0169] The server analyzes the calories and nutrients in oatmeal and bananas and determines that they are lacking in dietary fiber.
[0170] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[0171] 2. Medication Management:
[0172] The user says he takes his regular medication before going to bed.
[0173] The device takes a photo of the situation and sends it to a server to confirm the intake status.
[0174] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[0175] If there are no problems, the device will notify you with a "Good job!"
[0176] 3. Schedule Management:
[0177] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[0178] The terminal sends this information to the server and updates the schedule.
[0179] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[0180] This system will support the daily lives of the elderly and efficiently manage their health and schedules.
[0181] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0182] Dietary management
[0183] Step 1:
[0184] The user says to the device, "I'm going to start eating." The input voice data is collected by the device's microphone.
[0185] Step 2:
[0186] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[0187] Step 3:
[0188] The device recognizes the voice command and sends a signal to activate the camera module. The camera initializes and captures footage of the user's meal.
[0189] Step 4:
[0190] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[0191] Step 5:
[0192] The server analyzes the received video data using image processing software (e.g., OpenCV and TensorFlow). As a result of the analysis, the identified meal contents, calories, and nutrient information are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[0193] Step 6:
[0194] The server sends the analysis results to the device, which encodes them in JSON format and decodes them to use the data.
[0195] Step 7:
[0196] The analysis results received by the device are converted into voice using speech synthesis software (e.g., Google Text-to-Speech API), and feedback is given to the user, such as "Add some fruit containing vitamin C to your next meal." This is the final output.
[0197] Medication Management
[0198] Step 1:
[0199] The user says to the device, "I'm going to take my medicine." The input voice data is collected by the device's microphone.
[0200] Step 2:
[0201] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[0202] Step 3:
[0203] The device recognizes the voice command and sends a signal to activate the camera module. The camera is initialized and captures video of the user taking the medication.
[0204] Step 4:
[0205] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[0206] Step 5:
[0207] The server analyzes the received video data using image processing software (e.g., OpenCV and deep learning models). As a result of the analysis, the type of medication identified and the intake status are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[0208] Step 6:
[0209] The server determines whether the medication was taken properly and generates a "no problem" message if there is no problem, or an alert if there is a problem. This message becomes the input for the next step.
[0210] Step 7:
[0211] The device receives the message from the server and uses speech synthesis software (e.g., Google Text-to-Speech API) to notify the user. For example, it might say, "Did you forget your morning medicine?" This is the final output.
[0212] Schedule management
[0213] Step 1:
[0214] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 PM." The input voice data is collected by the terminal's microphone.
[0215] Step 2:
[0216] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[0217] Step 3:
[0218] The terminal sends the converted text data to the server, which becomes the input for the next step.
[0219] Step 4:
[0220] The server analyzes the received text data using a natural language processing engine (e.g., NLTK) to identify the reservation date and time. The identified reservation date and time becomes the input for the next step.
[0221] Step 5:
[0222] The server records the identified scheduled date and time in a database (e.g. MySQL) and updates the schedule.
[0223] Step 6:
[0224] 30 minutes before the scheduled time, the server generates a reminder and sends it to the device in JSON format. This reminder serves as input for the next step.
[0225] Step 7:
[0226] The device receives a reminder and notifies the user using speech synthesis software (e.g., Google Text-to-Speech API). For example, it might say, "You have a doctor's appointment at 1 p.m." This is the final output.
[0227] (Application example 1)
[0228] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0229] There is a need for systems that enable users, such as the elderly, to smoothly shop in brick-and-mortar stores and support health management and planned purchases. In particular, there is a lack of systems that integrate product information gathering, in-store navigation, and reminder functions. Therefore, the challenge is to provide a system that improves the safety and convenience of elderly people living independently.
[0230] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0231] In this invention, the server includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, means for receiving and analyzing audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the server, means for allowing the user to specify by voice where they want to go in the store and for recognizing the voice to provide navigation, means for the navigation to scan products with a camera and provide product information by voice, and means for the navigation to remind the user of products to be purchased in the store and when to next visit. This improves the convenience and safety of shopping in stores for elderly people and enables health management and planned purchases.
[0232] A "microphone" is a device that converts sound into an electrical signal and acquires sound data.
[0233] A "camera" is a device that captures images and videos and acquires them as digital data.
[0234] A "speaker" is a device that converts electrical signals into sound to notify or guide the user.
[0235] "Location information acquisition means" refers to devices or technologies for acquiring the current location of a terminal using GPS, Wi-Fi, etc.
[0236] A "sensor device" is a device that detects the environment and the user's situation and collects it as data.
[0237] A "terminal" is a device equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and used for direct interaction with a user.
[0238] A "server" is a computer system that receives and analyzes audio data, video data, location data, and the like sent from a terminal.
[0239] "Navigation" is a function that allows the user to specify the location or product they want to go to, and then recognizes the voice and provides guidance within the store.
[0240] "Product information" refers to detailed data such as the product name, price, ingredients, and nutrients.
[0241] "Remind" is a function that notifies users of schedules and important matters so that they do not forget.
[0242] This invention is a system that enables users, such as elderly people, to smoothly shop in brick-and-mortar stores and supports health management and planned purchases. The details of the system are described below.
[0243] System configuration
[0244] The system consists of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data sent from these terminals. The terminals interact directly with users, while the server provides feedback based on the analysis results.
[0245] Hardware and software used
[0246] Hardware: Smartphones, smart glasses, head-mounted displays
[0247] software:
[0248] Speech recognition: sr(SpeechRecognition instance), pyttsx3
[0249] Geolocation: geopy library
[0250] Image analysis: OpenCV
[0251] Program processing
[0252] The server receives the audio data, video data, and location data sent from the terminal and analyzes it as follows.
[0253] 1. Speech Recognition:
[0254] The device collects audio as the user speaks into the microphone, and the audio data is converted to text using a SpeechRecognition instance.
[0255] 2. Navigation provided:
[0256] When a user speaks the name of a place or product they want to go to, that information is sent as text data to the server, which then compares it with store map data and generates navigation instructions. The instructions are then provided to the user as voice feedback using pyttsx3.
[0257] 3.Product information provided:
[0258] When a user scans a product with a camera, the image data is sent to the server and analyzed using OpenCV. The analysis results, including detailed product information and nutritional information, are sent from the server to the device and notified to the user via voice.
[0259] 4. Reminder function:
[0260] The server stores the user's planned purchases and the next visit date in a database and sends reminders as appropriate. When a reminder is issued, the terminal notifies the user by voice.
[0261] Specific examples
[0262] For example, if a user says, "I want to go to the vegetable section," the system will use the camera and location information acquisition means to determine the user's current location and provide guidance to the desired vegetable section. An example of a prompt sentence in this case is, "Please tell us where you want to go. If the user says, 'Vegetable section,' the system will guide you, 'The vegetable section is in this direction.'"
[0263] Also, when a user scans an item with the camera saying, "Tell me the nutritional information of this apple," the system uses OpenCV to analyze the image of the apple and provides the nutritional information of the apple by voice. An example of a prompt sentence in this case is, "Please scan the item with the camera. When the user scans the item, the system will provide the nutritional information of that item by voice."
[0264] In this way, the entire system can provide multifaceted support for elderly people's shopping in brick-and-mortar stores, improving the quality of their daily lives.
[0265] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0266] Step 1:
[0267] The user speaks voice commands into the microphone regarding the places and products they want to go to.
[0268] Input: Audio data
[0269] Output: Audio file
[0270] Specific behavior:
[0271] The device collects audio data through a microphone and creates an audio file.
[0272] Step 2:
[0273] The voice data collected by the device is converted into text data using a SpeechRecognition instance.
[0274] Input: Audio file
[0275] Output: Text data
[0276] Specific behavior:
[0277] The audio file is analyzed and the audio data is converted into text, thereby obtaining the user's instructions as text.
[0278] Step 3:
[0279] The terminal transmits the text data to the server.
[0280] Input: Text data
[0281] Output: Sending text data
[0282] Specific behavior:
[0283] The terminal transmits the text data to the server via the network.
[0284] Step 4:
[0285] The server analyzes the text data and recognizes the user's instructions.
[0286] Input: Text data
[0287] Output: Analysis results
[0288] Specific behavior:
[0289] The server analyzes the text data to understand the location the user wants to go to and the products they are looking for, and then determines the next action based on the results of this analysis.
[0290] Step 5:
[0291] The server generates navigation instructions and sends them to the terminal.
[0292] Input: Analysis results
[0293] Output: Navigation instructions
[0294] Specific behavior:
[0295] The server then uses the analysis results to reference the store's map data and generates navigation instructions to the user's desired destination. These instructions are sent to the device in text format.
[0296] Step 6:
[0297] The terminal notifies the user of the navigation instructions received from the server using voice synthesis (pyttsx3).
[0298] Input: Navigation instructions
[0299] Output: Audio feedback
[0300] Specific behavior:
[0301] Based on the navigation instructions, the device uses voice synthesis technology to notify the user of the instructions aloud, for example, "The vegetable section is in this direction."
[0302] Step 7:
[0303] The user scans the item with the camera.
[0304] Input: Video data
[0305] Output: Video file
[0306] Specific behavior:
[0307] A camera is used to take a picture of a product designated by a user, and a video file is created.
[0308] Step 8:
[0309] The video data captured by the device is sent to the server.
[0310] Input: Video file
[0311] Output: Sending video data
[0312] Specific behavior:
[0313] The terminal transmits the video file to the server via the network.
[0314] Step 9:
[0315] The server uses OpenCV to analyze the video data and obtain detailed product information and nutritional information.
[0316] Input: Video data
[0317] Output: Product information
[0318] Specific behavior:
[0319] The server uses OpenCV to analyze the video data, identify the product, and retrieve detailed information from a database, such as the nutritional information of an apple.
[0320] Step 10:
[0321] The server sends the acquired product information to the terminal, and the terminal notifies the user using voice synthesis (pyttsx3).
[0322] Input: Product information
[0323] Output: Audio feedback
[0324] Specific behavior:
[0325] The device uses voice synthesis technology to notify the user of product information, such as "The vitamin C content of this apple is..."
[0326] Step 11:
[0327] The server stores the user's planned purchases and the next visit date in a database and sends reminders.
[0328] Input: Items to be purchased, time of visit
[0329] Output: Reminder notification
[0330] Specific behavior:
[0331] The server stores the user's planned purchases and the next visit date in a database, and generates appropriate reminders to notify the device. The device then issues a voice message, such as "The next time you visit is on this date."
[0332] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0333] This system provides multifaceted support for the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data. Furthermore, by combining it with an emotion engine, it has the function of recognizing the user's emotions and providing appropriate feedback and support based on that information.
[0334] Program processing
[0335] Dietary management
[0336] When a user begins eating, they say "I'm going to start eating" to the camera. The device recognizes the voice, activates the camera, captures video of the meal, and sends the data to the server. The server analyzes the video and calculates the calories and nutrients of the food. Based on the analysis results, the device gives the user audio feedback such as "Please add a fruit containing vitamin C to your next meal."
[0337] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[0338] Medication Management
[0339] When a user is about to take medicine, they say "I'm going to take my medicine" to the device. The device activates the camera, captures a video of the medicine, and sends it to the server. The server analyzes the video and checks the type of medicine and the intake status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user.
[0340] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[0341] Schedule management
[0342] When a user enters an appointment, they say to the device, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends it to the server. The server analyzes the text data and records the appointment on the calendar. 30 minutes before the scheduled time, the server generates reminder information, and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM."
[0343] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[0344] Emotion Recognition and Feedback
[0345] When a user expresses emotions in various everyday situations, the device captures their voice and facial expressions using a microphone and camera. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the device provides the user with suggestions for stress relief and relaxation.
[0346] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[0347] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[0348] The processing flow will be explained below.
[0349] Dietary management
[0350] Step 1:
[0351] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[0352] Step 2:
[0353] The device recognizes the voice and activates the camera to capture footage of the meal.
[0354] Step 3:
[0355] The device transmits the captured video data to the server.
[0356] Step 4:
[0357] The server analyzes the received video data.
[0358] The server uses computer vision technology to identify the food.
[0359] The server retrieves the calorie and nutrient information of the food from a database.
[0360] Step 5:
[0361] The server evaluates the diet based on the analysis results, identifies any nutrient deficiencies, and generates recommended menus.
[0362] Step 6:
[0363] The server sends the analysis results and suggestions to the terminal.
[0364] Step 7:
[0365] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[0366] Medication Management
[0367] Step 1:
[0368] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[0369] Step 2:
[0370] The device will recognize the voice and activate the camera.
[0371] Step 3:
[0372] The device captures a video of the medicine and sends the data to a server.
[0373] Step 4:
[0374] The server receives the video data and analyzes it.
[0375] The server identifies the type of drug and the intake situation from the video.
[0376] Step 5:
[0377] The server checks whether the medication is being taken properly and generates an alert if there is a shortage or abnormality.
[0378] Step 6:
[0379] The server sends the alert information to the terminal.
[0380] Step 7:
[0381] The device will notify the user by voice, "Did you forget your morning medicine?"
[0382] Schedule management
[0383] Step 1:
[0384] The user says to the device, "I have a dentist appointment tomorrow at 3:00 PM."
[0385] Step 2:
[0386] The device converts the speech into text and sends the data to the server.
[0387] Step 3:
[0388] The server receives the text data and analyzes it.
[0389] The server uses natural language processing technology to identify the scheduled date, time, and content.
[0390] Step 4:
[0391] The server records the event on the calendar based on the analysis results.
[0392] Step 5:
[0393] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[0394] Step 6:
[0395] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[0396] Emotion Recognition and Feedback
[0397] Step 1:
[0398] When a user expresses emotions in daily life, the device uses a microphone and a camera to capture their voice and facial expressions.
[0399] Step 2:
[0400] The terminal transmits the captured audio and video data to the server.
[0401] Step 3:
[0402] The server analyzes the received data and identifies the user's emotions using an emotion engine.
[0403] The server analyzes the emotional state from voice and facial expression data.
[0404] Step 4:
[0405] The server generates appropriate feedback and suggestions based on the sentiment analysis results.
[0406] Step 5:
[0407] The server sends the feedback data to the terminal.
[0408] Step 6:
[0409] The device will provide the user with appropriate voice feedback, such as "How about listening to some music for relaxation?"
[0410] Step 7:
[0411] The server accumulates emotional data and analyzes emotional trends over time.
[0412] Step 8:
[0413] If necessary, the server sends suggestions for counseling or mental support to the terminal, which then notifies the user.
[0414] summary
[0415] Through these detailed processing steps, the "Overly Nosy Dog" system, which combines an emotion engine, provides multifaceted support for the physical and mental health of elderly people, providing an environment in which they can live with peace of mind.
[0416] Example 2
[0417] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0418] The goal is to realize a system that provides multifaceted support for the lives of the elderly and an environment in which they can live with peace of mind. In particular, it is necessary to comprehensively support the user's health and daily life through managing food calories and nutrients, checking medication intake, and even analyzing emotions and providing feedback.
[0419] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0420] In this invention, the server includes means for analyzing video data and calculating food calories and nutrients, means for analyzing medication intake and generating an alert if an abnormality is detected, and means for analyzing the user's emotions using an emotion engine and providing appropriate feedback, thereby enabling the user to manage their diet, medication, and emotions.
[0421] A "microphone" is a device that converts sound into an electrical signal.
[0422] A "camera" is a device that captures images and stores them as digital data.
[0423] A "speaker" is a device that converts electrical signals into sound and plays it back.
[0424] The "location information acquisition means" is a device that measures the current location using a GPS or the like and acquires the data.
[0425] A "sensor device" is a device that includes a sensor for detecting the surrounding environment and the state of the user.
[0426] A "terminal" is an information processing device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[0427] A "server" is an information processing device that receives and analyzes data sent from terminals via a network.
[0428] "Analysis results" are information obtained by processing the data received by the server.
[0429] "Calories" is an indicator of the amount of energy contained in food.
[0430] Nutrients are substances contained in food that are necessary for the growth and maintenance of the body.
[0431] "Medicine intake status" is information indicating whether the user has taken the medicine correctly.
[0432] An "emotion engine" is an algorithm or software that analyzes audio and video to estimate a user's emotions.
[0433] "Feedback" refers to information or advice that the server provides to the user based on the analysis results.
[0434] An "alert" is a warning that notifies the user of an abnormality or a situation that requires attention.
[0435] This invention is a system that provides multifaceted support for the daily lives of the elderly, and is primarily composed of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that analyzes this data and provides feedback.
[0436] Dietary management
[0437] When starting a meal, the user speaks to the device, saying "I'm going to start eating." The device then recognizes the voice and activates the camera. The camera captures video of the meal, and the data is sent to the server. The server analyzes the received video data and calculates the calories and nutrients of the food. Based on the analysis results, the device provides audible feedback such as "Please add a fruit containing vitamin C to your next meal." The software used is Google Speech-to-Text API for speech recognition, OpenCV and TensorFlow for video analysis, and Google Text-to-Speech API for text-to-speech.
[0438] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[0439] Example prompt sentence:
[0440] When you say "start eating" and turn to the camera, your dietary management begins. What nutrients might be missing from your next meal?
[0441] Medication Management
[0442] When a user takes medicine, they say "take medicine" to the device, which activates the device's camera and captures the ingestion of the medicine on video. This video data is sent to a server, which analyzes the video to confirm the type of medicine and the ingestion status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user. The software used is Google Speech-to-Text API for voice recognition, and OpenCV and TensorFlow for video analysis.
[0443] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[0444] Example prompt sentence:
[0445] When you say "I'm going to take my medicine" and face the camera, the medication administration will begin. Please let us know if you have taken the medicine.
[0446] Schedule management
[0447] When a user enters an appointment, they speak to the device, saying, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends the text data to the server. The server analyzes the text and records the appointment on the calendar. 30 minutes before the scheduled time, a reminder is generated and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM." The software used is the Google Speech-to-Text API and natural language processing algorithms.
[0448] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[0449] Example prompt sentence:
[0450] Say, "I have a dentist appointment tomorrow at 3 PM," and we'll record the appointment for you. Send you a reminder 30 minutes before.
[0451] Emotion Recognition and Feedback
[0452] When a user expresses emotions in various everyday situations, the device uses a microphone and camera to capture their voice and facial expressions. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the system provides the user with suggestions for stress relief and relaxation. The software used is an emotion engine (e.g., IBM Watson Tone Analyzer).
[0453] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[0454] Example prompt sentence:
[0455] If you say, "I feel depressed today," you will receive relaxation suggestions. What methods do you recommend?
[0456] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[0457] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0458] Dietary management
[0459] Step 1:
[0460] The user says to the terminal, "I'm going to start eating."
[0461] Input: User's voice command
[0462] Output: Audio data
[0463] Action: The user says "Start eating" to the device.
[0464] Step 2:
[0465] The device performs voice recognition.
[0466] Input: Audio data
[0467] Output: Text data
[0468] What it does: Converts speech to text using speech recognition software (e.g., Google Speech-to-Text API).
[0469] Step 3:
[0470] The device will activate the camera.
[0471] Input: Text data
[0472] Output: Camera activation signal
[0473] Action: Activates the camera sensor and switches it into video capture mode.
[0474] Step 4:
[0475] The device captures video of the meal.
[0476] Input: Camera image
[0477] Output: Video data
[0478] How it works: The camera buffers your meal in real time.
[0479] Step 5:
[0480] The terminal transmits the video data to the server.
[0481] Input: Video data
[0482] Output: Upload to the server
[0483] How it works: Encoded video data is sent to the server via Wi-Fi or mobile data.
[0484] Step 6:
[0485] The server analyzes the video data.
[0486] Input: Video data
[0487] Output: Recognition data
[0488] How it works: Video analysis is performed using OpenCV and TensorFlow to identify the type and quantity of food.
[0489] Step 7:
[0490] The server calculates calories and nutrients.
[0491] Input: Recognition data
[0492] Output: Calorie and nutrient data
[0493] How it works: Looks up the FDA's food database and adds up the calories and nutrients for each food.
[0494] Step 8:
[0495] The server sends the calculation results to the terminal.
[0496] Input: Calorie and nutrient data
[0497] Output: Sending data to the terminal
[0498] Operation: The calculation results are returned to the terminal in JSON format or similar.
[0499] Step 9:
[0500] The device will provide audio feedback.
[0501] Input: Calorie and nutrient data
[0502] Output: Audio feedback
[0503] What it does: Uses the Google Text-to-Speech API to tell you to "Include a fruit containing vitamin C in your next meal."
[0504] Medication Management
[0505] Step 1:
[0506] The user says to the terminal, "I'm going to take my medicine."
[0507] Input: User's voice command
[0508] Output: Audio data
[0509] Action: The user says "take medicine."
[0510] Step 2:
[0511] The device performs voice recognition.
[0512] Input: Audio data
[0513] Output: Text data
[0514] What it does: Uses speech recognition software to convert speech to text.
[0515] Step 3:
[0516] The device will activate the camera.
[0517] Input: Text data
[0518] Output: Camera activation signal
[0519] Action: Activates the camera.
[0520] Step 4:
[0521] The device captures the drug intake as a video.
[0522] Input: Camera image
[0523] Output: Video data
[0524] How it works: The camera collects video data and stores it in a buffer.
[0525] Step 5:
[0526] The terminal transmits the video data to the server.
[0527] Input: Video data
[0528] Output: Upload to server
[0529] Operation: Encodes video data and sends it to the server.
[0530] Step 6:
[0531] The server analyzes the video data.
[0532] Input: Video data
[0533] Output: Recognition data
[0534] How it works: Uses video analysis algorithms to identify medication types and intake.
[0535] Step 7:
[0536] The server checks the medication intake.
[0537] Input: Recognition data
[0538] Output: Intake confirmation data
[0539] What it does: Checks a medication database to assess whether it was taken correctly.
[0540] Step 8:
[0541] The server sends the results to the terminal.
[0542] Input: Intake confirmation data
[0543] Output: Sending data to the terminal
[0544] Behavior: Sends the verification result to the device in JSON format.
[0545] Step 9:
[0546] The device will provide audio feedback.
[0547] Input: Intake confirmation data
[0548] Output: Audio feedback
[0549] What it does: Uses the Google Text-to-Speech API to say "Your medication was taken successfully."
[0550] Schedule management
[0551] Step 1:
[0552] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 p.m."
[0553] Input: User's voice command
[0554] Output: Audio data
[0555] Action: The user speaks the command "dentist appointment."
[0556] Step 2:
[0557] The device performs voice recognition.
[0558] Input: Audio data
[0559] Output: Text data
[0560] What it does: Uses speech recognition software to convert speech into text.
[0561] Step 3:
[0562] The terminal transmits the text data to the server.
[0563] Input: Text data
[0564] Output: Send data to the server
[0565] Action: Uploads the converted text data to the server.
[0566] Step 4:
[0567] The server parses the text.
[0568] Input: Text data
[0569] Output: Analysis data
[0570] How it works: Uses NLP algorithms to extract the event date, time, and content.
[0571] Step 5:
[0572] The server records the event in the calendar.
[0573] Input: Analysis data
[0574] Output: Schedule record data
[0575] Action: Adds new appointment information to the database.
[0576] Step 6:
[0577] The server generates the reminder information.
[0578] Input: Schedule record data
[0579] Output: Remind data
[0580] What it does: Generates a reminder 30 minutes before the scheduled time.
[0581] Step 7:
[0582] The device will remind you by voice.
[0583] Input: Remind data
[0584] Output: Voice reminder
[0585] What it does: Uses the Google Text-to-Speech API to notify you that "It's almost time for your 3pm dentist appointment."
[0586] Emotion Recognition and Feedback
[0587] Step 1:
[0588] The user expresses their feelings.
[0589] Input: Voice and facial expression data
[0590] Output: Emotion data
[0591] How it works: The camera and microphone capture the user's words and facial expressions that express their emotions.
[0592] Step 2:
[0593] The device captures audio and video simultaneously.
[0594] Input: Voice and facial expression data
[0595] Output: Capture data
[0596] Action: Activates microphone and camera to collect data.
[0597] Step 3:
[0598] The device sends the data to the server.
[0599] Input: Capture data
[0600] Output: Send data to the server
[0601] What it does: Encodes audio and video data and sends it to the server.
[0602] Step 4:
[0603] The server analyzes the data.
[0604] Input: Capture data
[0605] Output: Sentiment analysis data
[0606] What it does: Analyzes data using an emotion engine.
[0607] Step 5:
[0608] The server evaluates the emotion.
[0609] Input: Sentiment analysis data
[0610] Output: Emotion rating data
[0611] What it does: Classifies the results of the emotion engine and assigns emotion labels.
[0612] Step 6:
[0613] The server generates the feedback.
[0614] Input: Emotion rating data
[0615] Output: Feedback data
[0616] What it does: Automatically generate stress relief and relaxation suggestions.
[0617] Step 7:
[0618] The device will provide audio feedback.
[0619] Input: Feedback data
[0620] Output: Audio feedback
[0621] What it does: Uses the Google Text-to-Speech API to suggest, "Would you like to listen to some music for relaxation?"
[0622] (Application example 2)
[0623] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0624] Nutritional management, emotional support, and reminder functions are important for elderly people to live their daily lives with peace of mind. However, there are limited means to provide these functions in a single integrated system, and it has been common for individual services or devices to be used separately. This has resulted in problems of poor usability and a heavy burden on users. Furthermore, there are few systems that provide more appropriate support by linking dietary management and emotional support, and there are no systems that link with food delivery services to support the entire process from meal suggestions to ordering. The objective of this invention is to solve these problems.
[0625] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing video data of meals and calculating calories and nutrients, means for recognizing emotional states and suggesting relaxation, and means for suggesting and ordering meals in cooperation with food delivery services. This makes it possible to consistently perform everything from dietary management for elderly people to emotional support, and even meal suggestions and ordering, all in one system.
[0626] A "terminal" is a device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[0627] A "server" is a device that receives and analyzes audio and video data sent from a terminal.
[0628] "Analysis results" refers to information obtained by the server when it analyzes audio and video data.
[0629] "Notification" refers to the transmission of information to the user based on the analysis results.
[0630] "Nutritional information of meals" is data related to the calories and nutrients of meal contents.
[0631] "Feedback" refers to advice and suggestions to the user based on the server's analysis results.
[0632] "Emotional state" refers to the mental and emotional state of the user.
[0633] "Relaxation suggestions" are advice or suggestions regarding stress reduction and relaxation based on emotional state.
[0634] A "food delivery service" is a service that delivers meals.
[0635] "Meal suggestions" are suggestions for meal menus based on the user's health condition and required nutrients.
[0636] "Ordering" refers to the act of a user purchasing a suggested meal through a food delivery service.
[0637] This invention is a multi-functional system for supporting the lives of elderly people. Each element and its function will be explained below.
[0638] System configuration
[0639] The device is equipped with a microphone, camera, speaker, location information acquisition means, and sensor devices. The server has the ability to analyze audio and video data and notify the user based on the analysis results. It also has the ability to calculate nutritional information for meals and provide feedback, recognize emotional states and suggest relaxation activities, and connect with food delivery services to suggest and order meals.
[0640] Hardware and Software Configuration
[0641] Hardware used:
[0642] Smartphone (camera, microphone)
[0643] Server (high performance computer)
[0644] Software used:
[0645] Python
[0646] OpenCV (image processing library)
[0647] SpeechRecognition (speech recognition library)
[0648] Flask (server-side framework)
[0649] Emotion engine (AI model for emotion recognition)
[0650] System Operation
[0651] 1. Speech Recognition:
[0652] - When the user says "I'm going to start eating" to the device, the device's microphone captures the voice and converts it into a string using the SpeechRecognition library.
[0653] 2. Image capture:
[0654] - Once voice recognition is complete, the device's camera will activate and take a picture of the food, which will then be immediately sent to the server.
[0655] 3. Video analysis:
[0656] - The server analyzes the received image data using OpenCV and calculates the calories and nutrients of the meal. The calculation results are stored on the server.
[0657] 4. Emotion recognition:
[0658] - The emotion engine analyzes the user's facial expressions and tone of voice to understand their emotional state. The results are stored on the server as analytical data.
[0659] 5. Feedback and relaxation suggestions:
[0660] - Based on the analysis results, the server notifies the device of the nutrients needed for the next meal. If the emotional state is biased towards stress, it will also suggest music or activities for relaxation.
[0661] 6. Meal Suggestions and Ordering:
[0662] - The server connects the user's nutritional data and the food delivery service to suggest the next meal menu, which can then be ordered directly through the delivery service.
[0663] Natural language description example
[0664] When an elderly person starts eating, this system recognizes the voice command "start eating" and automatically activates the camera to film the meal. The video data is sent to a server, where an AI model analyzes the calories and nutrients. Next, an emotion engine analyzes the user's emotional state and provides appropriate feedback and relaxation suggestions. It also suggests meal menus based on the necessary nutrients, which can then be ordered directly through a food delivery service.
[0665] As a concrete example:
[0666] For example, when a user says "I'm going to start eating" before lunch, the smartphone camera captures an image of the meal. The image is sent to a server, where the calories and nutrients are analyzed. As a result, feedback such as "Add some fruit containing vitamin C to your next meal" is provided. Also, if the user is feeling stressed, suggestions such as "How about listening to some music to relax?" are made.
[0667] Prompt Sentence Examples
[0668] When you say "start eating," the camera automatically activates and takes a picture of your meal. The image data is sent to the server, where it is analyzed by an AI engine, which calculates calories and nutrients. Nutritional advice for your next meal is then provided via voice.
[0669] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0670] Step 1:
[0671] Voice Recognition:
[0672] When a user says "I'm going to start eating" to the device, the device's microphone captures the voice. The captured voice data is converted into text using Python's SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device's microphone records the voice, and the SpeechRecognition library analyzes the voice file and converts it into text format.
[0673] Step 2:
[0674] Image Capture:
[0675] Once voice recognition is complete and the text "Start eating" is confirmed, the device's camera will automatically start up and take an image of the meal. The input is video data from the camera, and the output is the captured image file. Specifically, the camera starts up, frames the user's meal, takes a photo, and saves it as an image file.
[0676] Step 3:
[0677] Video Analysis:
[0678] The image captured by the device is sent to the server. The server analyzes the received image data using the OpenCV library and calculates the type of food, calories, and nutrients. The input is image data, and the output is calorie and nutrient data. Specifically, the server runs an image analysis algorithm to recognize ingredients and output their nutritional information in a table format.
[0679] Step 4:
[0680] Emotion recognition:
[0681] The user's facial expression and tone of voice are also sent to the server and analyzed using the emotion engine. The input is facial image and voice data, and the output is emotional state data. Specifically, the server runs an emotion recognition model to determine the user's emotional state, such as stress, happiness, or fatigue.
[0682] Step 5:
[0683] Feedback and relaxation suggestions:
[0684] Based on the analysis results, the server provides nutritional advice and relaxation suggestions appropriate to the user. The input is calorie and nutrient data and emotional state data, and the output is a feedback message. Specifically, the server generates a text message with recommended meal contents and relaxation methods and sends it to the device.
[0685] Step 6:
[0686] Meal Suggestions and Ordering:
[0687] The server executes a function that links suggested meal menus to food delivery services based on the user's nutritional data. The input is nutritional balance data, and the output is meal suggestions and a confirmation message for the order. Specifically, the server generates an appropriate menu and executes the linking process to allow the user to easily order through the food delivery service.
[0688] Through these steps, the system will be able to smoothly carry out a series of steps from dietary management for the elderly to emotional support, and meal suggestions and ordering.
[0689] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0690] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0691] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0692] [Second embodiment]
[0693] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0694] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0695] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0696] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0697] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0698] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0699] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0700] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0701] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0702] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0703] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0704] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0705] This invention is a system that includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device to support the daily lives of elderly people, and a server that receives and analyzes data from these devices. Specific embodiments of the system program and its processing are described below.
[0706] Program processing
[0707] Dietary management
[0708] When a user wants to eat, they say "I'm going to start eating" to the camera. The device then activates the camera and captures the video of the meal. The captured video is then sent from the device to the server.
[0709] The server analyzes the received video to recognize the meal contents and retrieves the calories and nutrients obtained from the database. The analysis results are then sent from the server to the device, which then verbally informs the user of the next action to be taken and an evaluation of the meal.
[0710] For example, suppose a user is eating salad and chicken for lunch. The device captures the video, which is then received and analyzed by the server. As a result of the analysis, the calories and nutrients of the salad and chicken are calculated, and it is determined that they are lacking in vitamin C. Based on this information, the device provides feedback to the user, such as "Add some fruit containing vitamin C to your next meal."
[0711] Medication Management
[0712] When it's time for the user to take their medicine, they say "I'll take my medicine" to the device. In response, the device activates its camera, captures the video, and sends it to the server. The server analyzes the video data to confirm the type of medicine and the intake status. It checks whether the medicine has been taken properly and generates an alert if there is a problem.
[0713] For example, if a user is about to take their morning medicine, the device will record the action and send it to the server, which will then use the video to check whether the medicine was taken properly. If the user forgets to take their morning medicine, the device will notify the user, asking, "Did you forget to take your morning medicine?"
[0714] Schedule management
[0715] The user says to the device, "I have a dentist appointment tomorrow at 3 p.m." The device converts this speech into text and sends it to the server, where it analyzes the text data to identify the appointment date and time and records it on the calendar.
[0716] Thirty minutes before the scheduled time, the server generates a reminder and sends it to the device, which then issues a voice reminder saying, "It's almost time for your 3:00 PM dentist appointment."
[0717] Specific examples
[0718] 1. Dietary Management:
[0719] A user eats oatmeal and a banana for breakfast.
[0720] The device captures the video and sends it to the server.
[0721] The server analyzed the calories and nutrients in the oatmeal and banana and determined that they were lacking in dietary fiber.
[0722] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[0723] 2. Medication Management:
[0724] The user takes regular medication before going to bed.
[0725] The device takes a photo of the situation and sends it to a server to confirm the intake status.
[0726] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[0727] No problem: Under normal circumstances, the device will notify you that it was successful.
[0728] 3. Schedule Management:
[0729] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[0730] The device sends this information to the server and records it in the calendar.
[0731] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[0732] In this way, the entire system can provide multifaceted support for users' daily lives and ensure the safety and health of the elderly.
[0733] The processing flow will be explained below.
[0734] Dietary management
[0735] Step 1:
[0736] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[0737] Step 2:
[0738] The device recognizes the voice and activates the camera to capture footage of the meal.
[0739] Step 3:
[0740] The device sends the captured video to the server.
[0741] Step 4:
[0742] The server analyzes the received video data.
[0743] The server uses computer vision technology to identify the food.
[0744] The server retrieves calorie and nutrient information about the food from a database.
[0745] Step 5:
[0746] The server evaluates the diet based on the analysis results, identifies missing nutrients, and generates recommended menus.
[0747] Step 6:
[0748] The server sends the analysis results and suggestions to the terminal.
[0749] Step 7:
[0750] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[0751] Medication Management
[0752] Step 1:
[0753] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[0754] Step 2:
[0755] The device will recognize the voice and activate the camera.
[0756] Step 3:
[0757] The device captures a video of the medicine and sends the data to a server.
[0758] Step 4:
[0759] The server receives the video data and analyzes it.
[0760] The server identifies the type of drug and the intake situation from the video.
[0761] Step 5:
[0762] The server checks whether proper intake has been achieved and generates an alert if there is a shortage or abnormality.
[0763] Step 6:
[0764] The server sends the alert information to the terminal.
[0765] Step 7:
[0766] The device will notify the user by voice, "Did you forget your morning medicine?"
[0767] Schedule management
[0768] Step 1:
[0769] The user says to the device, "I have a dentist appointment tomorrow at 3:00 PM."
[0770] Step 2:
[0771] The device converts the speech into text and sends the data to the server.
[0772] Step 3:
[0773] The server receives the text data and analyzes it.
[0774] The server uses natural language processing technology to identify the scheduled date, time, and content.
[0775] Step 4:
[0776] The server records the event on the calendar based on the analysis results.
[0777] Step 5:
[0778] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[0779] Step 6:
[0780] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[0781] summary
[0782] Through these detailed processing steps, the "Overly Nosy Dog" system provides multifaceted support for the lives of the elderly, allowing users to live with peace of mind.
[0783] Example 1
[0784] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0785] There is a need for systems that support the lives of elderly people and efficiently manage their health and daily schedules. However, existing systems often lack the precision of analyzing voice input and camera footage, and provide inaccurate feedback to users. Furthermore, important tasks such as medication intake and dietary management are often not automated. This can make it difficult for elderly people to properly take medication and manage their nutrition, potentially increasing health risks. Furthermore, schedule management must be done manually, which can lead to forgetfulness.
[0786] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0787] In this invention, the server includes a terminal equipped with a microphone, a camera, a speaker, location information acquisition means, and a sensor device, an information processing device that receives and analyzes audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the information processing device, means for the terminal to photograph meal contents, the information processing device to analyze the images to calculate calories and nutrients, and the terminal to provide feedback based on the analysis results, means for the terminal to photograph medication intake status, the information processing device to analyze the images to check medication intake, and if there is an abnormality, generate an alert and notify the user, means for the terminal to convert audio to text, the information processing device to analyze the text data to identify an appointment and record it on a calendar, and means for the information processing device to generate a reminder a certain time before the appointment time and notify the user, thereby improving the quality of life of elderly people, reducing health risks, and enabling efficient management of daily life.
[0788] A "microphone" is a device that converts sound into an electrical signal.
[0789] A "camera" is a device that converts light into an electrical signal and captures images.
[0790] A "speaker" is a device that reproduces electrical signals as sound.
[0791] A "location information acquisition means" is a device that has the function of determining the current location of a device using GPS or other location information technology.
[0792] A "sensor device" is a device that detects changes in the physical environment and outputs them as an electrical signal.
[0793] A "terminal" is a device that transmits and receives information between a user and a system, and is a multifunction device that includes a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[0794] An "information processing device" is a computer device that receives and analyzes audio and video data transmitted from a terminal.
[0795] The "analysis results" are data obtained from the audio and video data analyzed by the information processing device, and are information that is fed back to the user.
[0796] "Feedback" refers to advice, prompts for action, and notifications provided to users based on the analysis results.
[0797] A "calorie" is a unit that indicates the amount of energy contained in food.
[0798] "Nutrients" are components contained in food and are substances necessary for the growth and maintenance of health of the body.
[0799] "Medicine intake status" refers to whether the user is taking prescribed medication appropriately.
[0800] "Abnormal" refers to a state that is not an expected normal state, and includes, for example, failure to take medication or not following schedules.
[0801] An "alert" is a notification that notifies the user when an abnormality or a condition requiring attention occurs.
[0802] "Speech-to-text" is the process of converting spoken words into text using speech recognition technology.
[0803] A "calendar" is a schedule management tool that records date information for managing appointments and important events.
[0804] "Reminder" is a function that notifies and reminds the user of pre-set schedules and tasks.
[0805] This invention is a system designed to support the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and an information processing device that receives and analyzes audio and video data transmitted from these devices. This system notifies the user based on the results of analyzing various types of data.
[0806] Hardware and Software Configuration
[0807] 1. Device:
[0808] Microphone: Used to collect audio data.
[0809] Camera: Used to collect video data. Example: Raspberry Pi camera module.
[0810] Speaker: Used to provide audio notifications to the user.
[0811] Location information acquisition means: Includes GPS modules, etc.
[0812] Sensor devices: Used to collect information about the user's surrounding environment.
[0813] 2. Information processing equipment:
[0814] Server: A central control device that receives and analyzes audio and video data sent from terminals.
[0815] Database: A database such as MySQL is used to store meal and medication information.
[0816] Image processing software: Analyzes received video data using OpenCV, TensorFlow, etc.
[0817] Speech recognition and speech synthesis software, such as Google Speech-to-Text API and Google Text-to-Speech API, to analyze and generate voice data.
[0818] Program processing
[0819] Dietary management
[0820] When a user starts eating, they say "I'm going to start eating" to the device. The device detects this voice command, activates the camera to capture video of the meal, and sends it to the server. The server analyzes the received video and identifies the meal contents. Based on the identified meal contents, it retrieves calorie and nutrient information from the database and sends the analysis results to the device. The device then notifies the user of the results by voice, giving advice such as "Add a fruit containing vitamin C to your next meal."
[0821] Medication Management
[0822] When the user wants to take their medicine, they say, "I'm taking my medicine." The device activates the camera and sends a video of the medicine being taken to the server. The server analyzes the video data and checks the type of medicine and the state of ingestion. If there is a problem, an alert is generated and the device notifies the user, "Did you forget to take your morning medicine?"
[0823] Schedule management
[0824] When a user says, "I have a dentist appointment tomorrow at 3 p.m.", the device converts the speech to text and sends it to the server. The server analyzes the text data to identify the appointment date and time and records it on the calendar. 30 minutes before the scheduled time, the server generates a reminder and the device notifies the user, "I have a doctor's appointment at 1 p.m."
[0825] Specific examples
[0826] 1. Dietary Management:
[0827] The user says they eat oatmeal and a banana for breakfast.
[0828] The device captures the video and sends it to the server.
[0829] The server analyzes the calories and nutrients in oatmeal and bananas and determines that they are lacking in dietary fiber.
[0830] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[0831] 2. Medication Management:
[0832] The user says he takes his regular medication before going to bed.
[0833] The device takes a photo of the situation and sends it to a server to confirm the intake status.
[0834] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[0835] If there are no problems, the device will notify you with a "Good job!"
[0836] 3. Schedule Management:
[0837] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[0838] The terminal sends this information to the server and updates the schedule.
[0839] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[0840] This system will support the daily lives of the elderly and efficiently manage their health and schedules.
[0841] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0842] Dietary management
[0843] Step 1:
[0844] The user says to the device, "I'm going to start eating." The input voice data is collected by the device's microphone.
[0845] Step 2:
[0846] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[0847] Step 3:
[0848] The device recognizes the voice command and sends a signal to activate the camera module. The camera initializes and captures footage of the user's meal.
[0849] Step 4:
[0850] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[0851] Step 5:
[0852] The server analyzes the received video data using image processing software (e.g., OpenCV and TensorFlow). As a result of the analysis, the identified meal contents, calories, and nutrient information are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[0853] Step 6:
[0854] The server sends the analysis results to the device, which encodes them in JSON format and decodes them to use the data.
[0855] Step 7:
[0856] The analysis results received by the device are converted into voice using speech synthesis software (e.g., Google Text-to-Speech API), and feedback is given to the user, such as "Add some fruit containing vitamin C to your next meal." This is the final output.
[0857] Medication Management
[0858] Step 1:
[0859] The user says to the device, "I'm going to take my medicine." The input voice data is collected by the device's microphone.
[0860] Step 2:
[0861] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[0862] Step 3:
[0863] The device recognizes the voice command and sends a signal to activate the camera module. The camera is initialized and captures video of the user taking the medication.
[0864] Step 4:
[0865] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[0866] Step 5:
[0867] The server analyzes the received video data using image processing software (e.g., OpenCV and deep learning models). As a result of the analysis, the type of medication identified and the intake status are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[0868] Step 6:
[0869] The server determines whether the medication was taken properly and generates a "no problem" message if there is no problem, or an alert if there is a problem. This message becomes the input for the next step.
[0870] Step 7:
[0871] The device receives the message from the server and uses speech synthesis software (e.g., Google Text-to-Speech API) to notify the user. For example, it might say, "Did you forget your morning medicine?" This is the final output.
[0872] Schedule management
[0873] Step 1:
[0874] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 PM." The input voice data is collected by the terminal's microphone.
[0875] Step 2:
[0876] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[0877] Step 3:
[0878] The terminal sends the converted text data to the server, which becomes the input for the next step.
[0879] Step 4:
[0880] The server analyzes the received text data using a natural language processing engine (e.g., NLTK) to identify the reservation date and time. The identified reservation date and time becomes the input for the next step.
[0881] Step 5:
[0882] The server records the identified scheduled date and time in a database (e.g. MySQL) and updates the schedule.
[0883] Step 6:
[0884] 30 minutes before the scheduled time, the server generates a reminder and sends it to the device in JSON format. This reminder serves as input for the next step.
[0885] Step 7:
[0886] The device receives a reminder and notifies the user using speech synthesis software (e.g., Google Text-to-Speech API). For example, it might say, "You have a doctor's appointment at 1 p.m." This is the final output.
[0887] (Application example 1)
[0888] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0889] There is a need for systems that enable users, such as the elderly, to smoothly shop in brick-and-mortar stores and support health management and planned purchases. In particular, there is a lack of systems that integrate product information gathering, in-store navigation, and reminder functions. Therefore, the challenge is to provide a system that improves the safety and convenience of elderly people living independently.
[0890] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0891] In this invention, the server includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, means for receiving and analyzing audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the server, means for allowing the user to specify by voice where they want to go in the store and for recognizing the voice to provide navigation, means for the navigation to scan products with a camera and provide product information by voice, and means for the navigation to remind the user of products to be purchased in the store and when to next visit. This improves the convenience and safety of shopping in stores for elderly people and enables health management and planned purchases.
[0892] A "microphone" is a device that converts sound into an electrical signal and acquires sound data.
[0893] A "camera" is a device that captures images and videos and acquires them as digital data.
[0894] A "speaker" is a device that converts electrical signals into sound to notify or guide the user.
[0895] "Location information acquisition means" refers to devices or technologies for acquiring the current location of a terminal using GPS, Wi-Fi, etc.
[0896] A "sensor device" is a device that detects the environment and the user's situation and collects it as data.
[0897] A "terminal" is a device equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and used for direct interaction with a user.
[0898] A "server" is a computer system that receives and analyzes audio data, video data, location data, and the like sent from a terminal.
[0899] "Navigation" is a function that allows the user to specify the location or product they want to go to, and then recognizes the voice and provides guidance within the store.
[0900] "Product information" refers to detailed data such as the product name, price, ingredients, and nutrients.
[0901] "Remind" is a function that notifies users of schedules and important matters so that they do not forget.
[0902] This invention is a system that enables users, such as elderly people, to smoothly shop in brick-and-mortar stores and supports health management and planned purchases. The details of the system are described below.
[0903] System configuration
[0904] The system consists of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data sent from these terminals. The terminals interact directly with users, while the server provides feedback based on the analysis results.
[0905] Hardware and software used
[0906] Hardware: Smartphones, smart glasses, head-mounted displays
[0907] software:
[0908] Speech recognition: sr(SpeechRecognition instance), pyttsx3
[0909] Geolocation: geopy library
[0910] Image analysis: OpenCV
[0911] Program processing
[0912] The server receives the audio data, video data, and location data sent from the terminal and analyzes it as follows.
[0913] 1. Speech Recognition:
[0914] The device collects audio as the user speaks into the microphone, and the audio data is converted to text using a SpeechRecognition instance.
[0915] 2. Navigation provided:
[0916] When a user speaks the name of a place or product they want to go to, that information is sent as text data to the server, which then compares it with store map data and generates navigation instructions. The instructions are then provided to the user as voice feedback using pyttsx3.
[0917] 3.Product information provided:
[0918] When a user scans a product with a camera, the image data is sent to the server and analyzed using OpenCV. The analysis results, including detailed product information and nutritional information, are sent from the server to the device and notified to the user via voice.
[0919] 4. Reminder function:
[0920] The server stores the user's planned purchases and the next visit date in a database and sends reminders as appropriate. When a reminder is issued, the terminal notifies the user by voice.
[0921] Specific examples
[0922] For example, if a user says, "I want to go to the vegetable section," the system will use the camera and location information acquisition means to determine the user's current location and provide guidance to the desired vegetable section. An example of a prompt sentence in this case is, "Please tell us where you want to go. If the user says, 'Vegetable section,' the system will guide you, 'The vegetable section is in this direction.'"
[0923] Also, when a user scans an item with the camera saying, "Tell me the nutritional information of this apple," the system uses OpenCV to analyze the image of the apple and provides the nutritional information of the apple by voice. An example of a prompt sentence in this case is, "Please scan the item with the camera. When the user scans the item, the system will provide the nutritional information of that item by voice."
[0924] In this way, the entire system can provide multifaceted support for elderly people's shopping in brick-and-mortar stores, improving the quality of their daily lives.
[0925] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0926] Step 1:
[0927] The user speaks voice commands into the microphone regarding the places and products they want to go to.
[0928] Input: Audio data
[0929] Output: Audio file
[0930] Specific behavior:
[0931] The device collects audio data through a microphone and creates an audio file.
[0932] Step 2:
[0933] The voice data collected by the device is converted into text data using a SpeechRecognition instance.
[0934] Input: Audio file
[0935] Output: Text data
[0936] Specific behavior:
[0937] The audio file is analyzed and the audio data is converted into text, thereby obtaining the user's instructions as text.
[0938] Step 3:
[0939] The terminal transmits the text data to the server.
[0940] Input: Text data
[0941] Output: Sending text data
[0942] Specific behavior:
[0943] The terminal transmits the text data to the server via the network.
[0944] Step 4:
[0945] The server analyzes the text data and recognizes the user's instructions.
[0946] Input: Text data
[0947] Output: Analysis results
[0948] Specific behavior:
[0949] The server analyzes the text data to understand the location the user wants to go to and the products they are looking for, and then determines the next action based on the results of this analysis.
[0950] Step 5:
[0951] The server generates navigation instructions and sends them to the terminal.
[0952] Input: Analysis results
[0953] Output: Navigation instructions
[0954] Specific behavior:
[0955] The server then uses the analysis results to reference the store's map data and generates navigation instructions to the user's desired destination. These instructions are sent to the device in text format.
[0956] Step 6:
[0957] The terminal notifies the user of the navigation instructions received from the server using voice synthesis (pyttsx3).
[0958] Input: Navigation instructions
[0959] Output: Audio feedback
[0960] Specific behavior:
[0961] Based on the navigation instructions, the device uses voice synthesis technology to notify the user of the instructions aloud, for example, "The vegetable section is in this direction."
[0962] Step 7:
[0963] The user scans the item with the camera.
[0964] Input: Video data
[0965] Output: Video file
[0966] Specific behavior:
[0967] A camera is used to take a picture of a product designated by a user, and a video file is created.
[0968] Step 8:
[0969] The video data captured by the device is sent to the server.
[0970] Input: Video file
[0971] Output: Sending video data
[0972] Specific behavior:
[0973] The terminal transmits the video file to the server via the network.
[0974] Step 9:
[0975] The server uses OpenCV to analyze the video data and obtain detailed product information and nutritional information.
[0976] Input: Video data
[0977] Output: Product information
[0978] Specific behavior:
[0979] The server uses OpenCV to analyze the video data, identify the product, and retrieve detailed information from a database, such as the nutritional information of an apple.
[0980] Step 10:
[0981] The server sends the acquired product information to the terminal, and the terminal notifies the user using voice synthesis (pyttsx3).
[0982] Input: Product information
[0983] Output: Audio feedback
[0984] Specific behavior:
[0985] The device uses voice synthesis technology to notify the user of product information, such as "The vitamin C content of this apple is..."
[0986] Step 11:
[0987] The server stores the user's planned purchases and the next visit date in a database and sends reminders.
[0988] Input: Items to be purchased, time of visit
[0989] Output: Reminder notification
[0990] Specific behavior:
[0991] The server stores the user's planned purchases and the next visit date in a database, and generates appropriate reminders to notify the device. The device then issues a voice message, such as "The next time you visit is on this date."
[0992] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0993] This system provides multifaceted support for the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data. Furthermore, by combining it with an emotion engine, it has the function of recognizing the user's emotions and providing appropriate feedback and support based on that information.
[0994] Program processing
[0995] Dietary management
[0996] When a user begins eating, they say "I'm going to start eating" to the camera. The device recognizes the voice, activates the camera, captures video of the meal, and sends the data to the server. The server analyzes the video and calculates the calories and nutrients of the food. Based on the analysis results, the device gives the user audio feedback such as "Please add a fruit containing vitamin C to your next meal."
[0997] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[0998] Medication Management
[0999] When a user is about to take medicine, they say "I'm going to take my medicine" to the device. The device activates the camera, captures a video of the medicine, and sends it to the server. The server analyzes the video and checks the type of medicine and the intake status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user.
[1000] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[1001] Schedule management
[1002] When a user enters an appointment, they say to the device, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends it to the server. The server analyzes the text data and records the appointment on the calendar. 30 minutes before the scheduled time, the server generates reminder information, and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM."
[1003] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[1004] Emotion Recognition and Feedback
[1005] When a user expresses emotions in various everyday situations, the device captures their voice and facial expressions using a microphone and camera. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the device provides the user with suggestions for stress relief and relaxation.
[1006] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[1007] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[1008] The processing flow will be explained below.
[1009] Dietary management
[1010] Step 1:
[1011] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[1012] Step 2:
[1013] The device recognizes the voice and activates the camera to capture footage of the meal.
[1014] Step 3:
[1015] The device transmits the captured video data to the server.
[1016] Step 4:
[1017] The server analyzes the received video data.
[1018] The server uses computer vision technology to identify the food.
[1019] The server retrieves the calorie and nutrient information of the food from a database.
[1020] Step 5:
[1021] The server evaluates the diet based on the analysis results, identifies any nutrient deficiencies, and generates recommended menus.
[1022] Step 6:
[1023] The server sends the analysis results and suggestions to the terminal.
[1024] Step 7:
[1025] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[1026] Medication Management
[1027] Step 1:
[1028] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[1029] Step 2:
[1030] The device will recognize the voice and activate the camera.
[1031] Step 3:
[1032] The device captures a video of the medicine and sends the data to a server.
[1033] Step 4:
[1034] The server receives the video data and analyzes it.
[1035] The server identifies the type of drug and the intake situation from the video.
[1036] Step 5:
[1037] The server checks whether the medication is being taken properly and generates an alert if there is a shortage or abnormality.
[1038] Step 6:
[1039] The server sends the alert information to the terminal.
[1040] Step 7:
[1041] The device will notify the user by voice, "Did you forget your morning medicine?"
[1042] Schedule management
[1043] Step 1:
[1044] The user says to the device, "I have a dentist appointment tomorrow at 3:00 PM."
[1045] Step 2:
[1046] The device converts the speech into text and sends the data to the server.
[1047] Step 3:
[1048] The server receives the text data and analyzes it.
[1049] The server uses natural language processing technology to identify the scheduled date, time, and content.
[1050] Step 4:
[1051] The server records the event on the calendar based on the analysis results.
[1052] Step 5:
[1053] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[1054] Step 6:
[1055] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[1056] Emotion Recognition and Feedback
[1057] Step 1:
[1058] When a user expresses emotions in daily life, the device uses a microphone and a camera to capture their voice and facial expressions.
[1059] Step 2:
[1060] The terminal transmits the captured audio and video data to the server.
[1061] Step 3:
[1062] The server analyzes the received data and identifies the user's emotions using an emotion engine.
[1063] The server analyzes the emotional state from voice and facial expression data.
[1064] Step 4:
[1065] The server generates appropriate feedback and suggestions based on the sentiment analysis results.
[1066] Step 5:
[1067] The server sends the feedback data to the terminal.
[1068] Step 6:
[1069] The device will provide the user with appropriate voice feedback, such as "How about listening to some music for relaxation?"
[1070] Step 7:
[1071] The server accumulates emotional data and analyzes emotional trends over time.
[1072] Step 8:
[1073] If necessary, the server sends suggestions for counseling or mental support to the terminal, which then notifies the user.
[1074] summary
[1075] Through these detailed processing steps, the "Overly Nosy Dog" system, which combines an emotion engine, provides multifaceted support for the physical and mental health of elderly people, providing an environment in which they can live with peace of mind.
[1076] Example 2
[1077] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1078] The goal is to realize a system that provides multifaceted support for the lives of the elderly and an environment in which they can live with peace of mind. In particular, it is necessary to comprehensively support the user's health and daily life through managing food calories and nutrients, checking medication intake, and even analyzing emotions and providing feedback.
[1079] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1080] In this invention, the server includes means for analyzing video data and calculating food calories and nutrients, means for analyzing medication intake and generating an alert if an abnormality is detected, and means for analyzing the user's emotions using an emotion engine and providing appropriate feedback, thereby enabling the user to manage their diet, medication, and emotions.
[1081] A "microphone" is a device that converts sound into an electrical signal.
[1082] A "camera" is a device that captures images and stores them as digital data.
[1083] A "speaker" is a device that converts electrical signals into sound and plays it back.
[1084] The "location information acquisition means" is a device that measures the current location using a GPS or the like and acquires the data.
[1085] A "sensor device" is a device that includes a sensor for detecting the surrounding environment and the state of the user.
[1086] A "terminal" is an information processing device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[1087] A "server" is an information processing device that receives and analyzes data sent from terminals via a network.
[1088] "Analysis results" are information obtained by processing the data received by the server.
[1089] "Calories" is an indicator of the amount of energy contained in food.
[1090] Nutrients are substances contained in food that are necessary for the growth and maintenance of the body.
[1091] "Medicine intake status" is information indicating whether the user has taken the medicine correctly.
[1092] An "emotion engine" is an algorithm or software that analyzes audio and video to estimate a user's emotions.
[1093] "Feedback" refers to information or advice that the server provides to the user based on the analysis results.
[1094] An "alert" is a warning that notifies the user of an abnormality or a situation that requires attention.
[1095] This invention is a system that provides multifaceted support for the daily lives of the elderly, and is primarily composed of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that analyzes this data and provides feedback.
[1096] Dietary management
[1097] When starting a meal, the user speaks to the device, saying "I'm going to start eating." The device then recognizes the voice and activates the camera. The camera captures video of the meal, and the data is sent to the server. The server analyzes the received video data and calculates the calories and nutrients of the food. Based on the analysis results, the device provides audible feedback such as "Please add a fruit containing vitamin C to your next meal." The software used is Google Speech-to-Text API for speech recognition, OpenCV and TensorFlow for video analysis, and Google Text-to-Speech API for text-to-speech.
[1098] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[1099] Example prompt sentence:
[1100] When you say "start eating" and turn to the camera, your dietary management begins. What nutrients might be missing from your next meal?
[1101] Medication Management
[1102] When a user takes medicine, they say "take medicine" to the device, which activates the device's camera and captures the ingestion of the medicine on video. This video data is sent to a server, which analyzes the video to confirm the type of medicine and the ingestion status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user. The software used is Google Speech-to-Text API for voice recognition, and OpenCV and TensorFlow for video analysis.
[1103] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[1104] Example prompt sentence:
[1105] When you say "I'm going to take my medicine" and face the camera, the medication administration will begin. Please let us know if you have taken the medicine.
[1106] Schedule management
[1107] When a user enters an appointment, they speak to the device, saying, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends the text data to the server. The server analyzes the text and records the appointment on the calendar. 30 minutes before the scheduled time, a reminder is generated and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM." The software used is the Google Speech-to-Text API and natural language processing algorithms.
[1108] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[1109] Example prompt sentence:
[1110] Say, "I have a dentist appointment tomorrow at 3 PM," and we'll record the appointment for you. Send you a reminder 30 minutes before.
[1111] Emotion Recognition and Feedback
[1112] When a user expresses emotions in various everyday situations, the device uses a microphone and camera to capture their voice and facial expressions. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the system provides the user with suggestions for stress relief and relaxation. The software used is an emotion engine (e.g., IBM Watson Tone Analyzer).
[1113] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[1114] Example prompt sentence:
[1115] If you say, "I feel depressed today," you will receive relaxation suggestions. What methods do you recommend?
[1116] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[1117] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1118] Dietary management
[1119] Step 1:
[1120] The user says to the terminal, "I'm going to start eating."
[1121] Input: User's voice command
[1122] Output: Audio data
[1123] Action: The user says "Start eating" to the device.
[1124] Step 2:
[1125] The device performs voice recognition.
[1126] Input: Audio data
[1127] Output: Text data
[1128] What it does: Converts speech to text using speech recognition software (e.g., Google Speech-to-Text API).
[1129] Step 3:
[1130] The device will activate the camera.
[1131] Input: Text data
[1132] Output: Camera activation signal
[1133] Action: Activates the camera sensor and switches it into video capture mode.
[1134] Step 4:
[1135] The device captures video of the meal.
[1136] Input: Camera image
[1137] Output: Video data
[1138] How it works: The camera buffers your meal in real time.
[1139] Step 5:
[1140] The terminal transmits the video data to the server.
[1141] Input: Video data
[1142] Output: Upload to the server
[1143] How it works: Encoded video data is sent to the server via Wi-Fi or mobile data.
[1144] Step 6:
[1145] The server analyzes the video data.
[1146] Input: Video data
[1147] Output: Recognition data
[1148] How it works: Video analysis is performed using OpenCV and TensorFlow to identify the type and quantity of food.
[1149] Step 7:
[1150] The server calculates calories and nutrients.
[1151] Input: Recognition data
[1152] Output: Calorie and nutrient data
[1153] How it works: Looks up the FDA's food database and adds up the calories and nutrients for each food.
[1154] Step 8:
[1155] The server sends the calculation results to the terminal.
[1156] Input: Calorie and nutrient data
[1157] Output: Sending data to the terminal
[1158] Operation: The calculation results are returned to the terminal in JSON format or similar.
[1159] Step 9:
[1160] The device will provide audio feedback.
[1161] Input: Calorie and nutrient data
[1162] Output: Audio feedback
[1163] What it does: Uses the Google Text-to-Speech API to tell you to "Include a fruit containing vitamin C in your next meal."
[1164] Medication Management
[1165] Step 1:
[1166] The user says to the terminal, "I'm going to take my medicine."
[1167] Input: User's voice command
[1168] Output: Audio data
[1169] Action: The user says "take medicine."
[1170] Step 2:
[1171] The device performs voice recognition.
[1172] Input: Audio data
[1173] Output: Text data
[1174] What it does: Uses speech recognition software to convert speech to text.
[1175] Step 3:
[1176] The device will activate the camera.
[1177] Input: Text data
[1178] Output: Camera activation signal
[1179] Action: Activates the camera.
[1180] Step 4:
[1181] The device captures the drug intake as a video.
[1182] Input: Camera image
[1183] Output: Video data
[1184] How it works: The camera collects video data and stores it in a buffer.
[1185] Step 5:
[1186] The terminal transmits the video data to the server.
[1187] Input: Video data
[1188] Output: Upload to server
[1189] Operation: Encodes video data and sends it to the server.
[1190] Step 6:
[1191] The server analyzes the video data.
[1192] Input: Video data
[1193] Output: Recognition data
[1194] How it works: Uses video analysis algorithms to identify medication types and intake.
[1195] Step 7:
[1196] The server checks the medication intake.
[1197] Input: Recognition data
[1198] Output: Intake confirmation data
[1199] What it does: Checks a medication database to assess whether it was taken correctly.
[1200] Step 8:
[1201] The server sends the results to the terminal.
[1202] Input: Intake confirmation data
[1203] Output: Sending data to the terminal
[1204] Behavior: Sends the verification result to the device in JSON format.
[1205] Step 9:
[1206] The device will provide audio feedback.
[1207] Input: Intake confirmation data
[1208] Output: Audio feedback
[1209] What it does: Uses the Google Text-to-Speech API to say "Your medication was taken successfully."
[1210] Schedule management
[1211] Step 1:
[1212] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 p.m."
[1213] Input: User's voice command
[1214] Output: Audio data
[1215] Action: The user speaks the command "dentist appointment."
[1216] Step 2:
[1217] The device performs voice recognition.
[1218] Input: Audio data
[1219] Output: Text data
[1220] What it does: Uses speech recognition software to convert speech into text.
[1221] Step 3:
[1222] The terminal transmits the text data to the server.
[1223] Input: Text data
[1224] Output: Send data to the server
[1225] Action: Uploads the converted text data to the server.
[1226] Step 4:
[1227] The server parses the text.
[1228] Input: Text data
[1229] Output: Analysis data
[1230] How it works: Uses NLP algorithms to extract the event date, time, and content.
[1231] Step 5:
[1232] The server records the event in the calendar.
[1233] Input: Analysis data
[1234] Output: Schedule record data
[1235] Action: Adds new appointment information to the database.
[1236] Step 6:
[1237] The server generates the reminder information.
[1238] Input: Schedule record data
[1239] Output: Remind data
[1240] What it does: Generates a reminder 30 minutes before the scheduled time.
[1241] Step 7:
[1242] The device will remind you by voice.
[1243] Input: Remind data
[1244] Output: Voice reminder
[1245] What it does: Uses the Google Text-to-Speech API to notify you that "It's almost time for your 3pm dentist appointment."
[1246] Emotion Recognition and Feedback
[1247] Step 1:
[1248] The user expresses their feelings.
[1249] Input: Voice and facial expression data
[1250] Output: Emotion data
[1251] How it works: The camera and microphone capture the user's words and facial expressions that express their emotions.
[1252] Step 2:
[1253] The device captures audio and video simultaneously.
[1254] Input: Voice and facial expression data
[1255] Output: Capture data
[1256] Action: Activates microphone and camera to collect data.
[1257] Step 3:
[1258] The device sends the data to the server.
[1259] Input: Capture data
[1260] Output: Send data to the server
[1261] What it does: Encodes audio and video data and sends it to the server.
[1262] Step 4:
[1263] The server analyzes the data.
[1264] Input: Capture data
[1265] Output: Sentiment analysis data
[1266] What it does: Analyzes data using an emotion engine.
[1267] Step 5:
[1268] The server evaluates the emotion.
[1269] Input: Sentiment analysis data
[1270] Output: Emotion rating data
[1271] What it does: Classifies the results of the emotion engine and assigns emotion labels.
[1272] Step 6:
[1273] The server generates the feedback.
[1274] Input: Emotion rating data
[1275] Output: Feedback data
[1276] What it does: Automatically generate stress relief and relaxation suggestions.
[1277] Step 7:
[1278] The device will provide audio feedback.
[1279] Input: Feedback data
[1280] Output: Audio feedback
[1281] What it does: Uses the Google Text-to-Speech API to suggest, "Would you like to listen to some music for relaxation?"
[1282] (Application example 2)
[1283] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1284] Nutritional management, emotional support, and reminder functions are important for elderly people to live their daily lives with peace of mind. However, there are limited means to provide these functions in a single integrated system, and it has been common for individual services or devices to be used separately. This has resulted in problems of poor usability and a heavy burden on users. Furthermore, there are few systems that provide more appropriate support by linking dietary management and emotional support, and there are no systems that link with food delivery services to support the entire process from meal suggestions to ordering. The objective of this invention is to solve these problems.
[1285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing video data of meals and calculating calories and nutrients, means for recognizing emotional states and suggesting relaxation, and means for suggesting and ordering meals in cooperation with food delivery services. This makes it possible to consistently perform everything from dietary management for elderly people to emotional support, and even meal suggestions and ordering, all in one system.
[1286] A "terminal" is a device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[1287] A "server" is a device that receives and analyzes audio and video data sent from a terminal.
[1288] "Analysis results" refers to information obtained by the server when it analyzes audio and video data.
[1289] "Notification" refers to the transmission of information to the user based on the analysis results.
[1290] "Nutritional information of meals" is data related to the calories and nutrients of meal contents.
[1291] "Feedback" refers to advice and suggestions to the user based on the server's analysis results.
[1292] "Emotional state" refers to the mental and emotional state of the user.
[1293] "Relaxation suggestions" are advice or suggestions regarding stress reduction and relaxation based on emotional state.
[1294] A "food delivery service" is a service that delivers meals.
[1295] "Meal suggestions" are suggestions for meal menus based on the user's health condition and required nutrients.
[1296] "Ordering" refers to the act of a user purchasing a suggested meal through a food delivery service.
[1297] This invention is a multi-functional system for supporting the lives of elderly people. Each element and its function will be explained below.
[1298] System configuration
[1299] The device is equipped with a microphone, camera, speaker, location information acquisition means, and sensor devices. The server has the ability to analyze audio and video data and notify the user based on the analysis results. It also has the ability to calculate nutritional information for meals and provide feedback, recognize emotional states and suggest relaxation activities, and connect with food delivery services to suggest and order meals.
[1300] Hardware and Software Configuration
[1301] Hardware used:
[1302] Smartphone (camera, microphone)
[1303] Server (high performance computer)
[1304] Software used:
[1305] Python
[1306] OpenCV (image processing library)
[1307] SpeechRecognition (speech recognition library)
[1308] Flask (server-side framework)
[1309] Emotion engine (AI model for emotion recognition)
[1310] System Operation
[1311] 1. Speech Recognition:
[1312] - When the user says "I'm going to start eating" to the device, the device's microphone captures the voice and converts it into a string using the SpeechRecognition library.
[1313] 2. Image capture:
[1314] - Once voice recognition is complete, the device's camera will activate and take a picture of the food, which will then be immediately sent to the server.
[1315] 3. Video analysis:
[1316] - The server analyzes the received image data using OpenCV and calculates the calories and nutrients of the meal. The calculation results are stored on the server.
[1317] 4. Emotion recognition:
[1318] - The emotion engine analyzes the user's facial expressions and tone of voice to understand their emotional state. The results are stored on the server as analytical data.
[1319] 5. Feedback and relaxation suggestions:
[1320] - Based on the analysis results, the server notifies the device of the nutrients needed for the next meal. If the emotional state is biased towards stress, it will also suggest music or activities for relaxation.
[1321] 6. Meal Suggestions and Ordering:
[1322] - The server connects the user's nutritional data and the food delivery service to suggest the next meal menu, which can then be ordered directly through the delivery service.
[1323] Natural language description example
[1324] When an elderly person starts eating, this system recognizes the voice command "start eating" and automatically activates the camera to film the meal. The video data is sent to a server, where an AI model analyzes the calories and nutrients. Next, an emotion engine analyzes the user's emotional state and provides appropriate feedback and relaxation suggestions. It also suggests meal menus based on the necessary nutrients, which can then be ordered directly through a food delivery service.
[1325] As a concrete example:
[1326] For example, when a user says "I'm going to start eating" before lunch, the smartphone camera captures an image of the meal. The image is sent to a server, where the calories and nutrients are analyzed. As a result, feedback such as "Add some fruit containing vitamin C to your next meal" is provided. Also, if the user is feeling stressed, suggestions such as "How about listening to some music to relax?" are made.
[1327] Prompt Sentence Examples
[1328] When you say "start eating," the camera automatically activates and takes a picture of your meal. The image data is sent to the server, where it is analyzed by an AI engine, which calculates calories and nutrients. Nutritional advice for your next meal is then provided via voice.
[1329] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1330] Step 1:
[1331] Voice Recognition:
[1332] When a user says "I'm going to start eating" to the device, the device's microphone captures the voice. The captured voice data is converted into text using Python's SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device's microphone records the voice, and the SpeechRecognition library analyzes the voice file and converts it into text format.
[1333] Step 2:
[1334] Image Capture:
[1335] Once voice recognition is complete and the text "Start eating" is confirmed, the device's camera will automatically start up and take an image of the meal. The input is video data from the camera, and the output is the captured image file. Specifically, the camera starts up, frames the user's meal, takes a photo, and saves it as an image file.
[1336] Step 3:
[1337] Video Analysis:
[1338] The image captured by the device is sent to the server. The server analyzes the received image data using the OpenCV library and calculates the type of food, calories, and nutrients. The input is image data, and the output is calorie and nutrient data. Specifically, the server runs an image analysis algorithm to recognize ingredients and output their nutritional information in a table format.
[1339] Step 4:
[1340] Emotion recognition:
[1341] The user's facial expression and tone of voice are also sent to the server and analyzed using the emotion engine. The input is facial image and voice data, and the output is emotional state data. Specifically, the server runs an emotion recognition model to determine the user's emotional state, such as stress, happiness, or fatigue.
[1342] Step 5:
[1343] Feedback and relaxation suggestions:
[1344] Based on the analysis results, the server provides nutritional advice and relaxation suggestions appropriate to the user. The input is calorie and nutrient data and emotional state data, and the output is a feedback message. Specifically, the server generates a text message with recommended meal contents and relaxation methods and sends it to the device.
[1345] Step 6:
[1346] Meal Suggestions and Ordering:
[1347] The server executes a function that links suggested meal menus to food delivery services based on the user's nutritional data. The input is nutritional balance data, and the output is meal suggestions and a confirmation message for the order. Specifically, the server generates an appropriate menu and executes the linking process to allow the user to easily order through the food delivery service.
[1348] Through these steps, the system will be able to smoothly carry out a series of steps from dietary management for the elderly to emotional support, and meal suggestions and ordering.
[1349] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1350] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1351] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1352] [Third embodiment]
[1353] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1354] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1355] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1356] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1357] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1358] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1359] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1360] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1361] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1362] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1363] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1364] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1365] This invention is a system that includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device to support the daily lives of elderly people, and a server that receives and analyzes data from these devices. Specific embodiments of the system program and its processing are described below.
[1366] Program processing
[1367] Dietary management
[1368] When a user wants to eat, they say "I'm going to start eating" to the camera. The device then activates the camera and captures the video of the meal. The captured video is then sent from the device to the server.
[1369] The server analyzes the received video to recognize the meal contents and retrieves the calories and nutrients obtained from the database. The analysis results are then sent from the server to the device, which then verbally informs the user of the next action to be taken and an evaluation of the meal.
[1370] For example, suppose a user is eating salad and chicken for lunch. The device captures the video, which is then received and analyzed by the server. As a result of the analysis, the calories and nutrients of the salad and chicken are calculated, and it is determined that they are lacking in vitamin C. Based on this information, the device provides feedback to the user, such as "Add some fruit containing vitamin C to your next meal."
[1371] Medication Management
[1372] When it's time for the user to take their medicine, they say "I'll take my medicine" to the device. In response, the device activates its camera, captures the video, and sends it to the server. The server analyzes the video data to confirm the type of medicine and the intake status. It checks whether the medicine has been taken properly and generates an alert if there is a problem.
[1373] For example, if a user is about to take their morning medicine, the device will record the action and send it to the server, which will then use the video to check whether the medicine was taken properly. If the user forgets to take their morning medicine, the device will notify the user, asking, "Did you forget to take your morning medicine?"
[1374] Schedule management
[1375] The user speaks to the device, "I have a dentist appointment tomorrow at 3 PM." The device converts this speech into text and sends it to the server. The server analyzes the text data to identify the appointment date and time and records it on the calendar.
[1376] Thirty minutes before the scheduled time, the server generates a reminder and sends it to the device, which then issues a voice reminder saying, "It's almost time for your 3:00 PM dentist appointment."
[1377] Specific examples
[1378] 1. Dietary Management:
[1379] A user eats oatmeal and a banana for breakfast.
[1380] The device captures the video and sends it to the server.
[1381] The server analyzed the calories and nutrients in the oatmeal and banana and determined that they were lacking in dietary fiber.
[1382] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[1383] 2. Medication Management:
[1384] The user takes regular medication before going to bed.
[1385] The device takes a picture of the situation and sends it to a server to confirm the intake status.
[1386] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[1387] No problem: Under normal circumstances, the device will notify you that it was successful.
[1388] 3. Schedule Management:
[1389] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[1390] The device sends this information to the server and records it in the calendar.
[1391] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[1392] In this way, the entire system can provide multifaceted support for users' daily lives and ensure the safety and health of the elderly.
[1393] The processing flow will be explained below.
[1394] Dietary management
[1395] Step 1:
[1396] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[1397] Step 2:
[1398] The device recognizes the voice and activates the camera to capture footage of the meal.
[1399] Step 3:
[1400] The device sends the captured video to the server.
[1401] Step 4:
[1402] The server analyzes the received video data.
[1403] The server uses computer vision technology to identify the food.
[1404] The server retrieves calorie and nutrient information about the food from a database.
[1405] Step 5:
[1406] The server evaluates the dietary content based on the analysis results, identifies missing nutrients, and generates recommended menus.
[1407] Step 6:
[1408] The server sends the analysis results and suggestions to the terminal.
[1409] Step 7:
[1410] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[1411] Medication Management
[1412] Step 1:
[1413] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[1414] Step 2:
[1415] The device will recognize the voice and activate the camera.
[1416] Step 3:
[1417] The device captures a video of the medicine and sends the data to a server.
[1418] Step 4:
[1419] The server receives the video data and analyzes it.
[1420] The server identifies the type of drug and the intake situation from the video.
[1421] Step 5:
[1422] The server checks whether proper intake has been achieved and generates an alert if there is a shortage or abnormality.
[1423] Step 6:
[1424] The server sends the alert information to the terminal.
[1425] Step 7:
[1426] The device will notify the user by voice, "Did you forget your morning medicine?"
[1427] Schedule management
[1428] Step 1:
[1429] The user says to the device, "I have a dentist appointment tomorrow at 3:00 PM."
[1430] Step 2:
[1431] The device converts the speech into text and sends the data to the server.
[1432] Step 3:
[1433] The server receives the text data and analyzes it.
[1434] The server uses natural language processing technology to identify the scheduled date, time, and content.
[1435] Step 4:
[1436] The server records the event on the calendar based on the analysis results.
[1437] Step 5:
[1438] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[1439] Step 6:
[1440] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[1441] summary
[1442] Through these detailed processing steps, the "Overly Nosy Dog" system provides multifaceted support for the lives of the elderly, allowing users to live with peace of mind.
[1443] Example 1
[1444] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1445] There is a need for systems that support the lives of elderly people and efficiently manage their health and daily schedules. However, existing systems often lack the precision of analyzing voice input and camera footage, and provide inaccurate feedback to users. Furthermore, important tasks such as medication intake and dietary management are often not automated. This can make it difficult for elderly people to properly take medication and manage their nutrition, potentially increasing health risks. Furthermore, schedule management must be done manually, which can lead to forgetfulness.
[1446] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1447] In this invention, the server includes a terminal equipped with a microphone, a camera, a speaker, location information acquisition means, and a sensor device, an information processing device that receives and analyzes audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the information processing device, means for the terminal to photograph meal contents, the information processing device to analyze the images to calculate calories and nutrients, and the terminal to provide feedback based on the analysis results, means for the terminal to photograph medication intake status, the information processing device to analyze the images to check medication intake, and if there is an abnormality, generate an alert and notify the user, means for the terminal to convert audio to text, the information processing device to analyze the text data to identify an appointment and record it on a calendar, and means for the information processing device to generate a reminder a certain time before the appointment time and notify the user, thereby improving the quality of life of elderly people, reducing health risks, and enabling efficient management of daily life.
[1448] A "microphone" is a device that converts sound into an electrical signal.
[1449] A "camera" is a device that converts light into an electrical signal and captures images.
[1450] A "speaker" is a device that reproduces electrical signals as sound.
[1451] A "location information acquisition means" is a device that has the function of determining the current location of a device using GPS or other location information technology.
[1452] A "sensor device" is a device that detects changes in the physical environment and outputs them as an electrical signal.
[1453] A "terminal" is a device that transmits and receives information between a user and a system, and is a multifunction device that includes a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[1454] An "information processing device" is a computer device that receives and analyzes audio and video data transmitted from a terminal.
[1455] The "analysis results" are data obtained from the audio and video data analyzed by the information processing device, and are information that is fed back to the user.
[1456] "Feedback" refers to advice, prompts for action, and notifications provided to users based on the analysis results.
[1457] A "calorie" is a unit that indicates the amount of energy contained in food.
[1458] "Nutrients" are components contained in food and are substances necessary for the growth and maintenance of health of the body.
[1459] "Medicine intake status" refers to whether the user is taking prescribed medication appropriately.
[1460] "Abnormal" refers to a state that is not an expected normal state, and includes, for example, failure to take medication or not following schedules.
[1461] An "alert" is a notification that notifies the user when an abnormality or a condition requiring attention occurs.
[1462] "Speech-to-text" is the process of converting spoken words into text using speech recognition technology.
[1463] A "calendar" is a schedule management tool that records date information for managing appointments and important events.
[1464] "Reminder" is a function that notifies and reminds the user of pre-set schedules and tasks.
[1465] This invention is a system designed to support the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and an information processing device that receives and analyzes audio and video data transmitted from these devices. This system notifies the user based on the results of analyzing various types of data.
[1466] Hardware and Software Configuration
[1467] 1. Device:
[1468] Microphone: Used to collect audio data.
[1469] Camera: Used to collect video data. Example: Raspberry Pi camera module.
[1470] Speaker: Used to provide audio notifications to the user.
[1471] Location information acquisition means: Includes GPS modules, etc.
[1472] Sensor devices: Used to collect information about the user's surrounding environment.
[1473] 2. Information processing equipment:
[1474] Server: A central control device that receives and analyzes audio and video data sent from terminals.
[1475] Database: A database such as MySQL is used to store meal and medication information.
[1476] Image processing software: Analyzes received video data using OpenCV, TensorFlow, etc.
[1477] Speech recognition and speech synthesis software, such as Google Speech-to-Text API and Google Text-to-Speech API, to analyze and generate voice data.
[1478] Program processing
[1479] Dietary management
[1480] When a user starts eating, they say "I'm going to start eating" to the device. The device detects this voice command, activates the camera to capture video of the meal, and sends it to the server. The server analyzes the received video and identifies the meal contents. Based on the identified meal contents, it retrieves calorie and nutrient information from the database and sends the analysis results to the device. The device then notifies the user of the results by voice, giving advice such as "Add a fruit containing vitamin C to your next meal."
[1481] Medication Management
[1482] When the user wants to take their medicine, they say, "I'm taking my medicine." The device activates the camera and sends a video of the medicine being taken to the server. The server analyzes the video data and checks the type of medicine and the state of ingestion. If there is a problem, an alert is generated and the device notifies the user, "Did you forget to take your morning medicine?"
[1483] Schedule management
[1484] When a user says, "I have a dentist appointment tomorrow at 3 p.m.", the device converts the speech to text and sends it to the server. The server analyzes the text data to identify the appointment date and time and records it on the calendar. 30 minutes before the scheduled time, the server generates a reminder and the device notifies the user, "I have a doctor's appointment at 1 p.m."
[1485] Specific examples
[1486] 1. Dietary Management:
[1487] The user says they eat oatmeal and a banana for breakfast.
[1488] The device captures the video and sends it to the server.
[1489] The server analyzes the calories and nutrients in oatmeal and bananas and determines that they are lacking in dietary fiber.
[1490] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[1491] 2. Medication Management:
[1492] The user says he takes his regular medication before going to bed.
[1493] The device takes a photo of the situation and sends it to a server to confirm the intake status.
[1494] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[1495] If there are no problems, the device will notify you with a "Good job!"
[1496] 3. Schedule Management:
[1497] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[1498] The terminal sends this information to the server and updates the schedule.
[1499] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[1500] This system will support the daily lives of the elderly and efficiently manage their health and schedules.
[1501] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1502] Dietary management
[1503] Step 1:
[1504] The user says to the device, "I'm going to start eating." The input voice data is collected by the device's microphone.
[1505] Step 2:
[1506] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[1507] Step 3:
[1508] The device recognizes the voice command and sends a signal to activate the camera module. The camera initializes and captures footage of the user's meal.
[1509] Step 4:
[1510] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[1511] Step 5:
[1512] The server analyzes the received video data using image processing software (e.g., OpenCV and TensorFlow). As a result of the analysis, the identified meal contents, calories, and nutrient information are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[1513] Step 6:
[1514] The server sends the analysis results to the device, which encodes them in JSON format and decodes them to use the data.
[1515] Step 7:
[1516] The analysis results received by the device are converted into voice using speech synthesis software (e.g., Google Text-to-Speech API), and feedback is given to the user, such as "Add some fruit containing vitamin C to your next meal." This is the final output.
[1517] Medication Management
[1518] Step 1:
[1519] The user says to the device, "I'm going to take my medicine." The input voice data is collected by the device's microphone.
[1520] Step 2:
[1521] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[1522] Step 3:
[1523] The device recognizes the voice command and sends a signal to activate the camera module. The camera is initialized and captures video of the user taking the medication.
[1524] Step 4:
[1525] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[1526] Step 5:
[1527] The server analyzes the received video data using image processing software (e.g., OpenCV and deep learning models). As a result of the analysis, the type of medication identified and the intake status are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[1528] Step 6:
[1529] The server determines whether the medication was taken properly and generates a "no problem" message if there is no problem, or an alert if there is a problem. This message becomes the input for the next step.
[1530] Step 7:
[1531] The device receives the message from the server and uses speech synthesis software (e.g., Google Text-to-Speech API) to notify the user. For example, it might say, "Did you forget your morning medicine?" This is the final output.
[1532] Schedule management
[1533] Step 1:
[1534] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 PM." The input voice data is collected by the terminal's microphone.
[1535] Step 2:
[1536] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[1537] Step 3:
[1538] The terminal sends the converted text data to the server, which becomes the input for the next step.
[1539] Step 4:
[1540] The server analyzes the received text data using a natural language processing engine (e.g., NLTK) to identify the reservation date and time. The identified reservation date and time becomes the input for the next step.
[1541] Step 5:
[1542] The server records the identified scheduled date and time in a database (e.g. MySQL) and updates the schedule.
[1543] Step 6:
[1544] 30 minutes before the scheduled time, the server generates a reminder and sends it to the device in JSON format. This reminder serves as input for the next step.
[1545] Step 7:
[1546] The device receives a reminder and notifies the user using speech synthesis software (e.g., Google Text-to-Speech API). For example, it might say, "You have a doctor's appointment at 1 p.m." This is the final output.
[1547] (Application example 1)
[1548] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1549] There is a need for systems that enable users, such as the elderly, to smoothly shop in brick-and-mortar stores and support health management and planned purchases. In particular, there is a lack of systems that integrate product information gathering, in-store navigation, and reminder functions. Therefore, the challenge is to provide a system that improves the safety and convenience of elderly people living independently.
[1550] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1551] In this invention, the server includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, means for receiving and analyzing audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the server, means for allowing the user to specify by voice where they want to go in the store and for recognizing the voice to provide navigation, means for the navigation to scan products with a camera and provide product information by voice, and means for the navigation to remind the user of products to be purchased in the store and when to next visit. This improves the convenience and safety of shopping in stores for elderly people and enables health management and planned purchases.
[1552] A "microphone" is a device that converts sound into an electrical signal and acquires sound data.
[1553] A "camera" is a device that captures images and videos and acquires them as digital data.
[1554] A "speaker" is a device that converts electrical signals into sound to notify or guide the user.
[1555] "Location information acquisition means" refers to devices or technologies for acquiring the current location of a terminal using GPS, Wi-Fi, etc.
[1556] A "sensor device" is a device that detects the environment and the user's situation and collects it as data.
[1557] A "terminal" is a device equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and used for direct interaction with a user.
[1558] A "server" is a computer system that receives and analyzes audio data, video data, location data, and the like sent from a terminal.
[1559] "Navigation" is a function that allows the user to specify the location or product they want to go to, and then recognizes the voice and provides guidance within the store.
[1560] "Product information" refers to detailed data such as the product name, price, ingredients, and nutrients.
[1561] "Remind" is a function that notifies users of schedules and important matters so that they do not forget.
[1562] This invention is a system that enables users, such as elderly people, to smoothly shop in brick-and-mortar stores and supports health management and planned purchases. The details of the system are described below.
[1563] System configuration
[1564] The system consists of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data sent from these terminals. The terminals interact directly with users, while the server provides feedback based on the analysis results.
[1565] Hardware and software used
[1566] Hardware: Smartphones, smart glasses, head-mounted displays
[1567] software:
[1568] Speech recognition: sr(SpeechRecognition instance), pyttsx3
[1569] Geolocation: geopy library
[1570] Image analysis: OpenCV
[1571] Program processing
[1572] The server receives the audio data, video data, and location data sent from the terminal and analyzes it as follows.
[1573] 1. Speech Recognition:
[1574] The device collects audio as the user speaks into the microphone, and the audio data is converted to text using a SpeechRecognition instance.
[1575] 2. Navigation provided:
[1576] When a user speaks the name of a place or product they want to go to, that information is sent as text data to the server, which then compares it with store map data and generates navigation instructions. The instructions are then provided to the user as voice feedback using pyttsx3.
[1577] 3.Product information provided:
[1578] When a user scans a product with a camera, the image data is sent to the server and analyzed using OpenCV. The analysis results, including detailed product information and nutritional information, are sent from the server to the device and notified to the user via voice.
[1579] 4. Reminder function:
[1580] The server stores the user's planned purchases and the next visit date in a database and sends reminders as appropriate. When a reminder is issued, the terminal notifies the user by voice.
[1581] Specific examples
[1582] For example, if a user says, "I want to go to the vegetable section," the system will use the camera and location information acquisition means to determine the user's current location and provide guidance to the desired vegetable section. An example of a prompt sentence in this case is, "Please tell us where you want to go. If the user says, 'Vegetable section,' the system will guide you, 'The vegetable section is in this direction.'"
[1583] Also, when a user scans an item with the camera saying, "Tell me the nutritional information of this apple," the system uses OpenCV to analyze the image of the apple and provides the nutritional information of the apple by voice. An example of a prompt sentence in this case is, "Please scan the item with the camera. When the user scans the item, the system will provide the nutritional information of that item by voice."
[1584] In this way, the entire system can provide multifaceted support for elderly people's shopping in brick-and-mortar stores, improving the quality of their daily lives.
[1585] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1586] Step 1:
[1587] The user speaks voice commands into the microphone regarding the places and products they want to go to.
[1588] Input: Audio data
[1589] Output: Audio file
[1590] Specific behavior:
[1591] The device collects audio data through a microphone and creates an audio file.
[1592] Step 2:
[1593] The voice data collected by the device is converted into text data using a SpeechRecognition instance.
[1594] Input: Audio file
[1595] Output: Text data
[1596] Specific behavior:
[1597] The audio file is analyzed and the audio data is converted into text, thereby obtaining the user's instructions as text.
[1598] Step 3:
[1599] The terminal transmits the text data to the server.
[1600] Input: Text data
[1601] Output: Sending text data
[1602] Specific behavior:
[1603] The terminal transmits the text data to the server via the network.
[1604] Step 4:
[1605] The server analyzes the text data and recognizes the user's instructions.
[1606] Input: Text data
[1607] Output: Analysis results
[1608] Specific behavior:
[1609] The server analyzes the text data to understand the location the user wants to go to and the products they are looking for, and then determines the next action based on the results of this analysis.
[1610] Step 5:
[1611] The server generates navigation instructions and sends them to the terminal.
[1612] Input: Analysis results
[1613] Output: Navigation instructions
[1614] Specific behavior:
[1615] The server then uses the analysis results to reference the store's map data and generates navigation instructions to the user's desired destination. These instructions are sent to the device in text format.
[1616] Step 6:
[1617] The terminal notifies the user of the navigation instructions received from the server using voice synthesis (pyttsx3).
[1618] Input: Navigation instructions
[1619] Output: Audio feedback
[1620] Specific behavior:
[1621] Based on the navigation instructions, the device uses voice synthesis technology to notify the user of the instructions aloud, for example, "The vegetable section is in this direction."
[1622] Step 7:
[1623] The user scans the item with the camera.
[1624] Input: Video data
[1625] Output: Video file
[1626] Specific behavior:
[1627] A camera is used to take a picture of a product designated by a user, and a video file is created.
[1628] Step 8:
[1629] The video data captured by the device is sent to the server.
[1630] Input: Video file
[1631] Output: Sending video data
[1632] Specific behavior:
[1633] The terminal transmits the video file to the server via the network.
[1634] Step 9:
[1635] The server uses OpenCV to analyze the video data and obtain detailed product information and nutritional information.
[1636] Input: Video data
[1637] Output: Product information
[1638] Specific behavior:
[1639] The server uses OpenCV to analyze the video data, identify the product, and retrieve detailed information from a database, such as the nutritional information of an apple.
[1640] Step 10:
[1641] The server sends the acquired product information to the terminal, and the terminal notifies the user using voice synthesis (pyttsx3).
[1642] Input: Product information
[1643] Output: Audio feedback
[1644] Specific behavior:
[1645] The device uses voice synthesis technology to notify the user of product information, such as "The vitamin C content of this apple is..."
[1646] Step 11:
[1647] The server stores the user's planned purchases and the next visit date in a database and sends reminders.
[1648] Input: Items to be purchased, time of visit
[1649] Output: Reminder notification
[1650] Specific behavior:
[1651] The server stores the user's planned purchases and the next visit date in a database, and generates appropriate reminders to notify the device. The device then issues a voice message, such as "The next time you visit is on this date."
[1652] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1653] This system provides multifaceted support for the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data. Furthermore, by combining it with an emotion engine, it has the function of recognizing the user's emotions and providing appropriate feedback and support based on that information.
[1654] Program processing
[1655] Dietary management
[1656] When a user begins eating, they say "I'm going to start eating" to the camera. The device recognizes the voice, activates the camera, captures video of the meal, and sends the data to the server. The server analyzes the video and calculates the calories and nutrients of the food. Based on the analysis results, the device gives the user audio feedback such as "Please add a fruit containing vitamin C to your next meal."
[1657] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[1658] Medication Management
[1659] When a user is about to take medicine, they say "I'm going to take my medicine" to the device. The device activates the camera, captures a video of the medicine, and sends it to the server. The server analyzes the video and checks the type of medicine and the intake status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user.
[1660] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[1661] Schedule management
[1662] When a user enters an appointment, they say to the device, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends it to the server. The server analyzes the text data and records the appointment on the calendar. 30 minutes before the scheduled time, the server generates reminder information, and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM."
[1663] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[1664] Emotion Recognition and Feedback
[1665] When a user expresses emotions in various everyday situations, the device captures their voice and facial expressions using a microphone and camera. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the device provides the user with suggestions for stress relief and relaxation.
[1666] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[1667] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[1668] The processing flow will be explained below.
[1669] Dietary management
[1670] Step 1:
[1671] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[1672] Step 2:
[1673] The device recognizes the voice and activates the camera to capture footage of the meal.
[1674] Step 3:
[1675] The device transmits the captured video data to the server.
[1676] Step 4:
[1677] The server analyzes the received video data.
[1678] The server uses computer vision technology to identify the food.
[1679] The server retrieves the calorie and nutrient information of the food from a database.
[1680] Step 5:
[1681] The server evaluates the diet based on the analysis results, identifies any nutrient deficiencies, and generates recommended menus.
[1682] Step 6:
[1683] The server sends the analysis results and suggestions to the terminal.
[1684] Step 7:
[1685] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[1686] Medication Management
[1687] Step 1:
[1688] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[1689] Step 2:
[1690] The device will recognize the voice and activate the camera.
[1691] Step 3:
[1692] The device captures a video of the medicine and sends the data to a server.
[1693] Step 4:
[1694] The server receives the video data and analyzes it.
[1695] The server identifies the type of drug and the intake situation from the video.
[1696] Step 5:
[1697] The server checks whether the medication is being taken properly and generates an alert if there is a shortage or abnormality.
[1698] Step 6:
[1699] The server sends the alert information to the terminal.
[1700] Step 7:
[1701] The device will notify the user by voice, "Did you forget your morning medicine?"
[1702] Schedule management
[1703] Step 1:
[1704] The user says to the device, "I have a dentist appointment tomorrow at 3:00 PM."
[1705] Step 2:
[1706] The device converts the speech into text and sends the data to the server.
[1707] Step 3:
[1708] The server receives the text data and analyzes it.
[1709] The server uses natural language processing technology to identify the scheduled date, time, and content.
[1710] Step 4:
[1711] The server records the event on the calendar based on the analysis results.
[1712] Step 5:
[1713] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[1714] Step 6:
[1715] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[1716] Emotion Recognition and Feedback
[1717] Step 1:
[1718] When a user expresses emotions in daily life, the device uses a microphone and a camera to capture their voice and facial expressions.
[1719] Step 2:
[1720] The terminal transmits the captured audio and video data to the server.
[1721] Step 3:
[1722] The server analyzes the received data and identifies the user's emotions using an emotion engine.
[1723] The server analyzes the emotional state from voice and facial expression data.
[1724] Step 4:
[1725] The server generates appropriate feedback and suggestions based on the sentiment analysis results.
[1726] Step 5:
[1727] The server sends the feedback data to the terminal.
[1728] Step 6:
[1729] The device will provide the user with appropriate voice feedback, such as "How about listening to some music for relaxation?"
[1730] Step 7:
[1731] The server accumulates emotional data and analyzes emotional trends over time.
[1732] Step 8:
[1733] If necessary, the server sends suggestions for counseling or mental support to the terminal, which then notifies the user.
[1734] summary
[1735] Through these detailed processing steps, the "Overly Nosy Dog" system, which combines an emotion engine, provides multifaceted support for the physical and mental health of elderly people, providing an environment in which they can live with peace of mind.
[1736] Example 2
[1737] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1738] The goal is to realize a system that provides multifaceted support for the lives of the elderly and an environment in which they can live with peace of mind. In particular, it is necessary to comprehensively support the user's health and daily life through managing food calories and nutrients, checking medication intake, and even analyzing emotions and providing feedback.
[1739] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1740] In this invention, the server includes means for analyzing video data and calculating food calories and nutrients, means for analyzing medication intake and generating an alert if an abnormality is detected, and means for analyzing the user's emotions using an emotion engine and providing appropriate feedback, thereby enabling the user to manage their diet, medication, and emotions.
[1741] A "microphone" is a device that converts sound into an electrical signal.
[1742] A "camera" is a device that captures images and stores them as digital data.
[1743] A "speaker" is a device that converts electrical signals into sound and plays it back.
[1744] The "location information acquisition means" is a device that measures the current location using a GPS or the like and acquires the data.
[1745] A "sensor device" is a device that includes a sensor for detecting the surrounding environment and the state of the user.
[1746] A "terminal" is an information processing device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[1747] A "server" is an information processing device that receives and analyzes data sent from terminals via a network.
[1748] "Analysis results" are information obtained by processing the data received by the server.
[1749] "Calories" is an indicator of the amount of energy contained in food.
[1750] Nutrients are substances contained in food that are necessary for the growth and maintenance of the body.
[1751] "Medicine intake status" is information indicating whether the user has taken the medicine correctly.
[1752] An "emotion engine" is an algorithm or software that analyzes audio and video to estimate a user's emotions.
[1753] "Feedback" refers to information or advice that the server provides to the user based on the analysis results.
[1754] An "alert" is a warning that notifies the user of an abnormality or a situation that requires attention.
[1755] This invention is a system that provides multifaceted support for the daily lives of the elderly, and is primarily composed of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that analyzes this data and provides feedback.
[1756] Dietary management
[1757] When starting a meal, the user speaks to the device, saying "I'm going to start eating." The device then recognizes the voice and activates the camera. The camera captures video of the meal, and the data is sent to the server. The server analyzes the received video data and calculates the calories and nutrients of the food. Based on the analysis results, the device provides audible feedback such as "Please add a fruit containing vitamin C to your next meal." The software used is Google Speech-to-Text API for speech recognition, OpenCV and TensorFlow for video analysis, and Google Text-to-Speech API for text-to-speech.
[1758] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[1759] Example prompt sentence:
[1760] When you say "start eating" and turn to the camera, your dietary management begins. What nutrients might be missing from your next meal?
[1761] Medication Management
[1762] When a user takes medicine, they say "take medicine" to the device, which activates the device's camera and captures the ingestion of the medicine on video. This video data is sent to a server, which analyzes the video to confirm the type of medicine and the ingestion status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user. The software used is Google Speech-to-Text API for voice recognition, and OpenCV and TensorFlow for video analysis.
[1763] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[1764] Example prompt sentence:
[1765] When you say "I'm going to take my medicine" and face the camera, the medication administration will begin. Please let us know if you have taken the medicine.
[1766] Schedule management
[1767] When a user enters an appointment, they speak to the device, saying, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends the text data to the server. The server analyzes the text and records the appointment on the calendar. 30 minutes before the scheduled time, a reminder is generated and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM." The software used is the Google Speech-to-Text API and natural language processing algorithms.
[1768] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[1769] Example prompt sentence:
[1770] Say, "I have a dentist appointment tomorrow at 3 PM," and we'll record the appointment for you. Send you a reminder 30 minutes before.
[1771] Emotion Recognition and Feedback
[1772] When a user expresses emotions in various everyday situations, the device uses a microphone and camera to capture their voice and facial expressions. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the system provides the user with suggestions for stress relief and relaxation. The software used is an emotion engine (e.g., IBM Watson Tone Analyzer).
[1773] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[1774] Example prompt sentence:
[1775] If you say, "I feel depressed today," you will receive relaxation suggestions. What methods do you recommend?
[1776] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[1777] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1778] Dietary management
[1779] Step 1:
[1780] The user says to the terminal, "I'm going to start eating."
[1781] Input: User's voice command
[1782] Output: Audio data
[1783] Action: The user says "Start eating" to the device.
[1784] Step 2:
[1785] The device performs voice recognition.
[1786] Input: Audio data
[1787] Output: Text data
[1788] What it does: Converts speech to text using speech recognition software (e.g., Google Speech-to-Text API).
[1789] Step 3:
[1790] The device will activate the camera.
[1791] Input: Text data
[1792] Output: Camera activation signal
[1793] Action: Activates the camera sensor and switches it into video capture mode.
[1794] Step 4:
[1795] The device captures video of the meal.
[1796] Input: Camera image
[1797] Output: Video data
[1798] How it works: The camera buffers your meal in real time.
[1799] Step 5:
[1800] The terminal transmits the video data to the server.
[1801] Input: Video data
[1802] Output: Upload to the server
[1803] How it works: Encoded video data is sent to the server via Wi-Fi or mobile data.
[1804] Step 6:
[1805] The server analyzes the video data.
[1806] Input: Video data
[1807] Output: Recognition data
[1808] How it works: Video analysis is performed using OpenCV and TensorFlow to identify the type and quantity of food.
[1809] Step 7:
[1810] The server calculates calories and nutrients.
[1811] Input: Recognition data
[1812] Output: Calorie and nutrient data
[1813] How it works: Looks up the FDA's food database and adds up the calories and nutrients for each food.
[1814] Step 8:
[1815] The server sends the calculation results to the terminal.
[1816] Input: Calorie and nutrient data
[1817] Output: Sending data to the terminal
[1818] Operation: The calculation results are returned to the terminal in JSON format or similar.
[1819] Step 9:
[1820] The device will provide audio feedback.
[1821] Input: Calorie and nutrient data
[1822] Output: Audio feedback
[1823] What it does: Uses the Google Text-to-Speech API to tell you to "Include a fruit containing vitamin C in your next meal."
[1824] Medication Management
[1825] Step 1:
[1826] The user says to the terminal, "I'm going to take my medicine."
[1827] Input: User's voice command
[1828] Output: Audio data
[1829] Action: The user says "take medicine."
[1830] Step 2:
[1831] The device performs voice recognition.
[1832] Input: Audio data
[1833] Output: Text data
[1834] What it does: Uses speech recognition software to convert speech to text.
[1835] Step 3:
[1836] The device will activate the camera.
[1837] Input: Text data
[1838] Output: Camera activation signal
[1839] Action: Activates the camera.
[1840] Step 4:
[1841] The device captures the drug intake as a video.
[1842] Input: Camera image
[1843] Output: Video data
[1844] How it works: The camera collects video data and stores it in a buffer.
[1845] Step 5:
[1846] The terminal transmits the video data to the server.
[1847] Input: Video data
[1848] Output: Upload to server
[1849] Operation: Encodes video data and sends it to the server.
[1850] Step 6:
[1851] The server analyzes the video data.
[1852] Input: Video data
[1853] Output: Recognition data
[1854] How it works: Uses video analysis algorithms to identify medication types and intake.
[1855] Step 7:
[1856] The server checks the medication intake.
[1857] Input: Recognition data
[1858] Output: Intake confirmation data
[1859] What it does: Checks a medication database to assess whether it was taken correctly.
[1860] Step 8:
[1861] The server sends the results to the terminal.
[1862] Input: Intake confirmation data
[1863] Output: Sending data to the terminal
[1864] Behavior: Sends the verification result to the device in JSON format.
[1865] Step 9:
[1866] The device will provide audio feedback.
[1867] Input: Intake confirmation data
[1868] Output: Audio feedback
[1869] What it does: Uses the Google Text-to-Speech API to say "Your medication was taken successfully."
[1870] Schedule management
[1871] Step 1:
[1872] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 p.m."
[1873] Input: User's voice command
[1874] Output: Audio data
[1875] Action: The user speaks the command "dentist appointment."
[1876] Step 2:
[1877] The device performs voice recognition.
[1878] Input: Audio data
[1879] Output: Text data
[1880] What it does: Uses speech recognition software to convert speech into text.
[1881] Step 3:
[1882] The terminal transmits the text data to the server.
[1883] Input: Text data
[1884] Output: Send data to the server
[1885] Action: Uploads the converted text data to the server.
[1886] Step 4:
[1887] The server parses the text.
[1888] Input: Text data
[1889] Output: Analysis data
[1890] How it works: Uses NLP algorithms to extract the event date, time, and content.
[1891] Step 5:
[1892] The server records the event in the calendar.
[1893] Input: Analysis data
[1894] Output: Schedule record data
[1895] Action: Adds new appointment information to the database.
[1896] Step 6:
[1897] The server generates the reminder information.
[1898] Input: Schedule record data
[1899] Output: Remind data
[1900] What it does: Generates a reminder 30 minutes before the scheduled time.
[1901] Step 7:
[1902] The device will remind you by voice.
[1903] Input: Remind data
[1904] Output: Voice reminder
[1905] What it does: Uses the Google Text-to-Speech API to notify you that "It's almost time for your 3pm dentist appointment."
[1906] Emotion Recognition and Feedback
[1907] Step 1:
[1908] The user expresses their feelings.
[1909] Input: Voice and facial expression data
[1910] Output: Emotion data
[1911] How it works: The camera and microphone capture the user's words and facial expressions that express their emotions.
[1912] Step 2:
[1913] The device captures audio and video simultaneously.
[1914] Input: Voice and facial expression data
[1915] Output: Capture data
[1916] Action: Activates microphone and camera to collect data.
[1917] Step 3:
[1918] The device sends the data to the server.
[1919] Input: Capture data
[1920] Output: Send data to the server
[1921] What it does: Encodes audio and video data and sends it to the server.
[1922] Step 4:
[1923] The server analyzes the data.
[1924] Input: Capture data
[1925] Output: Sentiment analysis data
[1926] What it does: Analyzes data using an emotion engine.
[1927] Step 5:
[1928] The server evaluates the emotion.
[1929] Input: Sentiment analysis data
[1930] Output: Emotion rating data
[1931] What it does: Classifies the results of the emotion engine and assigns emotion labels.
[1932] Step 6:
[1933] The server generates the feedback.
[1934] Input: Emotion rating data
[1935] Output: Feedback data
[1936] What it does: Automatically generate stress relief and relaxation suggestions.
[1937] Step 7:
[1938] The device will provide audio feedback.
[1939] Input: Feedback data
[1940] Output: Audio feedback
[1941] What it does: Uses the Google Text-to-Speech API to suggest, "Would you like to listen to some music for relaxation?"
[1942] (Application example 2)
[1943] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1944] Nutritional management, emotional support, and reminder functions are important for elderly people to live their daily lives with peace of mind. However, there are limited means to provide these functions in a single integrated system, and it has been common for individual services or devices to be used separately. This has resulted in problems of poor usability and a heavy burden on users. Furthermore, there are few systems that provide more appropriate support by linking dietary management and emotional support, and there are no systems that link with food delivery services to support the entire process from meal suggestions to ordering. The objective of this invention is to solve these problems.
[1945] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing video data of meals and calculating calories and nutrients, means for recognizing emotional states and suggesting relaxation, and means for suggesting and ordering meals in cooperation with food delivery services. This makes it possible to consistently perform everything from dietary management for elderly people to emotional support, and even meal suggestions and ordering, all in one system.
[1946] A "terminal" is a device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[1947] A "server" is a device that receives and analyzes audio and video data sent from a terminal.
[1948] "Analysis results" refers to information obtained by the server when it analyzes audio and video data.
[1949] "Notification" refers to the transmission of information to the user based on the analysis results.
[1950] "Nutritional information of meals" is data related to the calories and nutrients of meal contents.
[1951] "Feedback" refers to advice and suggestions to the user based on the server's analysis results.
[1952] "Emotional state" refers to the mental and emotional state of the user.
[1953] "Relaxation suggestions" are advice or suggestions regarding stress reduction and relaxation based on emotional state.
[1954] A "food delivery service" is a service that delivers meals.
[1955] "Meal suggestions" are suggestions for meal menus based on the user's health condition and required nutrients.
[1956] "Ordering" refers to the act of a user purchasing a suggested meal through a food delivery service.
[1957] This invention is a multi-functional system for supporting the lives of elderly people. Each element and its function will be explained below.
[1958] System configuration
[1959] The device is equipped with a microphone, camera, speaker, location information acquisition means, and sensor devices. The server has the ability to analyze audio and video data and notify the user based on the analysis results. It also has the ability to calculate nutritional information for meals and provide feedback, recognize emotional states and suggest relaxation activities, and connect with food delivery services to suggest and order meals.
[1960] Hardware and Software Configuration
[1961] Hardware used:
[1962] Smartphone (camera, microphone)
[1963] Server (high performance computer)
[1964] Software used:
[1965] Python
[1966] OpenCV (image processing library)
[1967] SpeechRecognition (speech recognition library)
[1968] Flask (server-side framework)
[1969] Emotion engine (AI model for emotion recognition)
[1970] System Operation
[1971] 1. Speech Recognition:
[1972] - When the user says "I'm going to start eating" to the device, the device's microphone captures the voice and converts it into a string using the SpeechRecognition library.
[1973] 2. Image capture:
[1974] - Once voice recognition is complete, the device's camera will activate and take a picture of the food, which will then be immediately sent to the server.
[1975] 3. Video analysis:
[1976] - The server analyzes the received image data using OpenCV and calculates the calories and nutrients of the meal. The calculation results are stored on the server.
[1977] 4. Emotion recognition:
[1978] - The emotion engine analyzes the user's facial expressions and tone of voice to understand their emotional state. The results are stored on the server as analytical data.
[1979] 5. Feedback and relaxation suggestions:
[1980] - Based on the analysis results, the server notifies the device of the nutrients needed for the next meal. If the emotional state is biased towards stress, it will also suggest music or activities for relaxation.
[1981] 6. Meal Suggestions and Ordering:
[1982] - The server connects the user's nutritional data and the food delivery service to suggest the next meal menu, which can then be ordered directly through the delivery service.
[1983] Natural language description example
[1984] When an elderly person starts eating, this system recognizes the voice command "start eating" and automatically activates the camera to film the meal. The video data is sent to a server, where an AI model analyzes the calories and nutrients. Next, an emotion engine analyzes the user's emotional state and provides appropriate feedback and relaxation suggestions. It also suggests meal menus based on the necessary nutrients, which can then be ordered directly through a food delivery service.
[1985] As a concrete example:
[1986] For example, when a user says "I'm going to start eating" before lunch, the smartphone camera captures an image of the meal. The image is sent to a server, where the calories and nutrients are analyzed. As a result, feedback such as "Add some fruit containing vitamin C to your next meal" is provided. Also, if the user is feeling stressed, suggestions such as "How about listening to some music to relax?" are made.
[1987] Prompt Sentence Examples
[1988] When you say "start eating," the camera automatically activates and takes a picture of your meal. The image data is sent to the server, where it is analyzed by an AI engine, which calculates calories and nutrients. Nutritional advice for your next meal is then provided via voice.
[1989] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1990] Step 1:
[1991] Voice Recognition:
[1992] When a user says "I'm going to start eating" to the device, the device's microphone captures the voice. The captured voice data is converted into text using Python's SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device's microphone records the voice, and the SpeechRecognition library analyzes the voice file and converts it into text format.
[1993] Step 2:
[1994] Image Capture:
[1995] Once voice recognition is complete and the text "Start eating" is confirmed, the device's camera will automatically start up and take an image of the meal. The input is video data from the camera, and the output is the captured image file. Specifically, the camera starts up, frames the user's meal, takes a photo, and saves it as an image file.
[1996] Step 3:
[1997] Video Analysis:
[1998] The image captured by the device is sent to the server. The server analyzes the received image data using the OpenCV library and calculates the type of food, calories, and nutrients. The input is image data, and the output is calorie and nutrient data. Specifically, the server runs an image analysis algorithm to recognize ingredients and output their nutritional information in a table format.
[1999] Step 4:
[2000] Emotion recognition:
[2001] The user's facial expression and tone of voice are also sent to the server and analyzed using the emotion engine. The input is facial image and voice data, and the output is emotional state data. Specifically, the server runs an emotion recognition model to determine the user's emotional state, such as stress, happiness, or fatigue.
[2002] Step 5:
[2003] Feedback and relaxation suggestions:
[2004] Based on the analysis results, the server provides nutritional advice and relaxation suggestions appropriate to the user. The input is calorie and nutrient data and emotional state data, and the output is a feedback message. Specifically, the server generates a text message with recommended meal contents and relaxation methods and sends it to the device.
[2005] Step 6:
[2006] Meal Suggestions and Ordering:
[2007] The server executes a function that links suggested meal menus to food delivery services based on the user's nutritional data. The input is nutritional balance data, and the output is meal suggestions and a confirmation message for the order. Specifically, the server generates an appropriate menu and executes the linking process to allow the user to easily order through the food delivery service.
[2008] Through these steps, the system will be able to smoothly carry out a series of steps from dietary management for the elderly to emotional support, and meal suggestions and ordering.
[2009] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2010] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2011] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2012] [Fourth embodiment]
[2013] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2014] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2015] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2016] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2017] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2018] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2019] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2020] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2021] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2022] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2023] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2024] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2025] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2026] This invention is a system that includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device to support the daily lives of elderly people, and a server that receives and analyzes data from these devices. Specific embodiments of the system program and its processing are described below.
[2027] Program processing
[2028] Dietary management
[2029] When a user wants to eat, they say "I'm going to start eating" to the camera. The device then activates the camera and captures the video of the meal. The captured video is then sent from the device to the server.
[2030] The server analyzes the received video to recognize the meal contents and retrieves the calories and nutrients obtained from the database. The analysis results are then sent from the server to the device, which then verbally informs the user of the next action to be taken and an evaluation of the meal.
[2031] For example, suppose a user is eating salad and chicken for lunch. The device captures the video, which is then received and analyzed by the server. As a result of the analysis, the calories and nutrients of the salad and chicken are calculated, and it is determined that they are lacking in vitamin C. Based on this information, the device provides feedback to the user, such as "Add some fruit containing vitamin C to your next meal."
[2032] Medication Management
[2033] When it's time for the user to take their medicine, they say "I'll take my medicine" to the device. In response, the device activates its camera, captures the video, and sends it to the server. The server analyzes the video data to confirm the type of medicine and the intake status. It checks whether the medicine has been taken properly and generates an alert if there is a problem.
[2034] For example, if a user is about to take their morning medicine, the device will record the action and send it to the server, which will then use the video to check whether the medicine was taken properly. If the user forgets to take their morning medicine, the device will notify the user, asking, "Did you forget to take your morning medicine?"
[2035] Schedule management
[2036] The user speaks to the device, "I have a dentist appointment tomorrow at 3 PM." The device converts this speech into text and sends it to the server. The server analyzes the text data to identify the appointment date and time and records it on the calendar.
[2037] Thirty minutes before the scheduled time, the server generates a reminder and sends it to the device, which then issues a voice reminder saying, "It's almost time for your 3:00 PM dentist appointment."
[2038] Specific examples
[2039] 1. Dietary Management:
[2040] A user eats oatmeal and a banana for breakfast.
[2041] The device captures the video and sends it to the server.
[2042] The server analyzed the calories and nutrients in the oatmeal and banana and determined that they were lacking in dietary fiber.
[2043] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[2044] 2. Medication Management:
[2045] The user takes regular medication before going to bed.
[2046] The device takes a picture of the situation and sends it to a server to confirm the intake status.
[2047] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[2048] No problem: Under normal circumstances, the device will notify you that it was successful.
[2049] 3. Schedule Management:
[2050] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[2051] The device sends this information to the server and records it in the calendar.
[2052] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[2053] In this way, the entire system can provide multifaceted support for users' daily lives and ensure the safety and health of the elderly.
[2054] The processing flow will be explained below.
[2055] Dietary management
[2056] Step 1:
[2057] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[2058] Step 2:
[2059] The device recognizes the voice and activates the camera to capture footage of the meal.
[2060] Step 3:
[2061] The device sends the captured video to the server.
[2062] Step 4:
[2063] The server analyzes the received video data.
[2064] The server uses computer vision technology to identify the food.
[2065] The server retrieves calorie and nutrient information about the food from a database.
[2066] Step 5:
[2067] The server evaluates the dietary content based on the analysis results, identifies missing nutrients, and generates recommended menus.
[2068] Step 6:
[2069] The server sends the analysis results and suggestions to the terminal.
[2070] Step 7:
[2071] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[2072] Medication Management
[2073] Step 1:
[2074] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[2075] Step 2:
[2076] The device will recognize the voice and activate the camera.
[2077] Step 3:
[2078] The device captures a video of the medicine and sends the data to a server.
[2079] Step 4:
[2080] The server receives the video data and analyzes it.
[2081] The server identifies the type of drug and the intake situation from the video.
[2082] Step 5:
[2083] The server checks whether proper intake has been achieved and generates an alert if there is a shortage or abnormality.
[2084] Step 6:
[2085] The server sends the alert information to the terminal.
[2086] Step 7:
[2087] The device will notify the user by voice, "Did you forget your morning medicine?"
[2088] Schedule management
[2089] Step 1:
[2090] The user says to the terminal, "I have a dentist appointment tomorrow at 3:00 PM."
[2091] Step 2:
[2092] The device converts the speech into text and sends the data to the server.
[2093] Step 3:
[2094] The server receives the text data and analyzes it.
[2095] The server uses natural language processing technology to identify the scheduled date, time, and content.
[2096] Step 4:
[2097] The server records the event on the calendar based on the analysis results.
[2098] Step 5:
[2099] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[2100] Step 6:
[2101] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[2102] summary
[2103] Through these detailed processing steps, the "Overly Nosy Dog" system provides multifaceted support for the lives of the elderly, allowing users to live with peace of mind.
[2104] Example 1
[2105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2106] There is a need to support the lives of the elderly and efficiently manage their health and daily schedules. However, existing systems often lack the precision of analyzing voice input and camera footage, and provide inaccurate feedback to users. Furthermore, important tasks such as medication intake and dietary management are often not automated. This can make it difficult for elderly people to properly take their medication and manage their nutrition, potentially increasing health risks. Furthermore, schedule management must be done manually, which can lead to forgetfulness.
[2107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2108] In this invention, the server includes a terminal equipped with a microphone, a camera, a speaker, location information acquisition means, and a sensor device, an information processing device that receives and analyzes audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the information processing device, means for the terminal to photograph meal contents, the information processing device to analyze the images to calculate calories and nutrients, and the terminal to provide feedback based on the analysis results, means for the terminal to photograph medication intake status, the information processing device to analyze the images to check medication intake, and if there is an abnormality, generate an alert and notify the user, means for the terminal to convert audio to text, the information processing device to analyze the text data to identify an appointment and record it on a calendar, and means for the information processing device to generate a reminder a certain time before the appointment time and notify the user, thereby improving the quality of life of elderly people, reducing health risks, and enabling efficient management of daily life.
[2109] A "microphone" is a device that converts sound into an electrical signal.
[2110] A "camera" is a device that converts light into an electrical signal and captures images.
[2111] A "speaker" is a device that reproduces electrical signals as sound.
[2112] A "location information acquisition means" is a device that has the function of determining the current location of a device using GPS or other location information technology.
[2113] A "sensor device" is a device that detects changes in the physical environment and outputs them as an electrical signal.
[2114] A "terminal" is a device that transmits and receives information between a user and a system, and is a multifunction device that includes a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[2115] An "information processing device" is a computer device that receives and analyzes audio and video data transmitted from a terminal.
[2116] The "analysis results" are data obtained from the audio and video data analyzed by the information processing device, and are information that is fed back to the user.
[2117] "Feedback" refers to advice, prompts for action, and notifications provided to users based on the analysis results.
[2118] A "calorie" is a unit that indicates the amount of energy contained in food.
[2119] "Nutrients" are components contained in food and are substances necessary for the growth and maintenance of health of the body.
[2120] "Medicine intake status" refers to whether the user is taking prescribed medication appropriately.
[2121] "Abnormal" refers to a state that is not an expected normal state, and includes, for example, failure to take medication or not following schedules.
[2122] An "alert" is a notification that notifies the user when an abnormality or a condition requiring attention occurs.
[2123] "Speech-to-text" is the process of converting spoken words into text using speech recognition technology.
[2124] A "calendar" is a schedule management tool that records date information for managing appointments and important events.
[2125] "Reminder" is a function that notifies and reminds the user of pre-set schedules and tasks.
[2126] This invention is a system designed to support the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and an information processing device that receives and analyzes audio and video data transmitted from these devices. This system notifies the user based on the results of analyzing various types of data.
[2127] Hardware and Software Configuration
[2128] 1. Device:
[2129] Microphone: Used to collect audio data.
[2130] Camera: Used to collect video data. Example: Raspberry Pi camera module.
[2131] Speaker: Used to provide audio notifications to the user.
[2132] Location information acquisition means: Includes GPS modules, etc.
[2133] Sensor devices: Used to collect information about the user's surrounding environment.
[2134] 2. Information processing equipment:
[2135] Server: A central control device that receives and analyzes audio and video data sent from terminals.
[2136] Database: A database such as MySQL is used to store meal and medication information.
[2137] Image processing software: Analyzes received video data using OpenCV, TensorFlow, etc.
[2138] Speech recognition and speech synthesis software, such as Google Speech-to-Text API and Google Text-to-Speech API, to analyze and generate voice data.
[2139] Program processing
[2140] Dietary management
[2141] When a user starts eating, they say "I'm going to start eating" to the device. The device detects this voice command, activates the camera to capture video of the meal, and sends it to the server. The server analyzes the received video and identifies the meal contents. Based on the identified meal contents, it retrieves calorie and nutrient information from the database and sends the analysis results to the device. The device then notifies the user of the results by voice, giving advice such as "Add a fruit containing vitamin C to your next meal."
[2142] Medication Management
[2143] When the user wants to take their medicine, they say, "I'm taking my medicine." The device activates the camera and sends a video of the medicine being taken to the server. The server analyzes the video data and checks the type of medicine and the state of ingestion. If there is a problem, an alert is generated and the device notifies the user, "Did you forget to take your morning medicine?"
[2144] Schedule management
[2145] When a user says, "I have a dentist appointment tomorrow at 3 p.m.", the device converts the speech to text and sends it to the server. The server analyzes the text data to identify the appointment date and time and records it on the calendar. 30 minutes before the scheduled time, the server generates a reminder and the device notifies the user, "I have a doctor's appointment at 1 p.m."
[2146] Specific examples
[2147] 1. Dietary Management:
[2148] The user says they eat oatmeal and a banana for breakfast.
[2149] The device captures the video and sends it to the server.
[2150] The server analyzes the calories and nutrients in oatmeal and bananas and determines that they are lacking in dietary fiber.
[2151] The device will provide feedback such as, "Add more vegetables, which are high in fiber, to your lunch."
[2152] 2. Medication Management:
[2153] The user says he takes his regular medication before going to bed.
[2154] The device takes a picture of the situation and sends it to a server to confirm the intake status.
[2155] The server checks whether the medication is being taken properly and sends an alert if there is a problem.
[2156] If there are no problems, the device will notify you with a "Good job!"
[2157] 3. Schedule Management:
[2158] The user says to the terminal, "I have a doctor's appointment next Monday at 1:00 p.m."
[2159] The terminal sends this information to the server and updates the schedule.
[2160] At 12:30 p.m. on Monday, your device reminds you that you have a doctor's appointment at 1 p.m.
[2161] This system will support the daily lives of the elderly and efficiently manage their health and schedules.
[2162] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2163] Dietary management
[2164] Step 1:
[2165] The user says to the device, "I'm going to start eating." The input voice data is collected by the device's microphone.
[2166] Step 2:
[2167] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[2168] Step 3:
[2169] The device recognizes the voice command and sends a signal to activate the camera module. The camera initializes and captures footage of the user's meal.
[2170] Step 4:
[2171] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[2172] Step 5:
[2173] The server analyzes the received video data using image processing software (e.g., OpenCV and TensorFlow). As a result of the analysis, the identified meal contents, calorie, and nutrient information are extracted from a database (e.g., MySQL). The analysis results are used as input for the next step.
[2174] Step 6:
[2175] The server sends the analysis results to the device, which encodes them in JSON format and decodes them for use by the device.
[2176] Step 7:
[2177] The analysis results received by the device are converted into voice using speech synthesis software (e.g., Google Text-to-Speech API), and feedback is given to the user, such as "Add some fruit containing vitamin C to your next meal." This is the final output.
[2178] Medication Management
[2179] Step 1:
[2180] The user says to the device, "I'm going to take my medicine." The input voice data is collected by the device's microphone.
[2181] Step 2:
[2182] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[2183] Step 3:
[2184] The device identifies the voice command and sends a signal to activate the camera module. The camera is initialized and captures video of the user taking the medication.
[2185] Step 4:
[2186] The device temporarily stores the captured video data and sends it to a server via the Internet. The sent video data becomes the input for the next step.
[2187] Step 5:
[2188] The server analyzes the received video data using image processing software (e.g., OpenCV and deep learning models). As a result of the analysis, the type of medication identified and the intake status are retrieved from a database (e.g., MySQL). The analysis results are used as input for the next step.
[2189] Step 6:
[2190] The server determines whether the medication was taken properly and generates a "no problem" message if there is no problem, or an alert if there is a problem. This message becomes the input for the next step.
[2191] Step 7:
[2192] The device receives the message from the server and uses speech synthesis software (e.g., Google Text-to-Speech API) to notify the user. For example, it might say, "Did you forget your morning medicine?" This is the final output.
[2193] Schedule management
[2194] Step 1:
[2195] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 PM." The input voice data is collected by the terminal's microphone.
[2196] Step 2:
[2197] The device detects the voice command and uses speech recognition software (e.g., Google Speech-to-Text API) to convert the speech into text data, which serves as input for the next step.
[2198] Step 3:
[2199] The terminal sends the converted text data to the server, which becomes the input for the next step.
[2200] Step 4:
[2201] The server analyzes the received text data using a natural language processing engine (e.g., NLTK) to identify the reservation date and time. The identified reservation date and time becomes the input for the next step.
[2202] Step 5:
[2203] The server records the identified scheduled date and time in a database (e.g. MySQL) and updates the schedule.
[2204] Step 6:
[2205] 30 minutes before the scheduled time, the server generates a reminder and sends it to the device in JSON format. This reminder serves as input for the next step.
[2206] Step 7:
[2207] The device receives a reminder and notifies the user using speech synthesis software (e.g., Google Text-to-Speech API). For example, it might say, "You have a doctor's appointment at 1 p.m." This is the final output.
[2208] (Application example 1)
[2209] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2210] There is a need for systems that enable users, such as the elderly, to smoothly shop in brick-and-mortar stores and support health management and planned purchases. In particular, there is a lack of systems that integrate product information gathering, in-store navigation, and reminder functions. Therefore, the challenge is to provide a system that improves the safety and convenience of elderly people living independently.
[2211] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2212] In this invention, the server includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, means for receiving and analyzing audio and video data transmitted from the terminal, means for notifying the user based on the analysis results from the server, means for allowing the user to specify by voice where they want to go in the store and for recognizing the voice to provide navigation, means for the navigation to scan products with a camera and provide product information by voice, and means for the navigation to remind the user of products to be purchased in the store and when to next visit. This improves the convenience and safety of shopping in stores for elderly people and enables health management and planned purchases.
[2213] A "microphone" is a device that converts sound into an electrical signal and acquires sound data.
[2214] A "camera" is a device that captures images and videos and acquires them as digital data.
[2215] A "speaker" is a device that converts electrical signals into sound to notify or guide the user.
[2216] "Location information acquisition means" refers to devices or technologies for acquiring the current location of a terminal using GPS, Wi-Fi, etc.
[2217] A "sensor device" is a device that detects the environment and the user's situation and collects it as data.
[2218] A "terminal" is a device equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and used for direct interaction with a user.
[2219] A "server" is a computer system that receives and analyzes audio data, video data, location data, and the like sent from a terminal.
[2220] "Navigation" is a function that allows the user to specify the location or product they want to go to, and then recognizes the voice and provides guidance within the store.
[2221] "Product information" refers to detailed data such as the product name, price, ingredients, and nutrients.
[2222] "Remind" is a function that notifies users of schedules and important matters so that they do not forget.
[2223] This invention is a system that enables users, such as elderly people, to smoothly shop in brick-and-mortar stores and supports health management and planned purchases. The details of the system are described below.
[2224] System configuration
[2225] The system consists of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data sent from these terminals. The terminals interact directly with users, while the server provides feedback based on the analysis results.
[2226] Hardware and software used
[2227] Hardware: Smartphones, smart glasses, head-mounted displays
[2228] software:
[2229] Speech recognition: sr(SpeechRecognition instance), pyttsx3
[2230] Geolocation: geopy library
[2231] Image analysis: OpenCV
[2232] Program processing
[2233] The server receives the audio data, video data, and location data sent from the terminal and analyzes them as follows.
[2234] 1. Speech Recognition:
[2235] The device collects audio as the user speaks into the microphone, and the audio data is converted to text using a SpeechRecognition instance.
[2236] 2. Navigation provided:
[2237] When a user speaks the name of a place or product they want to go to, that information is sent as text data to the server, which then compares it with store map data and generates navigation instructions. The instructions are then provided to the user as voice feedback using pyttsx3.
[2238] 3.Product information provided:
[2239] When a user scans a product with a camera, the image data is sent to the server and analyzed using OpenCV. The analysis results, including detailed product information and nutritional information, are sent from the server to the device and notified to the user via voice.
[2240] 4. Reminder function:
[2241] The server stores the user's planned purchases and the next visit date in a database and sends reminders as appropriate. When a reminder is issued, the terminal notifies the user by voice.
[2242] Specific examples
[2243] For example, if a user says, "I want to go to the vegetable section," the system will use the camera and location information acquisition means to determine the user's current location and provide guidance to the desired vegetable section. An example of a prompt sentence in this case is, "Please tell us where you want to go. If the user says, 'Vegetable section,' the system will guide you, 'The vegetable section is in this direction.'"
[2244] Also, when a user scans an item with the camera saying, "Tell me the nutritional information of this apple," the system uses OpenCV to analyze the image of the apple and provides the nutritional information of the apple by voice. An example of a prompt sentence in this case is, "Please scan the item with the camera. When the user scans the item, the system will provide the nutritional information of that item by voice."
[2245] In this way, the entire system can provide multifaceted support for elderly people's shopping in brick-and-mortar stores, improving the quality of their daily lives.
[2246] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2247] Step 1:
[2248] The user speaks voice commands into the microphone regarding the places and products they want to go to.
[2249] Input: Audio data
[2250] Output: Audio file
[2251] Specific behavior:
[2252] The device collects audio data through a microphone and creates an audio file.
[2253] Step 2:
[2254] The voice data collected by the device is converted into text data using a SpeechRecognition instance.
[2255] Input: Audio file
[2256] Output: Text data
[2257] Specific behavior:
[2258] The audio file is analyzed and the audio data is converted into text, thereby obtaining the user's instructions as text.
[2259] Step 3:
[2260] The terminal transmits the text data to the server.
[2261] Input: Text data
[2262] Output: Sending text data
[2263] Specific behavior:
[2264] The terminal transmits the text data to the server via the network.
[2265] Step 4:
[2266] The server analyzes the text data and recognizes the user's instructions.
[2267] Input: Text data
[2268] Output: Analysis results
[2269] Specific behavior:
[2270] The server analyzes the text data to understand the location the user wants to go to and the products they are looking for, and then determines the next action based on the results of this analysis.
[2271] Step 5:
[2272] The server generates navigation instructions and sends them to the terminal.
[2273] Input: Analysis results
[2274] Output: Navigation instructions
[2275] Specific behavior:
[2276] The server then uses the analysis results to reference the store's map data and generates navigation instructions to the user's desired destination. These instructions are sent to the device in text format.
[2277] Step 6:
[2278] The terminal notifies the user of the navigation instructions received from the server using voice synthesis (pyttsx3).
[2279] Input: Navigation instructions
[2280] Output: Audio feedback
[2281] Specific behavior:
[2282] Based on the navigation instructions, the device uses voice synthesis technology to notify the user of the instructions aloud, for example, "The vegetable section is in this direction."
[2283] Step 7:
[2284] The user scans the item with the camera.
[2285] Input: Video data
[2286] Output: Video file
[2287] Specific behavior:
[2288] A camera is used to take a picture of a product designated by a user, and a video file is created.
[2289] Step 8:
[2290] The video data captured by the device is sent to the server.
[2291] Input: Video file
[2292] Output: Sending video data
[2293] Specific behavior:
[2294] The terminal transmits the video file to the server via the network.
[2295] Step 9:
[2296] The server uses OpenCV to analyze the video data and obtain detailed product information and nutritional information.
[2297] Input: Video data
[2298] Output: Product information
[2299] Specific behavior:
[2300] The server uses OpenCV to analyze the video data, identify the product, and retrieve detailed information from a database, such as the nutritional information of an apple.
[2301] Step 10:
[2302] The server sends the acquired product information to the terminal, and the terminal notifies the user using voice synthesis (pyttsx3).
[2303] Input: Product information
[2304] Output: Audio feedback
[2305] Specific behavior:
[2306] The device uses voice synthesis technology to notify the user of product information, such as "The vitamin C content of this apple is..."
[2307] Step 11:
[2308] The server stores the user's planned purchases and the next visit date in a database and sends reminders.
[2309] Input: Items to be purchased, time of visit
[2310] Output: Reminder notification
[2311] Specific behavior:
[2312] The server stores the user's planned purchases and the next visit date in a database, and generates appropriate reminders to notify the device. The device then issues a voice message, such as "The next time you visit is on this date."
[2313] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2314] This system provides multifaceted support for the lives of the elderly, and includes a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that receives and analyzes the data. Furthermore, by combining it with an emotion engine, it has the function of recognizing the user's emotions and providing appropriate feedback and support based on that information.
[2315] Program processing
[2316] Dietary management
[2317] When a user begins eating, they say "I'm going to start eating" to the camera. The device recognizes the voice, activates the camera, captures video of the meal, and sends the data to the server. The server analyzes the video and calculates the calories and nutrients of the food. Based on the analysis results, the device gives the user audio feedback such as "Please add a fruit containing vitamin C to your next meal."
[2318] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, the server will recognize the food, calculate the calories and nutrients, and send the results to the device, which will then suggest the nutrients needed for the next meal.
[2319] Medication Management
[2320] When a user is about to take medicine, they say "I'm going to take my medicine" to the device. The device activates the camera, captures a video of the medicine, and sends it to the server. The server analyzes the video and checks the type of medicine and the intake status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user.
[2321] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, saying, "Did you forget to take your nighttime pill?"
[2322] Schedule management
[2323] When a user enters an appointment, they say to the device, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends it to the server. The server analyzes the text data and records the appointment on the calendar. 30 minutes before the scheduled time, the server generates reminder information, and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM."
[2324] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[2325] Emotion Recognition and Feedback
[2326] When a user expresses emotions in various everyday situations, the device captures their voice and facial expressions using a microphone and camera. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the device provides the user with suggestions for stress relief and relaxation.
[2327] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[2328] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[2329] The processing flow will be explained below.
[2330] Dietary management
[2331] Step 1:
[2332] When the user starts eating, he / she says "I'm going to start eating" to the terminal.
[2333] Step 2:
[2334] The device recognizes the voice and activates the camera to capture footage of the meal.
[2335] Step 3:
[2336] The device transmits the captured video data to the server.
[2337] Step 4:
[2338] The server analyzes the received video data.
[2339] The server uses computer vision technology to identify the food.
[2340] The server retrieves the calorie and nutrient information of the food from a database.
[2341] Step 5:
[2342] The server evaluates the diet based on the analysis results, identifies any nutrient deficiencies, and generates recommended menus.
[2343] Step 6:
[2344] The server sends the analysis results and suggestions to the terminal.
[2345] Step 7:
[2346] The device will provide the user with audible feedback such as, "Add a fruit containing vitamin C to your next meal."
[2347] Medication Management
[2348] Step 1:
[2349] When it is time for the user to take their medicine, they say to the device, "I'm going to take my medicine."
[2350] Step 2:
[2351] The device will recognize the voice and activate the camera.
[2352] Step 3:
[2353] The device captures a video of the medicine and sends the data to a server.
[2354] Step 4:
[2355] The server receives the video data and analyzes it.
[2356] The server identifies the type of drug and the intake situation from the video.
[2357] Step 5:
[2358] The server checks whether the medication is being taken properly and generates an alert if there is a shortage or abnormality.
[2359] Step 6:
[2360] The server sends the alert information to the terminal.
[2361] Step 7:
[2362] The device will notify the user by voice, "Did you forget your morning medicine?"
[2363] Schedule management
[2364] Step 1:
[2365] The user says to the terminal, "I have a dentist appointment tomorrow at 3:00 PM."
[2366] Step 2:
[2367] The device converts the speech into text and sends the data to the server.
[2368] Step 3:
[2369] The server receives the text data and analyzes it.
[2370] The server uses natural language processing technology to identify the scheduled date, time, and content.
[2371] Step 4:
[2372] The server records the event on the calendar based on the analysis results.
[2373] Step 5:
[2374] The server generates a reminder 30 minutes before the scheduled time and sends the reminder data to the device.
[2375] Step 6:
[2376] The device will then audibly remind the user that "it's almost time for your 3pm dentist appointment."
[2377] Emotion Recognition and Feedback
[2378] Step 1:
[2379] When a user expresses emotions in daily life, the device uses a microphone and a camera to capture their voice and facial expressions.
[2380] Step 2:
[2381] The terminal transmits the captured audio and video data to the server.
[2382] Step 3:
[2383] The server analyzes the received data and identifies the user's emotions using an emotion engine.
[2384] The server analyzes the emotional state from voice and facial expression data.
[2385] Step 4:
[2386] The server generates appropriate feedback and suggestions based on the sentiment analysis results.
[2387] Step 5:
[2388] The server sends the feedback data to the terminal.
[2389] Step 6:
[2390] The device will provide the user with appropriate voice feedback, such as "How about listening to some music for relaxation?"
[2391] Step 7:
[2392] The server accumulates emotional data and analyzes emotional trends over time.
[2393] Step 8:
[2394] If necessary, the server sends suggestions for counseling or mental support to the terminal, which then notifies the user.
[2395] summary
[2396] Through these detailed processing steps, the "Overly Nosy Dog" system, which combines an emotion engine, provides multifaceted support for the physical and mental health of elderly people, providing an environment in which they can live with peace of mind.
[2397] Example 2
[2398] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2399] The goal is to realize a system that provides multifaceted support for the lives of the elderly and an environment in which they can live with peace of mind. In particular, it is necessary to comprehensively support the user's health and daily life through managing food calories and nutrients, checking medication intake, and even analyzing emotions and providing feedback.
[2400] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2401] In this invention, the server includes means for analyzing video data and calculating food calories and nutrients, means for analyzing medication intake and generating an alert if an abnormality is detected, and means for analyzing the user's emotions using an emotion engine and providing appropriate feedback, thereby enabling the user to manage their diet, medication, and emotions.
[2402] A "microphone" is a device that converts sound into an electrical signal.
[2403] A "camera" is a device that captures images and stores them as digital data.
[2404] A "speaker" is a device that converts electrical signals into sound and plays it back.
[2405] The "location information acquisition means" is a device that measures the current location using a GPS or the like and acquires the data.
[2406] A "sensor device" is a device that includes a sensor for detecting the surrounding environment and the state of the user.
[2407] A "terminal" is an information processing device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[2408] A "server" is an information processing device that receives and analyzes data sent from terminals via a network.
[2409] "Analysis results" are information obtained by processing the data received by the server.
[2410] "Calories" is an indicator of the amount of energy contained in food.
[2411] Nutrients are substances contained in food that are necessary for the growth and maintenance of the body.
[2412] "Medicine intake status" is information indicating whether the user has taken the medicine correctly.
[2413] An "emotion engine" is an algorithm or software that analyzes audio and video to estimate a user's emotions.
[2414] "Feedback" refers to information or advice that the server provides to the user based on the analysis results.
[2415] An "alert" is a warning that notifies the user of an abnormality or a situation that requires attention.
[2416] This invention is a system that provides multifaceted support for the daily lives of the elderly, and is primarily composed of a terminal equipped with a microphone, camera, speaker, location information acquisition means, and sensor device, and a server that analyzes this data and provides feedback.
[2417] Dietary management
[2418] When starting a meal, the user speaks to the device, saying "I'm going to start eating." The device then recognizes the voice and activates the camera. The camera captures video of the meal, and the data is sent to the server. The server analyzes the received video data and calculates the calories and nutrients of the food. Based on the analysis results, the device provides audible feedback such as "Please add a fruit containing vitamin C to your next meal." The software used is Google Speech-to-Text API for speech recognition, OpenCV and TensorFlow for video analysis, and Google Text-to-Speech API for text-to-speech.
[2419] For example, if a user eats a sandwich and fruit for lunch, the device will take a video of the meal, and the server will recognize the food and calculate the calories and nutrients. The server will then send the results to the device, which will then suggest the nutrients needed for the next meal.
[2420] Example prompt sentence:
[2421] When you say "start eating" and turn to the camera, your dietary management begins. What nutrients might be missing from your next meal?
[2422] Medication Management
[2423] When a user takes medicine, they say "take medicine" to the device, which activates the device's camera and captures the ingestion of the medicine on video. This video data is sent to a server, which analyzes the video to confirm the type of medicine and the ingestion status. It checks whether the medicine has been taken properly, and if there is an abnormality, it generates an alert and notifies the user. The software used is Google Speech-to-Text API for voice recognition, and OpenCV and TensorFlow for video analysis.
[2424] For example, if a user takes a sleeping pill at night, the device will take a picture of the user taking the pill, and the server will analyze the information to confirm that the pill was taken. If the user forgets to take the pill, the device will notify the user, "Did you forget to take your nighttime pill?"
[2425] Example prompt sentence:
[2426] When you say "I'm going to take my medicine" and face the camera, the medication administration will begin. Please let us know if you have taken the medicine.
[2427] Schedule management
[2428] When a user enters an appointment, they speak to the device, saying, "I have a dentist appointment tomorrow at 3 PM." The device converts the speech into text and sends the text data to the server. The server analyzes the text and records the appointment on the calendar. 30 minutes before the scheduled time, a reminder is generated and the device issues a voice reminder saying, "It's almost time for your dentist appointment at 3 PM." The software used is the Google Speech-to-Text API and natural language processing algorithms.
[2429] As a specific example, if a user enters "I have a doctor's appointment on Wednesday at 10:00 AM," the server records that information in the calendar and the device sends a reminder 30 minutes before the scheduled time.
[2430] Example prompt sentence:
[2431] Say, "I have a dentist appointment tomorrow at 3 PM," and we'll record the appointment for you. Send you a reminder 30 minutes before.
[2432] Emotion Recognition and Feedback
[2433] When a user expresses emotions in various everyday situations, the device uses a microphone and camera to capture their voice and facial expressions. The device then sends this data to a server, which then uses an emotion engine to analyze the user's emotions. Based on the analysis results, the system provides the user with suggestions for stress relief and relaxation. The software used is an emotion engine (e.g., IBM Watson Tone Analyzer).
[2434] As a specific example, if a user shows signs of depression, the server will analyze their emotions and the device will suggest, "Why not listen to some music for relaxation?" The server will also accumulate emotional data, analyze emotional trends over time, and suggest counseling if necessary.
[2435] Example prompt sentence:
[2436] If you say, "I feel depressed today," you will receive relaxation suggestions. What methods do you recommend?
[2437] In this way, the present invention provides multifaceted support for the user's overall lifestyle, and in particular provides an environment in which the elderly can live with peace of mind through feedback based on their emotional state.
[2438] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2439] Dietary management
[2440] Step 1:
[2441] The user says to the terminal, "I'm going to start eating."
[2442] Input: User's voice command
[2443] Output: Audio data
[2444] Action: The user says "Start eating" to the device.
[2445] Step 2:
[2446] The device performs voice recognition.
[2447] Input: Audio data
[2448] Output: Text data
[2449] What it does: Converts speech to text using speech recognition software (e.g., Google Speech-to-Text API).
[2450] Step 3:
[2451] The device will activate the camera.
[2452] Input: Text data
[2453] Output: Camera activation signal
[2454] Action: Activates the camera sensor and switches it into video capture mode.
[2455] Step 4:
[2456] The device captures video of the meal.
[2457] Input: Camera image
[2458] Output: Video data
[2459] How it works: The camera buffers your meal in real time.
[2460] Step 5:
[2461] The terminal transmits the video data to the server.
[2462] Input: Video data
[2463] Output: Upload to the server
[2464] How it works: Encoded video data is sent to the server via Wi-Fi or mobile data.
[2465] Step 6:
[2466] The server analyzes the video data.
[2467] Input: Video data
[2468] Output: Recognition data
[2469] How it works: Video analysis is performed using OpenCV and TensorFlow to identify the type and quantity of food.
[2470] Step 7:
[2471] The server calculates calories and nutrients.
[2472] Input: Recognition data
[2473] Output: Calorie and nutrient data
[2474] How it works: Looks up the FDA's food database and adds up the calories and nutrients for each food.
[2475] Step 8:
[2476] The server sends the calculation results to the terminal.
[2477] Input: Calorie and nutrient data
[2478] Output: Sending data to the terminal
[2479] Operation: The calculation results are returned to the terminal in JSON format or similar.
[2480] Step 9:
[2481] The device will provide audio feedback.
[2482] Input: Calorie and nutrient data
[2483] Output: Audio feedback
[2484] What it does: Uses the Google Text-to-Speech API to tell you to "Include a fruit containing vitamin C in your next meal."
[2485] Medication Management
[2486] Step 1:
[2487] The user says to the terminal, "I'm going to take my medicine."
[2488] Input: User's voice command
[2489] Output: Audio data
[2490] Action: The user says "take medicine."
[2491] Step 2:
[2492] The device performs voice recognition.
[2493] Input: Audio data
[2494] Output: Text data
[2495] What it does: Uses speech recognition software to convert speech to text.
[2496] Step 3:
[2497] The device will activate the camera.
[2498] Input: Text data
[2499] Output: Camera activation signal
[2500] Action: Activates the camera.
[2501] Step 4:
[2502] The device captures the drug intake as a video.
[2503] Input: Camera image
[2504] Output: Video data
[2505] How it works: The camera collects video data and stores it in a buffer.
[2506] Step 5:
[2507] The terminal transmits the video data to the server.
[2508] Input: Video data
[2509] Output: Upload to server
[2510] Operation: Encodes video data and sends it to the server.
[2511] Step 6:
[2512] The server analyzes the video data.
[2513] Input: Video data
[2514] Output: Recognition data
[2515] How it works: Uses video analysis algorithms to identify medication types and intake.
[2516] Step 7:
[2517] The server checks the medication intake.
[2518] Input: Recognition data
[2519] Output: Intake confirmation data
[2520] What it does: Checks a medication database to assess whether it was taken correctly.
[2521] Step 8:
[2522] The server sends the results to the terminal.
[2523] Input: Intake confirmation data
[2524] Output: Sending data to the terminal
[2525] Behavior: Sends the verification result to the device in JSON format.
[2526] Step 9:
[2527] The device will provide audio feedback.
[2528] Input: Intake confirmation data
[2529] Output: Audio feedback
[2530] What it does: Uses the Google Text-to-Speech API to say "Your medication was taken successfully."
[2531] Schedule management
[2532] Step 1:
[2533] The user speaks into the terminal, "I have a dentist appointment tomorrow at 3:00 p.m."
[2534] Input: User's voice command
[2535] Output: Audio data
[2536] Action: The user speaks the command "dentist appointment."
[2537] Step 2:
[2538] The device performs voice recognition.
[2539] Input: Audio data
[2540] Output: Text data
[2541] What it does: Uses speech recognition software to convert speech into text.
[2542] Step 3:
[2543] The terminal transmits the text data to the server.
[2544] Input: Text data
[2545] Output: Send data to the server
[2546] Action: Uploads the converted text data to the server.
[2547] Step 4:
[2548] The server parses the text.
[2549] Input: Text data
[2550] Output: Analysis data
[2551] How it works: Uses NLP algorithms to extract the event date, time, and content.
[2552] Step 5:
[2553] The server records the event in the calendar.
[2554] Input: Analysis data
[2555] Output: Schedule record data
[2556] Action: Adds new appointment information to the database.
[2557] Step 6:
[2558] The server generates the reminder information.
[2559] Input: Schedule record data
[2560] Output: Remind data
[2561] What it does: Generates a reminder 30 minutes before the scheduled time.
[2562] Step 7:
[2563] The device will remind you by voice.
[2564] Input: Remind data
[2565] Output: Voice reminder
[2566] What it does: Uses the Google Text-to-Speech API to notify you that "It's almost time for your 3pm dentist appointment."
[2567] Emotion Recognition and Feedback
[2568] Step 1:
[2569] The user expresses their feelings.
[2570] Input: Voice and facial expression data
[2571] Output: Emotion data
[2572] How it works: The camera and microphone capture the user's words and facial expressions that express their emotions.
[2573] Step 2:
[2574] The device captures audio and video simultaneously.
[2575] Input: Voice and facial expression data
[2576] Output: Capture data
[2577] Action: Activates microphone and camera to collect data.
[2578] Step 3:
[2579] The device sends the data to the server.
[2580] Input: Capture data
[2581] Output: Send data to the server
[2582] What it does: Encodes audio and video data and sends it to the server.
[2583] Step 4:
[2584] The server analyzes the data.
[2585] Input: Capture data
[2586] Output: Sentiment analysis data
[2587] What it does: Analyzes data using an emotion engine.
[2588] Step 5:
[2589] The server evaluates the emotion.
[2590] Input: Sentiment analysis data
[2591] Output: Emotion rating data
[2592] What it does: Classifies the results of the emotion engine and assigns emotion labels.
[2593] Step 6:
[2594] The server generates the feedback.
[2595] Input: Emotion rating data
[2596] Output: Feedback data
[2597] What it does: Automatically generate stress relief and relaxation suggestions.
[2598] Step 7:
[2599] The device will provide audio feedback.
[2600] Input: Feedback data
[2601] Output: Audio feedback
[2602] What it does: Uses the Google Text-to-Speech API to suggest, "Would you like to listen to some music for relaxation?"
[2603] (Application example 2)
[2604] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2605] Nutritional management, emotional support, and reminder functions are important for elderly people to live their daily lives with peace of mind. However, there are limited means to provide these functions in a single integrated system, and it has been common for individual services or devices to be used separately. This has resulted in problems of poor usability and a heavy burden on users. Furthermore, there are few systems that provide more appropriate support by linking dietary management and emotional support, and there are no systems that link with food delivery services to support the entire process from meal suggestions to ordering. The objective of this invention is to solve these problems.
[2606] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing video data of meals and calculating calories and nutrients, means for recognizing emotional states and suggesting relaxation, and means for suggesting and ordering meals in cooperation with food delivery services. This makes it possible to consistently perform everything from dietary management for elderly people to emotional support, and even meal suggestions and ordering, all in one system.
[2607] A "terminal" is a device equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device.
[2608] A "server" is a device that receives and analyzes audio and video data sent from a terminal.
[2609] "Analysis results" refers to information obtained by the server when it analyzes audio and video data.
[2610] "Notification" refers to the transmission of information to the user based on the analysis results.
[2611] "Nutritional information of meals" is data related to the calories and nutrients of meal contents.
[2612] "Feedback" refers to advice and suggestions to the user based on the server's analysis results.
[2613] "Emotional state" refers to the mental and emotional state of the user.
[2614] "Relaxation suggestions" are advice or suggestions regarding stress reduction and relaxation based on emotional state.
[2615] A "food delivery service" is a service that delivers meals.
[2616] "Meal suggestions" are suggestions for meal menus based on the user's health condition and required nutrients.
[2617] "Ordering" refers to the act of a user purchasing a suggested meal through a food delivery service.
[2618] This invention is a multi-functional system for supporting the lives of elderly people. Each element and its function will be explained below.
[2619] System configuration
[2620] The device is equipped with a microphone, camera, speaker, location information acquisition means, and sensor devices. The server has the ability to analyze audio and video data and notify the user based on the analysis results. It also has the ability to calculate nutritional information for meals and provide feedback, recognize emotional states and suggest relaxation activities, and connect with food delivery services to suggest and order meals.
[2621] Hardware and Software Configuration
[2622] Hardware used:
[2623] Smartphone (camera, microphone)
[2624] Server (high performance computer)
[2625] Software used:
[2626] Python
[2627] OpenCV (image processing library)
[2628] SpeechRecognition (speech recognition library)
[2629] Flask (server-side framework)
[2630] Emotion engine (AI model for emotion recognition)
[2631] System Operation
[2632] 1. Speech Recognition:
[2633] - When the user says "I'm going to start eating" to the device, the device's microphone captures the audio and converts it into a string using the SpeechRecognition library.
[2634] 2. Image capture:
[2635] - Once voice recognition is complete, the device's camera will activate and take a picture of the food, which will then be immediately sent to the server.
[2636] 3. Video analysis:
[2637] - The server analyzes the received image data using OpenCV and calculates the calories and nutrients of the meal. The calculation results are stored on the server.
[2638] 4. Emotion recognition:
[2639] - The emotion engine analyzes the user's facial expressions and tone of voice to understand their emotional state. The results are stored on the server as analytical data.
[2640] 5. Feedback and relaxation suggestions:
[2641] - Based on the analysis results, the server notifies the device of the nutrients needed for the next meal. If the emotional state is biased towards stress, it will also suggest music or activities for relaxation.
[2642] 6. Meal Suggestions and Ordering:
[2643] - The server connects the user's nutritional data and the food delivery service to suggest the next meal menu, which can then be ordered directly through the delivery service.
[2644] Natural language description example
[2645] When an elderly person starts eating, this system recognizes the voice command "start eating" and automatically activates the camera to film the meal. The video data is sent to a server, where an AI model analyzes the calories and nutrients. Next, an emotion engine analyzes the user's emotional state and provides appropriate feedback and relaxation suggestions. It also suggests meal menus based on the necessary nutrients, which can then be ordered directly through a food delivery service.
[2646] As a concrete example:
[2647] For example, when a user says "I'm going to start eating" before lunch, the smartphone camera captures an image of the meal. The image is sent to a server, where the calories and nutrients are analyzed. As a result, feedback such as "Add some fruit containing vitamin C to your next meal" is provided. Also, if the user is feeling stressed, suggestions such as "How about listening to some music to relax?" are made.
[2648] Prompt Sentence Examples
[2649] When you say "start eating," the camera automatically activates and takes a picture of your meal. The image data is sent to the server, where it is analyzed by an AI engine, which calculates calories and nutrients. Nutritional advice for your next meal is then provided via voice.
[2650] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2651] Step 1:
[2652] Voice Recognition:
[2653] When a user says "I'm going to start eating" to the device, the device's microphone captures the voice. The captured voice data is converted into text using Python's SpeechRecognition library. The input is voice data, and the output is text data. Specifically, the device's microphone records the voice, and the SpeechRecognition library analyzes the voice file and converts it into text format.
[2654] Step 2:
[2655] Image Capture:
[2656] Once voice recognition is complete and the text "Start eating" is confirmed, the device's camera will automatically start up and take an image of the meal. The input is video data from the camera, and the output is the captured image file. Specifically, the camera starts up, frames the user's meal, takes a photo, and saves it as an image file.
[2657] Step 3:
[2658] Video Analysis:
[2659] The image captured by the device is sent to the server. The server analyzes the received image data using the OpenCV library and calculates the type of food, calories, and nutrients. The input is image data, and the output is calorie and nutrient data. Specifically, the server runs an image analysis algorithm to recognize ingredients and output their nutritional information in a table format.
[2660] Step 4:
[2661] Emotion recognition:
[2662] The user's facial expression and tone of voice are also sent to the server and analyzed using the emotion engine. The input is facial image and voice data, and the output is emotional state data. Specifically, the server runs an emotion recognition model to determine the user's emotional state, such as stress, happiness, or fatigue.
[2663] Step 5:
[2664] Feedback and relaxation suggestions:
[2665] Based on the analysis results, the server provides nutritional advice and relaxation suggestions appropriate to the user. The input is calorie and nutrient data and emotional state data, and the output is a feedback message. Specifically, the server generates a text message with recommended meal contents and relaxation methods and sends it to the device.
[2666] Step 6:
[2667] Meal Suggestions and Ordering:
[2668] The server executes a function that links suggested meal menus to food delivery services based on the user's nutritional data. The input is nutritional balance data, and the output is meal suggestions and a confirmation message for the order. Specifically, the server generates an appropriate menu and executes the linking process to allow the user to easily order through the food delivery service.
[2669] Through these steps, the system will be able to smoothly carry out a series of steps from dietary management for the elderly to emotional support, and meal suggestions and ordering.
[2670] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2671] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2672] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2673] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2674] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2675] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2676] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2677] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycl...
Claims
1. a terminal equipped with a microphone, a camera, a speaker, a location information acquisition means, and a sensor device; a server that receives and analyzes the audio and video data transmitted from the terminal; A system including a means for notifying a user based on the analysis results from the server.
2. The system according to claim 1 , further comprising means for the terminal to take a photograph of the meal contents, the server to analyze the image to calculate calories and nutrients, and the terminal to provide the user with feedback based on the analysis results.
3. The system according to claim 1, further comprising a means for the terminal to photograph the medication intake status, the server to analyze the video to confirm the medication intake, and the terminal to generate an alert to notify the user if any abnormalities are found.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A