System

A system addressing loneliness and health risks in elderly individuals by converting their voice to text, synthesizing family member voices, and alerting family members to abnormalities, effectively managing their health and safety.

JP2026035376APending Publication Date: 2026-03-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138219
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Elderly people living alone face increased risks of developing serious conditions like dementia and depression due to loneliness and health risks, and family members living far away struggle to provide daily care and monitor their health and safety effectively.

Method used

A system that receives and converts the elderly's voice into text, analyzes it for appropriate responses, synthesizes the response with a family member's voice, records conversations, and alerts family members of abnormalities, while managing health and safety through speech recognition, natural language processing, and speech synthesis technologies.

Benefits of technology

The system supports the elderly's daily life, reduces feelings of loneliness, and provides family members with real-time updates on their condition, enhancing health management and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026035376000001_ABST
    Figure 2026035376000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving voice of an elderly person; means for converting the received voice into text; means for analyzing the converted text and generating an appropriate response; means for synthesizing the response with voice of a registered family member; and means for providing the synthesized voice to the elderly person.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Elderly people living alone face the challenge of increasing their risk of developing serious conditions such as dementia and depression due to health risks and mental stress caused by increased feelings of loneliness. Furthermore, family members who live far away are constantly concerned about the health and safety of their elderly loved ones, but the physical distance makes it difficult to provide daily care. The present invention aims to solve these challenges by providing a system that supports the elderly in their daily lives while reducing the anxiety of their families. [Means for solving the problem]

[0005] The present invention solves the problems using the following means. The first means is a means for receiving the elderly's voice, thereby collecting information from the elderly. The second means includes a function for converting the received voice into text. The third means is a function for analyzing the text data and generating an appropriate response, and the fourth means is a means for synthesizing the response with the voice of a family member, thereby providing a response in a voice that is friendly to the elderly. The present invention also includes a means for providing the synthesized voice to the elderly. The present invention also includes a means for continuously recording conversations and sending the data to a server, and adds a function for notifying the family member of the analysis results from the server. The present invention also provides a system for efficiently managing the health and safety of the elderly by including a means for detecting signs of dementia and emotional abnormalities and a function for sending an alert if an abnormality is detected. These means make it possible to support the elderly's health management and communication, and reduce family anxiety.

[0006] "Elderly" refers to individuals who are older, generally aged 65 or older, and who often require assistance with daily living.

[0007] "Means for receiving voice" refers to devices or functions for capturing and processing voice data produced by the elderly person.

[0008] "Means for converting to text" refers to speech recognition technology or algorithms for converting received voice data into text information.

[0009] "Means for generating appropriate responses" refers to natural language processing technology and response generation algorithms that derive appropriate replies based on the content of speech from elderly people.

[0010] "Means for synthesizing with a family member's voice" refers to speech synthesis technology that uses a pre-recorded family member's voice to synthesize the generated response with that voice.

[0011] "Means for providing audio" refers to speakers or playback equipment that allows the elderly to hear the synthesized audio data through a physical device.

[0012] "Means for continuous recording" refers to devices or software functions that continuously record conversations with elderly people and store them as data.

[0013] "Means for transmitting to a server" refers to the communication technology and protocols used to transmit recorded data via a network to a remote server.

[0014] "Means for notifying family members of the analysis results" refers to a system or method for notifying family members of the data and information analyzed on the server in an appropriate manner on their devices.

[0015] "Means for detecting signs of dementia and emotional abnormalities" refers to algorithms and technologies that analyze the content and speech patterns of elderly people's conversations to determine whether they have dementia or emotional abnormalities.

[0016] "Means for sending alerts" refers to communication functions and systems for sending emergency notifications to family members and relevant organizations when an abnormality is detected. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time.

[0039] Feature Description

[0040] 1. Ability to talk to the elderly using the voices of children and grandchildren

[0041] Subject: Server

[0042] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[0043] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[0044] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[0045] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[0046] The device plays the received audio to the elderly person through a speaker.

[0047] Specific examples

[0048] An elderly person speaks to the device, asking, "What should I eat today?"

[0049] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[0050] 2. Ability to record conversations with elderly people and provide information to their families

[0051] Subject: Device

[0052] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[0053] The server converts the received conversation data into text and analyzes the content.

[0054] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[0055] Specific examples

[0056] An elderly person says, "My shoulder has been hurting lately."

[0057] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[0058] 3. Alerts for dementia and depression

[0059] Subject: Server

[0060] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[0061] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[0062] Specific examples

[0063] Elderly people say, "I don't know what I'm living for."

[0064] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[0065] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[0066] Subject: Device

[0067] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[0068] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[0069] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[0070] Specific examples

[0071] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0072] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[0073] 5. Fall detection and emergency contact function

[0074] Subject: Camera

[0075] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[0076] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[0077] Specific examples

[0078] An elderly person falls in the living room.

[0079] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[0080] As described above, the present invention is a system that comprehensively supports the health management of elderly people, reduces feelings of loneliness, and alleviates anxiety for their families. This system allows elderly people to live their daily lives with peace of mind, and allows their families to follow up in a timely manner.

[0081] The processing flow will be explained below.

[0082] A function that allows you to talk to the elderly using the voices of children and grandchildren

[0083] Step 1:

[0084] The user (elderly person) speaks into the terminal.

[0085] As a concrete example, say, "What shall we eat today?"

[0086] Step 2:

[0087] The terminal receives the elderly person's voice and transmits the voice data to the server.

[0088] The audio data is transferred to the server in real time.

[0089] Step 3:

[0090] The server converts the received voice data into text.

[0091] A speech recognition algorithm converts the speech data into a string of characters.

[0092] Step 4:

[0093] The server analyzes the converted text data and generates an appropriate response.

[0094] Use natural language processing (NLP) techniques to understand the context of text data.

[0095] Step 5:

[0096] The server synthesizes a response based on voice samples from family members.

[0097] A response message is synthesized using the voice of a pre-registered family member.

[0098] Step 6:

[0099] The server transmits the synthesized voice data to the terminal.

[0100] The generated voice data is sent back to the elderly person.

[0101] Step 7:

[0102] The device plays the synthesized speech through the speaker.

[0103] The elderly person hears a family member respond, "Maybe fish would be good?"

[0104] A function that records conversations with elderly people and provides information to their families

[0105] Step 1:

[0106] The device constantly records conversations with the elderly.

[0107] Record all your everyday conversations.

[0108] Step 2:

[0109] The recorded conversation data is periodically sent to a server.

[0110] For example, a data packet is sent every hour.

[0111] Step 3:

[0112] The server converts the received voice data into text.

[0113] Convert the speech into text using voice recognition technology.

[0114] Step 4:

[0115] The server analyzes the converted text data.

[0116] Natural language processing technology is used to extract important information and anomalies.

[0117] Step 5:

[0118] Based on the analysis results, the information to be notified to the family is determined.

[0119] If an anomaly is detected, detailed information about it will also be included.

[0120] Step 6:

[0121] The server sends the analysis results to the family.

[0122] Notifications are given in real time.

[0123] Examples:

[0124] An elderly person says, "My shoulder has been hurting lately."

[0125] The server analyzes the information and sends a notification to the family about the "shoulder pain."

[0126] A feature that sends alerts for dementia and depression

[0127] Step 1:

[0128] The server periodically analyzes conversation data with the elderly.

[0129] The analysis targets daily conversation logs.

[0130] Step 2:

[0131] The server detects changes in grammatical structure and sentiment.

[0132] It uses natural language processing techniques and sentiment analysis algorithms.

[0133] Step 3:

[0134] If an anomaly is detected, the server generates an alert.

[0135] Identify signs of dementia and emotional abnormalities.

[0136] Step 4:

[0137] The server generates alerts and sends them to family members and medical institutions.

[0138] The alert is sent as an emergency notification.

[0139] Examples:

[0140] Elderly people say, "I don't know what I'm living for."

[0141] The server performs emotion analysis, detects signs of depression, and sends emergency notifications.

[0142] Functions to prevent forgetting to take medicine and assist with meals and shopping

[0143] Step 1:

[0144] The server sets medication times and meal schedules.

[0145] Set a daily reminder.

[0146] Step 2:

[0147] The device will send a reminder notification at the set time.

[0148] A notification will play saying "Time for your medicine."

[0149] Step 3:

[0150] When the elderly person responds with a voice of confirmation, the device records the information and sends it to the server.

[0151] "Thanks, I drank it," he replied.

[0152] Step 4:

[0153] The server notifies the family of the reminder information.

[0154] Notifies you that confirmation of missed dose has been completed.

[0155] Step 5:

[0156] The server learns the elderly person's lifestyle patterns and generates a shopping list.

[0157] Make a list of the items you need.

[0158] Step 6:

[0159] The device notifies the elderly person of the shopping list.

[0160] Announce the list contents by voice.

[0161] Examples:

[0162] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0163] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[0164] Fall detection and emergency contact function

[0165] Step 1:

[0166] The camera monitors the footage at all times.

[0167] Monitor overall behavior.

[0168] Step 2:

[0169] The camera detects an abnormality.

[0170] For example, a falling motion is detected.

[0171] Step 3:

[0172] Video data is sent to the server in real time.

[0173] Send data immediately.

[0174] Step 4:

[0175] The server identifies the fall and generates an emergency alert.

[0176] Prepare notifications for medical institutions and families.

[0177] Step 5:

[0178] The server sends emergency alerts to family and medical institutions.

[0179] Encourage a rapid response.

[0180] Examples:

[0181] An elderly person falls in the living room.

[0182] The camera detects an abnormality and notifies the server.

[0183] The server sends emergency alerts to family and medical institutions.

[0184] These are the specific steps of the program processing, which will enable health management for elderly people living alone, reduce loneliness, and provide peace of mind to their families.

[0185] Example 1

[0186] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0187] There is a need for health management for the elderly, reducing loneliness, and improving family security. However, systems that enable appropriate communication and health monitoring for elderly people living alone have not yet been fully developed. When elderly people live in isolated situations, the lack of daily communication and the inability to detect health abnormalities early are problems. In particular, there is a lack of means to detect early signs of dementia and depression and provide appropriate treatment.

[0188] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0189] In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for saving voice samples provided by the family member and extracting features, and natural language processing means for analyzing the elderly person's speech and generating a response related to life support. This enables natural communication with the elderly, health management, reduction of loneliness, and early detection of signs of dementia and depression.

[0190] The "means for receiving the voice of the elderly person" is a function for acquiring the voice uttered by the elderly person using an input device such as a microphone.

[0191] The "means for converting received voice into text" is a function that converts voice data into corresponding text data using voice recognition technology.

[0192] "Means for analyzing the converted text and generating an appropriate response" refers to a function that uses natural language processing (NLP) technology to analyze text data and generate an appropriate response based on the content.

[0193] The "means for synthesizing the response with the voice of a registered family member" is a technology for converting text data into voice data using the voice characteristics of a designated family member.

[0194] The "means for providing synthesized voice to the elderly" is a function for playing back synthesized voice to the elderly using a voice output device such as a speaker.

[0195] The "means for saving voice samples provided by family members and extracting features" is a function for saving voice data provided by family members and extracting voice features.

[0196] "Natural language processing means for analyzing elderly people's statements and generating responses related to life support" is a technology that analyzes the content of elderly people's statements and generates information and actions necessary for life support based on the analysis results.

[0197] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. This system is realized by integrating speech recognition, natural language processing, and speech synthesis technologies.

[0198] Hardware and software used

[0199] Hardware

[0200] High-Precision Microphone

[0201] speaker

[0202] camera

[0203] server

[0204] Devices (smartphones, tablets, PCs, etc.)

[0205] software

[0206] Speech recognition: A speech recognition API for converting voice data into text (e.g., Google® Cloud Speech-to-Text API)

[0207] Natural Language Processing (NLP): NLP models (e.g., AWS® Comprehend) for analyzing text data and generating appropriate responses

[0208] Speech synthesis: A speech synthesis API for converting text to speech (e.g., IBM Watson® Text-to-Speech API)

[0209] Database: A database for storing audio samples and analysis data (e.g., Google Cloud Storage)

[0210] Explanation of the specific functions of the system

[0211] 1. Ability to talk to the elderly using the voices of children and grandchildren

[0212] Subject: Server, Terminal, User

[0213] The server stores voice samples provided by family members in a database. It also extracts voice features using a machine learning model. When the user (elderly person) speaks into the device, the device captures the voice data and sends it to the server. The server converts the received voice data into text using the Google Cloud Speech-to-Text API and analyzes the text data using a natural language processing (NLP) model. An appropriate response is generated, and speech is synthesized using the family member's voice using the IBM Watson Text-to-Speech API, and the speech data is sent to the device. The device then plays the received audio to the elderly person through a speaker.

[0214] Specific examples

[0215] An elderly person speaks to the device, asking, "What should I eat today?"

[0216] The server synthesizes the family member's voice saying, "Maybe fish would be good?" and plays it back to the elderly person via the device.

[0217] Example prompt

[0218] "When you say to a device, 'What should I eat today?' how does the server process the data and generate the voice response?"

[0219] 2. Ability to record conversations with elderly people and provide information to their families

[0220] Subject: Terminal, Server

[0221] The device continuously records conversations with the elderly. The recorded conversation data is periodically sent to a server, which converts the voice data into text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing (NLP) model, and if particularly important information or anomalies are detected, the server notifies the family using the Twilio API.

[0222] Specific examples

[0223] An elderly person says, "My shoulder has been hurting lately."

[0224] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[0225] Example prompt

[0226] "When the audio data recorded on the device is sent to the server, how will family members be notified if an abnormality is detected?"

[0227] 3. Alerts for dementia and depression

[0228] Subject: Server

[0229] The server periodically converts conversation data with the elderly person into text using the Google Cloud Speech-to-Text API, then analyzes the text data using a natural language processing (NLP) model. If signs of dementia or abnormal emotions are detected, an alert is generated and notifies family members and medical institutions using the Twilio API.

[0230] Specific examples

[0231] Elderly people say, "I don't know what I'm living for."

[0232] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[0233] Example prompt

[0234] "If an elderly person makes a statement that indicates signs of depression, what analysis does the server perform and how does it generate a notification?"

[0235] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[0236] Subject: Terminal, Server

[0237] The device uses the Google Calendar API to manage medication times and meal schedules, and sends a reminder at the set time: "It's time for your medicine." When the user (elderly person) responds with a voice confirmation, the voice is sent to the server and converted into text using the Google Cloud Speech-to-Text API. The server then notifies the family of the confirmation information. The server also learns the elderly person's lifestyle patterns, generates a shopping list, and provides it via voice from the device.

[0238] Specific examples

[0239] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0240] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[0241] Example prompt

[0242] "How can we remind and confirm elderly people to prevent them from forgetting to take their medication, and notify their families of this information?"

[0243] 5. Fall detection and emergency contact function

[0244] Subject: camera, server

[0245] The camera constantly monitors the video and runs an algorithm to detect anomalies. If an elderly person falls, the camera detects the movement and sends the video data in real time to a server. The server then confirms the fall and uses the Twilio API to send an emergency notification to family members and medical institutions.

[0246] Specific examples

[0247] An elderly person falls in the living room.

[0248] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[0249] Example prompt

[0250] "How does the fall detection feature work and how does it send emergency notifications after detection?"

[0251] As described above, this system recognizes the voice of the elderly person, generates an appropriate response using natural language processing, and then uses voice synthesis technology to respond in the voice of a family member, thereby providing comprehensive support for managing the elderly person's health, reducing feelings of loneliness, and providing a sense of security to their family.

[0252] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0253] Specific processing steps of the system program

[0254] Feature 1: Ability to talk to the elderly using the voices of children and grandchildren

[0255] Subject: Server, Terminal, User

[0256] Step 1:

[0257] Subject: Server

[0258] Voice samples provided by family members are stored in a database.

[0259] (Input) Voice samples recorded by family members

[0260] (Processing) The server receives the audio samples and stores them in a database.

[0261] (Output) Voice samples of family members stored in the database

[0262] Specific operation: The server receives the voice data and stores it in a database.

[0263] Step 2:

[0264] Subject: Server

[0265] The voice sample is analyzed using a machine learning model to extract features.

[0266] (Input) Audio samples stored in the database

[0267] (Processing) Apply a voice feature extraction algorithm to extract features

[0268] (Output) Extracted feature data

[0269] Specific operation: The server applies a feature extraction algorithm to the voice sample and stores the features in a database.

[0270] Step 3:

[0271] Subject: User

[0272] The user speaks into the terminal.

[0273] (Input) User's voice

[0274] (Processing) High-precision microphone captures audio

[0275] (Output) Captured audio data

[0276] Specific operation: The user speaks to the device, "What should I eat today?"

[0277] Step 4:

[0278] Subject: Device

[0279] The terminal transmits the voice data to the server.

[0280] (Input) Captured audio data

[0281] (Processing) Send the audio data to the server

[0282] (Output) Audio data received by the server

[0283] Specific operation: The device sends the captured audio data to a server via the Internet.

[0284] Step 5:

[0285] Subject: Server

[0286] Converts audio data into text.

[0287] (Input) Audio data received by the server

[0288] (Processing) Convert speech to text using a speech recognition API (e.g., Google Cloud Speech-to-Text API)

[0289] (Output) Converted text data

[0290] What happens: The server uses the Google Cloud Speech-to-Text API to convert the speech "What should we eat today?" into text.

[0291] Step 6:

[0292] Subject: Server

[0293] Analyze text data using an NLP model and generate appropriate responses.

[0294] (Input) Converted text data

[0295] (Processing) Analyze the text using an NLP model (e.g. AWS Comprehend) and generate an appropriate response

[0296] (Output) The generated response text

[0297] Specific behavior: The server generates a response to the question "What should we eat?" with "Maybe fish would be good?"

[0298] Step 7:

[0299] Subject: Server

[0300] The response text is converted into speech using the speech synthesis API.

[0301] (Input) Generated response text

[0302] (Processing) Convert text to speech using a speech synthesis API (e.g. IBM Watson Text-to-Speech API)

[0303] (Output) Synthesized voice data

[0304] Specific behavior: The server converts the text "Maybe fish would be good?" into speech.

[0305] Step 8:

[0306] Subject: Server

[0307] The synthesized voice data is transmitted to the terminal.

[0308] (Input) Synthesized voice data

[0309] (Processing) Send audio data to the terminal

[0310] (Output) Audio data received on the device

[0311] Specific operation: The server sends the synthesized voice data to the terminal.

[0312] Step 9:

[0313] Subject: Device

[0314] The device plays the audio.

[0315] (Input) Audio data received by the device

[0316] (Processing) Playing audio through speakers

[0317] (Output) Audio played to the elderly

[0318] Specific operation: The device plays a voice message from the speaker saying, "Maybe a fish would be good?"

[0319] You can add specific processing steps below for other features as well, but you can distinguish the steps for each feature based on this format.

[0320] If you have steps for other features below, please list them in a similar format.

[0321] (Application example 1)

[0322] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0323] In modern society, the number of elderly people living alone is increasing, and health management and reducing feelings of loneliness have become social issues. Furthermore, when elderly people living alone use self-driving vehicles, ensuring safety and communication with their families are important. However, conventional systems do not adequately provide comprehensive solutions to these issues. Therefore, there is a need for a system that integrates safe and secure transportation means for the elderly with support for their daily lives.

[0324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0325] In this invention, the server includes means for receiving the elderly person's voice, means for converting the voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the family member's voice, means for providing synthesized voice, means for setting a destination for the autonomous vehicle and providing route guidance, means for monitoring the driving situation and detecting abnormalities, and means for notifying the family member in real time when an abnormality occurs. This allows the elderly person to receive route guidance from the family member's voice when using the autonomous vehicle, improving safety and enabling the family member to understand the elderly person's situation in real time.

[0326] · An "elderly person" is an individual who, due to their advanced age, requires assistance in daily living.

[0327] "Means for receiving voice" refers to a device or method for recording and capturing the voice uttered by the elderly person and inputting it into the system.

[0328] "Means for converting speech to text" means the process or technology that converts received speech data into corresponding text data using speech recognition technology.

[0329] "Means of analyzing text" refers to natural language processing techniques used to understand meaning and intent based on converted text data.

[0330] - "Means for generating a response" refers to an algorithm or program that generates an appropriate response based on the analysis results.

[0331] "Means for synthesizing using family members' voices" refers to a method of reproducing the generated response using voice synthesis technology using pre-registered family members' voice data.

[0332] "Means for providing synthesized speech" means a speaker or other audio playback device that delivers synthesized speech to the senior citizen.

[0333] "Means for setting a destination in an autonomous vehicle" refers to the methods and technologies for inputting and setting the destination desired by the elderly person into an autonomous vehicle.

[0334] "Means for providing directions" refers to systems and technologies that allow autonomous vehicles to provide directions to seniors toward a set destination.

[0335] "Means for monitoring the driving situation" refers to sensors, cameras, and algorithms that monitor the driving situation in real time while the autonomous vehicle is driving to ensure safety.

[0336] "Means for detecting abnormalities" refers to techniques and methods for recognizing and identifying hazards and abnormal events that may occur during operation.

[0337] "Means of notifying family members in real time" refers to communication technologies and methods for immediately notifying family members of the situation when an abnormality is detected.

[0338] To implement this invention, it is necessary to build a system that receives the voice of the elderly person and synthesizes it with the voices of family members. This system uses the following hardware and software.

[0339] Hardware used

[0340] 1. Smartphone: Used to record the elderly person's voice and send the recording data to the server.

[0341] 2. Self-driving vehicles: Used as transportation for the elderly, with functions such as route guidance and safety monitoring.

[0342] 3. Speaker: Used as a device to transmit synthesized voices of family members to the elderly.

[0343] Software used

[0344] 1. pyaudio: A library for recording the voices of elderly people.

[0345] 2. wave: A library for saving and playing recorded audio.

[0346] 3. transformers (Wav2Vec2): A library for converting audio to text.

[0347] 4. gTTS: A library for synthesizing text in the voices of family members.

[0348] 5. smtplib: A library for implementing notification functions.

[0349] Data processing and calculation

[0350] The server first receives the elderly person's voice and records it using the pyaudio library. The recorded voice data is then temporarily saved using the wave library. The voice data is then converted to text using the Wav2Vec2 model in transformers. The converted text data is analyzed using natural language processing techniques to generate an appropriate response. The generated response is then synthesized in the voice of a family member using the gTTS library and provided to the elderly person through a speaker as synthesized speech.

[0351] If an anomaly is detected or if specific keywords are found, a real-time notification is sent to family members using the smtplib library, containing the analysis results and the elderly person's current condition.

[0352] Specific examples

[0353] For example, if an elderly person says, "Shall I confirm your next hospital appointment?", the system works as follows: First, the smartphone records this speech and sends it to the server. The server converts the recorded speech into text using the Wav2Vec2 model and analyzes its content. An appropriate response is generated, such as "Yes, Grandma. Your next hospital appointment is tomorrow at 10:00 AM," using a synthesized voice from a family member's voice, and this is played to the elderly through the speaker.

[0354] Prompt Sentence Examples

[0355] 1. "Grandpa, should I confirm your next doctor's appointment?"

[0356] 2. "Grandma, have you taken your morning medicine?"

[0357] 3. "Where is Grandpa going? Is he driving safely?"

[0358] This will allow elderly people to use self-driving vehicles with peace of mind, and will also enable family members to keep track of the elderly person's situation in real time.

[0359] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0360] Step 1:

[0361] The smartphone records the elderly person's voice. The device uses the pyaudio library to record 10 seconds of voice data uttered by the elderly person. The input here is the elderly person's voice, and the output is the recorded voice data (audio file).

[0362] Step 2:

[0363] The recorded audio data is sent to the server. The device uses the wave library to save the audio file and uploads it to the server. The input is the recorded audio file, and the output is the audio file saved on the server.

[0364] Step 3:

[0365] The server converts the audio data to text. The server uses the Wav2Vec2 model from transformers to convert the audio file to text. The input is the audio file, and the output is the converted text.

[0366] Step 4:

[0367] The server analyzes the text data and generates an appropriate response. Natural language processing technology is used to analyze the meaning and intent of the converted text data and generate an appropriate response text. The input is the text data, and the output is the generated response text.

[0368] Step 5:

[0369] The server synthesizes the response text in the voice of the family member. The server uses the gTTS library to synthesize the generated response text in the voice of the pre-registered family member and create an audio file. The input is the response text, and the output is the synthesized audio file.

[0370] Step 6:

[0371] The synthesized voice is provided to the elderly through a speaker. The server sends the synthesized voice to the terminal, and the terminal plays the voice through the speaker. The input is the synthesized voice file, and the output is the voice played to the elderly.

[0372] Step 7:

[0373] Monitors the driving situation and detects abnormalities. Autonomous vehicles use sensors and cameras to monitor the driving situation in real time to ensure safety. The input is real-time data during driving, and the output is analysis results and abnormality detection signals.

[0374] Step 8:

[0375] If an abnormality is detected, a notification is sent to family members in real time. The server uses the smtplib library to notify family members of the situation by email when it receives an abnormality detection signal. The input is the abnormality detection signal, and the output is the alert notification sent to family members.

[0376] Specific example prompts

[0377] "Grandpa, should I confirm your next doctor's appointment?"

[0378] "Grandma, have you taken your morning medicine?"

[0379] "Where is Grandpa going? Are you driving safely?"

[0380] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0381] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and reduce feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it achieves more effective responses and anomaly detection.

[0382] Feature Description

[0383] 1. Ability to talk to the elderly using the voices of children and grandchildren

[0384] Subject: Server

[0385] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[0386] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[0387] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[0388] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[0389] The device plays the received audio to the elderly person through a speaker.

[0390] Specific examples

[0391] An elderly person speaks to the device, asking, "What should I eat today?"

[0392] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[0393] 2. Ability to record conversations with elderly people and provide information to their families

[0394] Subject: Device

[0395] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[0396] The server converts the received conversation data into text and analyzes the content.

[0397] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[0398] Specific examples

[0399] An elderly person says, "My shoulder has been hurting lately."

[0400] The server analyzes the information and sends a notification about the "shoulder pain" to the family, who can then plan a visit.

[0401] 3. Alerts for dementia and depression

[0402] Subject: Server

[0403] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[0404] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[0405] Specific examples

[0406] Elderly people say, "I don't know what I'm living for."

[0407] The server performs emotion analysis, detects signs of depression, and sends emergency notifications to family members and medical institutions.

[0408] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[0409] Subject: Device

[0410] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[0411] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[0412] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[0413] Specific examples

[0414] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0415] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[0416] 5. Fall detection and emergency contact function

[0417] Subject: Camera

[0418] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[0419] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[0420] Specific examples

[0421] An elderly person falls in the living room.

[0422] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[0423] Features of inventions that combine emotion engines

[0424] 1. Emotion recognition and response generation using an emotion engine

[0425] Subject: Server

[0426] The server uses an emotion engine to recognize emotions from the elderly person's voice data, extracts the elderly person's emotions from the voice data, and generates a response based on that information.

[0427] Depending on the recognized emotion, the server generates an appropriate response and provides it to the elderly person as a synthesized voice.

[0428] Specific examples

[0429] An elderly person says, "I feel a little lonely today."

[0430] The server uses an emotion engine to recognize the emotion "loneliness" and generates a response such as "I see you're feeling lonely. Let's think of something we can do together."

[0431] 2. Anomaly detection and notification using emotion engine

[0432] Subject: Server

[0433] If the server detects negative emotions in an elderly person using the emotion engine, it generates an alert and notifies family members and relevant organizations.

[0434] If an anomaly is detected, a notification system will be activated to ensure a rapid response.

[0435] Specific examples

[0436] An elderly person says, "Nothing is fun anymore."

[0437] The server uses an emotion engine to detect emotions such as "despair" and "deep sadness" and sends emergency notifications to family members and medical institutions.

[0438] The above is a specific embodiment of the present invention that combines an emotion engine. This makes it possible to more accurately grasp the emotional state of elderly people living alone and respond appropriately. Furthermore, since it can respond sensitively to changes in emotions, it can also support the mental health of the elderly.

[0439] The processing flow will be explained below.

[0440] Emotion recognition and response generation using an emotion engine

[0441] Step 1:

[0442] The user (elderly person) speaks into the terminal.

[0443] As a specific example, say, "I feel a little lonely today."

[0444] Step 2:

[0445] The terminal receives the elderly person's voice and transmits the voice data to the server.

[0446] Step 3:

[0447] The server converts the received voice data into text.

[0448] A speech recognition algorithm converts the speech data into a string of characters.

[0449] Step 4:

[0450] The server passes the converted text data and voice data to the emotion engine.

[0451] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[0452] Step 5:

[0453] The server generates an appropriate response based on the emotional data recognized by the emotion engine.

[0454] Using natural language processing technology, it generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[0455] Step 6:

[0456] The server generates a response that is then synthesized into the family member's voice.

[0457] A response message is synthesized using the voice of a pre-registered family member.

[0458] Step 7:

[0459] The server transmits the synthesized voice data to the terminal.

[0460] The generated voice data is sent back to the elderly person.

[0461] Step 8:

[0462] The device plays the synthesized speech through the speaker.

[0463] The elderly person hears a family member respond, "I see you're feeling lonely. Let's think of something we can do together."

[0464] Anomaly detection and notification using emotion engine

[0465] Step 1:

[0466] The user (elderly person) speaks into the terminal.

[0467] For example, say, "Nothing is fun anymore."

[0468] Step 2:

[0469] The terminal receives the elderly person's voice and transmits the voice data to the server.

[0470] Step 3:

[0471] The server converts the received voice data into text.

[0472] A speech recognition algorithm converts the speech data into a string of characters.

[0473] Step 4:

[0474] The server passes the converted text data and voice data to the emotion engine.

[0475] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[0476] Step 5:

[0477] The emotion engine detects negative emotions such as "despair" and "deep sadness."

[0478] The emotion engine passes these abnormal emotion data to the server.

[0479] Step 6:

[0480] The server generates alerts based on the emotion engine data.

[0481] Prepare an abnormality detection message and notify family members and relevant organizations.

[0482] Step 7:

[0483] The server sends emergency notifications to family and medical institutions.

[0484] Alerts such as "elderly people are feeling deep sadness" will be sent via email or app notification.

[0485] Step 8:

[0486] Family and medical institutions will be notified and take appropriate action.

[0487] Schedule visits and medical consultations.

[0488] These are the processing steps of the system that combines the emotion engine, which makes it possible to grasp the emotional state of the elderly more accurately and take the necessary measures quickly.

[0489] Example 2

[0490] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0491] For elderly people living alone, reducing loneliness and managing their health are major challenges. It is particularly difficult to understand changes in the elderly's moods and health status when family members live far away. Early detection of dementia and depression and prevention of falls are also important issues. Furthermore, elderly people need assistance with daily activities such as remembering to take their medication, eating, and shopping. To solve these challenges, a voice-responsive system that includes emotion recognition is needed.

[0492] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for continuously recording conversations with the elderly person, means for transmitting the recorded conversation data to the server, means for analyzing the transmitted conversation data and notifying the family member of the analysis result, means for analyzing the conversation data with the elderly person and detecting signs of dementia or emotional abnormalities, means for transmitting an alert when signs of dementia or emotional abnormalities are detected, means for learning the elderly person's lifestyle patterns and setting reminders as necessary, means for notifying the family member based on the reminder, means for recording confirmation of the notified content and notifying the family member, means for detecting the elderly person's fall, means for transmitting an emergency notification when a fall is detected, means for recognizing the elderly person's emotion from the voice data, means for generating an appropriate response according to the recognized emotion, means for detecting the elderly person's negative emotion, and means for transmitting an alert to the family member or relevant organization when the emotion is detected. This reduces the elderly person's sense of loneliness and enables health management. It can also enable early detection of dementia and depression, prevent falls, prevent people from forgetting to take their medication, and provide assistance with meals and shopping.

[0493] The "means for receiving voice" is a device or system that captures the voice spoken by the elderly person and receives it as data.

[0494] A "speech-to-text means" is software or an algorithm that converts received speech data into text data.

[0495] The "means for analyzing text and generating an appropriate response" refers to a natural language processing engine or algorithm for analyzing text data converted from speech and generating an appropriate response.

[0496] The "means for synthesizing a response using the voice of a registered family member" refers to software or an algorithm for synthesizing the generated response using voice data of a registered family member in advance.

[0497] "Means for providing synthesized voices to the elderly" refers to devices or systems that transmit synthesized voices of family members to the elderly through speakers or communication devices.

[0498] "Means for continuously recording conversations" refers to equipment or systems that continuously record conversations with elderly people and store them as data.

[0499] The "means for transmitting recorded conversation data to a server" refers to a communication device or system for transmitting recorded conversation data to a server via a network.

[0500] "Means for analyzing the transmitted conversation data and notifying the family of the analysis results" refers to software or a communication system for analyzing the conversation data transmitted to the server and notifying the family of the analysis results.

[0501] "Means for detecting signs of dementia and emotional abnormalities" refers to algorithms and software that analyze conversation data with elderly people to detect signs of dementia and emotional abnormalities.

[0502] "Means for sending alerts when signs of dementia or abnormal emotions are detected" refers to a communication system or software for sending alerts to family members or relevant organizations when signs of dementia or abnormal emotions are detected.

[0503] "Means for learning lifestyle patterns and setting reminders" refers to algorithms or software that learn the lifestyle patterns of elderly people through data analysis and set reminders at the necessary times.

[0504] The "means for notifying based on a reminder" refers to a device or system for notifying the elderly based on the set reminder.

[0505] "Means for recording confirmation of the notified content and notifying the family" refers to a communication system or software for recording the notification confirmation from the elderly person and conveying the content to the family.

[0506] "Means for detecting falls in the elderly" refers to systems and algorithms that use cameras and sensors to detect falls in the elderly in real time.

[0507] The "means for sending an emergency notification when a fall is detected" is a communication system for sending an emergency notification to family members and medical institutions when a fall of an elderly person is detected.

[0508] "Means for recognizing the emotions of elderly people from voice data" refers to software or algorithms that use an emotion engine to extract and recognize emotions from the voice data of elderly people.

[0509] The "means for generating an appropriate response based on the recognized emotions" refers to a natural language processing engine or software that generates an appropriate response based on the emotions of the elderly person.

[0510] The "means for detecting negative emotions" refers to emotion analysis algorithms or software for identifying and detecting negative emotions from elderly people's voice data.

[0511] "Means for sending an alert to family members or relevant organizations when negative emotions are detected" refers to a communication system or software for quickly sending an alert to family members or relevant organizations when negative emotions are detected.

[0512] The present invention is a system that aims to manage the health of elderly people living alone and reduce feelings of loneliness through communication with family members. Furthermore, by combining it with an emotion engine, more effective responses and anomaly detection become possible. This system has multiple functions for connecting elderly people with their families. Specific embodiments of the system are described below.

[0513] System configuration

[0514] This system consists of three main components: a server, a terminal, and a user. The server processes and analyzes data, the terminal acts as an interface with the elderly, and the user refers to the elderly and their family members.

[0515] Hardware and software used

[0516] Hardware: The device can be a smart speaker such as Amazon Echo or GOOGLE HOME (registered trademark), and the camera can be a Nest Cam or similar.

[0517] Software: On the server side, we use Google Speech-to-Text for speech recognition, NLTK and Spacy for natural language processing, Google Text-to-Speech API for speech synthesis, and IBM Watson Tone Analyzer for sentiment analysis.

[0518] Program processing flow

[0519] 1. Saving and analyzing audio samples

[0520] The server receives voice samples provided by family members and stores them in a database. The stored voice samples are analyzed using machine learning models (e.g., TENSORFLOW®) to extract features. This analysis lays the foundation for delivering the voices of family members to the elderly.

[0521] 2. Audio capture and transmission

[0522] When an elderly person speaks into the device, the device captures the voice and transmits it to a server in real time. For example, Amazon Echo captures the elderly person's voice in high quality and transmits the data to a server over the Internet.

[0523] 3. Speech data text conversion and analysis

[0524] The server converts the received voice data into text using Google Speech-to-Text, which is then analyzed using natural language processing tools such as NLTK or Spacy. Based on this analysis, an appropriate response is generated.

[0525] 4. Speech synthesis and transmission

[0526] The generated response is synthesized using the Google Text-to-Speech API using the family member's voice. The synthesized voice data is sent to the device, which plays it back to the elderly person through the device's speaker. For example, if an elderly person asks, "What should we eat today?", the server synthesizes a family member's voice saying, "Maybe fish would be good?" and plays it back on the device.

[0527] 5. Recording and analyzing conversations

[0528] The device constantly records conversations with the elderly person and periodically sends the data to a server. The server then converts the conversation data into text using Google Speech-to-Text and detects important information and anomalies. The detection results are then notified to family members in real time. For example, if an elderly person says, "My shoulder has been hurting lately," that information is immediately notified to the family.

[0529] 6. Sentiment Analysis and Response

[0530] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize emotions from the elderly person's voice. Based on the recognized emotion, it generates an appropriate response and provides it to the elderly person as synthesized speech. For example, if an elderly person says, "I feel a little lonely today," the server generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[0531] 7. Fall detection and emergency notification

[0532] The camera constantly monitors the elderly person's movements and runs an algorithm (e.g., OpenCV) to detect abnormal behavior (e.g., falls). If a fall is detected, the server immediately sends an emergency notification to family members or medical institutions. For example, if an elderly person falls in the living room, the camera detects the movement and notifies the server in real time. The server confirms this and immediately sends an emergency alert.

[0533] Specific prompt examples

[0534] A specific example of a prompt sentence to input into the generative AI model using this system is as follows:

[0535] "An elderly person asks, 'What shall we eat today?' A child's voice responds, 'Maybe fish would be good?'"

[0536] "An elderly person says, 'I feel a little lonely today.' Use an emotion engine to generate an appropriate response."

[0537] The above is a specific embodiment of the present invention. This system can reduce the sense of loneliness felt by the elderly and help them manage their health. It can also provide early detection of dementia and depression, prevent falls, help with forgetting to take medication, and assist with meals and shopping.

[0538] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0539] Processing steps of this system's program

[0540] Step 1: Saving and analyzing audio samples

[0541] Subject: Server

[0542] Input: Voice sample provided by family member (e.g., audio file).

[0543] What it does: The server receives voice samples uploaded by family members and stores them in a database.

[0544] Data processing / computation: The server uses machine learning models (e.g., TensorFlow) to extract and analyze voice features.

[0545] Output: A feature-analyzed audio profile.

[0546] Step 2: Capture and send audio

[0547] Subject: Device

[0548] Input: Speech produced by an elderly person.

[0549] Specific operation: When an elderly person speaks into a device (e.g., a smart speaker), the device captures the voice.

[0550] Data processing / calculation: The device transmits the captured voice data to the server in real time.

[0551] Output: The audio data sent to the server.

[0552] Step 3: Converting audio data into text and analyzing it

[0553] Subject: Server

[0554] Input: The audio data sent to the server.

[0555] Specific operation: The server converts the received voice data into text using voice recognition software (e.g., Google Speech-to-Text).

[0556] Data processing / calculation: The converted text data is analyzed using natural language processing (NLP) tools (e.g., NLTK or Spacy).

[0557] Output: Parsed text data.

[0558] Step 4: Speech synthesis and transmission

[0559] Subject: Server

[0560] Input: Parsed text data.

[0561] What it does: The server synthesizes a response in the family member's voice based on the analyzed text (e.g., Google Text-to-Speech API).

[0562] Data processing / calculation: Voice synthesis using family voice profiles.

[0563] Output: Synthesized speech data.

[0564] Step 5: Play audio

[0565] Subject: Device

[0566] Input: Synthesized speech data.

[0567] Specific operation: The device plays the received audio to the elderly person through the speaker.

[0568] Output: Audio provided to the senior.

[0569] Step 6: Record and send the conversation

[0570] Subject: Device

[0571] Input: Conversation with an elderly person.

[0572] Specific operation: The device constantly records conversations with the elderly and periodically sends the data to a server.

[0573] Data processing / calculation: Converting the audio data into the format required to send it to the server.

[0574] Output: Conversation data sent to the server.

[0575] Step 7: Analyze conversation data and notify

[0576] Subject: Server

[0577] Input: Conversation data sent to the server.

[0578] What it does: The server uses Google Speech-to-Text to convert the conversation data into text and detects important information and anomalies.

[0579] Data processing / calculation: Generate information to notify family members based on the converted and analyzed data.

[0580] Output: Analysis results communicated to family members.

[0581] Step 8: Sentiment Analysis and Response Generation

[0582] Subject: Server

[0583] Input: Speech data of elderly people.

[0584] Specific operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize emotions from the elderly person's voice.

[0585] Data processing / computation: Generate appropriate responses based on recognized emotions and provide them as synthesized speech.

[0586] Output: Synthetic voice response provided to the senior.

[0587] Step 9: Fall Detection and Emergency Notification

[0588] Subject: Camera

[0589] Enter: Elderly Movement.

[0590] Specific operation: The camera constantly monitors the elderly person's movements and runs an algorithm (e.g., OpenCV) to detect abnormal movements (e.g., falls).

[0591] Data processing / calculation: When a fall is detected, data processing is performed to notify the server in real time.

[0592] Output: Emergency notification data sent to the server.

[0593] Step 10: Sending emergency alerts

[0594] Subject: Server

[0595] Input: Emergency notification data.

[0596] Specific operation: The server checks the fall information and immediately sends an emergency notification to family members and medical institutions.

[0597] Data manipulation / calculation: Create notification messages based on required contact information.

[0598] Output: Emergency alert sent to family and medical facilities.

[0599] The above are the detailed processing steps of this system. At each step, the voice data and behavioral data of the elderly person are analyzed and processed, enabling appropriate responses and notifications to be given.

[0600] (Application example 2)

[0601] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0602] In an aging society, safe transportation, health management, and reducing loneliness for the elderly are important issues. In particular, elderly people living alone need a means of safe transportation that reduces loneliness in their daily lives. Therefore, a system is needed that can grasp the emotions and health status of elderly people in real time and provide the necessary support.

[0603] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for analyzing the elderly person's voice data using a voice recognition system, means for recognizing the elderly person's emotion using an emotion engine, means for generating a response based on the emotion information, and means for operating the navigation system of the autonomous vehicle. This makes it possible to grasp the emotions and health status of the elderly person in real time and support safe and secure travel.

[0604] "Elderly" refers to individuals who, as they grow older, require some form of assistance or care for their daily lives and activities.

[0605] The "means for receiving voice" is a device such as a microphone for capturing the speech of the elderly person.

[0606] The "means for converting speech to text" is a speech recognition system that has the function of converting speech data into text information.

[0607] The "means for analyzing text and generating appropriate responses" is a system that analyzes the converted text data using natural language processing and generates responses appropriate for the dialogue.

[0608] The "means for synthesizing a response using the voice of a registered family member" is a system having a function for synthesizing the generated text-based response using the voice of a family member.

[0609] The "means for providing synthesized voice to the elderly" is a device that plays synthesized voice data to the elderly through a speaker or the like.

[0610] The "means for analyzing voice data of an elderly person using a voice recognition system" is a voice recognition technology for analyzing voice data of an elderly person and generating an appropriate response.

[0611] The "means for recognizing the emotions of the elderly using an emotion engine" is a system that analyzes and extracts the emotional state of the elderly from voice data and text data.

[0612] The "means for generating a response based on emotional information" is a system that generates an appropriate emotional response based on the recognized emotional information.

[0613] "Means for operating the navigation system of an autonomous vehicle" refers to a system that guides the autonomous vehicle to its destination and controls its driving.

[0614] A "means for continuous conversation recording" is a device or software feature that continuously records conversations between the senior and the system.

[0615] "Means for transmitting conversation data to a server" refers to a technology for transferring recorded conversation data to a central server.

[0616] The "means for notifying the family of the analysis results" is a system that promptly notifies the family of the results of the analysis performed by the server.

[0617] The "means for detecting signs of dementia and emotional abnormalities" is a system that analyzes conversation data and emotional states to find abnormal patterns and early signs of dementia.

[0618] "Means for sending alerts" refers to a system that promptly sends a warning to relevant parties when an abnormality or emergency is detected.

[0619] This invention relates to an autonomous vehicle system for assisting the elderly, aiming to provide safe transportation and health management for the elderly, as well as to reduce feelings of loneliness. This system is realized by combining cloud computing technology, voice recognition technology, emotion recognition engines, and autonomous driving technology.

[0620] In order to implement the present invention, the following hardware and software are used.

[0621] Hardware used:

[0622] 1. Self-driving vehicle: A vehicle equipped with self-driving technology and equipped with built-in sensors, speakers, microphones, and cameras.

[0623] 2. Microphone: A device for capturing the speech of elderly people in the vehicle.

[0624] 3. Speaker: A device that provides synthesized speech to the elderly.

[0625] 4. Camera: A device used to monitor the movements of elderly people and detect any abnormalities.

[0626] Software used:

[0627] 1. Speech recognition system: A platform for converting the elderly's speech into text (e.g., Google Speech-to-Text API).

[0628] 2. Emotion recognition engine: A machine learning model that recognizes the emotions of the elderly and generates appropriate responses (e.g., IBM Watson Tone Analyzer).

[0629] 3. Speech synthesis system: A platform for synthesizing responses in the voices of family members (e.g., Amazon Polly).

[0630] 4. Autonomous vehicle control systems: APIs for controlling navigation in autonomous vehicles (e.g., Tesla Autopilot API).

[0631] A natural language description of the program:

[0632] Audio capture and transmission:

[0633] The server receives the voice data of the elderly person captured by the microphone in the autonomous vehicle, and the voice data is transmitted to the server in real time.

[0634] Voice Recognition:

[0635] The server uses a speech recognition system (Google Speech-to-Text API) to convert the received voice data into text data, which is a preparatory step for analyzing the conversation content.

[0636] Emotion Recognition and Analysis:

[0637] Using an emotion recognition engine (IBM Watson Tone Analyzer), the emotions of the elderly are recognized from the converted text data. Based on the recognized emotional information, an appropriate response is generated.

[0638] Response generation and serving:

[0639] The server uses the data necessary to generate a response and synthesizes a response based on the family member's voice using a speech synthesis system (Amazon Polly). This synthesized voice is then provided to the elderly person through the vehicle's speakers.

[0640] Navigation controls:

[0641] The system analyzes the elderly person's speech to identify destinations and instructions, and uses the autonomous vehicle control system (Tesla Autopilot API) to provide the necessary navigation. For example, if the system detects an utterance such as "I want to go to the hospital," it will set the navigation system to head to the nearest hospital.

[0642] Examples:

[0643] For example, if an elderly person gets into an autonomous vehicle and says, "I want to go to the park today," the following series of processes will take place.

[0644] 1. Voice capture: The microphone captures "I want to go to the park today."

[0645] 2. Speech recognition: The server converts the utterance into text: "I want to go to the park today."

[0646] 3. Emotion recognition: The emotion engine recognizes the emotion of "fun."

[0647] 4. Navigation: The vehicle sets the destination as the park, and a family member's voice guides them, saying, "We're heading to the park. Have fun."

[0648] 5. Anomaly detection: No emergency notification is required unless an anomaly is detected.

[0649] Example prompt sentence:

[0650] Please provide an example of creating a program for an autonomous driving system to assist the elderly, which converts voice data into text, recognizes emotions, and generates responses. Use a speech recognition system for speech recognition, an emotion recognition engine for emotion recognition, a speech synthesis system for speech synthesis, and an autonomous vehicle control system for autonomous driving control.

[0651] This invention uses an elderly assistance autonomous driving vehicle system to effectively realize safe and secure transportation, health management, and reduction of feelings of loneliness for the elderly.

[0652] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0653] Step 1:

[0654] The server receives the elderly person's voice data captured by the microphone in the autonomous vehicle. The input in this step is the elderly person's voice, and the output is the voice data stored on the server. Specifically, the microphone captures the voice as digital data and transmits it to the server in real time.

[0655] Step 2:

[0656] The server uses a speech recognition system (voice recognition system) to convert the received voice data into text data. The input in this step is the voice data received in step 1, and the output is the converted text data. In concrete terms, the server sends the voice data to the speech recognition system and receives the result as text.

[0657] Step 3:

[0658] The server uses an emotion recognition engine (emotion recognition engine) to recognize the emotions of the elderly from the text data. The input in this step is the text data obtained in step 2, and the output is the recognized emotion information. Specifically, the server inputs the text data into the emotion recognition engine and outputs the emotion information.

[0659] Step 4:

[0660] The server generates a response based on the emotional information and uses a speech synthesis system to synthesize the response in the voice of the family member. The input in this step is the emotional information obtained in step 3 and the generated response text, and the output is synthesized voice data. Specifically, the server generates a response text according to the emotional information, sends it to the speech synthesis system, and synthesizes it in the voice of the family member.

[0661] Step 5:

[0662] The server sends the synthesized voice data to the autonomous vehicle and provides it to the elderly through the speaker. The input in this step is the synthesized voice data obtained in step 4, and the output is the voice to be played to the elderly. Specifically, the server sends the synthesized voice data to the vehicle's audio system and plays it through the speaker.

[0663] Step 6:

[0664] The server analyzes the destination and instructions and sends them to the autonomous vehicle control system to operate the autonomous vehicle's navigation system. The input in this step is the text data obtained from step 2, and the output is the vehicle's destination setting and driving control. Specifically, the server extracts destination information from the elderly person's speech and sends instructions to the autonomous vehicle control system to perform navigation.

[0665] Step 7:

[0666] The server analyzes the elderly person's health condition and emotional information, and sends an emergency notification to family members or medical institutions if an abnormality is detected. The input in this step is the emotional information obtained in step 3 and other data collected in steps 1 and 2, and the output is the notification to be sent. Specifically, the server continuously monitors the data, and if an abnormality is detected, it triggers the notification system to send an emergency notification.

[0667] Through the above processing steps, the present invention can realize safe and secure transportation and health management as an elderly assistance autonomous driving vehicle system.

[0668] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0669] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0670] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0671] [Second embodiment]

[0672] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0673] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0674] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0675] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0676] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0677] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0678] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0679] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0680] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0681] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0682] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0683] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0684] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time.

[0685] Feature Description

[0686] 1. Ability to talk to the elderly using the voices of children and grandchildren

[0687] Subject: Server

[0688] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[0689] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[0690] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[0691] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[0692] The device plays the received audio to the elderly person through a speaker.

[0693] Specific examples

[0694] An elderly person speaks to the device, asking, "What should I eat today?"

[0695] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[0696] 2. Ability to record conversations with elderly people and provide information to their families

[0697] Subject: Device

[0698] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[0699] The server converts the received conversation data into text and analyzes the content.

[0700] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[0701] Specific examples

[0702] An elderly person says, "My shoulder has been hurting lately."

[0703] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[0704] 3. Alerts for dementia and depression

[0705] Subject: Server

[0706] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[0707] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[0708] Specific examples

[0709] Elderly people say, "I don't know what I'm living for."

[0710] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[0711] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[0712] Subject: Device

[0713] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[0714] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[0715] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[0716] Specific examples

[0717] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0718] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[0719] 5. Fall detection and emergency contact function

[0720] Subject: Camera

[0721] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[0722] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[0723] Specific examples

[0724] An elderly person falls in the living room.

[0725] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[0726] As described above, the present invention is a system that comprehensively supports the health management of elderly people, reduces feelings of loneliness, and alleviates anxiety for their families. This system allows elderly people to live their daily lives with peace of mind, and allows their families to follow up in a timely manner.

[0727] The processing flow will be explained below.

[0728] A function that allows you to talk to the elderly using the voices of children and grandchildren

[0729] Step 1:

[0730] The user (elderly person) speaks into the terminal.

[0731] As a concrete example, say, "What shall we eat today?"

[0732] Step 2:

[0733] The terminal receives the elderly person's voice and transmits the voice data to the server.

[0734] The audio data is transferred to the server in real time.

[0735] Step 3:

[0736] The server converts the received voice data into text.

[0737] A speech recognition algorithm converts the speech data into a string of characters.

[0738] Step 4:

[0739] The server analyzes the converted text data and generates an appropriate response.

[0740] Use natural language processing (NLP) techniques to understand the context of text data.

[0741] Step 5:

[0742] The server synthesizes a response based on voice samples from family members.

[0743] A response message is synthesized using the voice of a pre-registered family member.

[0744] Step 6:

[0745] The server transmits the synthesized voice data to the terminal.

[0746] The generated voice data is sent back to the elderly person.

[0747] Step 7:

[0748] The device plays the synthesized speech through the speaker.

[0749] The elderly person hears a family member respond, "Maybe fish would be good?"

[0750] A function that records conversations with elderly people and provides information to their families

[0751] Step 1:

[0752] The device constantly records conversations with the elderly.

[0753] Record all your everyday conversations.

[0754] Step 2:

[0755] The recorded conversation data is periodically sent to a server.

[0756] For example, a data packet is sent every hour.

[0757] Step 3:

[0758] The server converts the received voice data into text.

[0759] Convert the speech into text using voice recognition technology.

[0760] Step 4:

[0761] The server analyzes the converted text data.

[0762] Natural language processing technology is used to extract important information and anomalies.

[0763] Step 5:

[0764] Based on the analysis results, the information to be notified to the family is determined.

[0765] If an anomaly is detected, detailed information about it will also be included.

[0766] Step 6:

[0767] The server sends the analysis results to the family.

[0768] Notifications are given in real time.

[0769] Examples:

[0770] An elderly person says, "My shoulder has been hurting lately."

[0771] The server analyzes the information and sends a notification to the family about the "shoulder pain."

[0772] A feature that sends alerts for dementia and depression

[0773] Step 1:

[0774] The server periodically analyzes conversation data with the elderly.

[0775] The analysis targets daily conversation logs.

[0776] Step 2:

[0777] The server detects changes in grammatical structure and sentiment.

[0778] It uses natural language processing techniques and sentiment analysis algorithms.

[0779] Step 3:

[0780] If an anomaly is detected, the server generates an alert.

[0781] Identify signs of dementia and emotional abnormalities.

[0782] Step 4:

[0783] The server generates alerts and sends them to family members and medical institutions.

[0784] The alert is sent as an emergency notification.

[0785] Examples:

[0786] Elderly people say, "I don't know what I'm living for."

[0787] The server performs emotion analysis, detects signs of depression, and sends emergency notifications.

[0788] Functions to prevent forgetting to take medicine and assist with meals and shopping

[0789] Step 1:

[0790] The server sets medication times and meal schedules.

[0791] Set a daily reminder.

[0792] Step 2:

[0793] The device will send a reminder notification at the set time.

[0794] A notification will play saying "Time for your medicine."

[0795] Step 3:

[0796] When the elderly person responds with a voice of confirmation, the device records the information and sends it to the server.

[0797] "Thanks, I drank it," he replied.

[0798] Step 4:

[0799] The server notifies the family of the reminder information.

[0800] Notifies you that confirmation of missed dose has been completed.

[0801] Step 5:

[0802] The server learns the elderly person's lifestyle patterns and generates a shopping list.

[0803] Make a list of the items you need.

[0804] Step 6:

[0805] The device notifies the elderly person of the shopping list.

[0806] Announce the list contents by voice.

[0807] Examples:

[0808] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0809] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[0810] Fall detection and emergency contact function

[0811] Step 1:

[0812] The camera monitors the footage at all times.

[0813] Monitor overall behavior.

[0814] Step 2:

[0815] The camera detects an abnormality.

[0816] For example, a falling motion is detected.

[0817] Step 3:

[0818] Video data is sent to the server in real time.

[0819] Send data immediately.

[0820] Step 4:

[0821] The server identifies the fall and generates an emergency alert.

[0822] Prepare notifications for medical institutions and families.

[0823] Step 5:

[0824] The server sends emergency alerts to family and medical institutions.

[0825] Encourage a rapid response.

[0826] Examples:

[0827] An elderly person falls in the living room.

[0828] The camera detects an abnormality and notifies the server.

[0829] The server sends emergency alerts to family and medical institutions.

[0830] These are the specific steps of the program processing, which will enable health management for elderly people living alone, reduce loneliness, and provide peace of mind to their families.

[0831] Example 1

[0832] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0833] There is a need for health management for the elderly, reducing loneliness, and improving family security. However, systems that enable appropriate communication and health monitoring for elderly people living alone have not yet been fully developed. When elderly people live in isolated situations, the lack of daily communication and the inability to detect health abnormalities early are problems. In particular, there is a lack of means to detect early signs of dementia and depression and provide appropriate treatment.

[0834] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0835] In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for saving voice samples provided by the family member and extracting features, and natural language processing means for analyzing the elderly person's speech and generating a response related to life support. This enables natural communication with the elderly, health management, reduction of loneliness, and early detection of signs of dementia and depression.

[0836] The "means for receiving the voice of the elderly person" is a function for acquiring the voice uttered by the elderly person using an input device such as a microphone.

[0837] The "means for converting received voice into text" is a function that converts voice data into corresponding text data using voice recognition technology.

[0838] "Means for analyzing the converted text and generating an appropriate response" refers to a function that uses natural language processing (NLP) technology to analyze text data and generate an appropriate response based on the content.

[0839] The "means for synthesizing the response with the voice of a registered family member" is a technology for converting text data into voice data using the voice characteristics of a designated family member.

[0840] The "means for providing synthesized voice to the elderly" is a function for playing back synthesized voice to the elderly using a voice output device such as a speaker.

[0841] The "means for saving voice samples provided by family members and extracting features" is a function for saving voice data provided by family members and extracting voice features.

[0842] "Natural language processing means for analyzing elderly people's statements and generating responses related to life support" is a technology that analyzes the content of elderly people's statements and generates information and actions necessary for life support based on the analysis results.

[0843] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. This system is realized by integrating speech recognition, natural language processing, and speech synthesis technologies.

[0844] Hardware and software used

[0845] Hardware

[0846] High-Precision Microphone

[0847] speaker

[0848] camera

[0849] server

[0850] Devices (smartphones, tablets, PCs, etc.)

[0851] software

[0852] Speech recognition: A speech recognition API for converting voice data into text (e.g., Google Cloud Speech-to-Text API)

[0853] Natural Language Processing (NLP): NLP models (e.g., AWS Comprehend) to analyze text data and generate appropriate responses.

[0854] Speech synthesis: A speech synthesis API to convert text to speech (e.g. IBM Watson Text-to-Speech API)

[0855] Database: A database for storing audio samples and analysis data (e.g., Google Cloud Storage)

[0856] Explanation of the specific functions of the system

[0857] 1. Ability to talk to the elderly using the voices of children and grandchildren

[0858] Subject: Server, Terminal, User

[0859] The server stores voice samples provided by family members in a database. It also extracts voice features using a machine learning model. When the user (elderly person) speaks into the device, the device captures the voice data and sends it to the server. The server converts the received voice data into text using the Google Cloud Speech-to-Text API and analyzes the text data using a natural language processing (NLP) model. An appropriate response is generated, and speech is synthesized using the family member's voice using the IBM Watson Text-to-Speech API, and the speech data is sent to the device. The device then plays the received audio to the elderly person through a speaker.

[0860] Specific examples

[0861] An elderly person speaks to the device, asking, "What should I eat today?"

[0862] The server synthesizes the family member's voice saying, "Maybe fish would be good?" and plays it back to the elderly person via the device.

[0863] Example prompt

[0864] "When you say to a device, 'What should I eat today?' how does the server process the data and generate the voice response?"

[0865] 2. Ability to record conversations with elderly people and provide information to their families

[0866] Subject: Terminal, Server

[0867] The device continuously records conversations with the elderly. The recorded conversation data is periodically sent to a server, which converts the voice data into text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing (NLP) model, and if particularly important information or anomalies are detected, the server notifies the family using the Twilio API.

[0868] Specific examples

[0869] An elderly person says, "My shoulder has been hurting lately."

[0870] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[0871] Example prompt

[0872] "When the audio data recorded on the device is sent to the server, how will family members be notified if an abnormality is detected?"

[0873] 3. Alerts for dementia and depression

[0874] Subject: Server

[0875] The server periodically converts conversation data with the elderly person into text using the Google Cloud Speech-to-Text API, then analyzes the text data using a natural language processing (NLP) model. If signs of dementia or abnormal emotions are detected, an alert is generated and notifies family members and medical institutions using the Twilio API.

[0876] Specific examples

[0877] Elderly people say, "I don't know what I'm living for."

[0878] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[0879] Example prompt

[0880] "If an elderly person makes a statement that indicates signs of depression, what analysis does the server perform and how does it generate a notification?"

[0881] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[0882] Subject: Terminal, Server

[0883] The device uses the Google Calendar API to manage medication times and meal schedules, and sends a reminder at the set time: "It's time for your medicine." When the user (elderly person) responds with a voice confirmation, the voice is sent to the server and converted into text using the Google Cloud Speech-to-Text API. The server then notifies the family of the confirmation information. The server also learns the elderly person's lifestyle patterns, generates a shopping list, and provides it via voice from the device.

[0884] Specific examples

[0885] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[0886] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[0887] Example prompt

[0888] "How can we remind and confirm elderly people to prevent them from forgetting to take their medication, and notify their families of this information?"

[0889] 5. Fall detection and emergency contact function

[0890] Subject: camera, server

[0891] The camera constantly monitors the video and runs an algorithm to detect anomalies. If an elderly person falls, the camera detects the movement and sends the video data in real time to a server. The server then confirms the fall and uses the Twilio API to send an emergency notification to family members and medical institutions.

[0892] Specific examples

[0893] An elderly person falls in the living room.

[0894] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[0895] Example prompt

[0896] "How does the fall detection feature work and how does it send emergency notifications after detection?"

[0897] As described above, this system recognizes the voice of the elderly person, generates an appropriate response using natural language processing, and then uses voice synthesis technology to respond in the voice of a family member, thereby providing comprehensive support for managing the elderly person's health, reducing feelings of loneliness, and providing a sense of security to their family.

[0898] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0899] Specific processing steps of the system program

[0900] Feature 1: Ability to talk to the elderly using the voices of children and grandchildren

[0901] Subject: Server, Terminal, User

[0902] Step 1:

[0903] Subject: Server

[0904] Voice samples provided by family members are stored in a database.

[0905] (Input) Voice samples recorded by family members

[0906] (Processing) The server receives the audio samples and stores them in a database.

[0907] (Output) Voice samples of family members stored in the database

[0908] Specific operation: The server receives the voice data and stores it in a database.

[0909] Step 2:

[0910] Subject: Server

[0911] The voice sample is analyzed using a machine learning model to extract features.

[0912] (Input) Audio samples stored in the database

[0913] (Processing) Apply a voice feature extraction algorithm to extract features

[0914] (Output) Extracted feature data

[0915] Specific operation: The server applies a feature extraction algorithm to the voice sample and stores the features in a database.

[0916] Step 3:

[0917] Subject: User

[0918] The user speaks into the terminal.

[0919] (Input) User's voice

[0920] (Processing) High-precision microphone captures audio

[0921] (Output) Captured audio data

[0922] Specific operation: The user speaks to the device, "What should I eat today?"

[0923] Step 4:

[0924] Subject: Device

[0925] The terminal transmits the voice data to the server.

[0926] (Input) Captured audio data

[0927] (Processing) Send the audio data to the server

[0928] (Output) Audio data received by the server

[0929] Specific operation: The device sends the captured audio data to a server via the Internet.

[0930] Step 5:

[0931] Subject: Server

[0932] Converts audio data into text.

[0933] (Input) Audio data received by the server

[0934] (Processing) Convert speech to text using a speech recognition API (e.g., Google Cloud Speech-to-Text API)

[0935] (Output) Converted text data

[0936] What happens: The server uses the Google Cloud Speech-to-Text API to convert the speech "What should we eat today?" into text.

[0937] Step 6:

[0938] Subject: Server

[0939] Analyze text data using an NLP model and generate appropriate responses.

[0940] (Input) Converted text data

[0941] (Processing) Analyze the text using an NLP model (e.g. AWS Comprehend) and generate an appropriate response

[0942] (Output) The generated response text

[0943] Specific behavior: The server generates a response to the question "What should we eat?" with "Maybe fish would be good?"

[0944] Step 7:

[0945] Subject: Server

[0946] The response text is converted into speech using the speech synthesis API.

[0947] (Input) Generated response text

[0948] (Processing) Convert text to speech using a speech synthesis API (e.g. IBM Watson Text-to-Speech API)

[0949] (Output) Synthesized voice data

[0950] Specific behavior: The server converts the text "Maybe fish would be good?" into speech.

[0951] Step 8:

[0952] Subject: Server

[0953] The synthesized voice data is transmitted to the terminal.

[0954] (Input) Synthesized voice data

[0955] (Processing) Send audio data to the terminal

[0956] (Output) Audio data received on the device

[0957] Specific operation: The server sends the synthesized voice data to the terminal.

[0958] Step 9:

[0959] Subject: Device

[0960] The device plays the audio.

[0961] (Input) Audio data received by the device

[0962] (Processing) Playing audio through speakers

[0963] (Output) Audio played to the elderly

[0964] Specific operation: The device plays a voice message from the speaker saying, "Maybe a fish would be good?"

[0965] You can add specific processing steps below for other features as well, but you can distinguish the steps for each feature based on this format.

[0966] If you have steps for other features below, please list them in a similar format.

[0967] (Application example 1)

[0968] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0969] In modern society, the number of elderly people living alone is increasing, and health management and reducing feelings of loneliness have become social issues. Furthermore, when elderly people living alone use self-driving vehicles, ensuring safety and communication with their families are important. However, conventional systems do not adequately provide comprehensive solutions to these issues. Therefore, there is a need for a system that integrates safe and secure transportation means for the elderly with support for their daily lives.

[0970] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0971] In this invention, the server includes means for receiving the elderly person's voice, means for converting the voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the family member's voice, means for providing synthesized voice, means for setting a destination for the autonomous vehicle and providing route guidance, means for monitoring the driving situation and detecting abnormalities, and means for notifying the family member in real time when an abnormality occurs. This allows the elderly person to receive route guidance from the family member's voice when using the autonomous vehicle, improving safety and enabling the family member to understand the elderly person's situation in real time.

[0972] · An "elderly person" is an individual who, due to their advanced age, requires assistance in daily living.

[0973] "Means for receiving voice" refers to a device or method for recording and capturing the voice uttered by the elderly person and inputting it into the system.

[0974] "Means for converting speech to text" means the process or technology that converts received speech data into corresponding text data using speech recognition technology.

[0975] "Means of analyzing text" refers to natural language processing techniques used to understand meaning and intent based on converted text data.

[0976] - "Means for generating a response" refers to an algorithm or program that generates an appropriate response based on the analysis results.

[0977] "Means for synthesizing using family members' voices" refers to a method of reproducing the generated response using voice synthesis technology using pre-registered family members' voice data.

[0978] "Means for providing synthesized speech" means a speaker or other audio playback device that delivers synthesized speech to the senior citizen.

[0979] "Means for setting a destination in an autonomous vehicle" refers to the methods and technologies for inputting and setting the destination desired by the elderly person into an autonomous vehicle.

[0980] "Means for providing directions" refers to systems and technologies that allow autonomous vehicles to provide directions to seniors toward a set destination.

[0981] "Means for monitoring the driving situation" refers to sensors, cameras, and algorithms that monitor the driving situation in real time while the autonomous vehicle is driving to ensure safety.

[0982] "Means for detecting abnormalities" refers to techniques and methods for recognizing and identifying hazards and abnormal events that may occur during operation.

[0983] "Means of notifying family members in real time" refers to communication technologies and methods for immediately notifying family members of the situation when an abnormality is detected.

[0984] To implement this invention, it is necessary to build a system that receives the voice of the elderly person and synthesizes it with the voices of family members. This system uses the following hardware and software.

[0985] Hardware used

[0986] 1. Smartphone: Used to record the elderly person's voice and send the recording data to the server.

[0987] 2. Self-driving vehicles: Used as transportation for the elderly, with functions such as route guidance and safety monitoring.

[0988] 3. Speaker: Used as a device to transmit synthesized voices of family members to the elderly.

[0989] Software used

[0990] 1. pyaudio: A library for recording the voices of elderly people.

[0991] 2. wave: A library for saving and playing recorded audio.

[0992] 3. transformers (Wav2Vec2): A library for converting audio to text.

[0993] 4. gTTS: A library for synthesizing text in the voices of family members.

[0994] 5. smtplib: A library for implementing notification functions.

[0995] Data processing and calculation

[0996] The server first receives the elderly person's voice and records it using the pyaudio library. The recorded voice data is then temporarily saved using the wave library. The voice data is then converted to text using the Wav2Vec2 model in transformers. The converted text data is analyzed using natural language processing techniques to generate an appropriate response. The generated response is then synthesized in the voice of a family member using the gTTS library and provided to the elderly person through a speaker as synthesized speech.

[0997] If an anomaly is detected or if specific keywords are found, a real-time notification is sent to family members using the smtplib library, containing the analysis results and the elderly person's current condition.

[0998] Specific examples

[0999] For example, if an elderly person says, "Shall I confirm your next hospital appointment?", the system works as follows: First, the smartphone records this speech and sends it to the server. The server converts the recorded speech into text using the Wav2Vec2 model and analyzes its content. An appropriate response is generated, such as "Yes, Grandma. Your next hospital appointment is tomorrow at 10:00 AM," using a synthesized voice from a family member's voice, and this is played to the elderly through the speaker.

[1000] Prompt Sentence Examples

[1001] 1. "Grandpa, should I confirm your next doctor's appointment?"

[1002] 2. "Grandma, have you taken your morning medicine?"

[1003] 3. "Where is Grandpa going? Is he driving safely?"

[1004] This will allow elderly people to use self-driving vehicles with peace of mind, and will also enable family members to keep track of the elderly person's situation in real time.

[1005] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1006] Step 1:

[1007] The smartphone records the elderly person's voice. The device uses the pyaudio library to record 10 seconds of voice data uttered by the elderly person. The input here is the elderly person's voice, and the output is the recorded voice data (audio file).

[1008] Step 2:

[1009] The recorded audio data is sent to the server. The device uses the wave library to save the audio file and uploads it to the server. The input is the recorded audio file, and the output is the audio file saved on the server.

[1010] Step 3:

[1011] The server converts the audio data to text. The server uses the Wav2Vec2 model from transformers to convert the audio file to text. The input is the audio file, and the output is the converted text.

[1012] Step 4:

[1013] The server analyzes the text data and generates an appropriate response. Natural language processing technology is used to analyze the meaning and intent of the converted text data and generate an appropriate response text. The input is the text data, and the output is the generated response text.

[1014] Step 5:

[1015] The server synthesizes the response text in the voice of the family member. The server uses the gTTS library to synthesize the generated response text in the voice of the pre-registered family member and create an audio file. The input is the response text, and the output is the synthesized audio file.

[1016] Step 6:

[1017] The synthesized voice is provided to the elderly through a speaker. The server sends the synthesized voice to the terminal, and the terminal plays the voice through the speaker. The input is the synthesized voice file, and the output is the voice played to the elderly.

[1018] Step 7:

[1019] Monitors the driving situation and detects abnormalities. Autonomous vehicles use sensors and cameras to monitor the driving situation in real time to ensure safety. The input is real-time data during driving, and the output is analysis results and abnormality detection signals.

[1020] Step 8:

[1021] If an abnormality is detected, a notification is sent to family members in real time. The server uses the smtplib library to notify family members of the situation by email when it receives an abnormality detection signal. The input is the abnormality detection signal, and the output is the alert notification sent to family members.

[1022] Specific example prompts

[1023] "Grandpa, should I confirm your next doctor's appointment?"

[1024] "Grandma, have you taken your morning medicine?"

[1025] "Where is Grandpa going? Are you driving safely?"

[1026] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1027] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and reduce feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it achieves more effective responses and anomaly detection.

[1028] Feature Description

[1029] 1. Ability to talk to the elderly using the voices of children and grandchildren

[1030] Subject: Server

[1031] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[1032] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[1033] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[1034] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[1035] The device plays the received audio to the elderly person through a speaker.

[1036] Specific examples

[1037] An elderly person speaks to the device, asking, "What should I eat today?"

[1038] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[1039] 2. Ability to record conversations with elderly people and provide information to their families

[1040] Subject: Device

[1041] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[1042] The server converts the received conversation data into text and analyzes the content.

[1043] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[1044] Specific examples

[1045] An elderly person says, "My shoulder has been hurting lately."

[1046] The server analyzes the information and sends a notification about the "shoulder pain" to the family, who can then plan a visit.

[1047] 3. Alerts for dementia and depression

[1048] Subject: Server

[1049] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[1050] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[1051] Specific examples

[1052] Elderly people say, "I don't know what I'm living for."

[1053] The server performs emotion analysis, detects signs of depression, and sends emergency notifications to family members and medical institutions.

[1054] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[1055] Subject: Device

[1056] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[1057] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[1058] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[1059] Specific examples

[1060] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[1061] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[1062] 5. Fall detection and emergency contact function

[1063] Subject: Camera

[1064] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[1065] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[1066] Specific examples

[1067] An elderly person falls in the living room.

[1068] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[1069] Features of inventions that combine emotion engines

[1070] 1. Emotion recognition and response generation using an emotion engine

[1071] Subject: Server

[1072] The server uses an emotion engine to recognize emotions from the elderly person's voice data, extracts the elderly person's emotions from the voice data, and generates a response based on that information.

[1073] Depending on the recognized emotion, the server generates an appropriate response and provides it to the elderly person as a synthesized voice.

[1074] Specific examples

[1075] An elderly person says, "I feel a little lonely today."

[1076] The server uses an emotion engine to recognize the emotion "loneliness" and generates a response such as "I see you're feeling lonely. Let's think of something we can do together."

[1077] 2. Anomaly detection and notification using emotion engine

[1078] Subject: Server

[1079] If the server detects negative emotions in an elderly person using the emotion engine, it generates an alert and notifies family members and relevant organizations.

[1080] If an anomaly is detected, a notification system will be activated to ensure a rapid response.

[1081] Specific examples

[1082] An elderly person says, "Nothing is fun anymore."

[1083] The server uses an emotion engine to detect emotions such as "despair" and "deep sadness" and sends emergency notifications to family members and medical institutions.

[1084] The above is a specific embodiment of the present invention that combines an emotion engine. This makes it possible to more accurately grasp the emotional state of elderly people living alone and respond appropriately. Furthermore, since it can respond sensitively to changes in emotions, it can also support the mental health of the elderly.

[1085] The processing flow will be explained below.

[1086] Emotion recognition and response generation using an emotion engine

[1087] Step 1:

[1088] The user (elderly person) speaks into the terminal.

[1089] As a specific example, say, "I feel a little lonely today."

[1090] Step 2:

[1091] The terminal receives the elderly person's voice and transmits the voice data to the server.

[1092] Step 3:

[1093] The server converts the received voice data into text.

[1094] A speech recognition algorithm converts the speech data into a string of characters.

[1095] Step 4:

[1096] The server passes the converted text data and voice data to the emotion engine.

[1097] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[1098] Step 5:

[1099] The server generates an appropriate response based on the emotional data recognized by the emotion engine.

[1100] Using natural language processing technology, it generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[1101] Step 6:

[1102] The server generates a response that is then synthesized into the family member's voice.

[1103] A response message is synthesized using the voice of a pre-registered family member.

[1104] Step 7:

[1105] The server transmits the synthesized voice data to the terminal.

[1106] The generated voice data is sent back to the elderly person.

[1107] Step 8:

[1108] The device plays the synthesized speech through the speaker.

[1109] The elderly person hears a family member respond, "I see you're feeling lonely. Let's think of something we can do together."

[1110] Anomaly detection and notification using emotion engine

[1111] Step 1:

[1112] The user (elderly person) speaks into the terminal.

[1113] For example, say, "Nothing is fun anymore."

[1114] Step 2:

[1115] The terminal receives the elderly person's voice and transmits the voice data to the server.

[1116] Step 3:

[1117] The server converts the received voice data into text.

[1118] A speech recognition algorithm converts the speech data into a string of characters.

[1119] Step 4:

[1120] The server passes the converted text data and voice data to the emotion engine.

[1121] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[1122] Step 5:

[1123] The emotion engine detects negative emotions such as "despair" and "deep sadness."

[1124] The emotion engine passes these abnormal emotion data to the server.

[1125] Step 6:

[1126] The server generates alerts based on the emotion engine data.

[1127] Prepare an abnormality detection message and notify family members and relevant organizations.

[1128] Step 7:

[1129] The server sends emergency notifications to family and medical institutions.

[1130] Alerts such as "elderly people are feeling deep sadness" will be sent via email or app notification.

[1131] Step 8:

[1132] Family and medical institutions will be notified and take appropriate action.

[1133] Schedule visits and medical consultations.

[1134] These are the processing steps of the system that combines the emotion engine, which makes it possible to grasp the emotional state of the elderly more accurately and take the necessary measures quickly.

[1135] Example 2

[1136] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1137] For elderly people living alone, reducing loneliness and managing their health are major challenges. It is particularly difficult to understand changes in the elderly's moods and health status when family members live far away. Early detection of dementia and depression and prevention of falls are also important issues. Furthermore, elderly people need assistance with daily activities such as remembering to take their medication, eating, and shopping. To solve these challenges, a voice-responsive system that includes emotion recognition is needed.

[1138] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for continuously recording conversations with the elderly person, means for transmitting the recorded conversation data to the server, means for analyzing the transmitted conversation data and notifying the family member of the analysis result, means for analyzing the conversation data with the elderly person and detecting signs of dementia or emotional abnormalities, means for transmitting an alert when signs of dementia or emotional abnormalities are detected, means for learning the elderly person's lifestyle patterns and setting reminders as necessary, means for notifying the family member based on the reminder, means for recording confirmation of the notified content and notifying the family member, means for detecting the elderly person's fall, means for transmitting an emergency notification when a fall is detected, means for recognizing the elderly person's emotion from the voice data, means for generating an appropriate response according to the recognized emotion, means for detecting the elderly person's negative emotion, and means for transmitting an alert to the family member or relevant organization when the emotion is detected. This reduces the elderly person's sense of loneliness and enables health management. It can also enable early detection of dementia and depression, prevent falls, prevent people from forgetting to take their medication, and provide assistance with meals and shopping.

[1139] The "means for receiving voice" is a device or system that captures the voice spoken by the elderly person and receives it as data.

[1140] A "speech-to-text means" is software or an algorithm that converts received speech data into text data.

[1141] The "means for analyzing text and generating an appropriate response" refers to a natural language processing engine or algorithm for analyzing text data converted from speech and generating an appropriate response.

[1142] The "means for synthesizing a response using the voice of a registered family member" refers to software or an algorithm for synthesizing the generated response using voice data of a registered family member in advance.

[1143] "Means for providing synthesized voices to the elderly" refers to devices or systems that transmit synthesized voices of family members to the elderly through speakers or communication devices.

[1144] "Means for continuously recording conversations" refers to equipment or systems that continuously record conversations with elderly people and store them as data.

[1145] The "means for transmitting recorded conversation data to a server" refers to a communication device or system for transmitting recorded conversation data to a server via a network.

[1146] "Means for analyzing the transmitted conversation data and notifying the family of the analysis results" refers to software or a communication system for analyzing the conversation data transmitted to the server and notifying the family of the analysis results.

[1147] "Means for detecting signs of dementia and emotional abnormalities" refers to algorithms and software that analyze conversation data with elderly people to detect signs of dementia and emotional abnormalities.

[1148] "Means for sending alerts when signs of dementia or abnormal emotions are detected" refers to a communication system or software for sending alerts to family members or relevant organizations when signs of dementia or abnormal emotions are detected.

[1149] "Means for learning lifestyle patterns and setting reminders" refers to algorithms or software that learn the lifestyle patterns of elderly people through data analysis and set reminders at the necessary times.

[1150] The "means for notifying based on a reminder" refers to a device or system for notifying the elderly based on the set reminder.

[1151] "Means for recording confirmation of the notified content and notifying the family" refers to a communication system or software for recording the notification confirmation from the elderly person and conveying the content to the family.

[1152] "Means for detecting falls in the elderly" refers to systems and algorithms that use cameras and sensors to detect falls in the elderly in real time.

[1153] The "means for sending an emergency notification when a fall is detected" is a communication system for sending an emergency notification to family members and medical institutions when a fall of an elderly person is detected.

[1154] "Means for recognizing the emotions of elderly people from voice data" refers to software or algorithms that use an emotion engine to extract and recognize emotions from the voice data of elderly people.

[1155] The "means for generating an appropriate response based on the recognized emotions" refers to a natural language processing engine or software that generates an appropriate response based on the emotions of the elderly person.

[1156] The "means for detecting negative emotions" refers to emotion analysis algorithms or software for identifying and detecting negative emotions from elderly people's voice data.

[1157] "Means for sending an alert to family members or relevant organizations when negative emotions are detected" refers to a communication system or software for quickly sending an alert to family members or relevant organizations when negative emotions are detected.

[1158] The present invention is a system that aims to manage the health of elderly people living alone and reduce feelings of loneliness through communication with family members. Furthermore, by combining it with an emotion engine, more effective responses and anomaly detection become possible. This system has multiple functions for connecting elderly people with their families. Specific embodiments of the system are described below.

[1159] System configuration

[1160] This system consists of three main components: a server, a terminal, and a user. The server processes and analyzes data, the terminal acts as an interface with the elderly, and the user refers to the elderly and their family members.

[1161] Hardware and software used

[1162] Hardware: Smart speakers such as Amazon Echo and Google Home can be used as devices, and cameras such as Nest Cam can be used.

[1163] Software: On the server side, we use Google Speech-to-Text for speech recognition, NLTK and Spacy for natural language processing, Google Text-to-Speech API for speech synthesis, and IBM Watson Tone Analyzer for sentiment analysis.

[1164] Program processing flow

[1165] 1. Saving and analyzing audio samples

[1166] The server receives voice samples provided by family members and stores them in a database. The stored voice samples are analyzed using machine learning models (e.g., TensorFlow) to extract features. This analysis lays the foundation for delivering the voices of family members to the elderly.

[1167] 2. Audio capture and transmission

[1168] When an elderly person speaks into the device, the device captures the voice and transmits it to a server in real time. For example, Amazon Echo captures the elderly person's voice in high quality and transmits the data to a server over the Internet.

[1169] 3. Speech data text conversion and analysis

[1170] The server converts the received voice data into text using Google Speech-to-Text, which is then analyzed using natural language processing tools such as NLTK or Spacy. Based on this analysis, an appropriate response is generated.

[1171] 4. Speech synthesis and transmission

[1172] The generated response is synthesized using the Google Text-to-Speech API using the family member's voice. The synthesized voice data is sent to the device, which plays it back to the elderly person through the device's speaker. For example, if an elderly person asks, "What should we eat today?", the server synthesizes a family member's voice saying, "Maybe fish would be good?" and plays it back on the device.

[1173] 5. Recording and analyzing conversations

[1174] The device constantly records conversations with the elderly person and periodically sends the data to a server. The server then converts the conversation data into text using Google Speech-to-Text and detects important information and anomalies. The detection results are then notified to family members in real time. For example, if an elderly person says, "My shoulder has been hurting lately," that information is immediately notified to the family.

[1175] 6. Sentiment Analysis and Response

[1176] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize emotions from the elderly person's voice. Based on the recognized emotion, it generates an appropriate response and provides it to the elderly person as synthesized speech. For example, if an elderly person says, "I feel a little lonely today," the server generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[1177] 7. Fall detection and emergency notification

[1178] The camera constantly monitors the elderly person's movements and runs an algorithm (e.g., OpenCV) to detect abnormal behavior (e.g., falls). If a fall is detected, the server immediately sends an emergency notification to family members or medical institutions. For example, if an elderly person falls in the living room, the camera detects the movement and notifies the server in real time. The server confirms this and immediately sends an emergency alert.

[1179] Specific prompt examples

[1180] A specific example of a prompt sentence to input into the generative AI model using this system is as follows:

[1181] "An elderly person asks, 'What shall we eat today?' A child's voice responds, 'Maybe fish would be good?'"

[1182] "An elderly person says, 'I feel a little lonely today.' Use an emotion engine to generate an appropriate response."

[1183] The above is a specific embodiment of the present invention. This system can reduce the sense of loneliness felt by the elderly and help them manage their health. It can also provide early detection of dementia and depression, prevent falls, help with forgetting to take medication, and assist with meals and shopping.

[1184] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1185] Processing steps of this system's program

[1186] Step 1: Saving and analyzing audio samples

[1187] Subject: Server

[1188] Input: Voice sample provided by family member (e.g., audio file).

[1189] What it does: The server receives voice samples uploaded by family members and stores them in a database.

[1190] Data processing / computation: The server uses machine learning models (e.g., TensorFlow) to extract and analyze voice features.

[1191] Output: A feature-analyzed audio profile.

[1192] Step 2: Capture and send audio

[1193] Subject: Device

[1194] Input: Speech produced by an elderly person.

[1195] Specific operation: When an elderly person speaks into a device (e.g., a smart speaker), the device captures the voice.

[1196] Data processing / calculation: The device transmits the captured voice data to the server in real time.

[1197] Output: The audio data sent to the server.

[1198] Step 3: Converting audio data into text and analyzing it

[1199] Subject: Server

[1200] Input: The audio data sent to the server.

[1201] Specific operation: The server converts the received voice data into text using voice recognition software (e.g., Google Speech-to-Text).

[1202] Data processing / calculation: The converted text data is analyzed using natural language processing (NLP) tools (e.g., NLTK or Spacy).

[1203] Output: Parsed text data.

[1204] Step 4: Speech synthesis and transmission

[1205] Subject: Server

[1206] Input: Parsed text data.

[1207] What it does: The server synthesizes a response in the family member's voice based on the analyzed text (e.g., Google Text-to-Speech API).

[1208] Data processing / calculation: Voice synthesis using family voice profiles.

[1209] Output: Synthesized speech data.

[1210] Step 5: Play audio

[1211] Subject: Device

[1212] Input: Synthesized speech data.

[1213] Specific operation: The device plays the received audio to the elderly person through the speaker.

[1214] Output: Audio provided to the senior.

[1215] Step 6: Record and send the conversation

[1216] Subject: Device

[1217] Input: Conversation with an elderly person.

[1218] Specific operation: The device constantly records conversations with the elderly and periodically sends the data to a server.

[1219] Data processing / calculation: Converting the audio data into the format required to send it to the server.

[1220] Output: Conversation data sent to the server.

[1221] Step 7: Analyze conversation data and notify

[1222] Subject: Server

[1223] Input: Conversation data sent to the server.

[1224] What it does: The server uses Google Speech-to-Text to convert the conversation data into text and detects important information and anomalies.

[1225] Data processing / calculation: Generate information to notify family members based on the converted and analyzed data.

[1226] Output: Analysis results communicated to family members.

[1227] Step 8: Sentiment Analysis and Response Generation

[1228] Subject: Server

[1229] Input: Speech data of elderly people.

[1230] Specific operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize emotions from the elderly person's voice.

[1231] Data processing / computation: Generate appropriate responses based on recognized emotions and provide them as synthesized speech.

[1232] Output: Synthetic voice response provided to the senior.

[1233] Step 9: Fall Detection and Emergency Notification

[1234] Subject: Camera

[1235] Enter: Elderly Movement.

[1236] Specific operation: The camera constantly monitors the elderly person's movements and runs an algorithm (e.g., OpenCV) to detect abnormal movements (e.g., falls).

[1237] Data processing / calculation: When a fall is detected, data processing is performed to notify the server in real time.

[1238] Output: Emergency notification data sent to the server.

[1239] Step 10: Sending emergency alerts

[1240] Subject: Server

[1241] Input: Emergency notification data.

[1242] Specific operation: The server checks the fall information and immediately sends an emergency notification to family members and medical institutions.

[1243] Data manipulation / calculation: Create notification messages based on required contact information.

[1244] Output: Emergency alert sent to family and medical facilities.

[1245] The above are the detailed processing steps of this system. At each step, the voice data and behavioral data of the elderly person are analyzed and processed, enabling appropriate responses and notifications to be given.

[1246] (Application example 2)

[1247] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1248] In an aging society, safe transportation, health management, and reducing loneliness for the elderly are important issues. In particular, elderly people living alone need a means of safe transportation that reduces loneliness in their daily lives. Therefore, a system is needed that can grasp the emotions and health status of elderly people in real time and provide the necessary support.

[1249] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for analyzing the elderly person's voice data using a voice recognition system, means for recognizing the elderly person's emotion using an emotion engine, means for generating a response based on the emotion information, and means for operating the navigation system of the autonomous vehicle. This makes it possible to grasp the emotions and health status of the elderly person in real time and support safe and secure travel.

[1250] "Elderly" refers to individuals who, as they grow older, require some form of assistance or care for their daily lives and activities.

[1251] The "means for receiving voice" is a device such as a microphone for capturing the speech of the elderly person.

[1252] The "means for converting speech to text" is a speech recognition system that has the function of converting speech data into text information.

[1253] The "means for analyzing text and generating appropriate responses" is a system that analyzes the converted text data using natural language processing and generates responses appropriate for the dialogue.

[1254] The "means for synthesizing a response using the voice of a registered family member" is a system having a function for synthesizing the generated text-based response using the voice of a family member.

[1255] The "means for providing synthesized voice to the elderly" is a device that plays synthesized voice data to the elderly through a speaker or the like.

[1256] The "means for analyzing voice data of an elderly person using a voice recognition system" is a voice recognition technology for analyzing voice data of an elderly person and generating an appropriate response.

[1257] The "means for recognizing the emotions of the elderly using an emotion engine" is a system that analyzes and extracts the emotional state of the elderly from voice data and text data.

[1258] The "means for generating a response based on emotional information" is a system that generates an appropriate emotional response based on the recognized emotional information.

[1259] "Means for operating the navigation system of an autonomous vehicle" refers to a system that guides the autonomous vehicle to its destination and controls its driving.

[1260] A "means for continuous conversation recording" is a device or software feature that continuously records conversations between the senior and the system.

[1261] "Means for transmitting conversation data to a server" refers to a technology for transferring recorded conversation data to a central server.

[1262] The "means for notifying the family of the analysis results" is a system that promptly notifies the family of the results of the analysis performed by the server.

[1263] The "means for detecting signs of dementia and emotional abnormalities" is a system that analyzes conversation data and emotional states to find abnormal patterns and early signs of dementia.

[1264] "Means for sending alerts" refers to a system that promptly sends a warning to relevant parties when an abnormality or emergency is detected.

[1265] This invention relates to an autonomous vehicle system for assisting the elderly, aiming to provide safe transportation and health management for the elderly, as well as to reduce feelings of loneliness. This system is realized by combining cloud computing technology, voice recognition technology, emotion recognition engines, and autonomous driving technology.

[1266] In order to implement the present invention, the following hardware and software are used.

[1267] Hardware used:

[1268] 1. Self-driving vehicle: A vehicle equipped with self-driving technology and equipped with built-in sensors, speakers, microphones, and cameras.

[1269] 2. Microphone: A device for capturing the speech of elderly people in the vehicle.

[1270] 3. Speaker: A device that provides synthesized speech to the elderly.

[1271] 4. Camera: A device used to monitor the movements of elderly people and detect any abnormalities.

[1272] Software used:

[1273] 1. Speech recognition system: A platform for converting the elderly's speech into text (e.g., Google Speech-to-Text API).

[1274] 2. Emotion recognition engine: A machine learning model that recognizes the emotions of the elderly and generates appropriate responses (e.g., IBM Watson Tone Analyzer).

[1275] 3. Speech synthesis system: A platform for synthesizing responses in the voices of family members (e.g., Amazon Polly).

[1276] 4. Autonomous vehicle control systems: APIs for controlling navigation in autonomous vehicles (e.g., Tesla Autopilot API).

[1277] A natural language description of the program:

[1278] Audio capture and transmission:

[1279] The server receives the voice data of the elderly person captured by the microphone in the autonomous vehicle, and the voice data is transmitted to the server in real time.

[1280] Voice Recognition:

[1281] The server uses a speech recognition system (Google Speech-to-Text API) to convert the received voice data into text data, which is a preparatory step for analyzing the conversation content.

[1282] Emotion Recognition and Analysis:

[1283] Using an emotion recognition engine (IBM Watson Tone Analyzer), the emotions of the elderly are recognized from the converted text data. Based on the recognized emotional information, an appropriate response is generated.

[1284] Response generation and serving:

[1285] The server uses the data necessary to generate a response and synthesizes a response based on the family member's voice using a speech synthesis system (Amazon Polly). This synthesized voice is then provided to the elderly person through the vehicle's speakers.

[1286] Navigation controls:

[1287] The system analyzes the elderly person's speech to identify destinations and instructions, and uses the autonomous vehicle control system (Tesla Autopilot API) to provide the necessary navigation. For example, if the system detects an utterance such as "I want to go to the hospital," it will set the navigation system to head to the nearest hospital.

[1288] Examples:

[1289] For example, if an elderly person gets into an autonomous vehicle and says, "I want to go to the park today," the following series of processes will take place.

[1290] 1. Voice capture: The microphone captures "I want to go to the park today."

[1291] 2. Speech recognition: The server converts the utterance into text: "I want to go to the park today."

[1292] 3. Emotion recognition: The emotion engine recognizes the emotion of "fun."

[1293] 4. Navigation: The vehicle sets the destination as the park, and a family member's voice guides them, saying, "We're heading to the park. Have fun."

[1294] 5. Anomaly detection: No emergency notification is required unless an anomaly is detected.

[1295] Example prompt sentence:

[1296] Please provide an example of creating a program for an autonomous driving system to assist the elderly, which converts voice data into text, recognizes emotions, and generates responses. Use a speech recognition system for speech recognition, an emotion recognition engine for emotion recognition, a speech synthesis system for speech synthesis, and an autonomous vehicle control system for autonomous driving control.

[1297] This invention uses an elderly assistance autonomous driving vehicle system to effectively realize safe and secure transportation, health management, and reduction of feelings of loneliness for the elderly.

[1298] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1299] Step 1:

[1300] The server receives the elderly person's voice data captured by the microphone in the autonomous vehicle. The input in this step is the elderly person's voice, and the output is the voice data stored on the server. Specifically, the microphone captures the voice as digital data and transmits it to the server in real time.

[1301] Step 2:

[1302] The server uses a speech recognition system (voice recognition system) to convert the received voice data into text data. The input in this step is the voice data received in step 1, and the output is the converted text data. In concrete terms, the server sends the voice data to the speech recognition system and receives the result as text.

[1303] Step 3:

[1304] The server uses an emotion recognition engine (emotion recognition engine) to recognize the emotions of the elderly from the text data. The input in this step is the text data obtained in step 2, and the output is the recognized emotion information. Specifically, the server inputs the text data into the emotion recognition engine and outputs the emotion information.

[1305] Step 4:

[1306] The server generates a response based on the emotional information and uses a speech synthesis system to synthesize the response in the voice of the family member. The input in this step is the emotional information obtained in step 3 and the generated response text, and the output is synthesized voice data. Specifically, the server generates a response text according to the emotional information, sends it to the speech synthesis system, and synthesizes it in the voice of the family member.

[1307] Step 5:

[1308] The server sends the synthesized voice data to the autonomous vehicle and provides it to the elderly through the speaker. The input in this step is the synthesized voice data obtained in step 4, and the output is the voice to be played to the elderly. Specifically, the server sends the synthesized voice data to the vehicle's audio system and plays it through the speaker.

[1309] Step 6:

[1310] The server analyzes the destination and instructions and sends them to the autonomous vehicle control system to operate the autonomous vehicle's navigation system. The input in this step is the text data obtained from step 2, and the output is the vehicle's destination setting and driving control. Specifically, the server extracts destination information from the elderly person's speech and sends instructions to the autonomous vehicle control system to perform navigation.

[1311] Step 7:

[1312] The server analyzes the elderly person's health condition and emotional information, and sends an emergency notification to family members or medical institutions if an abnormality is detected. The input in this step is the emotional information obtained in step 3 and other data collected in steps 1 and 2, and the output is the notification to be sent. Specifically, the server continuously monitors the data, and if an abnormality is detected, it triggers the notification system to send an emergency notification.

[1313] Through the above processing steps, the present invention can realize safe and secure transportation and health management as an elderly assistance autonomous driving vehicle system.

[1314] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1315] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1316] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1317] [Third embodiment]

[1318] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1319] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1320] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1321] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1322] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1323] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1324] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1325] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1326] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1327] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1328] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1329] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1330] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time.

[1331] Feature Description

[1332] 1. Ability to talk to the elderly using the voices of children and grandchildren

[1333] Subject: Server

[1334] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[1335] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[1336] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[1337] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[1338] The device plays the received audio to the elderly person through a speaker.

[1339] Specific examples

[1340] An elderly person speaks to the device, asking, "What should I eat today?"

[1341] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[1342] 2. Ability to record conversations with elderly people and provide information to their families

[1343] Subject: Device

[1344] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[1345] The server converts the received conversation data into text and analyzes the content.

[1346] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[1347] Specific examples

[1348] An elderly person says, "My shoulder has been hurting lately."

[1349] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[1350] 3. Alerts for dementia and depression

[1351] Subject: Server

[1352] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[1353] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[1354] Specific examples

[1355] Elderly people say, "I don't know what I'm living for."

[1356] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[1357] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[1358] Subject: Device

[1359] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[1360] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[1361] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[1362] Specific examples

[1363] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[1364] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[1365] 5. Fall detection and emergency contact function

[1366] Subject: Camera

[1367] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[1368] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[1369] Specific examples

[1370] An elderly person falls in the living room.

[1371] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[1372] As described above, the present invention is a system that comprehensively supports the health management of elderly people, reduces feelings of loneliness, and alleviates anxiety for their families. This system allows elderly people to live their daily lives with peace of mind, and allows their families to follow up in a timely manner.

[1373] The processing flow will be explained below.

[1374] A function that allows you to talk to the elderly using the voices of children and grandchildren

[1375] Step 1:

[1376] The user (elderly person) speaks into the terminal.

[1377] As a concrete example, say, "What shall we eat today?"

[1378] Step 2:

[1379] The terminal receives the elderly person's voice and transmits the voice data to the server.

[1380] The audio data is transferred to the server in real time.

[1381] Step 3:

[1382] The server converts the received voice data into text.

[1383] A speech recognition algorithm converts the speech data into a string of characters.

[1384] Step 4:

[1385] The server analyzes the converted text data and generates an appropriate response.

[1386] Use natural language processing (NLP) techniques to understand the context of text data.

[1387] Step 5:

[1388] The server synthesizes a response based on voice samples from family members.

[1389] A response message is synthesized using the voice of a pre-registered family member.

[1390] Step 6:

[1391] The server transmits the synthesized voice data to the terminal.

[1392] The generated voice data is sent back to the elderly person.

[1393] Step 7:

[1394] The device plays the synthesized speech through the speaker.

[1395] The elderly person hears a family member respond, "Maybe fish would be good?"

[1396] A function that records conversations with elderly people and provides information to their families

[1397] Step 1:

[1398] The device constantly records conversations with the elderly.

[1399] Record all your everyday conversations.

[1400] Step 2:

[1401] The recorded conversation data is periodically sent to a server.

[1402] For example, a data packet is sent every hour.

[1403] Step 3:

[1404] The server converts the received voice data into text.

[1405] Convert the speech into text using voice recognition technology.

[1406] Step 4:

[1407] The server analyzes the converted text data.

[1408] Natural language processing technology is used to extract important information and anomalies.

[1409] Step 5:

[1410] Based on the analysis results, the information to be notified to the family is determined.

[1411] If an anomaly is detected, detailed information about it will also be included.

[1412] Step 6:

[1413] The server sends the analysis results to the family.

[1414] Notifications are given in real time.

[1415] Examples:

[1416] An elderly person says, "My shoulder has been hurting lately."

[1417] The server analyzes the information and sends a notification to the family about the "shoulder pain."

[1418] A feature that sends alerts for dementia and depression

[1419] Step 1:

[1420] The server periodically analyzes conversation data with the elderly.

[1421] The analysis targets daily conversation logs.

[1422] Step 2:

[1423] The server detects changes in grammatical structure and sentiment.

[1424] It uses natural language processing techniques and sentiment analysis algorithms.

[1425] Step 3:

[1426] If an anomaly is detected, the server generates an alert.

[1427] Identify signs of dementia and emotional abnormalities.

[1428] Step 4:

[1429] The server generates alerts and sends them to family members and medical institutions.

[1430] The alert is sent as an emergency notification.

[1431] Examples:

[1432] Elderly people say, "I don't know what I'm living for."

[1433] The server performs emotion analysis, detects signs of depression, and sends emergency notifications.

[1434] Functions to prevent forgetting to take medicine and assist with meals and shopping

[1435] Step 1:

[1436] The server sets medication times and meal schedules.

[1437] Set a daily reminder.

[1438] Step 2:

[1439] The device will send a reminder notification at the set time.

[1440] A notification will play saying "Time for your medicine."

[1441] Step 3:

[1442] When the elderly person responds with a voice of confirmation, the device records the information and sends it to the server.

[1443] "Thanks, I drank it," he replied.

[1444] Step 4:

[1445] The server notifies the family of the reminder information.

[1446] Notifies you that confirmation of missed dose has been completed.

[1447] Step 5:

[1448] The server learns the elderly person's lifestyle patterns and generates a shopping list.

[1449] Make a list of the items you need.

[1450] Step 6:

[1451] The device notifies the elderly person of the shopping list.

[1452] Announce the list contents by voice.

[1453] Examples:

[1454] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[1455] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[1456] Fall detection and emergency contact function

[1457] Step 1:

[1458] The camera monitors the footage at all times.

[1459] Monitor overall behavior.

[1460] Step 2:

[1461] The camera detects an abnormality.

[1462] For example, a falling motion is detected.

[1463] Step 3:

[1464] Video data is sent to the server in real time.

[1465] Send data immediately.

[1466] Step 4:

[1467] The server identifies the fall and generates an emergency alert.

[1468] Prepare notifications for medical institutions and families.

[1469] Step 5:

[1470] The server sends emergency alerts to family and medical institutions.

[1471] Encourage a rapid response.

[1472] Examples:

[1473] An elderly person falls in the living room.

[1474] The camera detects an abnormality and notifies the server.

[1475] The server sends emergency alerts to family and medical institutions.

[1476] These are the specific steps of the program processing, which will enable health management for elderly people living alone, reduce loneliness, and provide peace of mind to their families.

[1477] Example 1

[1478] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1479] There is a need for health management for the elderly, reducing loneliness, and improving family security. However, systems that enable appropriate communication and health monitoring for elderly people living alone have not yet been fully developed. When elderly people live in isolated situations, the lack of daily communication and the inability to detect health abnormalities early are problems. In particular, there is a lack of means to detect early signs of dementia and depression and provide appropriate treatment.

[1480] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1481] In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for saving voice samples provided by the family member and extracting features, and natural language processing means for analyzing the elderly person's speech and generating a response related to life support. This enables natural communication with the elderly, health management, reduction of loneliness, and early detection of signs of dementia and depression.

[1482] The "means for receiving the voice of the elderly person" is a function for acquiring the voice uttered by the elderly person using an input device such as a microphone.

[1483] The "means for converting received voice into text" is a function that converts voice data into corresponding text data using voice recognition technology.

[1484] "Means for analyzing the converted text and generating an appropriate response" refers to a function that uses natural language processing (NLP) technology to analyze text data and generate an appropriate response based on the content.

[1485] The "means for synthesizing the response with the voice of a registered family member" is a technology for converting text data into voice data using the voice characteristics of a designated family member.

[1486] The "means for providing synthesized voice to the elderly" is a function for playing back synthesized voice to the elderly using a voice output device such as a speaker.

[1487] The "means for saving voice samples provided by family members and extracting features" is a function for saving voice data provided by family members and extracting voice features.

[1488] "Natural language processing means for analyzing elderly people's statements and generating responses related to life support" is a technology that analyzes the content of elderly people's statements and generates information and actions necessary for life support based on the analysis results.

[1489] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. This system is realized by integrating speech recognition, natural language processing, and speech synthesis technologies.

[1490] Hardware and software used

[1491] Hardware

[1492] High-Precision Microphone

[1493] speaker

[1494] camera

[1495] server

[1496] Devices (smartphones, tablets, PCs, etc.)

[1497] software

[1498] Speech recognition: A speech recognition API for converting voice data into text (e.g., Google Cloud Speech-to-Text API)

[1499] Natural Language Processing (NLP): NLP models (e.g., AWS Comprehend) to analyze text data and generate appropriate responses.

[1500] Speech synthesis: A speech synthesis API to convert text to speech (e.g. IBM Watson Text-to-Speech API)

[1501] Database: A database for storing audio samples and analysis data (e.g., Google Cloud Storage)

[1502] Explanation of the specific functions of the system

[1503] 1. Ability to talk to the elderly using the voices of children and grandchildren

[1504] Subject: Server, Terminal, User

[1505] The server stores voice samples provided by family members in a database. It also extracts voice features using a machine learning model. When the user (elderly person) speaks into the device, the device captures the voice data and sends it to the server. The server converts the received voice data into text using the Google Cloud Speech-to-Text API and analyzes the text data using a natural language processing (NLP) model. An appropriate response is generated, and speech is synthesized using the family member's voice using the IBM Watson Text-to-Speech API, and the speech data is sent to the device. The device then plays the received audio to the elderly person through a speaker.

[1506] Specific examples

[1507] An elderly person speaks to the device, asking, "What should I eat today?"

[1508] The server synthesizes the family member's voice saying, "Maybe fish would be good?" and plays it back to the elderly person via the device.

[1509] Example prompt

[1510] "When you say to a device, 'What should I eat today?' how does the server process the data and generate the voice response?"

[1511] 2. Ability to record conversations with elderly people and provide information to their families

[1512] Subject: Terminal, Server

[1513] The device continuously records conversations with the elderly. The recorded conversation data is periodically sent to a server, which converts the voice data into text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing (NLP) model, and if particularly important information or anomalies are detected, the server notifies the family using the Twilio API.

[1514] Specific examples

[1515] An elderly person says, "My shoulder has been hurting lately."

[1516] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[1517] Example prompt

[1518] "When the audio data recorded on the device is sent to the server, how will family members be notified if an abnormality is detected?"

[1519] 3. Alerts for dementia and depression

[1520] Subject: Server

[1521] The server periodically converts conversation data with the elderly person into text using the Google Cloud Speech-to-Text API, then analyzes the text data using a natural language processing (NLP) model. If signs of dementia or abnormal emotions are detected, an alert is generated and notifies family members and medical institutions using the Twilio API.

[1522] Specific examples

[1523] Elderly people say, "I don't know what I'm living for."

[1524] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[1525] Example prompt

[1526] "If an elderly person makes a statement that indicates signs of depression, what analysis does the server perform and how does it generate a notification?"

[1527] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[1528] Subject: Terminal, Server

[1529] The device uses the Google Calendar API to manage medication times and meal schedules, and sends a reminder at the set time: "It's time for your medicine." When the user (elderly person) responds with a voice confirmation, the voice is sent to the server and converted into text using the Google Cloud Speech-to-Text API. The server then notifies the family of the confirmation information. The server also learns the elderly person's lifestyle patterns, generates a shopping list, and provides it via voice from the device.

[1530] Specific examples

[1531] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[1532] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[1533] Example prompt

[1534] "How can we remind and confirm elderly people to prevent them from forgetting to take their medication, and notify their families of this information?"

[1535] 5. Fall detection and emergency contact function

[1536] Subject: camera, server

[1537] The camera constantly monitors the video and runs an algorithm to detect anomalies. If an elderly person falls, the camera detects the movement and sends the video data in real time to a server. The server then confirms the fall and uses the Twilio API to send an emergency notification to family members and medical institutions.

[1538] Specific examples

[1539] An elderly person falls in the living room.

[1540] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[1541] Example prompt

[1542] "How does the fall detection feature work and how does it send emergency notifications after detection?"

[1543] As described above, this system recognizes the voice of the elderly person, generates an appropriate response using natural language processing, and then uses voice synthesis technology to respond in the voice of a family member, thereby providing comprehensive support for managing the elderly person's health, reducing feelings of loneliness, and providing a sense of security to their family.

[1544] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1545] Specific processing steps of the system program

[1546] Feature 1: Ability to talk to the elderly using the voices of children and grandchildren

[1547] Subject: Server, Terminal, User

[1548] Step 1:

[1549] Subject: Server

[1550] Voice samples provided by family members are stored in a database.

[1551] (Input) Voice samples recorded by family members

[1552] (Processing) The server receives the audio samples and stores them in a database.

[1553] (Output) Voice samples of family members stored in the database

[1554] Specific operation: The server receives the voice data and stores it in a database.

[1555] Step 2:

[1556] Subject: Server

[1557] The voice sample is analyzed using a machine learning model to extract features.

[1558] (Input) Audio samples stored in the database

[1559] (Processing) Apply a voice feature extraction algorithm to extract features

[1560] (Output) Extracted feature data

[1561] Specific operation: The server applies a feature extraction algorithm to the voice sample and stores the features in a database.

[1562] Step 3:

[1563] Subject: User

[1564] The user speaks into the terminal.

[1565] (Input) User's voice

[1566] (Processing) High-precision microphone captures audio

[1567] (Output) Captured audio data

[1568] Specific operation: The user speaks to the device, "What should I eat today?"

[1569] Step 4:

[1570] Subject: Device

[1571] The terminal transmits the voice data to the server.

[1572] (Input) Captured audio data

[1573] (Processing) Send the audio data to the server

[1574] (Output) Audio data received by the server

[1575] Specific operation: The device sends the captured audio data to a server via the Internet.

[1576] Step 5:

[1577] Subject: Server

[1578] Converts audio data into text.

[1579] (Input) Audio data received by the server

[1580] (Processing) Convert speech to text using a speech recognition API (e.g., Google Cloud Speech-to-Text API)

[1581] (Output) Converted text data

[1582] What happens: The server uses the Google Cloud Speech-to-Text API to convert the speech "What should we eat today?" into text.

[1583] Step 6:

[1584] Subject: Server

[1585] Analyze text data using an NLP model and generate appropriate responses.

[1586] (Input) Converted text data

[1587] (Processing) Analyze the text using an NLP model (e.g. AWS Comprehend) and generate an appropriate response

[1588] (Output) The generated response text

[1589] Specific behavior: The server generates a response to the question "What should we eat?" with "Maybe fish would be good?"

[1590] Step 7:

[1591] Subject: Server

[1592] The response text is converted into speech using the speech synthesis API.

[1593] (Input) Generated response text

[1594] (Processing) Convert text to speech using a speech synthesis API (e.g. IBM Watson Text-to-Speech API)

[1595] (Output) Synthesized voice data

[1596] Specific behavior: The server converts the text "Maybe fish would be good?" into speech.

[1597] Step 8:

[1598] Subject: Server

[1599] The synthesized voice data is transmitted to the terminal.

[1600] (Input) Synthesized voice data

[1601] (Processing) Send audio data to the terminal

[1602] (Output) Audio data received on the device

[1603] Specific operation: The server sends the synthesized voice data to the terminal.

[1604] Step 9:

[1605] Subject: Device

[1606] The device plays the audio.

[1607] (Input) Audio data received by the device

[1608] (Processing) Playing audio through speakers

[1609] (Output) Audio played to the elderly

[1610] Specific operation: The device plays a voice message from the speaker saying, "Maybe a fish would be good?"

[1611] You can add specific processing steps below for other features as well, but you can distinguish the steps for each feature based on this format.

[1612] If you have steps for other features below, please list them in a similar format.

[1613] (Application example 1)

[1614] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1615] In modern society, the number of elderly people living alone is increasing, and health management and reducing feelings of loneliness have become social issues. Furthermore, when elderly people living alone use self-driving vehicles, ensuring safety and communication with their families are important. However, conventional systems do not adequately provide comprehensive solutions to these issues. Therefore, there is a need for a system that integrates safe and secure transportation means for the elderly with support for their daily lives.

[1616] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1617] In this invention, the server includes means for receiving the elderly person's voice, means for converting the voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the family member's voice, means for providing synthesized voice, means for setting a destination for the autonomous vehicle and providing route guidance, means for monitoring the driving situation and detecting abnormalities, and means for notifying the family member in real time when an abnormality occurs. This allows the elderly person to receive route guidance from the family member's voice when using the autonomous vehicle, improving safety and enabling the family member to understand the elderly person's situation in real time.

[1618] · An "elderly person" is an individual who, due to their advanced age, requires assistance in daily living.

[1619] "Means for receiving voice" refers to a device or method for recording and capturing the voice uttered by the elderly person and inputting it into the system.

[1620] "Means for converting speech to text" means the process or technology that converts received speech data into corresponding text data using speech recognition technology.

[1621] "Means of analyzing text" refers to natural language processing techniques used to understand meaning and intent based on converted text data.

[1622] - "Means for generating a response" refers to an algorithm or program that generates an appropriate response based on the analysis results.

[1623] "Means for synthesizing using family members' voices" refers to a method of reproducing the generated response using voice synthesis technology using pre-registered family members' voice data.

[1624] "Means for providing synthesized speech" means a speaker or other audio playback device that delivers synthesized speech to the senior citizen.

[1625] "Means for setting a destination in an autonomous vehicle" refers to the methods and technologies for inputting and setting the destination desired by the elderly person into an autonomous vehicle.

[1626] "Means for providing directions" refers to systems and technologies that allow autonomous vehicles to provide directions to seniors toward a set destination.

[1627] "Means for monitoring the driving situation" refers to sensors, cameras, and algorithms that monitor the driving situation in real time while the autonomous vehicle is driving to ensure safety.

[1628] "Means for detecting abnormalities" refers to techniques and methods for recognizing and identifying hazards and abnormal events that may occur during operation.

[1629] "Means of notifying family members in real time" refers to communication technologies and methods for immediately notifying family members of the situation when an abnormality is detected.

[1630] To implement this invention, it is necessary to build a system that receives the voice of the elderly person and synthesizes it with the voices of family members. This system uses the following hardware and software.

[1631] Hardware used

[1632] 1. Smartphone: Used to record the elderly person's voice and send the recording data to the server.

[1633] 2. Self-driving vehicles: Used as transportation for the elderly, with functions such as route guidance and safety monitoring.

[1634] 3. Speaker: Used as a device to transmit synthesized voices of family members to the elderly.

[1635] Software used

[1636] 1. pyaudio: A library for recording the voices of elderly people.

[1637] 2. wave: A library for saving and playing recorded audio.

[1638] 3. transformers (Wav2Vec2): A library for converting audio to text.

[1639] 4. gTTS: A library for synthesizing text in the voices of family members.

[1640] 5. smtplib: A library for implementing notification functions.

[1641] Data processing and calculation

[1642] The server first receives the elderly person's voice and records it using the pyaudio library. The recorded voice data is then temporarily saved using the wave library. The voice data is then converted to text using the Wav2Vec2 model in transformers. The converted text data is analyzed using natural language processing techniques to generate an appropriate response. The generated response is then synthesized in the voice of a family member using the gTTS library and provided to the elderly person through a speaker as synthesized speech.

[1643] If an anomaly is detected or if specific keywords are found, a real-time notification is sent to family members using the smtplib library, containing the analysis results and the elderly person's current condition.

[1644] Specific examples

[1645] For example, if an elderly person says, "Shall I confirm your next hospital appointment?", the system works as follows: First, the smartphone records this speech and sends it to the server. The server converts the recorded speech into text using the Wav2Vec2 model and analyzes its content. An appropriate response is generated, such as "Yes, Grandma. Your next hospital appointment is tomorrow at 10:00 AM," using a synthesized voice from a family member's voice, and this is played to the elderly through the speaker.

[1646] Prompt Sentence Examples

[1647] 1. "Grandpa, should I confirm your next doctor's appointment?"

[1648] 2. "Grandma, have you taken your morning medicine?"

[1649] 3. "Where is Grandpa going? Is he driving safely?"

[1650] This will allow elderly people to use self-driving vehicles with peace of mind, and will also enable family members to keep track of the elderly person's situation in real time.

[1651] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1652] Step 1:

[1653] The smartphone records the elderly person's voice. The device uses the pyaudio library to record 10 seconds of voice data uttered by the elderly person. The input here is the elderly person's voice, and the output is the recorded voice data (audio file).

[1654] Step 2:

[1655] The recorded audio data is sent to the server. The device uses the wave library to save the audio file and uploads it to the server. The input is the recorded audio file, and the output is the audio file saved on the server.

[1656] Step 3:

[1657] The server converts the audio data to text. The server uses the Wav2Vec2 model from transformers to convert the audio file to text. The input is the audio file, and the output is the converted text.

[1658] Step 4:

[1659] The server analyzes the text data and generates an appropriate response. Natural language processing technology is used to analyze the meaning and intent of the converted text data and generate an appropriate response text. The input is the text data, and the output is the generated response text.

[1660] Step 5:

[1661] The server synthesizes the response text in the voice of the family member. The server uses the gTTS library to synthesize the generated response text in the voice of the pre-registered family member and create an audio file. The input is the response text, and the output is the synthesized audio file.

[1662] Step 6:

[1663] The synthesized voice is provided to the elderly through a speaker. The server sends the synthesized voice to the terminal, and the terminal plays the voice through the speaker. The input is the synthesized voice file, and the output is the voice played to the elderly.

[1664] Step 7:

[1665] Monitors the driving situation and detects abnormalities. Autonomous vehicles use sensors and cameras to monitor the driving situation in real time to ensure safety. The input is real-time data during driving, and the output is analysis results and abnormality detection signals.

[1666] Step 8:

[1667] If an abnormality is detected, a notification is sent to family members in real time. The server uses the smtplib library to notify family members of the situation by email when it receives an abnormality detection signal. The input is the abnormality detection signal, and the output is the alert notification sent to family members.

[1668] Specific example prompts

[1669] "Grandpa, should I confirm your next doctor's appointment?"

[1670] "Grandma, have you taken your morning medicine?"

[1671] "Where is Grandpa going? Are you driving safely?"

[1672] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1673] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and reduce feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it achieves more effective responses and anomaly detection.

[1674] Feature Description

[1675] 1. Ability to talk to the elderly using the voices of children and grandchildren

[1676] Subject: Server

[1677] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[1678] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[1679] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[1680] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[1681] The device plays the received audio to the elderly person through a speaker.

[1682] Specific examples

[1683] An elderly person speaks to the device, asking, "What should I eat today?"

[1684] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[1685] 2. Ability to record conversations with elderly people and provide information to their families

[1686] Subject: Device

[1687] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[1688] The server converts the received conversation data into text and analyzes the content.

[1689] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[1690] Specific examples

[1691] An elderly person says, "My shoulder has been hurting lately."

[1692] The server analyzes the information and sends a notification about the "shoulder pain" to the family, who can then plan a visit.

[1693] 3. Alerts for dementia and depression

[1694] Subject: Server

[1695] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[1696] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[1697] Specific examples

[1698] Elderly people say, "I don't know what I'm living for."

[1699] The server performs emotion analysis, detects signs of depression, and sends emergency notifications to family members and medical institutions.

[1700] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[1701] Subject: Device

[1702] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[1703] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[1704] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[1705] Specific examples

[1706] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[1707] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[1708] 5. Fall detection and emergency contact function

[1709] Subject: Camera

[1710] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[1711] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[1712] Specific examples

[1713] An elderly person falls in the living room.

[1714] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[1715] Features of inventions that combine emotion engines

[1716] 1. Emotion recognition and response generation using an emotion engine

[1717] Subject: Server

[1718] The server uses an emotion engine to recognize emotions from the elderly person's voice data, extracts the elderly person's emotions from the voice data, and generates a response based on that information.

[1719] Depending on the recognized emotion, the server generates an appropriate response and provides it to the elderly person as a synthesized voice.

[1720] Specific examples

[1721] An elderly person says, "I feel a little lonely today."

[1722] The server uses an emotion engine to recognize the emotion "loneliness" and generates a response such as "I see you're feeling lonely. Let's think of something we can do together."

[1723] 2. Anomaly detection and notification using emotion engine

[1724] Subject: Server

[1725] If the server detects negative emotions in an elderly person using the emotion engine, it generates an alert and notifies family members and relevant organizations.

[1726] If an anomaly is detected, a notification system will be activated to ensure a rapid response.

[1727] Specific examples

[1728] An elderly person says, "Nothing is fun anymore."

[1729] The server uses an emotion engine to detect emotions such as "despair" and "deep sadness" and sends emergency notifications to family members and medical institutions.

[1730] The above is a specific embodiment of the present invention that combines an emotion engine. This makes it possible to more accurately grasp the emotional state of elderly people living alone and respond appropriately. Furthermore, since it can respond sensitively to changes in emotions, it can also support the mental health of the elderly.

[1731] The processing flow will be explained below.

[1732] Emotion recognition and response generation using an emotion engine

[1733] Step 1:

[1734] The user (elderly person) speaks into the terminal.

[1735] As a specific example, say, "I feel a little lonely today."

[1736] Step 2:

[1737] The terminal receives the elderly person's voice and transmits the voice data to the server.

[1738] Step 3:

[1739] The server converts the received voice data into text.

[1740] A speech recognition algorithm converts the speech data into a string of characters.

[1741] Step 4:

[1742] The server passes the converted text data and voice data to the emotion engine.

[1743] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[1744] Step 5:

[1745] The server generates an appropriate response based on the emotional data recognized by the emotion engine.

[1746] Using natural language processing technology, it generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[1747] Step 6:

[1748] The server generates a response that is then synthesized into the family member's voice.

[1749] A response message is synthesized using the voice of a pre-registered family member.

[1750] Step 7:

[1751] The server transmits the synthesized voice data to the terminal.

[1752] The generated voice data is sent back to the elderly person.

[1753] Step 8:

[1754] The device plays the synthesized speech through the speaker.

[1755] The elderly person hears a family member respond, "I see you're feeling lonely. Let's think of something we can do together."

[1756] Anomaly detection and notification using emotion engine

[1757] Step 1:

[1758] The user (elderly person) speaks into the terminal.

[1759] For example, say, "Nothing is fun anymore."

[1760] Step 2:

[1761] The terminal receives the elderly person's voice and transmits the voice data to the server.

[1762] Step 3:

[1763] The server converts the received voice data into text.

[1764] A speech recognition algorithm converts the speech data into a string of characters.

[1765] Step 4:

[1766] The server passes the converted text data and voice data to the emotion engine.

[1767] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[1768] Step 5:

[1769] The emotion engine detects negative emotions such as "despair" and "deep sadness."

[1770] The emotion engine passes these abnormal emotion data to the server.

[1771] Step 6:

[1772] The server generates alerts based on the emotion engine data.

[1773] Prepare an abnormality detection message and notify family members and relevant organizations.

[1774] Step 7:

[1775] The server sends emergency notifications to family and medical institutions.

[1776] Alerts such as "elderly people are feeling deep sadness" will be sent via email or app notification.

[1777] Step 8:

[1778] Family and medical institutions will be notified and take appropriate action.

[1779] Schedule visits and medical consultations.

[1780] These are the processing steps of the system that combines the emotion engine, which makes it possible to grasp the emotional state of the elderly more accurately and take the necessary measures quickly.

[1781] Example 2

[1782] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1783] For elderly people living alone, reducing loneliness and managing their health are major challenges. It is particularly difficult to understand changes in the elderly's moods and health status when family members live far away. Early detection of dementia and depression and prevention of falls are also important issues. Furthermore, elderly people need assistance with daily activities such as remembering to take their medication, eating, and shopping. To solve these challenges, a voice-responsive system that includes emotion recognition is needed.

[1784] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for continuously recording conversations with the elderly person, means for transmitting the recorded conversation data to the server, means for analyzing the transmitted conversation data and notifying the family member of the analysis result, means for analyzing the conversation data with the elderly person and detecting signs of dementia or emotional abnormalities, means for transmitting an alert when signs of dementia or emotional abnormalities are detected, means for learning the elderly person's lifestyle patterns and setting reminders as necessary, means for notifying the family member based on the reminder, means for recording confirmation of the notified content and notifying the family member, means for detecting the elderly person's fall, means for transmitting an emergency notification when a fall is detected, means for recognizing the elderly person's emotion from the voice data, means for generating an appropriate response according to the recognized emotion, means for detecting the elderly person's negative emotion, and means for transmitting an alert to the family member or relevant organization when the emotion is detected. This reduces the elderly person's sense of loneliness and enables health management. It can also enable early detection of dementia and depression, prevent falls, prevent people from forgetting to take their medication, and provide assistance with meals and shopping.

[1785] The "means for receiving voice" is a device or system that captures the voice spoken by the elderly person and receives it as data.

[1786] A "speech-to-text means" is software or an algorithm that converts received speech data into text data.

[1787] The "means for analyzing text and generating an appropriate response" refers to a natural language processing engine or algorithm for analyzing text data converted from speech and generating an appropriate response.

[1788] The "means for synthesizing a response using the voice of a registered family member" refers to software or an algorithm for synthesizing the generated response using voice data of a registered family member in advance.

[1789] "Means for providing synthesized voices to the elderly" refers to devices or systems that transmit synthesized voices of family members to the elderly through speakers or communication devices.

[1790] "Means for continuously recording conversations" refers to equipment or systems that continuously record conversations with elderly people and store them as data.

[1791] The "means for transmitting recorded conversation data to a server" refers to a communication device or system for transmitting recorded conversation data to a server via a network.

[1792] "Means for analyzing the transmitted conversation data and notifying the family of the analysis results" refers to software or a communication system for analyzing the conversation data transmitted to the server and notifying the family of the analysis results.

[1793] "Means for detecting signs of dementia and emotional abnormalities" refers to algorithms and software that analyze conversation data with elderly people to detect signs of dementia and emotional abnormalities.

[1794] "Means for sending alerts when signs of dementia or abnormal emotions are detected" refers to a communication system or software for sending alerts to family members or relevant organizations when signs of dementia or abnormal emotions are detected.

[1795] "Means for learning lifestyle patterns and setting reminders" refers to algorithms or software that learn the lifestyle patterns of elderly people through data analysis and set reminders at the necessary times.

[1796] The "means for notifying based on a reminder" refers to a device or system for notifying the elderly based on the set reminder.

[1797] "Means for recording confirmation of the notified content and notifying the family" refers to a communication system or software for recording the notification confirmation from the elderly person and conveying the content to the family.

[1798] "Means for detecting falls in the elderly" refers to systems and algorithms that use cameras and sensors to detect falls in the elderly in real time.

[1799] The "means for sending an emergency notification when a fall is detected" is a communication system for sending an emergency notification to family members and medical institutions when a fall of an elderly person is detected.

[1800] "Means for recognizing the emotions of elderly people from voice data" refers to software or algorithms that use an emotion engine to extract and recognize emotions from the voice data of elderly people.

[1801] The "means for generating an appropriate response based on the recognized emotions" refers to a natural language processing engine or software that generates an appropriate response based on the emotions of the elderly person.

[1802] The "means for detecting negative emotions" refers to emotion analysis algorithms or software for identifying and detecting negative emotions from elderly people's voice data.

[1803] "Means for sending an alert to family members or relevant organizations when negative emotions are detected" refers to a communication system or software for quickly sending an alert to family members or relevant organizations when negative emotions are detected.

[1804] The present invention is a system that aims to manage the health of elderly people living alone and reduce feelings of loneliness through communication with family members. Furthermore, by combining it with an emotion engine, more effective responses and anomaly detection become possible. This system has multiple functions for connecting elderly people with their families. Specific embodiments of the system are described below.

[1805] System configuration

[1806] This system consists of three main components: a server, a terminal, and a user. The server processes and analyzes data, the terminal acts as an interface with the elderly, and the user refers to the elderly and their family members.

[1807] Hardware and software used

[1808] Hardware: Smart speakers such as Amazon Echo and Google Home can be used as devices, and cameras such as Nest Cam can be used.

[1809] Software: On the server side, we use Google Speech-to-Text for speech recognition, NLTK and Spacy for natural language processing, Google Text-to-Speech API for speech synthesis, and IBM Watson Tone Analyzer for sentiment analysis.

[1810] Program processing flow

[1811] 1. Saving and analyzing audio samples

[1812] The server receives voice samples provided by family members and stores them in a database. The stored voice samples are analyzed using machine learning models (e.g., TensorFlow) to extract features. This analysis lays the foundation for delivering the voices of family members to the elderly.

[1813] 2. Audio capture and transmission

[1814] When an elderly person speaks into the device, the device captures the voice and transmits it to a server in real time. For example, Amazon Echo captures the elderly person's voice in high quality and transmits the data to a server over the Internet.

[1815] 3. Speech data text conversion and analysis

[1816] The server converts the received voice data into text using Google Speech-to-Text, which is then analyzed using natural language processing tools such as NLTK or Spacy. Based on this analysis, an appropriate response is generated.

[1817] 4. Speech synthesis and transmission

[1818] The generated response is synthesized using the Google Text-to-Speech API using the family member's voice. The synthesized voice data is sent to the device, which plays it back to the elderly person through the device's speaker. For example, if an elderly person asks, "What should we eat today?", the server synthesizes a family member's voice saying, "Maybe fish would be good?" and plays it back on the device.

[1819] 5. Recording and analyzing conversations

[1820] The device constantly records conversations with the elderly person and periodically sends the data to a server. The server then converts the conversation data into text using Google Speech-to-Text and detects important information and anomalies. The detection results are then notified to family members in real time. For example, if an elderly person says, "My shoulder has been hurting lately," that information is immediately notified to the family.

[1821] 6. Sentiment Analysis and Response

[1822] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize emotions from the elderly person's voice. Based on the recognized emotion, it generates an appropriate response and provides it to the elderly person as synthesized speech. For example, if an elderly person says, "I feel a little lonely today," the server generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[1823] 7. Fall detection and emergency notification

[1824] The camera constantly monitors the elderly person's movements and runs an algorithm (e.g., OpenCV) to detect abnormal behavior (e.g., falls). If a fall is detected, the server immediately sends an emergency notification to family members or medical institutions. For example, if an elderly person falls in the living room, the camera detects the movement and notifies the server in real time. The server confirms this and immediately sends an emergency alert.

[1825] Specific prompt examples

[1826] A specific example of a prompt sentence to input into the generative AI model using this system is as follows:

[1827] "An elderly person asks, 'What shall we eat today?' A child's voice responds, 'Maybe fish would be good?'"

[1828] "An elderly person says, 'I feel a little lonely today.' Use an emotion engine to generate an appropriate response."

[1829] The above is a specific embodiment of the present invention. This system can reduce the sense of loneliness felt by the elderly and help them manage their health. It can also provide early detection of dementia and depression, prevent falls, help with forgetting to take medication, and assist with meals and shopping.

[1830] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1831] Processing steps of this system's program

[1832] Step 1: Saving and analyzing audio samples

[1833] Subject: Server

[1834] Input: Voice sample provided by family member (e.g., audio file).

[1835] What it does: The server receives voice samples uploaded by family members and stores them in a database.

[1836] Data processing / computation: The server uses machine learning models (e.g., TensorFlow) to extract and analyze voice features.

[1837] Output: A feature-analyzed audio profile.

[1838] Step 2: Capture and send audio

[1839] Subject: Device

[1840] Input: Speech produced by an elderly person.

[1841] Specific operation: When an elderly person speaks into a device (e.g., a smart speaker), the device captures the voice.

[1842] Data processing / calculation: The device transmits the captured voice data to the server in real time.

[1843] Output: The audio data sent to the server.

[1844] Step 3: Converting audio data into text and analyzing it

[1845] Subject: Server

[1846] Input: The audio data sent to the server.

[1847] Specific operation: The server converts the received voice data into text using voice recognition software (e.g., Google Speech-to-Text).

[1848] Data processing / calculation: The converted text data is analyzed using natural language processing (NLP) tools (e.g., NLTK or Spacy).

[1849] Output: Parsed text data.

[1850] Step 4: Speech synthesis and transmission

[1851] Subject: Server

[1852] Input: Parsed text data.

[1853] What it does: The server synthesizes a response in the family member's voice based on the analyzed text (e.g., Google Text-to-Speech API).

[1854] Data processing / calculation: Voice synthesis using family voice profiles.

[1855] Output: Synthesized speech data.

[1856] Step 5: Play audio

[1857] Subject: Device

[1858] Input: Synthesized speech data.

[1859] Specific operation: The device plays the received audio to the elderly person through the speaker.

[1860] Output: Audio provided to the senior.

[1861] Step 6: Record and send the conversation

[1862] Subject: Device

[1863] Input: Conversation with an elderly person.

[1864] Specific operation: The device constantly records conversations with the elderly and periodically sends the data to a server.

[1865] Data processing / calculation: Converting the audio data into the format required to send it to the server.

[1866] Output: Conversation data sent to the server.

[1867] Step 7: Analyze conversation data and notify

[1868] Subject: Server

[1869] Input: Conversation data sent to the server.

[1870] What it does: The server uses Google Speech-to-Text to convert the conversation data into text and detects important information and anomalies.

[1871] Data processing / calculation: Generate information to notify family members based on the converted and analyzed data.

[1872] Output: Analysis results communicated to family members.

[1873] Step 8: Sentiment Analysis and Response Generation

[1874] Subject: Server

[1875] Input: Speech data of elderly people.

[1876] Specific operation: The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to recognize emotions from the elderly person's voice.

[1877] Data processing / computation: Generate appropriate responses based on recognized emotions and provide them as synthesized speech.

[1878] Output: Synthetic voice response provided to the senior.

[1879] Step 9: Fall Detection and Emergency Notification

[1880] Subject: Camera

[1881] Enter: Elderly Movement.

[1882] Specific operation: The camera constantly monitors the elderly person's movements and runs an algorithm (e.g., OpenCV) to detect abnormal movements (e.g., falls).

[1883] Data processing / calculation: When a fall is detected, data processing is performed to notify the server in real time.

[1884] Output: Emergency notification data sent to the server.

[1885] Step 10: Sending emergency alerts

[1886] Subject: Server

[1887] Input: Emergency notification data.

[1888] Specific operation: The server checks the fall information and immediately sends an emergency notification to family members and medical institutions.

[1889] Data manipulation / calculation: Create notification messages based on required contact information.

[1890] Output: Emergency alert sent to family and medical facilities.

[1891] The above are the detailed processing steps of this system. At each step, the voice data and behavioral data of the elderly person are analyzed and processed, enabling appropriate responses and notifications to be given.

[1892] (Application example 2)

[1893] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1894] In an aging society, safe transportation, health management, and reducing loneliness for the elderly are important issues. In particular, elderly people living alone need a means of safe transportation that reduces loneliness in their daily lives. Therefore, a system is needed that can grasp the emotions and health status of elderly people in real time and provide the necessary support.

[1895] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for analyzing the elderly person's voice data using a voice recognition system, means for recognizing the elderly person's emotion using an emotion engine, means for generating a response based on the emotion information, and means for operating the navigation system of the autonomous vehicle. This makes it possible to grasp the emotions and health status of the elderly person in real time and support safe and secure travel.

[1896] "Elderly" refers to individuals who, as they grow older, require some form of assistance or care for their daily lives and activities.

[1897] The "means for receiving voice" is a device such as a microphone for capturing the speech of the elderly person.

[1898] The "means for converting speech to text" is a speech recognition system that has the function of converting speech data into text information.

[1899] The "means for analyzing text and generating appropriate responses" is a system that analyzes the converted text data using natural language processing and generates responses appropriate for the dialogue.

[1900] The "means for synthesizing a response using the voice of a registered family member" is a system having a function for synthesizing the generated text-based response using the voice of a family member.

[1901] The "means for providing synthesized voice to the elderly" is a device that plays synthesized voice data to the elderly through a speaker or the like.

[1902] The "means for analyzing voice data of an elderly person using a voice recognition system" is a voice recognition technology for analyzing voice data of an elderly person and generating an appropriate response.

[1903] The "means for recognizing the emotions of the elderly using an emotion engine" is a system that analyzes and extracts the emotional state of the elderly from voice data and text data.

[1904] The "means for generating a response based on emotional information" is a system that generates an appropriate emotional response based on the recognized emotional information.

[1905] "Means for operating the navigation system of an autonomous vehicle" refers to a system that guides the autonomous vehicle to its destination and controls its driving.

[1906] A "means for continuous conversation recording" is a device or software feature that continuously records conversations between the senior and the system.

[1907] "Means for transmitting conversation data to a server" refers to a technology for transferring recorded conversation data to a central server.

[1908] The "means for notifying the family of the analysis results" is a system that promptly notifies the family of the results of the analysis performed by the server.

[1909] The "means for detecting signs of dementia and emotional abnormalities" is a system that analyzes conversation data and emotional states to find abnormal patterns and early signs of dementia.

[1910] "Means for sending alerts" refers to a system that promptly sends a warning to relevant parties when an abnormality or emergency is detected.

[1911] This invention relates to an autonomous vehicle system for assisting the elderly, aiming to provide safe transportation and health management for the elderly, as well as to reduce feelings of loneliness. This system is realized by combining cloud computing technology, voice recognition technology, emotion recognition engines, and autonomous driving technology.

[1912] In order to implement the present invention, the following hardware and software are used.

[1913] Hardware used:

[1914] 1. Self-driving vehicle: A vehicle equipped with self-driving technology and equipped with built-in sensors, speakers, microphones, and cameras.

[1915] 2. Microphone: A device for capturing the speech of elderly people in the vehicle.

[1916] 3. Speaker: A device that provides synthesized speech to the elderly.

[1917] 4. Camera: A device used to monitor the movements of elderly people and detect any abnormalities.

[1918] Software used:

[1919] 1. Speech recognition system: A platform for converting the elderly's speech into text (e.g., Google Speech-to-Text API).

[1920] 2. Emotion recognition engine: A machine learning model that recognizes the emotions of the elderly and generates appropriate responses (e.g., IBM Watson Tone Analyzer).

[1921] 3. Speech synthesis system: A platform for synthesizing responses in the voices of family members (e.g., Amazon Polly).

[1922] 4. Autonomous vehicle control systems: APIs for controlling navigation in autonomous vehicles (e.g., Tesla Autopilot API).

[1923] A natural language description of the program:

[1924] Audio capture and transmission:

[1925] The server receives the voice data of the elderly person captured by the microphone in the autonomous vehicle, and the voice data is transmitted to the server in real time.

[1926] Voice Recognition:

[1927] The server uses a speech recognition system (Google Speech-to-Text API) to convert the received voice data into text data, which is a preparatory step for analyzing the conversation content.

[1928] Emotion Recognition and Analysis:

[1929] Using an emotion recognition engine (IBM Watson Tone Analyzer), the emotions of the elderly are recognized from the converted text data. Based on the recognized emotional information, an appropriate response is generated.

[1930] Response generation and serving:

[1931] The server uses the data necessary to generate a response and synthesizes a response based on the family member's voice using a speech synthesis system (Amazon Polly). This synthesized voice is then provided to the elderly person through the vehicle's speakers.

[1932] Navigation controls:

[1933] The system analyzes the elderly person's speech to identify destinations and instructions, and uses the autonomous vehicle control system (Tesla Autopilot API) to provide the necessary navigation. For example, if the system detects an utterance such as "I want to go to the hospital," it will set the navigation system to head to the nearest hospital.

[1934] Examples:

[1935] For example, if an elderly person gets into an autonomous vehicle and says, "I want to go to the park today," the following series of processes will take place.

[1936] 1. Voice capture: The microphone captures "I want to go to the park today."

[1937] 2. Speech recognition: The server converts the utterance into text: "I want to go to the park today."

[1938] 3. Emotion recognition: The emotion engine recognizes the emotion of "fun."

[1939] 4. Navigation: The vehicle sets the destination as the park, and a family member's voice guides them, saying, "We're heading to the park. Have fun."

[1940] 5. Anomaly detection: No emergency notification is required unless an anomaly is detected.

[1941] Example prompt sentence:

[1942] Please provide an example of creating a program for an autonomous driving system to assist the elderly, which converts voice data into text, recognizes emotions, and generates responses. Use a speech recognition system for speech recognition, an emotion recognition engine for emotion recognition, a speech synthesis system for speech synthesis, and an autonomous vehicle control system for autonomous driving control.

[1943] This invention uses an elderly assistance autonomous driving vehicle system to effectively realize safe and secure transportation, health management, and reduction of feelings of loneliness for the elderly.

[1944] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1945] Step 1:

[1946] The server receives the elderly person's voice data captured by the microphone in the autonomous vehicle. The input in this step is the elderly person's voice, and the output is the voice data stored on the server. Specifically, the microphone captures the voice as digital data and transmits it to the server in real time.

[1947] Step 2:

[1948] The server uses a speech recognition system (voice recognition system) to convert the received voice data into text data. The input in this step is the voice data received in step 1, and the output is the converted text data. In concrete terms, the server sends the voice data to the speech recognition system and receives the result as text.

[1949] Step 3:

[1950] The server uses an emotion recognition engine (emotion recognition engine) to recognize the emotions of the elderly from the text data. The input in this step is the text data obtained in step 2, and the output is the recognized emotion information. Specifically, the server inputs the text data into the emotion recognition engine and outputs the emotion information.

[1951] Step 4:

[1952] The server generates a response based on the emotional information and uses a speech synthesis system to synthesize the response in the voice of the family member. The input in this step is the emotional information obtained in step 3 and the generated response text, and the output is synthesized voice data. Specifically, the server generates a response text according to the emotional information, sends it to the speech synthesis system, and synthesizes it in the voice of the family member.

[1953] Step 5:

[1954] The server sends the synthesized voice data to the autonomous vehicle and provides it to the elderly through the speaker. The input in this step is the synthesized voice data obtained in step 4, and the output is the voice to be played to the elderly. Specifically, the server sends the synthesized voice data to the vehicle's audio system and plays it through the speaker.

[1955] Step 6:

[1956] The server analyzes the destination and instructions and sends them to the autonomous vehicle control system to operate the autonomous vehicle's navigation system. The input in this step is the text data obtained from step 2, and the output is the vehicle's destination setting and driving control. Specifically, the server extracts destination information from the elderly person's speech and sends instructions to the autonomous vehicle control system to perform navigation.

[1957] Step 7:

[1958] The server analyzes the elderly person's health condition and emotional information, and sends an emergency notification to family members or medical institutions if an abnormality is detected. The input in this step is the emotional information obtained in step 3 and other data collected in steps 1 and 2, and the output is the notification to be sent. Specifically, the server continuously monitors the data, and if an abnormality is detected, it triggers the notification system to send an emergency notification.

[1959] Through the above processing steps, the present invention can realize safe and secure transportation and health management as an elderly assistance autonomous driving vehicle system.

[1960] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1961] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1962] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1963] [Fourth embodiment]

[1964] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1965] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1966] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1967] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1968] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1969] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1970] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1971] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1972] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1973] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1974] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1975] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1976] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1977] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time.

[1978] Feature Description

[1979] 1. Ability to talk to the elderly using the voices of children and grandchildren

[1980] Subject: Server

[1981] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[1982] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[1983] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[1984] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[1985] The device plays the received audio to the elderly person through a speaker.

[1986] Specific examples

[1987] An elderly person speaks to the device, asking, "What should I eat today?"

[1988] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[1989] 2. Ability to record conversations with elderly people and provide information to their families

[1990] Subject: Device

[1991] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[1992] The server converts the received conversation data into text and analyzes the content.

[1993] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[1994] Specific examples

[1995] An elderly person says, "My shoulder has been hurting lately."

[1996] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[1997] 3. Alerts for dementia and depression

[1998] Subject: Server

[1999] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[2000] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[2001] Specific examples

[2002] Elderly people say, "I don't know what I'm living for."

[2003] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[2004] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[2005] Subject: Device

[2006] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[2007] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[2008] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[2009] Specific examples

[2010] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[2011] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[2012] 5. Fall detection and emergency contact function

[2013] Subject: Camera

[2014] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[2015] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[2016] Specific examples

[2017] An elderly person falls in the living room.

[2018] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[2019] As described above, the present invention is a system that comprehensively supports the health management of elderly people, reduces feelings of loneliness, and alleviates anxiety for their families. This system allows elderly people to live their daily lives with peace of mind, and allows their families to follow up in a timely manner.

[2020] The processing flow will be explained below.

[2021] A function that allows you to talk to the elderly using the voices of children and grandchildren

[2022] Step 1:

[2023] The user (elderly person) speaks into the terminal.

[2024] As a concrete example, say, "What shall we eat today?"

[2025] Step 2:

[2026] The terminal receives the elderly person's voice and transmits the voice data to the server.

[2027] The audio data is transferred to the server in real time.

[2028] Step 3:

[2029] The server converts the received voice data into text.

[2030] A speech recognition algorithm converts the speech data into a string of characters.

[2031] Step 4:

[2032] The server analyzes the converted text data and generates an appropriate response.

[2033] Use natural language processing (NLP) techniques to understand the context of text data.

[2034] Step 5:

[2035] The server synthesizes a response based on voice samples from family members.

[2036] A response message is synthesized using the voice of a pre-registered family member.

[2037] Step 6:

[2038] The server transmits the synthesized voice data to the terminal.

[2039] The generated voice data is sent back to the elderly person.

[2040] Step 7:

[2041] The device plays the synthesized speech through the speaker.

[2042] The elderly person hears a family member respond, "Maybe fish would be good?"

[2043] A function that records conversations with elderly people and provides information to their families

[2044] Step 1:

[2045] The device constantly records conversations with the elderly.

[2046] Record all your everyday conversations.

[2047] Step 2:

[2048] The recorded conversation data is periodically sent to a server.

[2049] For example, a data packet is sent every hour.

[2050] Step 3:

[2051] The server converts the received voice data into text.

[2052] Convert the speech into text using voice recognition technology.

[2053] Step 4:

[2054] The server analyzes the converted text data.

[2055] Natural language processing technology is used to extract important information and anomalies.

[2056] Step 5:

[2057] Based on the analysis results, the information to be notified to the family is determined.

[2058] If an anomaly is detected, detailed information about it will also be included.

[2059] Step 6:

[2060] The server sends the analysis results to the family.

[2061] Notifications are given in real time.

[2062] Examples:

[2063] An elderly person says, "My shoulder has been hurting lately."

[2064] The server analyzes the information and sends a notification to the family about the "shoulder pain."

[2065] A feature that sends alerts for dementia and depression

[2066] Step 1:

[2067] The server periodically analyzes conversation data with the elderly.

[2068] The analysis targets daily conversation logs.

[2069] Step 2:

[2070] The server detects changes in grammatical structure and sentiment.

[2071] It uses natural language processing techniques and sentiment analysis algorithms.

[2072] Step 3:

[2073] If an anomaly is detected, the server generates an alert.

[2074] Identify signs of dementia and emotional abnormalities.

[2075] Step 4:

[2076] The server generates alerts and sends them to family members and medical institutions.

[2077] The alert is sent as an emergency notification.

[2078] Examples:

[2079] Elderly people say, "I don't know what I'm living for."

[2080] The server performs emotion analysis, detects signs of depression, and sends emergency notifications.

[2081] Functions to prevent forgetting to take medicine and assist with meals and shopping

[2082] Step 1:

[2083] The server sets medication times and meal schedules.

[2084] Set a daily reminder.

[2085] Step 2:

[2086] The device will send a reminder notification at the set time.

[2087] A notification will play saying "Time for your medicine."

[2088] Step 3:

[2089] When the elderly person responds with a voice of confirmation, the device records the information and sends it to the server.

[2090] "Thanks, I drank it," he replied.

[2091] Step 4:

[2092] The server notifies the family of the reminder information.

[2093] Notifies you that confirmation of missed dose has been completed.

[2094] Step 5:

[2095] The server learns the elderly person's lifestyle patterns and generates a shopping list.

[2096] Make a list of the items you need.

[2097] Step 6:

[2098] The device notifies the elderly person of the shopping list.

[2099] Announce the list contents by voice.

[2100] Examples:

[2101] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[2102] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[2103] Fall detection and emergency contact function

[2104] Step 1:

[2105] The camera monitors the footage at all times.

[2106] Monitor overall behavior.

[2107] Step 2:

[2108] The camera detects an abnormality.

[2109] For example, a falling motion is detected.

[2110] Step 3:

[2111] Video data is sent to the server in real time.

[2112] Send data immediately.

[2113] Step 4:

[2114] The server identifies the fall and generates an emergency alert.

[2115] Prepare notifications for medical institutions and families.

[2116] Step 5:

[2117] The server sends emergency alerts to family and medical institutions.

[2118] Encourage a rapid response.

[2119] Examples:

[2120] An elderly person falls in the living room.

[2121] The camera detects an abnormality and notifies the server.

[2122] The server sends emergency alerts to family and medical institutions.

[2123] These are the specific steps of the program processing, which will enable health management for elderly people living alone, reduce loneliness, and provide peace of mind to their families.

[2124] Example 1

[2125] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2126] There is a need for health management for the elderly, reducing loneliness, and improving family security. However, systems that enable appropriate communication and health monitoring for elderly people living alone have not yet been fully developed. When elderly people live in isolated situations, the lack of daily communication and the inability to detect health abnormalities early are problems. In particular, there is a lack of means to detect early signs of dementia and depression and provide appropriate treatment.

[2127] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2128] In this invention, the server includes means for receiving the elderly person's voice, means for converting the received voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the voice of a registered family member, means for providing the synthesized voice to the elderly person, means for saving voice samples provided by the family member and extracting features, and natural language processing means for analyzing the elderly person's speech and generating a response related to life support. This enables natural communication with the elderly, health management, reduction of loneliness, and early detection of signs of dementia and depression.

[2129] The "means for receiving the voice of the elderly person" is a function for acquiring the voice uttered by the elderly person using an input device such as a microphone.

[2130] The "means for converting received voice into text" is a function that converts voice data into corresponding text data using voice recognition technology.

[2131] "Means for analyzing the converted text and generating an appropriate response" refers to a function that uses natural language processing (NLP) technology to analyze text data and generate an appropriate response based on the content.

[2132] The "means for synthesizing the response with the voice of a registered family member" is a technology for converting text data into voice data using the voice characteristics of a designated family member.

[2133] The "means for providing synthesized voice to the elderly" is a function for playing back synthesized voice to the elderly using a voice output device such as a speaker.

[2134] The "means for saving voice samples provided by family members and extracting features" is a function for saving voice data provided by family members and extracting voice features.

[2135] "Natural language processing means for analyzing elderly people's statements and generating responses related to life support" is a technology that analyzes the content of elderly people's statements and generates information and actions necessary for life support based on the analysis results.

[2136] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and alleviate feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. This system is realized by integrating speech recognition, natural language processing, and speech synthesis technologies.

[2137] Hardware and software used

[2138] Hardware

[2139] High-Precision Microphone

[2140] speaker

[2141] camera

[2142] server

[2143] Devices (smartphones, tablets, PCs, etc.)

[2144] software

[2145] Speech recognition: A speech recognition API for converting voice data into text (e.g., Google Cloud Speech-to-Text API)

[2146] Natural Language Processing (NLP): NLP models (e.g., AWS Comprehend) to analyze text data and generate appropriate responses.

[2147] Speech synthesis: A speech synthesis API to convert text to speech (e.g. IBM Watson Text-to-Speech API)

[2148] Database: A database for storing audio samples and analysis data (e.g., Google Cloud Storage)

[2149] Explanation of the specific functions of the system

[2150] 1. Ability to talk to the elderly using the voices of children and grandchildren

[2151] Subject: Server, Terminal, User

[2152] The server stores voice samples provided by family members in a database. It also extracts voice features using a machine learning model. When the user (elderly person) speaks into the device, the device captures the voice data and sends it to the server. The server converts the received voice data into text using the Google Cloud Speech-to-Text API and analyzes the text data using a natural language processing (NLP) model. An appropriate response is generated, and speech is synthesized using the family member's voice using the IBM Watson Text-to-Speech API, and the speech data is sent to the device. The device then plays the received audio to the elderly person through a speaker.

[2153] Specific examples

[2154] An elderly person speaks to the device, asking, "What should I eat today?"

[2155] The server synthesizes the family member's voice saying, "Maybe fish would be good?" and plays it back to the elderly person via the device.

[2156] Example prompt

[2157] "When you say to a device, 'What should I eat today?' how does the server process the data and generate the voice response?"

[2158] 2. Ability to record conversations with elderly people and provide information to their families

[2159] Subject: Terminal, Server

[2160] The device continuously records conversations with the elderly. The recorded conversation data is periodically sent to a server, which converts the voice data into text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing (NLP) model, and if particularly important information or anomalies are detected, the server notifies the family using the Twilio API.

[2161] Specific examples

[2162] An elderly person says, "My shoulder has been hurting lately."

[2163] The server analyzes the information and notifies the family of any abnormalities, allowing them to plan a visit.

[2164] Example prompt

[2165] "When the audio data recorded on the device is sent to the server, how will family members be notified if an abnormality is detected?"

[2166] 3. Alerts for dementia and depression

[2167] Subject: Server

[2168] The server periodically converts conversation data with the elderly person into text using the Google Cloud Speech-to-Text API, then analyzes the text data using a natural language processing (NLP) model. If signs of dementia or abnormal emotions are detected, an alert is generated and notifies family members and medical institutions using the Twilio API.

[2169] Specific examples

[2170] Elderly people say, "I don't know what I'm living for."

[2171] The server performs emotion analysis, detects signs of depression, and sends an emergency notification to family members and medical institutions.

[2172] Example prompt

[2173] "If an elderly person makes a statement that indicates signs of depression, what analysis does the server perform and how does it generate a notification?"

[2174] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[2175] Subject: Terminal, Server

[2176] The device uses the Google Calendar API to manage medication times and meal schedules, and sends a reminder at the set time: "It's time for your medicine." When the user (elderly person) responds with a voice confirmation, the voice is sent to the server and converted into text using the Google Cloud Speech-to-Text API. The server then notifies the family of the confirmation information. The server also learns the elderly person's lifestyle patterns, generates a shopping list, and provides it via voice from the device.

[2177] Specific examples

[2178] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[2179] When the elderly person replies, "Thank you, I drank it," the server records this and notifies the family.

[2180] Example prompt

[2181] "How can we remind and confirm elderly people to prevent them from forgetting to take their medication, and notify their families of this information?"

[2182] 5. Fall detection and emergency contact function

[2183] Subject: camera, server

[2184] The camera constantly monitors the video and runs an algorithm to detect anomalies. If an elderly person falls, the camera detects the movement and sends the video data in real time to a server. The server then confirms the fall and uses the Twilio API to send an emergency notification to family members and medical institutions.

[2185] Specific examples

[2186] An elderly person falls in the living room.

[2187] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[2188] Example prompt

[2189] "How does the fall detection feature work and how does it send emergency notifications after detection?"

[2190] As described above, this system recognizes the voice of the elderly person, generates an appropriate response using natural language processing, and then uses voice synthesis technology to respond in the voice of a family member, thereby providing comprehensive support for managing the elderly person's health, reducing feelings of loneliness, and providing a sense of security to their family.

[2191] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2192] Specific processing steps of the system program

[2193] Feature 1: Ability to talk to the elderly using the voices of children and grandchildren

[2194] Subject: Server, Terminal, User

[2195] Step 1:

[2196] Subject: Server

[2197] Voice samples provided by family members are stored in a database.

[2198] (Input) Voice samples recorded by family members

[2199] (Processing) The server receives the audio samples and stores them in a database.

[2200] (Output) Voice samples of family members stored in the database

[2201] Specific operation: The server receives the voice data and stores it in a database.

[2202] Step 2:

[2203] Subject: Server

[2204] The voice sample is analyzed using a machine learning model to extract features.

[2205] (Input) Audio samples stored in the database

[2206] (Processing) Apply a voice feature extraction algorithm to extract features

[2207] (Output) Extracted feature data

[2208] Specific operation: The server applies a feature extraction algorithm to the voice sample and stores the features in a database.

[2209] Step 3:

[2210] Subject: User

[2211] The user speaks into the terminal.

[2212] (Input) User's voice

[2213] (Processing) High-precision microphone captures audio

[2214] (Output) Captured audio data

[2215] Specific operation: The user speaks to the device, "What should I eat today?"

[2216] Step 4:

[2217] Subject: Device

[2218] The terminal transmits the voice data to the server.

[2219] (Input) Captured audio data

[2220] (Processing) Send the audio data to the server

[2221] (Output) Audio data received by the server

[2222] Specific operation: The device sends the captured audio data to a server via the Internet.

[2223] Step 5:

[2224] Subject: Server

[2225] Converts audio data into text.

[2226] (Input) Audio data received by the server

[2227] (Processing) Convert speech to text using a speech recognition API (e.g., Google Cloud Speech-to-Text API)

[2228] (Output) Converted text data

[2229] What happens: The server uses the Google Cloud Speech-to-Text API to convert the speech "What should we eat today?" into text.

[2230] Step 6:

[2231] Subject: Server

[2232] Analyze text data using an NLP model and generate appropriate responses.

[2233] (Input) Converted text data

[2234] (Processing) Analyze the text using an NLP model (e.g. AWS Comprehend) and generate an appropriate response

[2235] (Output) The generated response text

[2236] Specific behavior: The server generates a response to the question "What should we eat?" with "Maybe fish would be good?"

[2237] Step 7:

[2238] Subject: Server

[2239] The response text is converted into speech using the speech synthesis API.

[2240] (Input) Generated response text

[2241] (Processing) Convert text to speech using a speech synthesis API (e.g. IBM Watson Text-to-Speech API)

[2242] (Output) Synthesized voice data

[2243] Specific behavior: The server converts the text "Maybe fish would be good?" into speech.

[2244] Step 8:

[2245] Subject: Server

[2246] The synthesized voice data is transmitted to the terminal.

[2247] (Input) Synthesized voice data

[2248] (Processing) Send audio data to the terminal

[2249] (Output) Audio data received on the device

[2250] Specific operation: The server sends the synthesized voice data to the terminal.

[2251] Step 9:

[2252] Subject: Device

[2253] The device plays the audio.

[2254] (Input) Audio data received by the device

[2255] (Processing) Playing audio through speakers

[2256] (Output) Audio played to the elderly

[2257] Specific operation: The device plays a voice message from the speaker saying, "Maybe a fish would be good?"

[2258] You can add specific processing steps below for other features as well, but you can distinguish the steps for each feature based on this format.

[2259] If you have steps for other features below, please list them in a similar format.

[2260] (Application example 1)

[2261] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2262] In modern society, the number of elderly people living alone is increasing, and health management and reducing feelings of loneliness have become social issues. Furthermore, when elderly people living alone use self-driving vehicles, ensuring safety and communication with their families are important. However, conventional systems do not adequately provide comprehensive solutions to these issues. Therefore, there is a need for a system that integrates safe and secure transportation means for the elderly with support for their daily lives.

[2263] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2264] In this invention, the server includes means for receiving the elderly person's voice, means for converting the voice into text, means for analyzing the converted text and generating an appropriate response, means for synthesizing the response with the family member's voice, means for providing synthesized voice, means for setting a destination for the autonomous vehicle and providing route guidance, means for monitoring the driving situation and detecting abnormalities, and means for notifying the family member in real time when an abnormality occurs. This allows the elderly person to receive route guidance from the family member's voice when using the autonomous vehicle, improving safety and enabling the family member to understand the elderly person's situation in real time.

[2265] · An "elderly person" is an individual who, due to their advanced age, requires assistance in daily living.

[2266] "Means for receiving voice" refers to a device or method for recording and capturing the voice uttered by the elderly person and inputting it into the system.

[2267] "Means for converting speech to text" means the process or technology that converts received speech data into corresponding text data using speech recognition technology.

[2268] "Means of analyzing text" refers to natural language processing techniques used to understand meaning and intent based on converted text data.

[2269] - "Means for generating a response" refers to an algorithm or program that generates an appropriate response based on the analysis results.

[2270] "Means for synthesizing using family members' voices" refers to a method of reproducing the generated response using voice synthesis technology using pre-registered family members' voice data.

[2271] "Means for providing synthesized speech" means a speaker or other audio playback device that delivers synthesized speech to the senior citizen.

[2272] "Means for setting a destination in an autonomous vehicle" refers to the methods and technologies for inputting and setting the destination desired by the elderly person into an autonomous vehicle.

[2273] "Means for providing directions" refers to systems and technologies that allow autonomous vehicles to provide directions to seniors toward a set destination.

[2274] "Means for monitoring the driving situation" refers to sensors, cameras, and algorithms that monitor the driving situation in real time while the autonomous vehicle is driving to ensure safety.

[2275] "Means for detecting abnormalities" refers to techniques and methods for recognizing and identifying hazards and abnormal events that may occur during operation.

[2276] "Means of notifying family members in real time" refers to communication technologies and methods for immediately notifying family members of the situation when an abnormality is detected.

[2277] To implement this invention, it is necessary to build a system that receives the voice of the elderly person and synthesizes it with the voices of family members. This system uses the following hardware and software.

[2278] Hardware used

[2279] 1. Smartphone: Used to record the elderly person's voice and send the recording data to the server.

[2280] 2. Self-driving vehicles: Used as transportation for the elderly, with functions such as route guidance and safety monitoring.

[2281] 3. Speaker: Used as a device to transmit synthesized voices of family members to the elderly.

[2282] Software used

[2283] 1. pyaudio: A library for recording the voices of elderly people.

[2284] 2. wave: A library for saving and playing recorded audio.

[2285] 3. transformers (Wav2Vec2): A library for converting audio to text.

[2286] 4. gTTS: A library for synthesizing text in the voices of family members.

[2287] 5. smtplib: A library for implementing notification functions.

[2288] Data processing and calculation

[2289] The server first receives the elderly person's voice and records it using the pyaudio library. The recorded voice data is then temporarily saved using the wave library. The voice data is then converted to text using the Wav2Vec2 model in transformers. The converted text data is analyzed using natural language processing techniques to generate an appropriate response. The generated response is then synthesized in the voice of a family member using the gTTS library and provided to the elderly person through a speaker as synthesized speech.

[2290] If an anomaly is detected or if specific keywords are found, a real-time notification is sent to family members using the smtplib library, containing the analysis results and the elderly person's current condition.

[2291] Specific examples

[2292] For example, if an elderly person says, "Shall I confirm your next hospital appointment?", the system works as follows: First, the smartphone records this speech and sends it to the server. The server converts the recorded speech into text using the Wav2Vec2 model and analyzes its content. An appropriate response is generated, such as "Yes, Grandma. Your next hospital appointment is tomorrow at 10:00 AM," using a synthesized voice from a family member's voice, and this is played to the elderly through the speaker.

[2293] Prompt Sentence Examples

[2294] 1. "Grandpa, should I confirm your next doctor's appointment?"

[2295] 2. "Grandma, have you taken your morning medicine?"

[2296] 3. "Where is Grandpa going? Is he driving safely?"

[2297] This will allow elderly people to use self-driving vehicles with peace of mind, and will also enable family members to keep track of the elderly person's situation in real time.

[2298] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2299] Step 1:

[2300] The smartphone records the elderly person's voice. The device uses the pyaudio library to record 10 seconds of voice data uttered by the elderly person. The input here is the elderly person's voice, and the output is the recorded voice data (audio file).

[2301] Step 2:

[2302] The recorded audio data is sent to the server. The device uses the wave library to save the audio file and uploads it to the server. The input is the recorded audio file, and the output is the audio file saved on the server.

[2303] Step 3:

[2304] The server converts the audio data to text. The server uses the Wav2Vec2 model from transformers to convert the audio file to text. The input is the audio file, and the output is the converted text.

[2305] Step 4:

[2306] The server analyzes the text data and generates an appropriate response. Natural language processing technology is used to analyze the meaning and intent of the converted text data and generate an appropriate response text. The input is the text data, and the output is the generated response text.

[2307] Step 5:

[2308] The server synthesizes the response text in the voice of the family member. The server uses the gTTS library to synthesize the generated response text in the voice of the pre-registered family member and create an audio file. The input is the response text, and the output is the synthesized audio file.

[2309] Step 6:

[2310] The synthesized voice is provided to the elderly through a speaker. The server sends the synthesized voice to the terminal, and the terminal plays the voice through the speaker. The input is the synthesized voice file, and the output is the voice played to the elderly.

[2311] Step 7:

[2312] Monitors the driving situation and detects abnormalities. Autonomous vehicles use sensors and cameras to monitor the driving situation in real time to ensure safety. The input is real-time data during driving, and the output is analysis results and abnormality detection signals.

[2313] Step 8:

[2314] If an abnormality is detected, a notification is sent to family members in real time. The server uses the smtplib library to notify family members of the situation by email when it receives an abnormality detection signal. The input is the abnormality detection signal, and the output is the alert notification sent to family members.

[2315] Specific example prompts

[2316] "Grandpa, should I confirm your next doctor's appointment?"

[2317] "Grandma, have you taken your morning medicine?"

[2318] "Where is Grandpa going? Are you driving safely?"

[2319] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2320] This system uses the voices of family members to communicate with elderly people living alone, helping them manage their health and reduce feelings of loneliness. It also provides a sense of security to family members living far away by informing them of the elderly person's condition in real time. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it achieves more effective responses and anomaly detection.

[2321] Feature Description

[2322] 1. Ability to talk to the elderly using the voices of children and grandchildren

[2323] Subject: Server

[2324] The server stores voice samples provided by family members in a database, which is then analyzed with a machine learning model for feature extraction.

[2325] When a user (elderly person) speaks into the terminal, the terminal captures the voice data and sends it to the server.

[2326] The server converts the received voice data into text, analyzes the text data using natural language processing (NLP), and generates an appropriate response.

[2327] The generated response is synthesized with the family member's voice, and the voice data is sent to the terminal.

[2328] The device plays the received audio to the elderly person through a speaker.

[2329] Specific examples

[2330] An elderly person speaks to the device, asking, "What should I eat today?"

[2331] The server synthesizes a child's voice saying, "Wouldn't a fish be good?" and plays it to the elderly person via the device.

[2332] 2. Ability to record conversations with elderly people and provide information to their families

[2333] Subject: Device

[2334] The device constantly records conversations with the elderly person, and the conversation data is periodically sent to a server.

[2335] The server converts the received conversation data into text and analyzes the content.

[2336] If particularly important information or abnormalities are detected, the analysis results will be notified to the family.

[2337] Specific examples

[2338] An elderly person says, "My shoulder has been hurting lately."

[2339] The server analyzes the information and sends a notification about the "shoulder pain" to the family, who can then plan a visit.

[2340] 3. Alerts for dementia and depression

[2341] Subject: Server

[2342] The server periodically analyzes conversation data with the elderly to detect changes in grammatical structure and emotions.

[2343] If signs of dementia or emotional abnormalities are detected, an alert will be generated and sent to family and medical institutions.

[2344] Specific examples

[2345] Elderly people say, "I don't know what I'm living for."

[2346] The server performs emotion analysis, detects signs of depression, and sends emergency notifications to family members and medical institutions.

[2347] 4. Functions to prevent forgetting to take medicine and assist with meals and shopping

[2348] Subject: Device

[2349] The device notifies the elderly of pre-set times for taking medicine and meal schedules, and when the set times arrive, the device will remind them, "It's time to take your medicine."

[2350] When the elderly person responds with a voice confirmation, the device records it, sends it to the server, and notifies the family.

[2351] The server learns the elderly person's lifestyle patterns, generates shopping lists when needed, and provides them via voice from the terminal.

[2352] Specific examples

[2353] Every day at 8:00 a.m., the device will notify you, "Good morning. It's time for your medicine."

[2354] When the elderly person replies, "Thank you, I drank it," the server records it and notifies their family.

[2355] 5. Fall detection and emergency contact function

[2356] Subject: Camera

[2357] The camera constantly monitors the video and runs an algorithm to detect abnormalities. If an elderly person falls, the system detects the movement and sends the video data to a server in real time.

[2358] The server identifies the fall and sends an emergency notification to the family and medical authorities.

[2359] Specific examples

[2360] An elderly person falls in the living room.

[2361] The camera detects an abnormality and notifies the server, which then sends an emergency alert to family members and medical institutions.

[2362] Features of inventions that combine emotion engines

[2363] 1. Emotion recognition and response generation using an emotion engine

[2364] Subject: Server

[2365] The server uses an emotion engine to recognize emotions from the elderly person's voice data, extracts the elderly person's emotions from the voice data, and generates a response based on that information.

[2366] Depending on the recognized emotion, the server generates an appropriate response and provides it to the elderly person as a synthesized voice.

[2367] Specific examples

[2368] An elderly person says, "I feel a little lonely today."

[2369] The server uses an emotion engine to recognize the emotion "loneliness" and generates a response such as "I see you're feeling lonely. Let's think of something we can do together."

[2370] 2. Anomaly detection and notification using emotion engine

[2371] Subject: Server

[2372] If the server detects negative emotions in an elderly person using the emotion engine, it generates an alert and notifies family members and relevant organizations.

[2373] If an anomaly is detected, a notification system will be activated to ensure a rapid response.

[2374] Specific examples

[2375] An elderly person says, "Nothing is fun anymore."

[2376] The server uses an emotion engine to detect emotions such as "despair" and "deep sadness" and sends emergency notifications to family members and medical institutions.

[2377] The above is a specific embodiment of the present invention that combines an emotion engine. This makes it possible to more accurately grasp the emotional state of elderly people living alone and respond appropriately. Furthermore, since it can respond sensitively to changes in emotions, it can also support the mental health of the elderly.

[2378] The processing flow will be explained below.

[2379] Emotion recognition and response generation using an emotion engine

[2380] Step 1:

[2381] The user (elderly person) speaks into the terminal.

[2382] As a specific example, say, "I feel a little lonely today."

[2383] Step 2:

[2384] The terminal receives the elderly person's voice and transmits the voice data to the server.

[2385] Step 3:

[2386] The server converts the received voice data into text.

[2387] A speech recognition algorithm converts the speech data into a string of characters.

[2388] Step 4:

[2389] The server passes the converted text data and voice data to the emotion engine.

[2390] The emotion engine analyzes the tone and context of the voice to recognize emotions.

[2391] Step 5:

[2392] The server generates an appropriate response based on the emotional data recognized by the emotion engine.

[2393] Using natural language processing technology, it generates a response such as, "I see you're feeling lonely. Let's think of something we can do together."

[2394] Step 6:

[2395] The server generates a response that is then synthesized into the family member's voice.

[2396] A response message is synthesized using the voice of a pre-registered family member.

[2397] Step 7:

[2398] The server transmits the synthesized voice data to the termin...

Claims

1. a means for receiving the voice of the senior citizen; means for converting the received speech into text; means for analyzing the converted text and generating an appropriate response; means for synthesizing the response in the voice of a registered family member; a means for providing the synthesized speech to the senior citizen; A system including:

2. 10. The system of claim 1, A means of constantly recording conversations with the elderly, means for transmitting the recorded conversation data to a server; A means for analyzing the transmitted conversation data and notifying the family of the analysis result; The system further comprises:

3. 10. The system of claim 1, A method for analyzing conversation data with elderly people to detect signs of dementia and emotional abnormalities, A means of sending alerts if signs of dementia or abnormal emotions are detected; The system further comprises:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A