system

A voice-based system addresses communication challenges for elderly individuals by converting voice inputs to text for managing daily tasks and emergencies, enhancing their support and convenience.

JP2026062293APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Elderly people living alone face challenges in communication, especially during emergencies and daily support, due to limited means and unfamiliarity with digital technologies like online consultations and messaging.

Method used

A system that allows users to input information and requests via voice, which is converted to text, processed by a server, and executed to facilitate everyday conversations, message sending, medical appointments, and other operations.

Benefits of technology

Enables elderly individuals to easily manage daily life tasks and receive support, improving communication and convenience through a user-friendly voice-based interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062293000001_ABST
    Figure 2026062293000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】means for the user to input basic information, means for sending the basic information to the server, a server for storing the basic information in a database, means for notifying that the storage has been completed, means for voice-inputting daily conversations with the user, means for sending the voice data to the server, a server for converting the voice data into text, a server for generating a response based on the text, means for outputting the response to the user by voice, means for voice-inputting the recipient and the message content from the user, a server for sending the voice data to the server and sending a message to the recipient, means for voice-inputting the desired date and time of the user's medical appointment, means for sending the desired date and time to the server, a server for making a medical appointment based on the desired date and time, means for voice-inputting the user's requests, a server for sending the voice data to the server and executing the corresponding operations, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Elderly people living alone often have limited communication means in their daily lives, especially facing difficulties in case of emergencies or when daily support is needed. Also, they may be unfamiliar with digital technologies such as using online medical consultations and sending and receiving messages, and it has become a problem that it is difficult for them to enjoy the convenience of these technologies. There is a need for a comprehensive communication tool to solve these problems and easily support the daily lives of elderly people living alone.

Means for Solving the Problems

[0005] The present invention provides a system in which a user inputs basic information, which is then transmitted to a server and stored in a database. Furthermore, it includes means for inputting everyday conversations with the user via voice, transmitting the voice data to the server for conversion into text, generating a response, and outputting it to the user via voice. It also includes a function for inputting the recipient and message content via voice from the user, transmitting the voice data to the server, and sending the message to the recipient. In addition, it includes a function for inputting a desired date and time for a medical appointment via voice, transmitting the desired date and time to the server to make an online medical appointment, and a server that inputs user requests via voice, transmits the voice data to the server, and executes the corresponding operation. This provides a system that allows elderly people living alone to easily receive support in their daily lives and enables smooth support in emergencies and daily life.

[0006] A "user" refers to an elderly person living alone who uses the system to receive support for their daily life.

[0007] "Basic information" refers to personal information such as the user's name, contact information, and emergency contact information.

[0008] A "server" refers to a computer system that receives, processes, and stores data sent by users, and interacts with other systems as needed.

[0009] A "database" refers to a storage device where a server stores basic user information and other related data.

[0010] "Voice input" refers to the process of recording the voice spoken by the user and converting that voice data into a format that can be processed within the system.

[0011] "Audio data" refers to digital data generated from voice input.

[0012] "Converting to text" refers to the process of converting audio data into text data.

[0013] "Response" refers to the reply message generated by the system based on the user's voice input.

[0014] "Message content" refers to the text containing the information that the user wants to send to other recipients.

[0015] "Recipient" refers to an individual who receives messages or communications from the user.

[0016] "Medical appointment" refers to the date and time of the medical treatment agreed upon with a medical institution.

[0017] "Desired date and time" refers to the specific time when the user hopes to receive medical treatment or engage in other activities.

[0018] "Requests" refer to specific operations or requests that the user wants the system to perform.

[0019] "Operation" refers to the specific actions or procedures that the system performs according to the user's requests.

[0020] These definitions clarify the meanings of the important words used within the scope of the patent claims.

Brief Description of the Drawings

[0021] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6]It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0022] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0023] First, the language used in the following description will be explained.

[0024] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0025] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0026] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0027] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0029] [First Embodiment]

[0030] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0031] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0034] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0037] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0041] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0042] The communication tool for elderly people living alone according to the present invention is designed to support daily life and improve the convenience of communication. This system includes a user-friendly terminal and a server that processes data and provides necessary services.

[0043] Specifically, the user enters basic information through their device and sends that information to the server. The server then stores the information in a database and notifies the device that registration is complete. Through this process, the system can manage the user's basic information.

[0044] Furthermore, users can enjoy everyday conversations by speaking directly to the device. For example, if a user says, "The weather's nice today," the device records the voice and sends it to the server. The server converts the voice data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[0045] Furthermore, sending and receiving messages is easy. When a user says to the device, "Send a message to my son saying 'I'm doing well'," the device records the voice and sends it to the server. The server retrieves the recipient's (in this case, the son's) contact information from its database and sends the specified message to that contact. Once the server has finished processing, it notifies the device, and the device then informs the user by voice, "The message has been sent."

[0046] Making appointments is also easy for users. When a user says, "Please schedule an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data and makes an online appointment for the specified date and time. Once the appointment is complete, the server sends the information back to the device, which then notifies the user that "the appointment is complete."

[0047] Furthermore, it can also respond to user requests. For example, if a user says, "Add milk to my shopping list," the device sends voice data to the server, and the server adds the appropriate item to the shopping list. Once the server has finished processing, it notifies the device of the result, and the device informs the user of the result.

[0048] Through these processes, the system of the present invention provides elderly people living alone with the communication and support they need in their daily lives, significantly improving their convenience. Furthermore, since all operations are voice-based, it is designed to be easy to use even for users unfamiliar with digital technology.

[0049] The following describes the processing flow.

[0050] User Registration

[0051] Step 1:

[0052] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[0053] Step 2:

[0054] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[0055] Step 3:

[0056] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[0057] Step 4:

[0058] The server notifies the terminal that saving the user information is complete.

[0059] Step 5:

[0060] The device receives a notification and displays "Registration Complete".

[0061] Support for everyday conversation

[0062] Step 1:

[0063] The user speaks to the device. For example, they might say, "The weather's nice today."

[0064] Step 2:

[0065] The device records the user's voice and sends the audio data to the server.

[0066] Step 3:

[0067] The server uses a speech recognition engine to convert the audio data into text.

[0068] Step 4:

[0069] The server generates an appropriate response based on the converted text. Natural language processing is used in this process.

[0070] Step 5:

[0071] The server sends a response text to the terminal.

[0072] Step 6:

[0073] The device converts the received response into speech and plays it back to the user. For example, it might say, "What beautiful weather we have today."

[0074] Send message

[0075] Step 1:

[0076] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[0077] Step 2:

[0078] The device sends audio data to the server. The data is encoded in the appropriate format.

[0079] Step 3:

[0080] The server converts the audio data into text and analyzes the message content for the recipient.

[0081] Step 4:

[0082] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[0083] Step 5:

[0084] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[0085] Step 6:

[0086] The device informs the user of notifications it has received. For example, it might say, "A message has been sent."

[0087] Online medical consultation booking

[0088] Step 1:

[0089] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[0090] Step 2:

[0091] The device sends voice data to the server.

[0092] Step 3:

[0093] The server analyzes the voice data and converts the desired reservation date and time into text.

[0094] Step 4:

[0095] The server accesses the online medical appointment system and attempts to make a reservation for that date and time.

[0096] Step 5:

[0097] The server receives the reservation confirmation information and sends the result to the terminal.

[0098] Step 6:

[0099] The device plays a voice message to the user confirming that the reservation has been received. For example, it might say, "Your reservation is complete."

[0100] Implementing requested actions (such as updating the shopping list)

[0101] Step 1:

[0102] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[0103] Step 2:

[0104] The device sends voice data to the server.

[0105] Step 3:

[0106] The server analyzes the audio data and converts the request into text.

[0107] Step 4:

[0108] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list in the database.

[0109] Step 5:

[0110] The server confirms that the request has been completed and notifies the terminal.

[0111] Step 6:

[0112] The device will verbally inform the user of notifications it has received. For example, it might say, "Milk has been added to your shopping list."

[0113] (Example 1)

[0114] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0115] The objective of this invention is to address the lack of communication and support that elderly people living alone face in their daily lives. In particular, it aims to provide a voice-based interface that can be easily used even by elderly people who are unfamiliar with digital technology, enabling them to enjoy everyday conversations, obtain necessary information, send and receive messages, and easily make medical appointments.

[0116] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0117] In this invention, the server includes means for the user to input general information, means for transmitting the general information to an information processing device, and an information processing device for storing the general information in a storage device. This makes it possible for the user to easily register basic information, enjoy everyday conversations, send messages, make medical appointments, and more, all through voice.

[0118] A "user" is an individual who utilizes the system of the present invention.

[0119] "General information" refers to basic information such as name, address, and contact information entered by the user.

[0120] An "information processing device" is a device that processes information transmitted by a user and provides necessary services.

[0121] A "storage device" is a device used by an information processing device to permanently store information.

[0122] "Voice input" is a method of inputting information or instructions using voice.

[0123] "Speech" refers to sound information generated by the user's utterances.

[0124] A "document" is information in the form obtained when an information processing device converts speech into text.

[0125] A "response" is the answer generated by an information processing device in response to a user's inquiry.

[0126] A "recipient" is an individual or organization that receives a message from a user.

[0127] "Message content" refers to the information or notification that the user wants to send.

[0128] "Appointment scheduling" is the process of making a medical appointment for a date and time specified by the user.

[0129] "Desired date and time" refers to the specific date and time the user wishes to schedule an appointment.

[0130] A "request" is a specific demand or instruction that a user wants to perform through the system.

[0131] "Natural language processing" is a technology that enables information processing devices to understand human language and generate appropriate responses.

[0132] "Methods of confirming online" refer to methods of confirming procedures such as reservations via the internet.

[0133] "The relevant operation" refers to the specific actions or procedures that the system performs in response to the user's request.

[0134] This invention relates to a communication tool for elderly people living alone, and is a system that uses a voice-based interface to allow users to easily input basic information, enjoy daily conversations, send messages, make medical appointments, and perform other requests.

[0135] How users enter information

[0136] The user inputs general information by voice through the microphone into the terminal. The terminal has a recording function, and the input voice is saved as digital audio data. The terminal sends this digital audio data to an information processing device (server). Specifically, the data is transmitted securely using the HTTPS protocol.

[0137] The server receives the transmitted information and saves it to a storage device (e.g., a MySQL® database). It also notifies the terminal when saving is complete, and the terminal notifies the user that "registration is complete." A speech synthesis engine (e.g., Amazon Polly) is used for this notification.

[0138] Processing everyday conversations

[0139] When a user speaks to the device, for example, "The weather's nice today," the device records the audio and sends it to the server. The server uses Google® Cloud Speech-to-Text to convert the audio data into text. Then, it uses a generative AI model (e.g., OpenAI®'s GPT-3®) to generate an appropriate response and sends it to the device. The device plays the response using a speech synthesis engine, allowing the user to enjoy a natural conversation.

[0140] Send message

[0141] When the user says, "Send my son a message saying 'I'm doing well'," the device records the voice and sends it to the server. The server converts the voice data to text and retrieves the recipient's (the son's) contact information from its database. The server uses Twilio's SMS API to send the message to the recipient. Once the server notifies the user that processing is complete, the device informs the user that "the message has been sent."

[0142] Appointment

[0143] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the device sends the voice data to the server. The server converts the voice data into text and extracts the specified date and time information. The server uses the online medical consultation system's API to schedule the appointment. Once the appointment is complete, the server notifies the device that "the appointment is complete."

[0144] Responding to requests

[0145] When a user says, "Add milk to my shopping list," the device sends that voice data to the server. The server converts the voice data into text and adds the milk to the shopping list. Once the process is complete, the server notifies the device that "Milk has been added to your shopping list."

[0146] Examples of specific cases and prompt statements

[0147] For example, if a user says, "The weather's nice today," the following specific actions will be taken:

[0148] Voice input: "The weather is nice today."

[0149] Server processing: Converts audio data to text and generates responses using a generative AI model.

[0150] Response: "It really is a beautiful day today."

[0151] Example of a prompt:

[0152] "Please tell me what kind of response you should generate if a user says, 'The weather is nice today.'"

[0153] Hardware or software to use

[0154] Voice recording device: Terminal with microphone

[0155] Data transmission: HTTPS protocol

[0156] Speech recognition software: Google Cloud Speech-to-Text

[0157] Response generation model: OpenAI's GPT-3

[0158] Speech synthesis engine: Amazon Polly

[0159] Message sending API: Twilio

[0160] Data storage: MySQL database

[0161] The above describes the specific embodiments of the present invention. This system solves the lack of communication and support faced by elderly people living alone and supports their daily lives.

[0162] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0163] Basic Information Registration

[0164] Step 1:

[0165] The user inputs general information (name, address, contact information, etc.) by voice into the device. The device records this voice data as input.

[0166] Step 2:

[0167] The device converts recorded audio data into text data. This process uses speech recognition software (e.g., Google Cloud Speech-to-Text). Audio data is taken as input, and text data is obtained as output.

[0168] Step 3:

[0169] The terminal sends the converted text data to the server using the HTTPS protocol. Text data is used as input and sent to the server as output.

[0170] Step 4:

[0171] The server saves the received text data to a database. A storage device (e.g., a MySQL database) is used to permanently store the input text data. The output is a status indicating that saving is complete.

[0172] Step 5:

[0173] The server sends a notification to the device when saving is complete. The save completion status is used as input, and the notification is sent to the device as output.

[0174] Step 6:

[0175] The device converts the save completion notification into speech using a speech synthesis engine (e.g., Amazon Polly) and notifies the user that "Registration complete." The notification status is used as input, and the voice message is obtained as output.

[0176] Processing everyday conversations

[0177] Step 1:

[0178] The user speaks into the device saying, "The weather's nice today." The device receives the user's voice data as input. The device then records this voice data.

[0179] Step 2:

[0180] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0181] Step 3:

[0182] The server uses speech recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. Audio data is used as input, and text data is obtained as output.

[0183] Step 4:

[0184] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response based on the text. Text data is used as input, and the response text is obtained as output.

[0185] Step 5:

[0186] The server sends the generated response text to the terminal. The response text is used as input and sent to the terminal as output.

[0187] Step 6:

[0188] The device converts the response text into speech using a text-to-speech engine (e.g., Amazon Polly) and plays it back to the user. The response text is used as input, and the output is a voiced message.

[0189] Send message

[0190] Step 1:

[0191] The user says, "Send a message to my son saying, 'I'm doing well.'" The user's voice data is obtained as input. The device records this voice data.

[0192] Step 2:

[0193] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0194] Step 3:

[0195] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[0196] Step 4:

[0197] The server retrieves the recipient's contact information from the database. Text data is used as input, and the recipient's contact information is obtained as output.

[0198] Step 5:

[0199] The server uses Twilio's SMS API to send a message to the recipient. The recipient's contact information and message content are used as input, and the message sending status is obtained as output.

[0200] Step 6:

[0201] The server sends a status message to the terminal indicating that the transmission is complete. The message transmission status is used as input and sent to the terminal as output.

[0202] Step 7:

[0203] The device converts the message transmission completion notification into speech using a speech synthesis engine and notifies the user that "the message has been sent." The notification status is used as input, and the voice message is obtained as output.

[0204] Appointment

[0205] Step 1:

[0206] The user says, "Please schedule an appointment for me this Friday at 3 PM." The user's voice data is obtained as input. The device records this voice data.

[0207] Step 2:

[0208] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0209] Step 3:

[0210] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[0211] Step 4:

[0212] The server analyzes the text data and extracts information for the desired date and time. Text data is used as input, and the output is information for the desired date and time.

[0213] Step 5:

[0214] The server uses the online medical consultation system's API to make reservations. The desired date and time are used as input, and the output is a reservation completion status.

[0215] Step 6:

[0216] The server sends a reservation completion notification to the device. The reservation completion status is used as input and sent to the device as output.

[0217] Step 7:

[0218] The device converts the reservation completion notification into speech using a speech synthesis engine and notifies the user that "Your reservation is complete." The notification status is used as input, and the voice message is obtained as output.

[0219] Responding to requests

[0220] Step 1:

[0221] The user says, "Add milk to my shopping list." The device receives the user's voice data as input. The device records this voice data.

[0222] Step 2:

[0223] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0224] Step 3:

[0225] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[0226] Step 4:

[0227] The server analyzes the text data and identifies the relevant requests. Text data is used as input, and the requests are obtained as output.

[0228] Step 5:

[0229] The server performs the appropriate operation based on the requested information. The requested information is used as input, and the operation execution status is obtained as output.

[0230] Step 6:

[0231] The server sends a notification to the terminal indicating that the operation is complete. The operation execution status is used as input and sent to the terminal as output.

[0232] Step 7:

[0233] The device converts the operation completion notification into speech using a speech synthesis engine and notifies the user, "Milk has been added to your shopping list." The notification status is used as input, and the voice message is obtained as output.

[0234] The above is the specific processing flow of the system.

[0235] (Application Example 1)

[0236] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0237] Elderly people living alone face the problem of being unable to rely on others for daily life support, such as ordering food delivery, because traditional methods are cumbersome to use. Furthermore, users unfamiliar with digital technology require natural communication using voice recognition. Therefore, there is a need for a system that allows elderly people living alone to easily and naturally order food delivery and access other life support.

[0238] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0239] In this invention, the server is

[0240] A means for users to input basic information,

[0241] Means for transmitting the aforementioned basic information to a server,

[0242] A server that stores the aforementioned basic information in a database,

[0243] A means for notifying that the aforementioned storage has been completed,

[0244] A means of inputting everyday conversations with the user via voice,

[0245] Means for transmitting the aforementioned audio data to a server,

[0246] A server that converts the aforementioned audio data into text,

[0247] A server that generates a response based on the aforementioned text,

[0248] A means for outputting the aforementioned response to the user in voice,

[0249] A means of inputting the recipient and message content by the user,

[0250] A server that sends the aforementioned audio data to a server and sends a message to a recipient,

[0251] A method for users to input their preferred date and time for medical appointments via voice input,

[0252] A means for sending the aforementioned desired date and time to the server,

[0253] A server that makes medical appointments based on the aforementioned preferred date and time,

[0254] A means of inputting user requests by voice,

[0255] A server that transmits the aforementioned audio data to a server and performs the corresponding operation,

[0256] A means for users to input food delivery orders by voice,

[0257] Means for sending the food delivery order to the server,

[0258] A server that processes orders based on the aforementioned food delivery orders,

[0259] A means for notifying the completion of the aforementioned order processing,

[0260] Includes.

[0261] This will make it easy and natural for elderly people living alone to order food delivery and other forms of assistance for daily life.

[0262] "User" refers to any person who uses this system.

[0263] "Basic information" refers to necessary information about the user, such as their name, address, and contact information.

[0264] A "server" refers to a computer system that processes and stores data, and notifies users.

[0265] A "database" refers to an information system used to systematically store basic information and other data.

[0266] "Means of notification" refers to methods of informing users about the completion of information storage or processing.

[0267] "Voice input" refers to a method of inputting information by voice through a microphone or similar device.

[0268] "Audio data" refers to digital audio information obtained from voice input.

[0269] "Converting to text" refers to the process of converting audio data into written text.

[0270] "Response" refers to the reply or instruction generated by the server based on the user's voice input.

[0271] A "message" refers to information with specified content that a user sends to other recipients.

[0272] "Medical appointment booking" refers to making a reservation for a specific date and time to receive medical treatment at a medical institution or similar facility.

[0273] "Requests" refer to the operations or requests that users want to perform on the system.

[0274] "Food delivery" refers to the process of a user ordering food to be delivered to a specified location.

[0275] "Order processing" refers to the procedure for confirming the details of a food delivery order and handling it appropriately.

[0276] "Means of notifying completion" refers to methods of informing users that an order or operation has been completed.

[0277] This invention is a system that provides the communication and support that elderly people living alone need in their daily lives. This system includes a terminal that is easy for the user to operate and a server that processes data and provides the necessary services.

[0278] The system is configured as follows:

[0279] The user first enters basic information (name, address, contact information, etc.) into the terminal. This basic information is sent from the terminal to the server, which stores it in a database. Once the saving is complete, the server notifies the terminal, and the terminal informs the user that the saving is complete. In this way, the system can manage the user's basic information.

[0280] Regarding support for everyday conversation, when a user says, "The weather's nice today," the device records the audio and sends it to the server. The server converts the audio data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[0281] It is also possible to send and receive messages. For example, if a user says, "Send a message to my son saying, 'I'm doing well'," the device will record the voice and send it to the server. The server will retrieve the recipient's contact information from its database and send the specified message. Once the server has finished processing, it will notify the device, and the device will inform the user by voice, "The message has been sent."

[0282] Similarly, for medical appointments, when a user says "Reserve a medical appointment for this Friday at 3 PM", the terminal sends the voice to the server. The server analyzes the voice data and makes a medical appointment for the specified date and time. When the reservation is completed, the server notifies the terminal, and the terminal informs the user that "The reservation has been completed".

[0283] Also, this system supports food delivery orders. When a user says "Order curry for lunch today", the terminal records the voice and sends it to the server. The server converts the voice data into text, selects an appropriate food delivery service, and places the order. When the order is completed, the server notifies the terminal, and the terminal informs the user in voice that "The order has been completed".

[0284] As specific hardware for implementing this invention, smartphones and tablets can be used. For software, the "speech_recognition" library is used for speech recognition, the "pyttsx3" library is used for speech synthesis, and the "requests" library is used to process API requests. Also, a generative AI model is utilized for natural language processing (NLP).

[0285] Examples of specific prompt sentences are as follows.

[0286] "Order ~ for lunch today"

[0287] "Reserve a medical appointment for this Friday at 3 PM"

[0288] This system provides great convenience by enabling the elderly living alone to easily perform various operations necessary for daily life by voice. This reduces the barriers to digital technology and makes it possible to improve the quality of daily life.

[0289] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0290] Step 1:

[0291] The user enters basic information. The user enters their name, address, contact information, etc., on the screen of their smartphone or tablet. This constitutes input.

[0292] Step 2:

[0293] The terminal sends the entered basic information to the server. The basic information is converted into a digital format and sent as a request to the server. The server then receives the data.

[0294] Step 3:

[0295] The server saves the received basic information to the database. Data processing includes correctly formatting the input information and writing it to the database. The server then verifies that saving is complete.

[0296] Step 4:

[0297] The server notifies the user when the basic information has been saved. It sends the status of save completion to the terminal. The terminal informs the user of this information via voice or screen display.

[0298] Step 5:

[0299] The user inputs everyday conversations via voice. For example, they might say, "The weather's nice today." This becomes the input.

[0300] Step 6:

[0301] The terminal records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data. A process of converting it to a data format takes place here.

[0302] Step 7:

[0303] The server converts the received voice data into text. For the conversion, a generative AI model is used to perform the conversion from voice to text. The converted text is obtained as the output.

[0304] Step 8:

[0305] The server generates a response based on the converted text. It analyzes the text using natural language processing (NLP) to create appropriate response text, which is output as the response.

[0306] Step 9:

[0307] The server sends the generated response to the terminal. The text data of the response is sent to the terminal, and preparations for speech synthesis are made.

[0308] Step 10:

[0309] The terminal plays the response received from the server as voice. It uses a speech synthesis tool (pyttsx3) to convert the text data into voice and plays it for the user. This becomes the output to the user.

[0310] Step 11:

[0311] The user inputs the recipient and the message content by voice. For example, the user says example, the user says "Send a message 'I'm fine' to my son", which becomes the input.

[0312] Step 12:

[0313] The terminal records the voice data and sends it to the server. The recorded data is digitized and sent to the server as voice data.

[0314] Step 13:

[0315] The server sends a message based on the recipient and the message content. It retrieves the recipient's contact information from the database and sends the specified message.

[0316] Step 14:

[0317] The server notifies the terminal that the message has been sent. It sends a status message to the terminal indicating that the message has been sent, and the terminal then verbally informs the user that "the message has been sent."

[0318] Step 15:

[0319] Users can place food delivery orders by voice. For example, they might say, "Order curry for lunch today." This constitutes the input.

[0320] Step 16:

[0321] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[0322] Step 17:

[0323] The server converts voice data into text and processes the necessary food delivery orders. It uses a generative AI model to convert voice to text and then places orders with food delivery services based on that information.

[0324] Step 18:

[0325] The server notifies the terminal that the order processing is complete. It confirms that the order is complete and sends the status to the terminal. The terminal then informs the user by voice, "Your order is complete."

[0326] Through these steps, the system will be able to naturally use voice commands to order food deliveries and provide daily support for elderly people living alone.

[0327] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0328] The communication tool for elderly people living alone according to the present invention is designed to support the user's daily life and improve the convenience of communication. This system includes means for inputting and saving the user's basic information, support for daily conversation, sending and receiving messages, scheduling online medical appointments, and fulfilling requests, as well as a function that uses an emotion engine to recognize the user's emotions and generate responses.

[0329] First, the user enters their basic information into the device and sends that information to the server. The server saves the received information to a database and notifies the device when saving is complete. The device displays "Registration Complete," and the management of the basic information is finished.

[0330] Next, the user speaks to the device to enjoy a casual conversation. For example, if the user says, "The weather's nice today," the device records the voice and sends the voice data to the server. The server converts the voice data into text and recognizes the emotion using an emotion engine. Based on the recognized emotion, it generates the most appropriate response and sends the response text to the device. The device then converts the response back into voice and plays it back to the user, saying, "The weather's really nice today." In this way, the user can receive emotional support through a natural conversation.

[0331] Furthermore, sending and receiving messages is easy. When a user speaks to the device, saying, "Send a message to my son saying, 'I'm fine,'" the device records the voice and sends it to the server. The server converts the recipient and message content into text, recognizes the user's emotions using an emotion engine, and generates a message in an appropriate tone. For example, if the user sends a message in a sad tone, the emotion engine recognizes that emotion and generates a message in a tone such as, "I'm fine. I'm a little lonely, though." The server sends the message to the recipient and notifies the device when it is complete. The device then informs the user by voice, "Message sent."

[0332] Furthermore, booking an online consultation is also easy. When a user says, "Please book an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data, transcribes the desired date and time into text, and attempts to book the appointment at that time by accessing the online consultation booking system. Once the booking is complete, the server sends that information to the device, which then notifies the user that "the booking is complete."

[0333] Regarding requests, when a user says, "Add milk to my shopping list," the device sends voice data to the server. The server analyzes the request, recognizes the user's emotions using an emotion engine, and performs the appropriate action. If the user is stressed, the system responds in a tone such as, "Milk has been added to the list. Is there anything else I can help with?" The server notifies the device that the action is complete, and the device then informs the user, "Milk has been added to your shopping list."

[0334] Through these processes, the system of the present invention recognizes the emotions of elderly people living alone and generates responses based on them, thereby providing more humane support. Since all operations are voice-based, it can be easily used even by users unfamiliar with digital technology. This system is expected to support the user's daily life and reduce feelings of loneliness.

[0335] The following describes the processing flow.

[0336] User Registration

[0337] Step 1:

[0338] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[0339] Step 2:

[0340] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[0341] Step 3:

[0342] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[0343] Step 4:

[0344] The server notifies the terminal that saving the user information is complete.

[0345] Step 5:

[0346] The device receives a notification and displays "Registration Complete".

[0347] Support for everyday conversation

[0348] Step 1:

[0349] The user speaks to the device. For example, they might say, "The weather's nice today."

[0350] Step 2:

[0351] The device records the user's voice and sends the audio data to the server.

[0352] Step 3:

[0353] The server uses a speech recognition engine to convert the audio data into text.

[0354] Step 4:

[0355] The server uses an emotion engine to recognize the user's emotions based on the converted text. For example, if the user's voice has a bright tone, it is recognized as "positive," and if it has a somber tone, it is recognized as "negative."

[0356] Step 5:

[0357] The server generates the most appropriate response based on the emotions it perceives. For example, if someone says "The weather's nice today," it will respond with "It really is wonderful weather."

[0358] Step 6:

[0359] The server sends a response text to the terminal.

[0360] Step 7:

[0361] The device converts the received response into speech and plays it back to the user. For example, it might respond with, "What beautiful weather we have today."

[0362] Send message

[0363] Step 1:

[0364] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[0365] Step 2:

[0366] The device sends audio data to the server. The data is encoded in the appropriate format.

[0367] Step 3:

[0368] The server converts the audio data into text and analyzes the message content for the recipient.

[0369] Step 4:

[0370] The server uses an emotion engine to recognize the user's emotions. For example, if a message is sent in a sad tone, the server will recognize that emotion.

[0371] Step 5:

[0372] The server generates messages based on the emotions it perceives. For example, if a message saying "I'm fine" is sent in a sad tone, it will generate a message saying "I'm fine, but I'm a little lonely."

[0373] Step 6:

[0374] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[0375] Step 7:

[0376] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[0377] Step 8:

[0378] The device notifies the user, for example, saying, "A message has been sent."

[0379] Online medical consultation booking

[0380] Step 1:

[0381] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[0382] Step 2:

[0383] The device sends voice data to the server.

[0384] Step 3:

[0385] The server converts the voice data into text and analyzes the desired reservation date and time.

[0386] Step 4:

[0387] Based on these analysis results, the server accesses the online medical appointment system and makes an appointment for the desired date and time.

[0388] Step 5:

[0389] The server receives the reservation confirmation information and sends it to the terminal.

[0390] Step 6:

[0391] The device receives a notification that the reservation is complete and informs the user via voice, "Your reservation is complete."

[0392] Implementing requested actions (such as updating the shopping list)

[0393] Step 1:

[0394] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[0395] Step 2:

[0396] The device sends voice data to the server.

[0397] Step 3:

[0398] The server analyzes the audio data and converts the request into text.

[0399] Step 4:

[0400] The server uses an emotion engine to recognize the user's emotions. If the user is feeling stressed, it will recognize that.

[0401] Step 5:

[0402] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list.

[0403] Step 6:

[0404] The server notifies the terminal that the request has been completed.

[0405] Step 7:

[0406] The device receives a notification and informs the user by voice, "Milk has been added to your shopping list." It can also provide additional, emotionally-based responses such as, "Is there anything else I can help you with?"

[0407] This allows users to recognize their emotions and receive appropriate responses, enabling them to receive both life support and emotional care simultaneously.

[0408] (Example 2)

[0409] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0410] In modern society, elderly people living alone suffer from loneliness and a lack of support in their daily lives. This problem can have a significant impact on their mental health and quality of life. Furthermore, for elderly people who are unfamiliar with digital technology, the means to receive appropriate support are limited. Against this backdrop, there is a need for communication tools that are easy for elderly people living alone to use, understand their emotions, and respond appropriately, but existing solutions do not have sufficient accuracy in emotion recognition or natural dialogue.

[0411] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0412] In this invention, the server includes means for the user to input basic information, means for transmitting the basic information to the server, and means for storing the basic information in a database. This makes it possible to register the user's basic information smoothly.

[0413] Furthermore, the server includes means for voice input of everyday conversations with the user, means for transmitting the voice data to the server, means for converting the voice data into text, means for recognizing emotions based on the text, means for generating a response based on the recognized emotions, and means for outputting the response to the user in voice. This enables natural everyday conversations and emotional support with the user.

[0414] Furthermore, the server includes means for inputting voice from the user, including the recipient and message content; means for transmitting the voice data to the server and sending the message to the recipient; means for recognizing the emotion of the message and generating a response in an appropriate tone; and means for notifying that the message transmission is complete. This allows users to send emotionally charged messages to others, and the process is seamless.

[0415] The server also includes means for voice input of the user's preferred date and time for a medical appointment, means for transmitting the preferred date and time to the server, and means for making a medical appointment based on the preferred date and time. This makes online medical appointments intuitive and improves user convenience.

[0416] Finally, the server includes means for voice input of user requests, means for transmitting the voice data to the server and executing the corresponding operation, and means for generating a response that takes into account the emotions recognized based on the user's voice input. This allows the server to understand the user's emotions and then perform the request and notify the user of the result.

[0417] These methods provide a system that allows users, even those unfamiliar with digital technology, to receive emotional support through natural conversation and perform necessary daily tasks without stress.

[0418] "Basic information" refers to detailed information about the user, such as name, age, address, and health information, which is necessary to identify the individual and improve the quality of support.

[0419] A "server" is a computer system that stores basic information, manages voice data from users, and generates responses.

[0420] "Voice input means" refers to a device that includes a microphone and its control software for capturing user speech as digital voice data.

[0421] "Audio data" refers to data obtained by converting a user's speech into a digital format.

[0422] A "database" is a management system used by a server to store basic information and other user data.

[0423] "Means of converting to text" refers to software or algorithms that analyze audio data and convert it into text information.

[0424] "Means of recognizing emotions" refer to algorithms and software that analyze and determine emotions from the textual content of a user's speech.

[0425] A "means for generating responses" refers to a system that uses natural language processing techniques to produce appropriate responses based on recognized emotions and user input.

[0426] "Means of sending messages" refers to communication technology for sending messages generated from voice data to recipients specified by the user.

[0427] "Method for making medical appointments" refers to a system that allows users to confirm online appointments with medical institutions based on their preferred date and time.

[0428] "Means for executing operations" refers to a control system that receives requests from users and executes corresponding processing on the server.

[0429] "Notification methods" refer to functions that inform users of important information, such as the completion of a process or the sending of a message, via voice or text.

[0430] This invention is a communication tool for elderly people living alone, aiming to support users' daily lives and improve the convenience of communication. This system provides the following functions:

[0431] Basic Information Registration

[0432] The user enters their basic information, sends it to the server, and the server stores it in a database. This process allows the system to understand the user's individual information and use it to improve future service provision. For example, a user enters their name, age, address, and health information into their device and sends it to the server. The server receives it, saves it in the database, and then sends a notification to the device that the data has been saved.

[0433] Daily conversation support

[0434] When a user speaks into the device to enjoy a casual conversation, the device records the audio and sends it to a server. The server converts the audio to text, recognizes the emotion, and generates the most appropriate response. The device then converts that response back into audio and plays it back to the user. For example, if the user says, "The weather's nice today," the device might respond, "It really is wonderful weather."

[0435] Sending and receiving messages

[0436] When a user wants to send a message, they input their voice into the device. The device sends the voice data to a server, which converts the voice into text, recognizes the emotion, and generates an appropriate message. The generated message is sent to the recipient, and the device is notified when the message has been sent. For example, if a user says, "Send my son a message saying 'I'm fine'," the server recognizes that emotion and generates and sends the message, "I'm fine. I'm a little lonely, though."

[0437] Online medical consultation booking

[0438] When a user voice-inputs their desired date and time for a medical appointment, the device sends it to a server. The server analyzes the requested date and time, accesses the online medical appointment system, and makes the reservation. Once the reservation is complete, a notification is sent to the device. For example, if a user says, "Please schedule an appointment for me this Friday at 3pm," the server analyzes the information and confirms the online medical appointment.

[0439] Implementation of requested items

[0440] When a user speaks a request to the device, the device records the audio and sends it to the server. The server analyzes the request, recognizes the emotion, and performs the appropriate action. Once the action is complete, a notification is sent to the device. For example, if a user says, "Add milk to my shopping list," the server recognizes the request and adds the milk to the shopping list.

[0441] The following hardware and software are used to implement each function.

[0442] Speech recognition software: Google Speech-to-Text API

[0443] Text conversion software: IBM Watson® Natural Language Understanding

[0444] Natural Language Processing Technology: OpenAI's GPT-3

[0445] Speech synthesis software: Amazon Polly

[0446] Example of a prompt

[0447] "Based on what a user says to the system designed for elderly people living alone, please use an emotion engine to generate the most appropriate response. For example, please tell us what kind of response would be appropriate to a conversation like, 'The weather is nice today.'"

[0448] As described above, users can perform various operations using voice commands and receive emotional support. This system is expected to improve the quality of life for elderly people living alone and reduce feelings of loneliness.

[0449] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0450] Basic Information Registration

[0451] Step 1:

[0452] The user enters their basic information (name, age, address, health information, etc.) into the device. Specifically, the user enters the necessary information into the input form using the device's keyboard or touchscreen.

[0453] Input: User's basic information (name, age, address, health information)

[0454] Output: Input basic information data

[0455] Step 2:

[0456] The terminal sends the entered basic information to the server. Specifically, the terminal's application uses an HTTP POST request to send data to the server's information registration API.

[0457] Input: Entered basic information data

[0458] Output: Status of successful transmission to server

[0459] Step 3:

[0460] The server saves the basic information it receives to the database. Specifically, the server opens a database connection and saves the information using an SQL INSERT statement.

[0461] Input: Basic information data sent to the server

[0462] Output: Status of saving to database

[0463] Step 4:

[0464] The server sends a notification to the device when saving is complete. Specifically, the server returns a save completion message in an HTTP response.

[0465] Input: Database save complete status

[0466] Output: Notification that saving to the device is complete.

[0467] Step 5:

[0468] The device receives a notification that the save is complete and displays "Registration complete." Specifically, it updates the GUI and displays a pop-up message.

[0469] Input: Server save completion notification

[0470] Output: Display of the message "Registration complete".

[0471] Daily conversation support

[0472] Step 1:

[0473] The user speaks to the device, saying, "The weather's nice today." Specifically, this is done using the device's microphone function for voice input.

[0474] Input: User voice input ("The weather is nice today.")

[0475] Output: Audio data

[0476] Step 2:

[0477] The device records audio data. Specifically, it captures audio input using the microphone as digital data.

[0478] Input: User's voice

[0479] Output: Recorded audio data

[0480] Step 3:

[0481] The device sends audio data to the server. Specifically, it sends the audio data, which has been Base64 encoded, via an HTTP POST request.

[0482] Input: Recorded audio data

[0483] Output: Status of successful transmission to server

[0484] Step 4:

[0485] The server converts the audio data to text. Specifically, it uses the Google Speech-to-Text API to convert the audio data to text.

[0486] Input: Audio data sent to the server

[0487] Output: Text data ("The weather is nice today.")

[0488] Step 5:

[0489] The server recognizes emotions from text data. Specifically, it uses IBM Watson Natural Language Understanding to extract emotions from the text.

[0490] Input: Text data

[0491] Output: Recognized emotion (positive emotion)

[0492] Step 6:

[0493] The server generates responses based on emotions. Specifically, it uses OpenAI's GPT-3 model to generate appropriate reply text.

[0494] Input: Recognized emotions and text data

[0495] Output: Response text ("What lovely weather we have today!")

[0496] Step 7:

[0497] The server sends the response text to the terminal. Specifically, it returns the response text data as an HTTP response.

[0498] Input: Response text

[0499] Output: Sending status to terminal and response text

[0500] Step 8:

[0501] The device converts the response text into speech and plays it back to the user. Specifically, it uses Amazon Polly Text-to-Speech software to generate the speech and plays it through the device's speaker.

[0502] Input: Response text

[0503] Output: Voice response ("What beautiful weather we have today!")

[0504] Sending and receiving messages

[0505] Step 1:

[0506] The user speaks to the device, saying, "Send a message to my son saying, 'I'm doing well.'" Specifically, they use the voice input function to instruct the device to send the message.

[0507] Input: User's voice command to send a message

[0508] Output: Audio data

[0509] Step 2:

[0510] The device records audio. Specifically, it uses the microphone to capture audio as digital data.

[0511] Input: User's voice command to send a message

[0512] Output: Recording data

[0513] Step 3:

[0514] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[0515] Input: Recording data

[0516] Output: Status of successful transmission to server

[0517] Step 4:

[0518] The server converts the speech to text. Specifically, it uses the Google Speech-to-Text API to convert the speech data to text.

[0519] Input: Audio data

[0520] Output: Text data ("Send a message to your son saying 'I'm doing well.'")

[0521] Step 5:

[0522] The server recognizes emotions and generates messages in an appropriate tone. Specifically, it uses IBM Watson Natural Language Understanding to recognize emotions and OpenAI GPT-3 to generate messages.

[0523] Input: Text data, user sentiment

[0524] Output: Response text ("I'm fine. I'm a little lonely though.")

[0525] Step 6:

[0526] The server sends the message to the recipient. Specifically, it uses the SMS API or email sending API to send the message.

[0527] Input: Response text

[0528] Output: Message sent successfully status

[0529] Step 7:

[0530] The server notifies the terminal that the transmission is complete. Specifically, it sends a transmission completion notification as an HTTP response.

[0531] Input: Message sent status

[0532] Output: Notification of successful transmission to the terminal

[0533] Step 8:

[0534] The terminal notifies the user that the message has been sent. Specifically, it updates the GUI and displays "Message sent."

[0535] Input: Server transmission completion notification

[0536] Output: Displays "Message sent".

[0537] Online medical consultation booking

[0538] Step 1:

[0539] The user speaks into the terminal, saying, "Please schedule an appointment for me this Friday at 3 PM." Specifically, they use the voice input function to enter their desired appointment date and time.

[0540] Input: User voice commands

[0541] Output: Audio data

[0542] Step 2:

[0543] The device records audio. Specifically, it uses the microphone to capture the user's voice as digital data.

[0544] Input: User voice commands

[0545] Output: Recording data

[0546] Step 3:

[0547] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[0548] Input: Recording data

[0549] Output: Status of successful transmission to server

[0550] Step 4:

[0551] The server analyzes the audio and converts the desired date and time into text. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text data and extract the date and time information.

[0552] Input: Audio data

[0553] Output: Text data of the desired date and time ("This Friday at 3 PM")

[0554] Step 5:

[0555] The server accesses the online medical appointment system and attempts to book an appointment for the desired date and time. Specifically, it sends an appointment request using the medical appointment API.

[0556] Input: Text data of desired date and time

[0557] Output: Booking complete status

[0558] Step 6:

[0559] The server notifies the terminal that the reservation is complete. Specifically, it sends a reservation completion notification as an HTTP response.

[0560] Input: Booking complete status

[0561] Output: Reservation completion notification to the device

[0562] Step 7:

[0563] The device notifies the user that the reservation is complete. Specifically, it updates the GUI and displays "Reservation complete."

[0564] Input: Reservation completion notification from the server

[0565] Output: "Reservation complete" is displayed.

[0566] Implementation of requested items

[0567] Step 1:

[0568] The user speaks to the device, saying, "Add milk to my shopping list." Specifically, they use the voice input function to enter their request.

[0569] Input: User voice commands

[0570] Output: Audio data

[0571] Step 2:

[0572] The device records audio. Specifically, it uses the microphone to capture the user's voice as digital data.

[0573] Input: User voice commands

[0574] Output: Recording data

[0575] Step 3:

[0576] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[0577] Input: Recording data

[0578] Output: Status of successful transmission to server

[0579] Step 4:

[0580] The server analyzes the audio data and converts the request into text. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text data and extract the request.

[0581] Input: Audio data

[0582] Output: Text data of the request ("Add milk to the shopping list")

[0583] Step 5:

[0584] The server analyzes the request, recognizes the emotion using an emotion engine, and performs the appropriate action. Specifically, it uses a generative AI model (GPT-3) to generate an emotion-based response and perform the corresponding action. For example, adding an item to a shopping list.

[0585] Input: Text data of the request and the recognized emotion.

[0586] Output: Operation complete status ("Milk has been added to shopping list")

[0587] Step 6:

[0588] The server notifies the terminal that the operation is complete. Specifically, it sends an operation completion notification as an HTTP response.

[0589] Input: Operation completion status

[0590] Output: Operation completion notification to the terminal

[0591] Step 7:

[0592] The device notifies the user that the operation is complete. Specifically, it updates the GUI and displays a message such as "Milk has been added to your shopping list." Alternatively, it provides an audio notification.

[0593] Input: Operation completion notification from the server

[0594] Output: A message or notification saying "Milk has been added to your shopping list."

[0595] The above outlines the specific processing steps of the system's program. In this way, even users unfamiliar with digital technology can perform various daily tasks while receiving emotional support through natural conversation.

[0596] (Application Example 2)

[0597] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0598] When elderly people living alone or older seniors look for products in physical stores or have questions, it can be time-consuming to ask store staff for assistance. Furthermore, appropriate responses tailored to their emotional state are required, but not all store staff necessarily possess the ability or time to do so. In this situation, there is a need to improve the convenience and sense of security for elderly people living alone or older seniors when they are doing their daily shopping.

[0599] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input basic information, means for transmitting the basic information to the server, and means for storing the basic information in a database. This ensures that the user's basic information is securely stored on the server, providing a foundation for daily support. The system also includes means for voice input of daily conversations with the user, means for transmitting the voice data to the server, and means for converting the voice data into text. This allows the system to recognize and transcribe questions even when the user asks them by voice.

[0600] Furthermore, the system includes means for recognizing the user's emotions using an emotion engine, means for generating a response based on the text and adjusting it based on the user's emotions, means for outputting the response to the user as audio, and means for providing the optimal response to the user's questions in a physical store. This makes it possible to generate a response that matches the user's emotional state, enabling a more human-like interaction.

[0601] As a result, elderly people living alone and older adults can shop with peace of mind, reduce stress in stores, and be provided with an environment where they can easily find products and ask questions.

[0602] "Means for inputting basic user information" refers to devices or interfaces that allow users to input information about themselves (such as name, age, and address) into a terminal.

[0603] "Means of sending basic information to a server" refers to functions or protocols that transfer basic information entered by a user to a server via the internet or a local network.

[0604] A "server that stores basic information in a database" is a server equipped with a digital storage system that efficiently and securely stores the basic information of users that has been transmitted.

[0605] "Means of notifying that saving is complete" refers to functions or mechanisms that inform the user that basic information has been successfully saved to the database.

[0606] "Means for inputting everyday conversations by voice" refers to microphones or voice recognition devices that allow users to input everyday conversations as voice into a system.

[0607] "Means for sending audio data to a server" refers to the functions and protocols for transferring acquired audio data to a server via the internet or a local network.

[0608] A "server that converts audio data to text" is a server equipped with software and hardware to receive audio data and convert it into text format.

[0609] A "response-generating server" is a server equipped with algorithms and software that generate responses to users based on text data.

[0610] "Means of outputting a response to the user in audio" refers to speakers or headphones that play back the generated response as audio to the user using speech synthesis technology.

[0611] "Means for inputting the recipient and message content as voice" refers to a microphone or voice recognition device used by a user to input the content of a message they intend to send to a specific recipient as voice.

[0612] A "server that sends audio data to a server and sends a message to a recipient" is a server equipped with the function of converting audio data into text and sending a message to a designated recipient.

[0613] "A means of inputting desired date and time for medical appointments by voice" refers to a microphone or voice recognition device that allows the user to input their desired date and time for medical appointments into the system by voice.

[0614] "Means for sending desired date and time to the server" refers to a function or protocol that transmits the user's voice-inputted desired date and time for a medical consultation to a server via the internet or a local network.

[0615] A "server that makes medical appointments based on requested dates and times" is a server equipped with the function to confirm appointments in conjunction with a medical appointment system based on the received requested dates and times.

[0616] An "emotion engine" is an algorithm or software that estimates emotions from a user's words and actions and adjusts its response accordingly.

[0617] "Means of adjusting responses based on user emotions" refers to a function that appropriately modifies the tone and content of generated responses based on the user's emotional state recognized by the emotion engine.

[0618] "Means of providing optimal responses to user questions within a physical store" refers to interfaces and systems that provide appropriate information and instructions via voice in response to user questions and requests within a physical store.

[0619] This invention's system is designed to make it easier and safer for elderly people living alone or those in general to navigate situations such as shopping at physical stores. The system begins with the input of the user's basic information and then supports daily conversations, sending and receiving messages, scheduling online medical appointments, and fulfilling requests. In addition, it features an emotion engine that recognizes the user's emotions and adjusts its responses accordingly.

[0620] 1. Management of user basic information

[0621] The user first enters their basic information into their smartphone or smart glasses, and this information is sent to the server. The server stores the information in a database and sends a notification to the device when the saving is complete.

[0622] 2. Support for everyday conversation

[0623] Users can speak to a terminal in a physical store. For example, they might ask, "Where is this product?" The voice device records this, and the voice data is sent to a server. The server converts the voice data into text and uses an emotion engine to recognize emotions. Based on this information, the most appropriate response is generated and sent to the terminal. The terminal converts this response back into voice and provides the user with a response such as, "That product is in the second column from the right."

[0624] 3. Sending and receiving messages

[0625] For example, if a user says, "Send my son a message saying 'I'm doing well'," the voice is recorded and sent to the server. The server converts the voice data into text and uses an emotion engine to recognize the emotion. It then generates a message in an appropriate tone and sends it to the recipient. For example, it might provide a tailored message like, "I'm doing well. I'm a little lonely, though."

[0626] 4. Booking an online medical consultation

[0627] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the voice data is sent to the server. The server analyzes this data and enters the desired date and time into the online booking system. Once the booking is complete, the information is sent to the terminal, and the terminal notifies the user that "Your booking is complete."

[0628] 5. Implementation of requested items

[0629] User requests are entered via voice input and sent to the server. For example, if a user says, "Add milk to my shopping list," that voice data is sent to the server and the corresponding action is performed. An emotion engine is also included, and if the user is feeling stressed, a response in a more appropriate tone will be provided, such as, "Milk has been added to the list. Is there anything else I can help you with?"

[0630] Hardware / software to use

[0631] The primary hardware used will be smartphones and smart glasses, utilizing their built-in microphones and speakers. The following technologies and tools will be used in the software:

[0632] Speech recognition: Use the speech_recognition library to convert speech data into text.

[0633] Emotion Recognition: Implement an emotion recognition model using the transformers library.

[0634] Speech synthesis: Generate speech using gTTS or other speech synthesis libraries, and play it back using the playsound library.

[0635] Examples of specific cases and prompt statements

[0636] Specific example

[0637] Assume the user will be using smart glasses in the store:

[0638] User: "Where can I find this product?"

[0639] System: "That item is in the second row from the right. Please let us know if you need further information."

[0640] Example of a prompt

[0641] If the user wants to have a casual conversation:

[0642] "It's a beautiful day today."

[0643] System response: "It truly is a beautiful day!"

[0644] This system will improve the in-store shopping experience and allow elderly people living alone and seniors to shop with peace of mind.

[0645] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0646] Step 1:

[0647] The user enters basic information. The user uses the input interface of their smartphone or smart glasses to enter basic information such as their name, age, and address. This entered information is stored in a database, creating a baseline of user information.

[0648] Step 2:

[0649] The terminal sends the entered basic information to the server. The server receives this information and processes it for storage in the database. As output, a save completion notification is generated and sent to the terminal.

[0650] Step 3:

[0651] The system uses voice input for everyday conversations. When a user speaks into the device, the voice device (microphone) records it. The input voice data is sent to a server, which stores it in a database and processes it for impact analysis and recognition.

[0652] Step 4:

[0653] The device sends the recorded audio data to the server. This data is stored on the server and used in the next step.

[0654] Step 5:

[0655] The process converts audio data to text. The server uses speech recognition software (e.g., the speech_recognition library) to convert the transmitted audio data into text. The input is audio data, and the output is text data. During this conversion process, data calculations are performed to analyze the characteristics of the audio data and convert it into text format.

[0656] Step 6:

[0657] The system recognizes emotions and generates responses. The server uses a generative AI model to recognize user emotions from text data (e.g., the emotion recognition model in the transformers library). Input is text data, and output is response text with emotion labels. Based on this, the system generates responses in an appropriate emotional tone.

[0658] Step 7:

[0659] The response is output as audio. After the server recognizes the emotion, it uses speech synthesis software (e.g., gTTS library) to convert the response text into audio data. The terminal receives this audio data and plays it back to the user through speakers or headphones. The input is the response text, and the output is the audio data.

[0660] Step 8:

[0661] Sending a message. When the user issues a message sending command, the device sends it as audio data to the server. The server converts the audio data to text, performs emotion recognition, generates a message text in an appropriate tone, and sends it to the designated recipient. The input is the audio data, and the output is the message to be sent.

[0662] Step 9:

[0663] The process involves making a medical appointment. The user inputs their desired appointment date and time via voice, and the terminal sends this to the server. The server converts the requested date and time into text, accesses the online booking system, and confirms the appointment. The input is the voice data of the desired date and time, and the output is a booking confirmation notice.

[0664] Step 10:

[0665] The system executes the requested actions. The user inputs their specific request via voice, and the terminal sends this to the server. The server analyzes the request, recognizes the emotion, generates a response in an appropriate tone, and performs the corresponding action. The input is the voice data of the request, and the output is the appropriate response and the execution result.

[0666] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0667] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0668] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0669] [Second Embodiment]

[0670] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0671] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0672] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0673] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0674] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0675] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0676] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0677] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0678] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0679] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0680] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0681] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0682] The communication tool for elderly people living alone according to the present invention is designed to support daily life and improve the convenience of communication. This system includes a user-friendly terminal and a server that processes data and provides necessary services.

[0683] Specifically, the user enters basic information through their device and sends that information to the server. The server then stores the information in a database and notifies the device that registration is complete. Through this process, the system can manage the user's basic information.

[0684] Furthermore, users can enjoy everyday conversations by speaking directly to the device. For example, if a user says, "The weather's nice today," the device records the voice and sends it to the server. The server converts the voice data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[0685] Furthermore, sending and receiving messages is easy. When a user says to the device, "Send a message to my son saying 'I'm doing well'," the device records the voice and sends it to the server. The server retrieves the recipient's (in this case, the son's) contact information from its database and sends the specified message to that contact. Once the server has finished processing, it notifies the device, and the device then informs the user by voice, "The message has been sent."

[0686] Making appointments is also easy for users. When a user says, "Please schedule an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data and makes an online appointment for the specified date and time. Once the appointment is complete, the server sends the information back to the device, which then notifies the user that "the appointment is complete."

[0687] Furthermore, it can also respond to user requests. For example, if a user says, "Add milk to my shopping list," the device sends voice data to the server, and the server adds the appropriate item to the shopping list. Once the server has finished processing, it notifies the device of the result, and the device informs the user of the result.

[0688] Through these processes, the system of the present invention provides elderly people living alone with the communication and support they need in their daily lives, significantly improving their convenience. Furthermore, since all operations are voice-based, it is designed to be easy to use even for users unfamiliar with digital technology.

[0689] The following describes the processing flow.

[0690] User Registration

[0691] Step 1:

[0692] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[0693] Step 2:

[0694] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[0695] Step 3:

[0696] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[0697] Step 4:

[0698] The server notifies the terminal that saving the user information is complete.

[0699] Step 5:

[0700] The device receives a notification and displays "Registration Complete".

[0701] Support for everyday conversation

[0702] Step 1:

[0703] The user speaks to the device. For example, they might say, "The weather's nice today."

[0704] Step 2:

[0705] The device records the user's voice and sends the audio data to the server.

[0706] Step 3:

[0707] The server uses a speech recognition engine to convert the audio data into text.

[0708] Step 4:

[0709] The server generates an appropriate response based on the converted text. Natural language processing is used in this process.

[0710] Step 5:

[0711] The server sends a response text to the terminal.

[0712] Step 6:

[0713] The device converts the received response into speech and plays it back to the user. For example, it might say, "What beautiful weather we have today."

[0714] Send message

[0715] Step 1:

[0716] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[0717] Step 2:

[0718] The device sends audio data to the server. The data is encoded in the appropriate format.

[0719] Step 3:

[0720] The server converts the audio data into text and analyzes the message content for the recipient.

[0721] Step 4:

[0722] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[0723] Step 5:

[0724] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[0725] Step 6:

[0726] The device informs the user of notifications it has received. For example, it might say, "A message has been sent."

[0727] Online medical consultation booking

[0728] Step 1:

[0729] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[0730] Step 2:

[0731] The device sends voice data to the server.

[0732] Step 3:

[0733] The server analyzes the voice data and converts the desired reservation date and time into text.

[0734] Step 4:

[0735] The server accesses the online medical appointment system and attempts to make a reservation for that date and time.

[0736] Step 5:

[0737] The server receives the reservation confirmation information and sends the result to the terminal.

[0738] Step 6:

[0739] The device plays a voice message to the user confirming that the reservation has been received. For example, it might say, "Your reservation is complete."

[0740] Implementing requested actions (such as updating the shopping list)

[0741] Step 1:

[0742] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[0743] Step 2:

[0744] The device sends voice data to the server.

[0745] Step 3:

[0746] The server analyzes the audio data and converts the request into text.

[0747] Step 4:

[0748] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list in the database.

[0749] Step 5:

[0750] The server confirms that the request has been completed and notifies the terminal.

[0751] Step 6:

[0752] The device will verbally inform the user of notifications it has received. For example, it might say, "Milk has been added to your shopping list."

[0753] (Example 1)

[0754] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0755] The objective of this invention is to address the lack of communication and support that elderly people living alone face in their daily lives. In particular, it aims to provide a voice-based interface that can be easily used even by elderly people who are unfamiliar with digital technology, enabling them to enjoy everyday conversations, obtain necessary information, send and receive messages, and easily make medical appointments.

[0756] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0757] In this invention, the server includes means for the user to input general information, means for transmitting the general information to an information processing device, and an information processing device for storing the general information in a storage device. This makes it possible for the user to easily register basic information, enjoy everyday conversations, send messages, make medical appointments, and more, all through voice.

[0758] A "user" is an individual who utilizes the system of the present invention.

[0759] "General information" refers to basic information such as name, address, and contact information entered by the user.

[0760] An "information processing device" is a device that processes information transmitted by a user and provides necessary services.

[0761] A "storage device" is a device used by an information processing device to permanently store information.

[0762] "Voice input" is a method of inputting information or instructions using voice.

[0763] "Speech" refers to sound information generated by the user's utterances.

[0764] A "document" is information in the form obtained when an information processing device converts speech into text.

[0765] A "response" is the answer generated by an information processing device in response to a user's inquiry.

[0766] A "recipient" is an individual or organization that receives a message from a user.

[0767] "Message content" refers to the information or notification that the user wants to send.

[0768] "Appointment scheduling" is the process of making a medical appointment for a date and time specified by the user.

[0769] "Desired date and time" refers to the specific date and time the user wishes to schedule an appointment.

[0770] A "request" is a specific demand or instruction that a user wants to perform through the system.

[0771] "Natural language processing" is a technology that enables information processing devices to understand human language and generate appropriate responses.

[0772] "Methods of confirming online" refer to methods of confirming procedures such as reservations via the internet.

[0773] "The relevant operation" refers to the specific actions or procedures that the system performs in response to the user's request.

[0774] This invention relates to a communication tool for elderly people living alone, and is a system that uses a voice-based interface to allow users to easily input basic information, enjoy daily conversations, send messages, make medical appointments, and perform other requests.

[0775] How users enter information

[0776] The user inputs general information by voice through the microphone into the terminal. The terminal has a recording function, and the input voice is saved as digital audio data. The terminal sends this digital audio data to an information processing device (server). Specifically, the data is transmitted securely using the HTTPS protocol.

[0777] The server receives the transmitted information and saves it to a storage device (e.g., a MySQL database). It also notifies the terminal when saving is complete, and the terminal notifies the user that "registration is complete." A speech synthesis engine (e.g., Amazon Polly) is used for this notification.

[0778] Processing everyday conversations

[0779] When a user speaks to the device, for example, saying "The weather's nice today," the device records the audio and sends it to the server. The server uses Google Cloud Speech-to-Text to convert the audio data into text. Then, it uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response and sends it to the device. The device plays the response using a speech synthesis engine, allowing the user to enjoy a natural conversation.

[0780] Send message

[0781] When the user says, "Send my son a message saying 'I'm doing well'," the device records the voice and sends it to the server. The server converts the voice data to text and retrieves the recipient's (the son's) contact information from its database. The server uses Twilio's SMS API to send the message to the recipient. Once the server notifies the user that processing is complete, the device informs the user that "the message has been sent."

[0782] Appointment booking

[0783] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the device sends the voice data to the server. The server converts the voice data into text and extracts the specified date and time information. The server uses the online medical consultation system's API to schedule the appointment. Once the appointment is complete, the server notifies the device that "the appointment is complete."

[0784] Responding to requests

[0785] When a user says, "Add milk to my shopping list," the device sends the voice data to the server. The server converts the voice data to text and adds the milk to the shopping list. Once the process is complete, the server notifies the device that "Milk has been added to your shopping list."

[0786] Examples of specific cases and prompt statements

[0787] For example, if a user says, "The weather's nice today," the following specific actions will be taken:

[0788] Voice input: "The weather is nice today."

[0789] Server processing: Converts audio data to text and generates responses using a generative AI model.

[0790] Response: "It really is a beautiful day today."

[0791] Example of a prompt:

[0792] "Please tell me what kind of response you should generate if a user says, 'The weather is nice today.'"

[0793] Hardware or software to use

[0794] Voice recording device: Terminal with microphone

[0795] Data transmission: HTTPS protocol

[0796] Speech recognition software: Google Cloud Speech-to-Text

[0797] Response generation model: OpenAI's GPT-3

[0798] Speech synthesis engine: Amazon Polly

[0799] Message sending API: Twilio

[0800] Data storage: MySQL database

[0801] The above describes the specific embodiments of the present invention. This system solves the lack of communication and support faced by elderly people living alone and supports their daily lives.

[0802] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0803] Basic Information Registration

[0804] Step 1:

[0805] The user inputs general information (name, address, contact information, etc.) by voice into the device. The device records this voice data as input.

[0806] Step 2:

[0807] The device converts recorded audio data into text data. This process uses speech recognition software (e.g., Google Cloud Speech-to-Text). Audio data is taken as input, and text data is obtained as output.

[0808] Step 3:

[0809] The terminal sends the converted text data to the server using the HTTPS protocol. Text data is used as input and sent to the server as output.

[0810] Step 4:

[0811] The server saves the received text data to a database. A storage device (e.g., a MySQL database) is used to permanently store the input text data. The output is a status indicating that saving is complete.

[0812] Step 5:

[0813] The server sends a notification to the device when saving is complete. The save completion status is used as input, and the notification is sent to the device as output.

[0814] Step 6:

[0815] The device converts the save completion notification into speech using a speech synthesis engine (e.g., Amazon Polly) and notifies the user that "Registration complete." The notification status is used as input, and the output is a voice message.

[0816] Processing everyday conversations

[0817] Step 1:

[0818] The user speaks into the device saying, "The weather's nice today." The device receives the user's voice data as input. The device then records this voice data.

[0819] Step 2:

[0820] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0821] Step 3:

[0822] The server uses speech recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. Audio data is used as input, and text data is obtained as output.

[0823] Step 4:

[0824] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response based on the text. Text data is used as input, and the response text is obtained as output.

[0825] Step 5:

[0826] The server sends the generated response text to the terminal. The response text is used as input and sent to the terminal as output.

[0827] Step 6:

[0828] The device converts the response text into speech using a text-to-speech engine (e.g., Amazon Polly) and plays it back to the user. The response text is used as input, and the output is a voiced message.

[0829] Send message

[0830] Step 1:

[0831] The user says, "Send a message to my son saying, 'I'm doing well.'" The input is the user's voice data. The device records this voice data.

[0832] Step 2:

[0833] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0834] Step 3:

[0835] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[0836] Step 4:

[0837] The server retrieves the recipient's contact information from the database. Text data is used as input, and the recipient's contact information is obtained as output.

[0838] Step 5:

[0839] The server uses Twilio's SMS API to send a message to the recipient. The recipient's contact information and message content are used as input, and the message sending status is obtained as output.

[0840] Step 6:

[0841] The server sends a status message to the terminal indicating that the transmission is complete. The message transmission status is used as input and sent to the terminal as output.

[0842] Step 7:

[0843] The device converts the message transmission completion notification into speech using a speech synthesis engine and notifies the user that "the message has been sent." The notification status is used as input, and the voice message is obtained as output.

[0844] Appointment booking

[0845] Step 1:

[0846] The user says, "Please schedule an appointment for me this Friday at 3 PM." The user's voice data is obtained as input. The device records this voice data.

[0847] Step 2:

[0848] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0849] Step 3:

[0850] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[0851] Step 4:

[0852] The server analyzes the text data and extracts information about the desired date and time. Text data is used as input, and the output is information about the desired date and time.

[0853] Step 5:

[0854] The server uses the online medical consultation system's API to make reservations. The desired date and time are used as input, and the output is a reservation completion status.

[0855] Step 6:

[0856] The server sends a reservation completion notification to the device. The reservation completion status is used as input and sent to the device as output.

[0857] Step 7:

[0858] The device converts the reservation completion notification into speech using a speech synthesis engine and notifies the user that "Your reservation is complete." The notification status is used as input, and the voice message is obtained as output.

[0859] Responding to requests

[0860] Step 1:

[0861] The user says, "Add milk to my shopping list." The device receives the user's voice data as input. The device then records this voice data.

[0862] Step 2:

[0863] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[0864] Step 3:

[0865] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[0866] Step 4:

[0867] The server analyzes the text data and identifies the relevant requests. Text data is used as input, and the requests are obtained as output.

[0868] Step 5:

[0869] The server performs the appropriate operation based on the requested information. The requested information is used as input, and the operation execution status is obtained as output.

[0870] Step 6:

[0871] The server sends a notification to the terminal that the operation is complete. The operation execution status is used as input and sent to the terminal as output.

[0872] Step 7:

[0873] The device converts the completion notification into speech using a speech synthesis engine and notifies the user that "Milk has been added to your shopping list." The notification status is used as input, and the voice message is obtained as output.

[0874] The above is the specific processing flow of the system.

[0875] (Application Example 1)

[0876] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0877] Elderly people living alone face the problem of being unable to rely on others for daily life support, such as ordering food delivery or other assistance, because traditional methods are cumbersome to use. Furthermore, users unfamiliar with digital technology require natural communication using voice recognition. Therefore, there is a need for a system that allows elderly people living alone to easily and naturally order food delivery and other daily life support.

[0878] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0879] In this invention, the server is

[0880] A means for users to input basic information,

[0881] Means for transmitting the aforementioned basic information to a server,

[0882] A server that stores the aforementioned basic information in a database,

[0883] A means for notifying that the aforementioned storage has been completed,

[0884] A means of inputting everyday conversations with the user via voice,

[0885] Means for transmitting the aforementioned audio data to a server,

[0886] A server that converts the aforementioned audio data into text,

[0887] A server that generates a response based on the aforementioned text,

[0888] A means for outputting the aforementioned response to the user in voice,

[0889] A means of inputting the recipient and message content by the user,

[0890] A server that sends the aforementioned audio data to a server and sends a message to a recipient,

[0891] A method for users to input their preferred date and time for medical appointments via voice input,

[0892] A means for sending the aforementioned desired date and time to the server,

[0893] A server that makes medical appointments based on the aforementioned preferred date and time,

[0894] A means of inputting user requests by voice,

[0895] A server that transmits the aforementioned audio data to a server and performs the corresponding operation,

[0896] A means for users to input food delivery orders by voice,

[0897] Means for sending the food delivery order to the server,

[0898] A server that processes orders based on the aforementioned food delivery orders,

[0899] A means for notifying the completion of the aforementioned order processing,

[0900] Includes.

[0901] This will make it easy and natural for elderly people living alone to order food delivery and other forms of assistance for daily life.

[0902] "User" refers to any person who uses this system.

[0903] "Basic information" refers to necessary information about the user, such as their name, address, and contact information.

[0904] A "server" refers to a computer system that processes and stores data, and notifies users.

[0905] A "database" refers to an information system used to systematically store basic information and other data.

[0906] "Means of notification" refers to methods of informing users about the completion of information storage or processing.

[0907] "Voice input" refers to a method of inputting information by voice through a microphone or similar device.

[0908] "Audio data" refers to digital audio information obtained from voice input.

[0909] "Converting to text" refers to the process of converting audio data into written text.

[0910] "Response" refers to the reply or instruction generated by the server based on the user's voice input.

[0911] A "message" refers to information with specified content that a user sends to other recipients.

[0912] "Medical appointment booking" refers to making a reservation for a specific date and time to receive medical treatment at a medical institution or similar facility.

[0913] "Requests" refer to the operations or requests that users want to perform on the system.

[0914] "Food delivery" refers to the process of a user ordering food to be delivered to a specified location.

[0915] "Order processing" refers to the procedure for confirming the details of a food delivery order and handling it appropriately.

[0916] "Means of notifying completion" refers to methods of informing users that an order or operation has been completed.

[0917] This invention is a system that provides the communication and support that elderly people living alone need in their daily lives. This system includes a terminal that is easy for the user to operate and a server that processes data and provides necessary services.

[0918] The system is configured as follows:

[0919] The user first enters basic information (name, address, contact information, etc.) into the terminal. This basic information is sent from the terminal to the server, which stores it in a database. Once the saving is complete, the server notifies the terminal, and the terminal informs the user that the saving is complete. In this way, the system can manage the user's basic information.

[0920] Regarding support for everyday conversation, when a user says, "The weather's nice today," the device records the audio and sends it to the server. The server converts the audio data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[0921] It is also possible to send and receive messages. For example, if a user says, "Send a message to my son saying, 'I'm doing well'," the device will record the voice and send it to the server. The server will retrieve the recipient's contact information from its database and send the specified message. Once the server has finished processing, it will notify the device, and the device will inform the user by voice, "The message has been sent."

[0922] Similarly, when scheduling an appointment, if the user says, "Please schedule an appointment for me this Friday at 3pm," the device sends that voice message to the server. The server analyzes the voice data and schedules an appointment for the specified date and time. Once the appointment is complete, the server notifies the device, and the device informs the user that "the appointment is complete."

[0923] This system also supports ordering food delivery. When a user says, "Order curry for lunch today," the terminal records the voice and sends it to the server. The server converts the voice data into text, selects an appropriate food delivery service, and places the order. Once the order is complete, the server notifies the terminal of the result, and the terminal informs the user via voice, "Your order is complete."

[0924] The specific hardware used to realize this invention can be a smartphone or tablet. For the software, the "speech_recognition" library is used for speech recognition, the "pyttsx3" library for speech synthesis, and the "requests" library for processing API requests. Furthermore, a generative AI model is used for natural language processing (NLP).

[0925] Examples of specific prompt messages include the following:

[0926] "I'd like to order ~ for lunch today."

[0927] "Please schedule an appointment for this Friday at 3 PM."

[0928] This system offers significant convenience by allowing elderly people living alone to easily perform various daily tasks using voice commands. This reduces barriers to digital technology and improves the quality of daily life.

[0929] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0930] Step 1:

[0931] The user enters basic information. The user enters their name, address, contact information, etc., on the screen of their smartphone or tablet. This constitutes input.

[0932] Step 2:

[0933] The terminal sends the entered basic information to the server. The basic information is converted into a digital format and sent as a request to the server. The server then receives the data.

[0934] Step 3:

[0935] The server saves the received basic information to the database. Data processing includes correctly formatting the input information and writing it to the database. The server then verifies that saving is complete.

[0936] Step 4:

[0937] The server notifies the user when the basic information has been saved. It sends the status of save completion to the terminal. The terminal informs the user of this information via voice or screen display.

[0938] Step 5:

[0939] The user inputs everyday conversations via voice. For example, they might say, "The weather's nice today." This becomes the input.

[0940] Step 6:

[0941] The terminal records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data. A process of converting it to a data format takes place here.

[0942] Step 7:

[0943] The server converts the received audio data into text. A generative AI model is used for the conversion from speech to text. The converted text is obtained as output.

[0944] Step 8:

[0945] The server generates a response based on the converted text. It uses natural language processing (NLP) to analyze the text and create an appropriate response text, which is then output as the response.

[0946] Step 9:

[0947] The server sends the generated response to the terminal. The text data of the response is sent to the terminal, and preparations for speech synthesis are made.

[0948] Step 10:

[0949] The terminal plays back the response received from the server as audio. It uses a speech synthesis tool (pyttsx3) to convert text data into speech and plays it back to the user. This is the output to the user.

[0950] Step 11:

[0951] The user inputs the recipient and message content by voice. For example, they might say, "Send a message to my son saying, 'I'm doing well.'" This is the input.

[0952] Step 12:

[0953] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[0954] Step 13:

[0955] The server sends a message based on the recipient and message content. It retrieves the recipient's contact information from the database and sends the specified message.

[0956] Step 14:

[0957] The server notifies the terminal that the message has been sent. It sends a status message to the terminal indicating that the message has been sent, and the terminal then verbally informs the user that "the message has been sent."

[0958] Step 15:

[0959] Users can place food delivery orders by voice. For example, they might say, "Order curry for lunch today." This constitutes the input.

[0960] Step 16:

[0961] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[0962] Step 17:

[0963] The server converts voice data into text and processes the necessary food delivery orders. It uses a generative AI model to convert voice to text and then places orders with food delivery services based on that information.

[0964] Step 18:

[0965] The server notifies the terminal that the order processing is complete. It confirms that the order is complete and sends the status to the terminal. The terminal then informs the user by voice, "Your order is complete."

[0966] Through these steps, the system will be able to naturally use voice commands to order food deliveries and provide daily support for elderly people living alone.

[0967] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0968] The communication tool for elderly people living alone according to the present invention is designed to support the user's daily life and improve the convenience of communication. This system includes means for inputting and saving the user's basic information, support for daily conversation, sending and receiving messages, scheduling online medical appointments, and fulfilling requests, as well as a function that uses an emotion engine to recognize the user's emotions and generate responses.

[0969] First, the user enters their basic information into the device and sends that information to the server. The server saves the received information to a database and notifies the device when the saving is complete. The device displays "Registration Complete," and the management of the basic information is finished.

[0970] Next, the user speaks to the device to enjoy a casual conversation. For example, if the user says, "The weather's nice today," the device records the voice and sends the voice data to the server. The server converts the voice data into text and recognizes the emotion using an emotion engine. Based on the recognized emotion, it generates the most appropriate response and sends the response text to the device. The device then converts the response back into voice and plays it back to the user, saying, "The weather's really nice today." In this way, the user can receive emotional support through a natural conversation.

[0971] Furthermore, sending and receiving messages is easy. When a user speaks to the device, saying, "Send a message to my son saying, 'I'm fine,'" the device records the voice and sends it to the server. The server converts the recipient and message content into text, recognizes the user's emotions using an emotion engine, and generates a message in an appropriate tone. For example, if the user sends a message in a sad tone, the emotion engine recognizes that emotion and generates a message in a tone such as, "I'm fine. I'm a little lonely, though." The server sends the message to the recipient and notifies the device when it is complete. The device then informs the user by voice, "Message sent."

[0972] Furthermore, booking an online consultation is also easy. When a user says, "Please book an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data, transcribes the desired date and time into text, and attempts to book the appointment at that time by accessing the online consultation booking system. Once the booking is complete, the server sends that information to the device, which then notifies the user that "the booking is complete."

[0973] Regarding requests, when a user says, "Add milk to my shopping list," the device sends voice data to the server. The server analyzes the request, recognizes the user's emotions using an emotion engine, and performs the appropriate action. If the user is stressed, the system responds in a tone such as, "Milk has been added to the list. Is there anything else I can help with?" The server notifies the device that the action is complete, and the device then informs the user, "Milk has been added to your shopping list."

[0974] Through these processes, the system of the present invention recognizes the emotions of elderly people living alone and generates responses based on them, thereby providing more humane support. Since all operations are voice-based, it can be easily used even by users unfamiliar with digital technology. This system is expected to support the user's daily life and reduce feelings of loneliness.

[0975] The following describes the processing flow.

[0976] User Registration

[0977] Step 1:

[0978] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[0979] Step 2:

[0980] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[0981] Step 3:

[0982] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[0983] Step 4:

[0984] The server notifies the terminal that saving the user information is complete.

[0985] Step 5:

[0986] The device receives a notification and displays "Registration Complete".

[0987] Support for everyday conversation

[0988] Step 1:

[0989] The user speaks to the device. For example, they might say, "The weather's nice today."

[0990] Step 2:

[0991] The device records the user's voice and sends the audio data to the server.

[0992] Step 3:

[0993] The server uses a speech recognition engine to convert the audio data into text.

[0994] Step 4:

[0995] The server uses an emotion engine to recognize the user's emotions based on the converted text. For example, if the user's voice has a bright tone, it is recognized as "positive," and if it has a somber tone, it is recognized as "negative."

[0996] Step 5:

[0997] The server generates the most appropriate response based on the emotions it perceives. For example, if someone says "The weather's nice today," it will respond with "It really is wonderful weather."

[0998] Step 6:

[0999] The server sends a response text to the terminal.

[1000] Step 7:

[1001] The device converts the received response into speech and plays it back to the user. For example, it might respond with, "What beautiful weather we have today."

[1002] Send message

[1003] Step 1:

[1004] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[1005] Step 2:

[1006] The device sends audio data to the server. The data is encoded in the appropriate format.

[1007] Step 3:

[1008] The server converts the audio data into text and analyzes the message content for the recipient.

[1009] Step 4:

[1010] The server uses an emotion engine to recognize the user's emotions. For example, if a message is sent in a sad tone, the server will recognize that emotion.

[1011] Step 5:

[1012] The server generates messages based on the emotions it perceives. For example, if a message saying "I'm fine" is sent in a sad tone, it will generate a message saying "I'm fine, but I'm a little lonely."

[1013] Step 6:

[1014] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[1015] Step 7:

[1016] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[1017] Step 8:

[1018] The device notifies the user, for example, saying, "A message has been sent."

[1019] Online medical consultation booking

[1020] Step 1:

[1021] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[1022] Step 2:

[1023] The device sends voice data to the server.

[1024] Step 3:

[1025] The server converts the voice data into text and analyzes the desired reservation date and time.

[1026] Step 4:

[1027] Based on these analysis results, the server accesses the online medical appointment system and makes an appointment for the desired date and time.

[1028] Step 5:

[1029] The server receives the reservation confirmation information and sends it to the terminal.

[1030] Step 6:

[1031] The device receives a notification that the reservation is complete and informs the user via voice, "Your reservation is complete."

[1032] Implementing requested actions (such as updating the shopping list)

[1033] Step 1:

[1034] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[1035] Step 2:

[1036] The device sends voice data to the server.

[1037] Step 3:

[1038] The server analyzes the audio data and converts the request into text.

[1039] Step 4:

[1040] The server uses an emotion engine to recognize the user's emotions. If the user is feeling stressed, it will recognize that.

[1041] Step 5:

[1042] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list.

[1043] Step 6:

[1044] The server notifies the terminal that the request has been completed.

[1045] Step 7:

[1046] The device receives a notification and informs the user by voice, "Milk has been added to your shopping list." It can also provide additional, emotionally-based responses such as, "Is there anything else I can help you with?"

[1047] This allows users to recognize their emotions and receive appropriate responses, enabling them to receive both life support and emotional care simultaneously.

[1048] (Example 2)

[1049] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[1050] In modern society, elderly people living alone suffer from loneliness and a lack of support in their daily lives. This problem can have a significant impact on their mental health and quality of life. Furthermore, for elderly people who are unfamiliar with digital technology, the means to receive appropriate support are limited. Against this backdrop, there is a need for communication tools that are easy for elderly people living alone to use, understand their emotions, and respond appropriately, but existing solutions do not have sufficient accuracy in emotion recognition or natural dialogue.

[1051] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1052] In this invention, the server includes means for the user to input basic information, means for transmitting the basic information to the server, and means for storing the basic information in a database. This makes it possible to register the user's basic information smoothly.

[1053] Furthermore, the server includes means for voice input of everyday conversations with the user, means for transmitting the voice data to the server, means for converting the voice data into text, means for recognizing emotions based on the text, means for generating a response based on the recognized emotions, and means for outputting the response to the user in voice. This enables natural everyday conversations and emotional support with the user.

[1054] Furthermore, the server includes means for inputting voice from the user, including the recipient and message content; means for transmitting the voice data to the server and sending the message to the recipient; means for recognizing the emotion of the message and generating a response in an appropriate tone; and means for notifying that the message transmission is complete. This allows users to send emotionally charged messages to others, and the process is seamless.

[1055] The server also includes means for voice input of the user's preferred date and time for a medical appointment, means for transmitting the preferred date and time to the server, and means for making a medical appointment based on the preferred date and time. This makes online medical appointments intuitive and improves user convenience.

[1056] Finally, the server includes means for voice input of user requests, means for transmitting the voice data to the server and executing the corresponding operation, and means for generating a response that takes into account the emotions recognized based on the user's voice input. This allows the server to understand the user's emotions and then perform the request and notify the user of the result.

[1057] These methods provide a system that allows users, even those unfamiliar with digital technology, to receive emotional support through natural conversation and perform necessary daily tasks without stress.

[1058] "Basic information" refers to detailed information about the user, such as name, age, address, and health information, which is necessary to identify the individual and improve the quality of support.

[1059] A "server" is a computer system that stores basic information, manages voice data from users, and generates responses.

[1060] "Voice input means" refers to a device that includes a microphone and its control software for capturing user speech as digital voice data.

[1061] "Audio data" refers to data obtained by converting a user's speech into a digital format.

[1062] A "database" is a management system used by a server to store basic information and other user data.

[1063] "Means of converting to text" refers to software or algorithms that analyze audio data and convert it into text information.

[1064] "Means of recognizing emotions" refer to algorithms and software that analyze and determine emotions from the textual content of a user's speech.

[1065] A "means for generating responses" refers to a system that uses natural language processing techniques to produce appropriate responses based on recognized emotions and user input.

[1066] "Means of sending messages" refers to communication technology for sending messages generated from voice data to recipients specified by the user.

[1067] "Method for making medical appointments" refers to a system that allows users to confirm online appointments with medical institutions based on their preferred date and time.

[1068] "Means for executing operations" refers to a control system that receives requests from users and executes corresponding processing on the server.

[1069] "Notification methods" refer to functions that inform users of important information, such as the completion of a process or the sending of a message, via voice or text.

[1070] This invention is a communication tool for elderly people living alone, aiming to support users' daily lives and improve the convenience of communication. This system provides the following functions:

[1071] Basic Information Registration

[1072] The user enters their basic information, sends it to the server, and the server stores it in a database. This process allows the system to understand the user's individual information and use it to improve future service provision. For example, a user enters their name, age, address, and health information into their device and sends it to the server. The server receives it, saves it in the database, and then sends a notification to the device that the data has been saved.

[1073] Daily conversation support

[1074] When a user speaks into the device to enjoy a casual conversation, the device records the audio and sends it to a server. The server converts the audio to text, recognizes the emotion, and generates the most appropriate response. The device then converts that response back into audio and plays it back to the user. For example, if the user says, "The weather's nice today," the device might respond, "It really is wonderful weather."

[1075] Sending and receiving messages

[1076] When a user wants to send a message, they input their voice into the device. The device sends the voice data to a server, which converts the voice into text, recognizes the emotion, and generates an appropriate message. The generated message is sent to the recipient, and the device is notified when the message has been sent. For example, if a user says, "Send my son a message saying 'I'm fine'," the server recognizes that emotion and generates and sends the message, "I'm fine. I'm a little lonely, though."

[1077] Online medical consultation booking

[1078] When a user voice-inputs their desired date and time for a medical appointment, the device sends it to a server. The server analyzes the requested date and time, accesses the online medical appointment system, and makes the reservation. Once the reservation is complete, a notification is sent to the device. For example, if a user says, "Please schedule an appointment for me this Friday at 3pm," the server analyzes the information and confirms the online medical appointment.

[1079] Implementation of requested items

[1080] When a user speaks a request to the device, the device records the audio and sends it to the server. The server analyzes the request, recognizes the emotion, and performs the appropriate action. Once the action is complete, a notification is sent to the device. For example, if a user says, "Add milk to my shopping list," the server recognizes the request and adds the milk to the shopping list.

[1081] The following hardware and software are used to implement each function.

[1082] Speech recognition software: Google Speech-to-Text API

[1083] Text conversion software: IBM Watson Natural Language Understanding

[1084] Natural Language Processing Technology: OpenAI's GPT-3

[1085] Text-to-speech software: Amazon Polly

[1086] Example of a prompt

[1087] "Based on what a user says to the system designed for elderly people living alone, please use an emotion engine to generate the most appropriate response. For example, please tell us what kind of response would be appropriate to a conversation like, 'The weather is nice today.'"

[1088] As described above, users can perform various operations using voice commands and receive emotional support. This system is expected to improve the quality of life for elderly people living alone and reduce feelings of loneliness.

[1089] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1090] Basic Information Registration

[1091] Step 1:

[1092] The user enters their basic information (name, age, address, health information, etc.) into the device. Specifically, the user enters the necessary information into the input form using the device's keyboard or touchscreen.

[1093] Input: User's basic information (name, age, address, health information)

[1094] Output: Input basic information data

[1095] Step 2:

[1096] The terminal sends the entered basic information to the server. Specifically, the terminal's application uses an HTTP POST request to send data to the server's information registration API.

[1097] Input: Entered basic information data

[1098] Output: Status of successful transmission to server

[1099] Step 3:

[1100] The server saves the basic information it receives to the database. Specifically, the server opens a database connection and saves the information using an SQL INSERT statement.

[1101] Input: Basic information data sent to the server

[1102] Output: Status of saving to database

[1103] Step 4:

[1104] The server sends a notification to the device when saving is complete. Specifically, the server returns a save completion message in an HTTP response.

[1105] Input: Database save complete status

[1106] Output: Notification that saving to the device is complete.

[1107] Step 5:

[1108] The device receives a notification that the save is complete and displays "Registration complete." Specifically, it updates the GUI and displays a pop-up message.

[1109] Input: Server save completion notification

[1110] Output: Display of the message "Registration complete".

[1111] Daily conversation support

[1112] Step 1:

[1113] The user speaks to the device, saying, "The weather's nice today." Specifically, this is done using the device's microphone function for voice input.

[1114] Input: User voice input ("The weather is nice today.")

[1115] Output: Audio data

[1116] Step 2:

[1117] The device records audio data. Specifically, it captures audio input using the microphone as digital data.

[1118] Input: User's voice

[1119] Output: Recorded audio data

[1120] Step 3:

[1121] The device sends audio data to the server. Specifically, it sends the audio data, which has been Base64 encoded, via an HTTP POST request.

[1122] Input: Recorded audio data

[1123] Output: Status of successful transmission to server

[1124] Step 4:

[1125] The server converts the audio data to text. Specifically, it uses the Google Speech-to-Text API to convert the audio data to text.

[1126] Input: Audio data sent to the server

[1127] Output: Text data ("The weather is nice today.")

[1128] Step 5:

[1129] The server recognizes emotions from text data. Specifically, it uses IBM Watson Natural Language Understanding to extract emotions from the text.

[1130] Input: Text data

[1131] Output: Recognized emotion (positive emotion)

[1132] Step 6:

[1133] The server generates responses based on emotions. Specifically, it uses OpenAI's GPT-3 model to generate appropriate reply text.

[1134] Input: Recognized emotions and text data

[1135] Output: Response text ("What lovely weather we have today!")

[1136] Step 7:

[1137] The server sends the response text to the terminal. Specifically, it returns the response text data as an HTTP response.

[1138] Input: Response text

[1139] Output: Sending status to terminal and response text

[1140] Step 8:

[1141] The device converts the response text into speech and plays it back to the user. Specifically, it uses Amazon Polly Text-to-Speech software to generate the speech and plays it through the device's speaker.

[1142] Input: Response text

[1143] Output: Voice response ("What beautiful weather we have today!")

[1144] Sending and receiving messages

[1145] Step 1:

[1146] The user speaks to the device, saying, "Send a message to my son saying, 'I'm doing well.'" Specifically, they use the voice input function to instruct the device to send the message.

[1147] Input: User's voice command to send a message

[1148] Output: Audio data

[1149] Step 2:

[1150] The device records audio. Specifically, it uses the microphone to capture audio as digital data.

[1151] Input: User's voice command to send a message

[1152] Output: Recording data

[1153] Step 3:

[1154] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[1155] Input: Recording data

[1156] Output: Status of successful transmission to server

[1157] Step 4:

[1158] The server converts the speech to text. Specifically, it uses the Google Speech-to-Text API to convert the speech data to text.

[1159] Input: Audio data

[1160] Output: Text data ("Send a message to your son saying 'I'm doing well.'")

[1161] Step 5:

[1162] The server recognizes emotions and generates messages in an appropriate tone. Specifically, it uses IBM Watson Natural Language Understanding to recognize emotions and OpenAI GPT-3 to generate messages.

[1163] Input: Text data, user sentiment

[1164] Output: Response text ("I'm fine. I'm a little lonely though.")

[1165] Step 6:

[1166] The server sends the message to the recipient. Specifically, it uses the SMS API or email sending API to send the message.

[1167] Input: Response text

[1168] Output: Message sent successfully status

[1169] Step 7:

[1170] The server notifies the terminal that the transmission is complete. Specifically, it sends a transmission completion notification as an HTTP response.

[1171] Input: Message sent status

[1172] Output: Notification of successful transmission to the terminal

[1173] Step 8:

[1174] The terminal notifies the user that the message has been sent. Specifically, it updates the GUI and displays "Message sent."

[1175] Input: Server transmission completion notification

[1176] Output: Displays "Message sent".

[1177] Online medical consultation booking

[1178] Step 1:

[1179] The user speaks into the terminal, saying, "Please schedule an appointment for me this Friday at 3 PM." Specifically, they use the voice input function to enter their desired appointment date and time.

[1180] Input: User voice commands

[1181] Output: Audio data

[1182] Step 2:

[1183] The device records audio. Specifically, it uses the microphone to capture the user's voice as digital data.

[1184] Input: User voice commands

[1185] Output: Recording data

[1186] Step 3:

[1187] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[1188] Input: Recording data

[1189] Output: Status of successful transmission to server

[1190] Step 4:

[1191] The server analyzes the audio and converts the desired date and time into text. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text data and extract the date and time information.

[1192] Input: Audio data

[1193] Output: Text data of the desired date and time ("This Friday at 3 PM")

[1194] Step 5:

[1195] The server accesses the online medical appointment system and attempts to book an appointment for the desired date and time. Specifically, it sends an appointment request using the medical appointment API.

[1196] Input: Text data of desired date and time

[1197] Output: Booking complete status

[1198] Step 6:

[1199] The server notifies the terminal that the reservation is complete. Specifically, it sends a reservation completion notification as an HTTP response.

[1200] Input: Booking complete status

[1201] Output: Reservation completion notification to the device

[1202] Step 7:

[1203] The device notifies the user that the reservation is complete. Specifically, it updates the GUI and displays "Reservation complete."

[1204] Input: Reservation completion notification from the server

[1205] Output: "Reservation complete" is displayed.

[1206] Implementation of requested items

[1207] Step 1:

[1208] The user speaks to the device, saying, "Add milk to my shopping list." Specifically, they use the voice input function to enter their request.

[1209] Input: User voice commands

[1210] Output: Audio data

[1211] Step 2:

[1212] The device records audio. Specifically, it uses the microphone to capture the user's voice as digital data.

[1213] Input: User voice commands

[1214] Output: Recording data

[1215] Step 3:

[1216] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[1217] Input: Recording data

[1218] Output: Status of successful transmission to server

[1219] Step 4:

[1220] The server analyzes the audio data and converts the request into text. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text data and extract the request.

[1221] Input: Audio data

[1222] Output: Text data of the request ("Add milk to the shopping list")

[1223] Step 5:

[1224] The server analyzes the request, recognizes the emotion using an emotion engine, and performs the appropriate action. Specifically, it uses a generative AI model (GPT-3) to generate an emotion-based response and perform the corresponding action. For example, adding an item to a shopping list.

[1225] Input: Text data of the request and the recognized emotion.

[1226] Output: Operation complete status ("Milk has been added to shopping list")

[1227] Step 6:

[1228] The server notifies the terminal that the operation is complete. Specifically, it sends an operation completion notification as an HTTP response.

[1229] Input: Operation completion status

[1230] Output: Operation completion notification to the terminal

[1231] Step 7:

[1232] The device notifies the user that the operation is complete. Specifically, it updates the GUI and displays a message such as "Milk has been added to your shopping list." Alternatively, it provides an audio notification.

[1233] Input: Operation completion notification from the server

[1234] Output: A message or notification saying "Milk has been added to your shopping list."

[1235] The above outlines the specific processing steps of the system's program. In this way, even users unfamiliar with digital technology can perform various daily tasks while receiving emotional support through natural conversation.

[1236] (Application Example 2)

[1237] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1238] When elderly people living alone or older seniors look for products in physical stores or have questions, it can be time-consuming to ask store staff for assistance. Furthermore, appropriate responses tailored to their emotional state are required, but not all store staff necessarily possess the ability or time to do so. In this situation, there is a need to improve the convenience and sense of security for elderly people living alone or older seniors when they are doing their daily shopping.

[1239] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input basic information, means for transmitting the basic information to the server, and means for storing the basic information in a database. This ensures that the user's basic information is securely stored on the server, providing a foundation for daily support. The system also includes means for voice input of daily conversations with the user, means for transmitting the voice data to the server, and means for converting the voice data into text. This allows the system to recognize and transcribe questions even when the user asks them by voice.

[1240] Furthermore, the system includes means for recognizing the user's emotions using an emotion engine, means for generating a response based on the text and adjusting it based on the user's emotions, means for outputting the response to the user as audio, and means for providing the optimal response to the user's questions in a physical store. This makes it possible to generate a response that matches the user's emotional state, enabling a more human-like interaction.

[1241] As a result, elderly people living alone and older adults can shop with peace of mind, reduce stress in stores, and be provided with an environment where they can easily find products and ask questions.

[1242] "Means for inputting basic user information" refers to devices or interfaces that allow users to input information about themselves (such as name, age, and address) into a terminal.

[1243] "Means of sending basic information to a server" refers to functions or protocols that transfer basic information entered by a user to a server via the internet or a local network.

[1244] A "server that stores basic information in a database" is a server equipped with a digital storage system that efficiently and securely stores the basic information of users that has been transmitted.

[1245] "Means of notifying that saving is complete" refers to functions or mechanisms that inform the user that basic information has been successfully saved to the database.

[1246] "Means for inputting everyday conversations by voice" refers to microphones or voice recognition devices that allow users to input everyday conversations as voice into a system.

[1247] "Means for sending audio data to a server" refers to the functions and protocols for transferring acquired audio data to a server via the internet or a local network.

[1248] A "server that converts audio data to text" is a server equipped with software and hardware to receive audio data and convert it into text format.

[1249] A "response-generating server" is a server equipped with algorithms and software that generate responses to users based on text data.

[1250] "Means of outputting a response to the user in audio" refers to speakers or headphones that play back the generated response as audio to the user using speech synthesis technology.

[1251] "Means for inputting the recipient and message content as voice" refers to a microphone or voice recognition device used by a user to input the content of a message they intend to send to a specific recipient as voice.

[1252] A "server that sends audio data to a server and sends a message to a recipient" is a server equipped with the function of converting audio data into text and sending a message to a designated recipient.

[1253] "A means of inputting desired date and time for medical appointments by voice" refers to a microphone or voice recognition device that allows the user to input their desired date and time for medical appointments into the system by voice.

[1254] "Means for sending desired date and time to the server" refers to a function or protocol that transmits the user's voice-inputted desired date and time for a medical consultation to a server via the internet or a local network.

[1255] A "server that makes medical appointments based on requested dates and times" is a server equipped with the function to confirm appointments in conjunction with a medical appointment system based on the received requested dates and times.

[1256] An "emotion engine" is an algorithm or software that estimates emotions from a user's words and actions and adjusts its response accordingly.

[1257] "Means of adjusting responses based on user emotions" refers to a function that appropriately modifies the tone and content of generated responses based on the user's emotional state recognized by the emotion engine.

[1258] "Means of providing optimal responses to user questions within a physical store" refers to interfaces and systems that provide appropriate information and instructions via voice in response to user questions and requests within a physical store.

[1259] This invention's system is designed to make it easier and safer for elderly people living alone or those in general to navigate situations such as shopping at physical stores. The system begins with the input of the user's basic information and then supports daily conversations, sending and receiving messages, scheduling online medical appointments, and fulfilling requests. In addition, it features an emotion engine that recognizes the user's emotions and adjusts its responses accordingly.

[1260] 1. Management of user basic information

[1261] The user first enters their basic information into their smartphone or smart glasses, and this information is sent to the server. The server stores the information in a database and sends a notification to the device when the saving is complete.

[1262] 2. Support for everyday conversation

[1263] Users can speak to a terminal in a physical store. For example, they might ask, "Where is this product?" The voice device records this, and the voice data is sent to a server. The server converts the voice data into text and uses an emotion engine to recognize emotions. Based on this information, the most appropriate response is generated and sent to the terminal. The terminal converts this response back into voice and provides the user with a response such as, "That product is in the second column from the right."

[1264] 3. Sending and receiving messages

[1265] For example, if a user says, "Send my son a message saying 'I'm doing well'," the voice is recorded and sent to the server. The server converts the voice data into text and uses an emotion engine to recognize the emotion. It then generates a message in an appropriate tone and sends it to the recipient. For example, it might provide a tailored message like, "I'm doing well. I'm a little lonely, though."

[1266] 4. Booking an online medical consultation

[1267] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the voice data is sent to the server. The server analyzes this data and enters the desired date and time into the online booking system. Once the booking is complete, the information is sent to the terminal, and the terminal notifies the user that "Your booking is complete."

[1268] 5. Implementation of requested items

[1269] User requests are entered via voice input and sent to the server. For example, if a user says, "Add milk to my shopping list," that voice data is sent to the server and the corresponding action is performed. An emotion engine is also included, and if the user is feeling stressed, a response in a more appropriate tone will be provided, such as, "Milk has been added to the list. Is there anything else I can help you with?"

[1270] Hardware / software to use

[1271] The primary hardware used will be smartphones and smart glasses, utilizing their built-in microphones and speakers. The following technologies and tools will be used in the software:

[1272] Speech recognition: Use the speech_recognition library to convert speech data into text.

[1273] Emotion Recognition: Implement an emotion recognition model using the transformers library.

[1274] Speech synthesis: Generate speech using gTTS or other speech synthesis libraries, and play it back using the playsound library.

[1275] Examples of specific cases and prompt statements

[1276] Specific example

[1277] Assume the user will be using smart glasses in the store:

[1278] User: "Where can I find this product?"

[1279] System: "That item is in the second row from the right. Please let us know if you need further information."

[1280] Example of a prompt

[1281] If the user wants to have a casual conversation:

[1282] "It's a beautiful day today."

[1283] System response: "It truly is a beautiful day!"

[1284] This system will improve the in-store shopping experience and allow elderly people living alone and seniors to shop with peace of mind.

[1285] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1286] Step 1:

[1287] The user enters basic information. The user uses the input interface of their smartphone or smart glasses to enter basic information such as their name, age, and address. This entered information is stored in a database, creating a baseline of user information.

[1288] Step 2:

[1289] The terminal sends the entered basic information to the server. The server receives this information and processes it for storage in the database. As output, a save completion notification is generated and sent to the terminal.

[1290] Step 3:

[1291] The system uses voice input for everyday conversations. When a user speaks into the device, the voice device (microphone) records it. The input voice data is sent to a server, which stores it in a database and processes it for impact analysis and recognition.

[1292] Step 4:

[1293] The device sends the recorded audio data to the server. This data is stored on the server and used in the next step.

[1294] Step 5:

[1295] The process converts audio data to text. The server uses speech recognition software (e.g., the speech_recognition library) to convert the transmitted audio data into text. The input is audio data, and the output is text data. During this conversion process, data calculations are performed to analyze the characteristics of the audio data and convert it into text format.

[1296] Step 6:

[1297] The system recognizes emotions and generates responses. The server uses a generative AI model to recognize user emotions from text data (e.g., the emotion recognition model in the transformers library). Input is text data, and output is response text with emotion labels. Based on this, the system generates responses in an appropriate emotional tone.

[1298] Step 7:

[1299] The response is output as audio. After the server recognizes the emotion, it uses speech synthesis software (e.g., gTTS library) to convert the response text into audio data. The terminal receives this audio data and plays it back to the user through speakers or headphones. The input is the response text, and the output is the audio data.

[1300] Step 8:

[1301] Sending a message. When the user issues a message sending command, the device sends it as audio data to the server. The server converts the audio data to text, performs emotion recognition, generates a message text in an appropriate tone, and sends it to the designated recipient. The input is the audio data, and the output is the message to be sent.

[1302] Step 9:

[1303] The process involves making a medical appointment. The user inputs their desired appointment date and time via voice, and the terminal sends this to the server. The server converts the desired date and time into text, accesses the online booking system, and confirms the appointment. The input is the voice data of the desired date and time, and the output is a booking confirmation notification.

[1304] Step 10:

[1305] The system executes the requested actions. The user inputs their specific request via voice, and the terminal sends this to the server. The server analyzes the request, recognizes the emotion, generates a response in an appropriate tone, and performs the corresponding action. The input is the voice data of the request, and the output is the appropriate response and the execution result.

[1306] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1307] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1308] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1309] [Third Embodiment]

[1310] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1311] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1312] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1313] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1314] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1315] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1316] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1317] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1318] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1319] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1320] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1321] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1322] The communication tool for elderly people living alone according to the present invention is designed to support daily life and improve the convenience of communication. This system includes a user-friendly terminal and a server that processes data and provides necessary services.

[1323] Specifically, the user enters basic information through their device and sends that information to the server. The server then stores the information in a database and notifies the device that registration is complete. Through this process, the system can manage the user's basic information.

[1324] Furthermore, users can enjoy everyday conversations by speaking directly to the device. For example, if a user says, "The weather's nice today," the device records the voice and sends it to the server. The server converts the voice data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[1325] Furthermore, sending and receiving messages is easy. When a user says to the device, "Send a message to my son saying 'I'm doing well'," the device records the voice and sends it to the server. The server retrieves the recipient's (in this case, the son's) contact information from its database and sends the specified message to that contact. Once the server has finished processing, it notifies the device, and the device then informs the user by voice, "The message has been sent."

[1326] Making appointments is also easy for users. When a user says, "Please schedule an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data and makes an online appointment for the specified date and time. Once the appointment is complete, the server sends the information back to the device, which then notifies the user that "the appointment is complete."

[1327] Furthermore, it can also respond to user requests. For example, if a user says, "Add milk to my shopping list," the device sends voice data to the server, and the server adds the appropriate item to the shopping list. Once the server has finished processing, it notifies the device of the result, and the device informs the user of the result.

[1328] Through these processes, the system of the present invention provides elderly people living alone with the communication and support they need in their daily lives, significantly improving their convenience. Furthermore, since all operations are voice-based, it is designed to be easy to use even for users unfamiliar with digital technology.

[1329] The following describes the processing flow.

[1330] User Registration

[1331] Step 1:

[1332] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[1333] Step 2:

[1334] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[1335] Step 3:

[1336] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[1337] Step 4:

[1338] The server notifies the terminal that saving the user information is complete.

[1339] Step 5:

[1340] The device receives a notification and displays "Registration Complete".

[1341] Support for everyday conversation

[1342] Step 1:

[1343] The user speaks to the device. For example, they might say, "The weather's nice today."

[1344] Step 2:

[1345] The device records the user's voice and sends the audio data to the server.

[1346] Step 3:

[1347] The server uses a speech recognition engine to convert the audio data into text.

[1348] Step 4:

[1349] The server generates an appropriate response based on the converted text. Natural language processing is used in this process.

[1350] Step 5:

[1351] The server sends a response text to the terminal.

[1352] Step 6:

[1353] The device converts the received response into speech and plays it back to the user. For example, it might say, "What beautiful weather we have today."

[1354] Send message

[1355] Step 1:

[1356] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[1357] Step 2:

[1358] The device sends audio data to the server. The data is encoded in the appropriate format.

[1359] Step 3:

[1360] The server converts the audio data into text and analyzes the message content for the recipient.

[1361] Step 4:

[1362] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[1363] Step 5:

[1364] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[1365] Step 6:

[1366] The device informs the user of notifications it has received. For example, it might say, "A message has been sent."

[1367] Online medical consultation booking

[1368] Step 1:

[1369] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[1370] Step 2:

[1371] The device sends voice data to the server.

[1372] Step 3:

[1373] The server analyzes the voice data and converts the desired reservation date and time into text.

[1374] Step 4:

[1375] The server accesses the online medical appointment system and attempts to make a reservation for that date and time.

[1376] Step 5:

[1377] The server receives the reservation confirmation information and sends the result to the terminal.

[1378] Step 6:

[1379] The device plays a voice message to the user confirming that the reservation has been received. For example, it might say, "Your reservation is complete."

[1380] Implementing requested actions (such as updating the shopping list)

[1381] Step 1:

[1382] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[1383] Step 2:

[1384] The device sends voice data to the server.

[1385] Step 3:

[1386] The server analyzes the audio data and converts the request into text.

[1387] Step 4:

[1388] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list in the database.

[1389] Step 5:

[1390] The server confirms that the request has been completed and notifies the terminal.

[1391] Step 6:

[1392] The device will verbally inform the user of notifications it has received. For example, it might say, "Milk has been added to your shopping list."

[1393] (Example 1)

[1394] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1395] The objective of this invention is to address the lack of communication and support that elderly people living alone face in their daily lives. In particular, it aims to provide a voice-based interface that can be easily used even by elderly people who are unfamiliar with digital technology, enabling them to enjoy everyday conversations, obtain necessary information, send and receive messages, and easily make medical appointments.

[1396] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1397] In this invention, the server includes means for the user to input general information, means for transmitting the general information to an information processing device, and an information processing device for storing the general information in a storage device. This makes it possible for the user to easily register basic information, enjoy everyday conversations, send messages, make medical appointments, and more, all through voice.

[1398] A "user" is an individual who utilizes the system of the present invention.

[1399] "General information" refers to basic information such as name, address, and contact information entered by the user.

[1400] An "information processing device" is a device that processes information transmitted by a user and provides necessary services.

[1401] A "storage device" is a device used by an information processing device to permanently store information.

[1402] "Voice input" is a method of inputting information or instructions using voice.

[1403] "Speech" refers to sound information generated by the user's utterances.

[1404] A "document" is information in the form obtained when an information processing device converts speech into text.

[1405] A "response" is the answer generated by an information processing device in response to a user's inquiry.

[1406] A "recipient" is an individual or organization that receives a message from a user.

[1407] "Message content" refers to the information or notification that the user wants to send.

[1408] "Appointment scheduling" is the process of making a medical appointment for a date and time specified by the user.

[1409] "Desired date and time" refers to the specific date and time the user wishes to schedule an appointment.

[1410] A "request" is a specific demand or instruction that a user wants to perform through the system.

[1411] "Natural language processing" is a technology that enables information processing devices to understand human language and generate appropriate responses.

[1412] "Methods of confirming online" refer to methods of confirming procedures such as reservations via the internet.

[1413] "The relevant operation" refers to the specific actions or procedures that the system performs in response to the user's request.

[1414] This invention relates to a communication tool for elderly people living alone, and is a system that uses a voice-based interface to allow users to easily input basic information, enjoy daily conversations, send messages, make medical appointments, and perform other requests.

[1415] How users enter information

[1416] The user inputs general information by voice through the microphone into the terminal. The terminal has a recording function, and the input voice is saved as digital audio data. The terminal sends this digital audio data to an information processing device (server). Specifically, the data is transmitted securely using the HTTPS protocol.

[1417] The server receives the transmitted information and saves it to a storage device (e.g., a MySQL database). It also notifies the terminal when saving is complete, and the terminal notifies the user that "registration is complete." A speech synthesis engine (e.g., Amazon Polly) is used for this notification.

[1418] Processing everyday conversations

[1419] When a user speaks to the device, for example, saying "The weather's nice today," the device records the audio and sends it to the server. The server uses Google Cloud Speech-to-Text to convert the audio data into text. Then, it uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response and sends it to the device. The device plays the response using a speech synthesis engine, allowing the user to enjoy a natural conversation.

[1420] Send message

[1421] When the user says, "Send my son a message saying 'I'm doing well'," the device records the voice and sends it to the server. The server converts the voice data to text and retrieves the recipient's (the son's) contact information from its database. The server uses Twilio's SMS API to send the message to the recipient. Once the server notifies the user that processing is complete, the device informs the user that "the message has been sent."

[1422] Appointment booking

[1423] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the device sends the voice data to the server. The server converts the voice data into text and extracts the specified date and time information. The server uses the online medical consultation system's API to schedule the appointment. Once the appointment is complete, the server notifies the device that "the appointment is complete."

[1424] Responding to requests

[1425] When a user says, "Add milk to my shopping list," the device sends the voice data to the server. The server converts the voice data to text and adds the milk to the shopping list. Once the process is complete, the server notifies the device that "Milk has been added to your shopping list."

[1426] Examples of specific cases and prompt statements

[1427] For example, if a user says, "The weather's nice today," the following specific actions will be taken:

[1428] Voice input: "The weather is nice today."

[1429] Server processing: Converts audio data to text and generates responses using a generative AI model.

[1430] Response: "It really is a beautiful day today."

[1431] Example of a prompt:

[1432] "Please tell me what kind of response you should generate if a user says, 'The weather is nice today.'"

[1433] Hardware or software to use

[1434] Voice recording device: Terminal with microphone

[1435] Data transmission: HTTPS protocol

[1436] Speech recognition software: Google Cloud Speech-to-Text

[1437] Response generation model: OpenAI's GPT-3

[1438] Speech synthesis engine: Amazon Polly

[1439] Message sending API: Twilio

[1440] Data storage: MySQL database

[1441] The above describes the specific embodiments of the present invention. This system solves the lack of communication and support faced by elderly people living alone and supports their daily lives.

[1442] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1443] Basic Information Registration

[1444] Step 1:

[1445] The user inputs general information (name, address, contact information, etc.) by voice into the device. The device records this voice data as input.

[1446] Step 2:

[1447] The device converts recorded audio data into text data. This process uses speech recognition software (e.g., Google Cloud Speech-to-Text). Audio data is taken as input, and text data is obtained as output.

[1448] Step 3:

[1449] The terminal sends the converted text data to the server using the HTTPS protocol. Text data is used as input and sent to the server as output.

[1450] Step 4:

[1451] The server saves the received text data to a database. A storage device (e.g., a MySQL database) is used to permanently store the input text data. The output is a status indicating that saving is complete.

[1452] Step 5:

[1453] The server sends a notification to the device when saving is complete. The save completion status is used as input, and the notification is sent to the device as output.

[1454] Step 6:

[1455] The device converts the save completion notification into speech using a speech synthesis engine (e.g., Amazon Polly) and notifies the user that "Registration complete." The notification status is used as input, and the output is a voice message.

[1456] Processing everyday conversations

[1457] Step 1:

[1458] The user speaks into the device saying, "The weather's nice today." The device receives the user's voice data as input. The device then records this voice data.

[1459] Step 2:

[1460] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[1461] Step 3:

[1462] The server uses speech recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. Audio data is used as input, and text data is obtained as output.

[1463] Step 4:

[1464] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response based on the text. Text data is used as input, and the response text is obtained as output.

[1465] Step 5:

[1466] The server sends the generated response text to the terminal. The response text is used as input and sent to the terminal as output.

[1467] Step 6:

[1468] The device converts the response text into speech using a text-to-speech engine (e.g., Amazon Polly) and plays it back to the user. The response text is used as input, and the output is a voiced message.

[1469] Send message

[1470] Step 1:

[1471] The user says, "Send a message to my son saying, 'I'm doing well.'" The input is the user's voice data. The device records this voice data.

[1472] Step 2:

[1473] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[1474] Step 3:

[1475] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[1476] Step 4:

[1477] The server retrieves the recipient's contact information from the database. Text data is used as input, and the recipient's contact information is obtained as output.

[1478] Step 5:

[1479] The server uses Twilio's SMS API to send a message to the recipient. The recipient's contact information and message content are used as input, and the message sending status is obtained as output.

[1480] Step 6:

[1481] The server sends a status message to the terminal indicating that the transmission is complete. The message transmission status is used as input and sent to the terminal as output.

[1482] Step 7:

[1483] The device converts the message transmission completion notification into speech using a speech synthesis engine and notifies the user that "the message has been sent." The notification status is used as input, and the voice message is obtained as output.

[1484] Appointment booking

[1485] Step 1:

[1486] The user says, "Please schedule an appointment for me this Friday at 3 PM." The user's voice data is obtained as input. The device records this voice data.

[1487] Step 2:

[1488] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[1489] Step 3:

[1490] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[1491] Step 4:

[1492] The server analyzes the text data and extracts information about the desired date and time. Text data is used as input, and the output is information about the desired date and time.

[1493] Step 5:

[1494] The server uses the online medical consultation system's API to make reservations. The desired date and time are used as input, and the output is a reservation completion status.

[1495] Step 6:

[1496] The server sends a reservation completion notification to the device. The reservation completion status is used as input and sent to the device as output.

[1497] Step 7:

[1498] The device converts the reservation completion notification into speech using a speech synthesis engine and notifies the user that "Your reservation is complete." The notification status is used as input, and the voice message is obtained as output.

[1499] Responding to requests

[1500] Step 1:

[1501] The user says, "Add milk to my shopping list." The device receives the user's voice data as input. The device then records this voice data.

[1502] Step 2:

[1503] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[1504] Step 3:

[1505] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[1506] Step 4:

[1507] The server analyzes the text data and identifies the relevant requests. Text data is used as input, and the requests are obtained as output.

[1508] Step 5:

[1509] The server performs the appropriate operation based on the requested information. The requested information is used as input, and the operation execution status is obtained as output.

[1510] Step 6:

[1511] The server sends a notification to the terminal that the operation is complete. The operation execution status is used as input and sent to the terminal as output.

[1512] Step 7:

[1513] The device converts the completion notification into speech using a speech synthesis engine and notifies the user that "Milk has been added to your shopping list." The notification status is used as input, and the voice message is obtained as output.

[1514] The above is the specific processing flow of the system.

[1515] (Application Example 1)

[1516] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1517] Elderly people living alone face the problem of being unable to rely on others for daily life support, such as ordering food delivery or other assistance, because traditional methods are cumbersome to use. Furthermore, users unfamiliar with digital technology require natural communication using voice recognition. Therefore, there is a need for a system that allows elderly people living alone to easily and naturally order food delivery and other daily life support.

[1518] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1519] In this invention, the server is

[1520] A means for users to input basic information,

[1521] Means for transmitting the aforementioned basic information to a server,

[1522] A server that stores the aforementioned basic information in a database,

[1523] A means for notifying that the aforementioned storage has been completed,

[1524] A means of inputting everyday conversations with the user via voice,

[1525] Means for transmitting the aforementioned audio data to a server,

[1526] A server that converts the aforementioned audio data into text,

[1527] A server that generates a response based on the aforementioned text,

[1528] A means for outputting the aforementioned response to the user in voice,

[1529] A means of inputting the recipient and message content by the user,

[1530] A server that sends the aforementioned audio data to a server and sends a message to a recipient,

[1531] A method for users to input their preferred date and time for medical appointments via voice input,

[1532] A means for sending the aforementioned desired date and time to the server,

[1533] A server that makes medical appointments based on the aforementioned preferred date and time,

[1534] A means of inputting user requests by voice,

[1535] A server that transmits the aforementioned audio data to a server and performs the corresponding operation,

[1536] A means for users to input food delivery orders by voice,

[1537] Means for sending the food delivery order to the server,

[1538] A server that processes orders based on the aforementioned food delivery orders,

[1539] A means for notifying the completion of the aforementioned order processing,

[1540] Includes.

[1541] This will make it easy and natural for elderly people living alone to order food delivery and other forms of assistance for daily life.

[1542] "User" refers to any person who uses this system.

[1543] "Basic information" refers to necessary information about the user, such as their name, address, and contact information.

[1544] A "server" refers to a computer system that processes and stores data, and notifies users.

[1545] A "database" refers to an information system used to systematically store basic information and other data.

[1546] "Means of notification" refers to methods of informing users about the completion of information storage or processing.

[1547] "Voice input" refers to a method of inputting information by voice through a microphone or similar device.

[1548] "Audio data" refers to digital audio information obtained from voice input.

[1549] "Converting to text" refers to the process of converting audio data into written text.

[1550] "Response" refers to the reply or instruction generated by the server based on the user's voice input.

[1551] A "message" refers to information with specified content that a user sends to other recipients.

[1552] "Medical appointment booking" refers to making a reservation for a specific date and time to receive medical treatment at a medical institution or similar facility.

[1553] "Requests" refer to the operations or requests that users want to perform on the system.

[1554] "Food delivery" refers to the process of a user ordering food to be delivered to a specified location.

[1555] "Order processing" refers to the procedure for confirming the details of a food delivery order and handling it appropriately.

[1556] "Means of notifying completion" refers to methods of informing users that an order or operation has been completed.

[1557] This invention is a system that provides the communication and support that elderly people living alone need in their daily lives. This system includes a terminal that is easy for the user to operate and a server that processes data and provides the necessary services.

[1558] The system is configured as follows:

[1559] The user first enters basic information (name, address, contact information, etc.) into the terminal. This basic information is sent from the terminal to the server, which stores it in a database. Once the saving is complete, the server notifies the terminal, and the terminal informs the user that the saving is complete. In this way, the system can manage the user's basic information.

[1560] Regarding support for everyday conversation, when a user says, "The weather's nice today," the device records the audio and sends it to the server. The server converts the audio data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[1561] It is also possible to send and receive messages. For example, if a user says, "Send a message to my son saying, 'I'm doing well'," the device will record the voice and send it to the server. The server will retrieve the recipient's contact information from its database and send the specified message. Once the server has finished processing, it will notify the device, and the device will inform the user by voice, "The message has been sent."

[1562] Similarly, when scheduling an appointment, if the user says, "Please schedule an appointment for me this Friday at 3pm," the device sends that voice message to the server. The server analyzes the voice data and schedules an appointment for the specified date and time. Once the appointment is complete, the server notifies the device, and the device informs the user that "the appointment is complete."

[1563] This system also supports ordering food delivery. When a user says, "Order curry for lunch today," the terminal records the voice and sends it to the server. The server converts the voice data into text, selects an appropriate food delivery service, and places the order. Once the order is complete, the server notifies the terminal of the result, and the terminal informs the user via voice, "Your order is complete."

[1564] The specific hardware used to realize this invention can be a smartphone or tablet. For the software, the "speech_recognition" library is used for speech recognition, the "pyttsx3" library for speech synthesis, and the "requests" library for processing API requests. Furthermore, a generative AI model is used for natural language processing (NLP).

[1565] Examples of specific prompt messages include the following:

[1566] "I'd like to order ~ for lunch today."

[1567] "Please schedule an appointment for this Friday at 3 PM."

[1568] This system offers significant convenience by allowing elderly people living alone to easily perform various daily tasks using voice commands. This reduces barriers to digital technology and improves the quality of daily life.

[1569] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1570] Step 1:

[1571] The user enters basic information. The user enters their name, address, contact information, etc., on the screen of their smartphone or tablet. This constitutes input.

[1572] Step 2:

[1573] The terminal sends the entered basic information to the server. The basic information is converted into a digital format and sent as a request to the server. The server then receives the data.

[1574] Step 3:

[1575] The server saves the received basic information to the database. Data processing includes correctly formatting the input information and writing it to the database. The server then verifies that saving is complete.

[1576] Step 4:

[1577] The server notifies the user when the basic information has been saved. It sends the status of save completion to the terminal. The terminal informs the user of this information via voice or screen display.

[1578] Step 5:

[1579] The user inputs everyday conversations via voice. For example, they might say, "The weather's nice today." This becomes the input.

[1580] Step 6:

[1581] The terminal records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data. A process of converting it to a data format takes place here.

[1582] Step 7:

[1583] The server converts the received audio data into text. A generative AI model is used for the conversion from speech to text. The converted text is obtained as output.

[1584] Step 8:

[1585] The server generates a response based on the converted text. It uses natural language processing (NLP) to analyze the text and create an appropriate response text, which is then output as the response.

[1586] Step 9:

[1587] The server sends the generated response to the terminal. The text data of the response is sent to the terminal, and preparations for speech synthesis are made.

[1588] Step 10:

[1589] The terminal plays back the response received from the server as audio. It uses a speech synthesis tool (pyttsx3) to convert text data into speech and plays it back to the user. This is the output to the user.

[1590] Step 11:

[1591] The user inputs the recipient and message content by voice. For example, they might say, "Send a message to my son saying, 'I'm doing well.'" This is the input.

[1592] Step 12:

[1593] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[1594] Step 13:

[1595] The server sends a message based on the recipient and message content. It retrieves the recipient's contact information from the database and sends the specified message.

[1596] Step 14:

[1597] The server notifies the terminal that the message has been sent. It sends a status message to the terminal indicating that the message has been sent, and the terminal then verbally informs the user that "the message has been sent."

[1598] Step 15:

[1599] Users can place food delivery orders by voice. For example, they might say, "Order curry for lunch today." This constitutes the input.

[1600] Step 16:

[1601] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[1602] Step 17:

[1603] The server converts voice data into text and processes the necessary food delivery orders. It uses a generative AI model to convert voice to text and then places orders with food delivery services based on that information.

[1604] Step 18:

[1605] The server notifies the terminal that the order processing is complete. It confirms that the order is complete and sends the status to the terminal. The terminal then informs the user by voice, "Your order is complete."

[1606] Through these steps, the system will be able to naturally use voice commands to order food deliveries and provide daily support for elderly people living alone.

[1607] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1608] The communication tool for elderly people living alone according to the present invention is designed to support the user's daily life and improve the convenience of communication. This system includes means for inputting and saving the user's basic information, support for daily conversation, sending and receiving messages, scheduling online medical appointments, and fulfilling requests, as well as a function that uses an emotion engine to recognize the user's emotions and generate responses.

[1609] First, the user enters their basic information into the device and sends that information to the server. The server saves the received information to a database and notifies the device when the saving is complete. The device displays "Registration Complete," and the management of the basic information is finished.

[1610] Next, the user speaks to the device to enjoy a casual conversation. For example, if the user says, "The weather's nice today," the device records the voice and sends the voice data to the server. The server converts the voice data into text and recognizes the emotion using an emotion engine. Based on the recognized emotion, it generates the most appropriate response and sends the response text to the device. The device then converts the response back into voice and plays it back to the user, saying, "The weather's really nice today." In this way, the user can receive emotional support through a natural conversation.

[1611] Furthermore, sending and receiving messages is easy. When a user speaks to the device, saying, "Send a message to my son saying, 'I'm fine,'" the device records the voice and sends it to the server. The server converts the recipient and message content into text, recognizes the user's emotions using an emotion engine, and generates a message in an appropriate tone. For example, if the user sends a message in a sad tone, the emotion engine recognizes that emotion and generates a message in a tone such as, "I'm fine. I'm a little lonely, though." The server sends the message to the recipient and notifies the device when it is complete. The device then informs the user by voice, "Message sent."

[1612] Furthermore, booking an online consultation is also easy. When a user says, "Please book an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data, transcribes the desired date and time into text, and attempts to book the appointment at that time by accessing the online consultation booking system. Once the booking is complete, the server sends that information to the device, which then notifies the user that "the booking is complete."

[1613] Regarding requests, when a user says, "Add milk to my shopping list," the device sends voice data to the server. The server analyzes the request, recognizes the user's emotions using an emotion engine, and performs the appropriate action. If the user is stressed, the system responds in a tone such as, "Milk has been added to the list. Is there anything else I can help with?" The server notifies the device that the action is complete, and the device then informs the user, "Milk has been added to your shopping list."

[1614] Through these processes, the system of the present invention recognizes the emotions of elderly people living alone and generates responses based on them, thereby providing more humane support. Since all operations are voice-based, it can be easily used even by users unfamiliar with digital technology. This system is expected to support the user's daily life and reduce feelings of loneliness.

[1615] The following describes the processing flow.

[1616] User Registration

[1617] Step 1:

[1618] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[1619] Step 2:

[1620] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[1621] Step 3:

[1622] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[1623] Step 4:

[1624] The server notifies the terminal that saving the user information is complete.

[1625] Step 5:

[1626] The device receives a notification and displays "Registration Complete".

[1627] Support for everyday conversation

[1628] Step 1:

[1629] The user speaks to the device. For example, they might say, "The weather's nice today."

[1630] Step 2:

[1631] The device records the user's voice and sends the audio data to the server.

[1632] Step 3:

[1633] The server uses a speech recognition engine to convert the audio data into text.

[1634] Step 4:

[1635] The server uses an emotion engine to recognize the user's emotions based on the converted text. For example, if the user's voice has a bright tone, it is recognized as "positive," and if it has a somber tone, it is recognized as "negative."

[1636] Step 5:

[1637] The server generates the most appropriate response based on the emotions it perceives. For example, if someone says "The weather's nice today," it will respond with "It really is wonderful weather."

[1638] Step 6:

[1639] The server sends a response text to the terminal.

[1640] Step 7:

[1641] The device converts the received response into speech and plays it back to the user. For example, it might respond with, "What beautiful weather we have today."

[1642] Send message

[1643] Step 1:

[1644] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[1645] Step 2:

[1646] The device sends audio data to the server. The data is encoded in the appropriate format.

[1647] Step 3:

[1648] The server converts the audio data into text and analyzes the message content for the recipient.

[1649] Step 4:

[1650] The server uses an emotion engine to recognize the user's emotions. For example, if a message is sent in a sad tone, the server will recognize that emotion.

[1651] Step 5:

[1652] The server generates messages based on the emotions it perceives. For example, if a message saying "I'm fine" is sent in a sad tone, it will generate a message saying "I'm fine, but I'm a little lonely."

[1653] Step 6:

[1654] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[1655] Step 7:

[1656] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[1657] Step 8:

[1658] The device notifies the user, for example, saying, "A message has been sent."

[1659] Online medical consultation booking

[1660] Step 1:

[1661] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[1662] Step 2:

[1663] The device sends voice data to the server.

[1664] Step 3:

[1665] The server converts the voice data into text and analyzes the desired reservation date and time.

[1666] Step 4:

[1667] Based on these analysis results, the server accesses the online medical appointment system and makes an appointment for the desired date and time.

[1668] Step 5:

[1669] The server receives the reservation confirmation information and sends it to the terminal.

[1670] Step 6:

[1671] The device receives a notification that the reservation is complete and informs the user via voice, "Your reservation is complete."

[1672] Implementing requested actions (such as updating the shopping list)

[1673] Step 1:

[1674] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[1675] Step 2:

[1676] The device sends voice data to the server.

[1677] Step 3:

[1678] The server analyzes the audio data and converts the request into text.

[1679] Step 4:

[1680] The server uses an emotion engine to recognize the user's emotions. If the user is feeling stressed, it will recognize that.

[1681] Step 5:

[1682] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list.

[1683] Step 6:

[1684] The server notifies the terminal that the request has been completed.

[1685] Step 7:

[1686] The device receives a notification and informs the user by voice, "Milk has been added to your shopping list." It can also provide additional, emotionally-based responses such as, "Is there anything else I can help you with?"

[1687] This allows users to recognize their emotions and receive appropriate responses, enabling them to receive both life support and emotional care simultaneously.

[1688] (Example 2)

[1689] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1690] In modern society, elderly people living alone suffer from loneliness and a lack of support in their daily lives. This problem can have a significant impact on their mental health and quality of life. Furthermore, for elderly people who are unfamiliar with digital technology, the means to receive appropriate support are limited. Against this backdrop, there is a need for communication tools that are easy for elderly people living alone to use, understand their emotions, and respond appropriately, but existing solutions do not have sufficient accuracy in emotion recognition or natural dialogue.

[1691] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1692] In this invention, the server includes means for the user to input basic information, means for transmitting the basic information to the server, and means for storing the basic information in a database. This makes it possible to register the user's basic information smoothly.

[1693] Furthermore, the server includes means for voice input of everyday conversations with the user, means for transmitting the voice data to the server, means for converting the voice data into text, means for recognizing emotions based on the text, means for generating a response based on the recognized emotions, and means for outputting the response to the user in voice. This enables natural everyday conversations and emotional support with the user.

[1694] Furthermore, the server includes means for inputting voice from the user, including the recipient and message content; means for transmitting the voice data to the server and sending the message to the recipient; means for recognizing the emotion of the message and generating a response in an appropriate tone; and means for notifying that the message transmission is complete. This allows users to send emotionally charged messages to others, and the process is seamless.

[1695] The server also includes means for voice input of the user's preferred date and time for a medical appointment, means for transmitting the preferred date and time to the server, and means for making a medical appointment based on the preferred date and time. This makes online medical appointments intuitive and improves user convenience.

[1696] Finally, the server includes means for voice input of user requests, means for transmitting the voice data to the server and executing the corresponding operation, and means for generating a response that takes into account the emotions recognized based on the user's voice input. This allows the server to understand the user's emotions and then perform the request and notify the user of the result.

[1697] These methods provide a system that allows users, even those unfamiliar with digital technology, to receive emotional support through natural conversation and perform necessary daily tasks without stress.

[1698] "Basic information" refers to detailed information about the user, such as name, age, address, and health information, which is necessary to identify the individual and improve the quality of support.

[1699] A "server" is a computer system that stores basic information, manages voice data from users, and generates responses.

[1700] "Voice input means" refers to a device that includes a microphone and its control software for capturing user speech as digital voice data.

[1701] "Audio data" refers to data obtained by converting a user's speech into a digital format.

[1702] A "database" is a management system used by a server to store basic information and other user data.

[1703] "Means of converting to text" refers to software or algorithms that analyze audio data and convert it into text information.

[1704] "Means of recognizing emotions" refer to algorithms and software that analyze and determine emotions from the textual content of a user's speech.

[1705] A "means for generating responses" refers to a system that uses natural language processing techniques to produce appropriate responses based on recognized emotions and user input.

[1706] "Means of sending messages" refers to communication technology for sending messages generated from voice data to recipients specified by the user.

[1707] "Method for making medical appointments" refers to a system that allows users to confirm online appointments with medical institutions based on their preferred date and time.

[1708] "Means for executing operations" refers to a control system that receives requests from users and executes corresponding processing on the server.

[1709] "Notification methods" refer to functions that inform users of important information, such as the completion of a process or the sending of a message, via voice or text.

[1710] This invention is a communication tool for elderly people living alone, aiming to support users' daily lives and improve the convenience of communication. This system provides the following functions:

[1711] Basic Information Registration

[1712] The user enters their basic information, sends it to the server, and the server stores it in a database. This process allows the system to understand the user's individual information and use it to improve future service provision. For example, a user enters their name, age, address, and health information into their device and sends it to the server. The server receives it, saves it in the database, and then sends a notification to the device that the data has been saved.

[1713] Daily conversation support

[1714] When a user speaks into the device to enjoy a casual conversation, the device records the audio and sends it to a server. The server converts the audio to text, recognizes the emotion, and generates the most appropriate response. The device then converts that response back into audio and plays it back to the user. For example, if the user says, "The weather's nice today," the device might respond, "It really is wonderful weather."

[1715] Sending and receiving messages

[1716] When a user wants to send a message, they input their voice into the device. The device sends the voice data to a server, which converts the voice into text, recognizes the emotion, and generates an appropriate message. The generated message is sent to the recipient, and the device is notified when the message has been sent. For example, if a user says, "Send my son a message saying 'I'm fine'," the server recognizes that emotion and generates and sends the message, "I'm fine. I'm a little lonely, though."

[1717] Online medical consultation booking

[1718] When a user voice-inputs their desired date and time for a medical appointment, the device sends it to a server. The server analyzes the requested date and time, accesses the online medical appointment system, and makes the reservation. Once the reservation is complete, a notification is sent to the device. For example, if a user says, "Please schedule an appointment for me this Friday at 3pm," the server analyzes the information and confirms the online medical appointment.

[1719] Implementation of requested items

[1720] When a user speaks a request to the device, the device records the audio and sends it to the server. The server analyzes the request, recognizes the emotion, and performs the appropriate action. Once the action is complete, a notification is sent to the device. For example, if a user says, "Add milk to my shopping list," the server recognizes the request and adds the milk to the shopping list.

[1721] The following hardware and software are used to implement each function.

[1722] Speech recognition software: Google Speech-to-Text API

[1723] Text conversion software: IBM Watson Natural Language Understanding

[1724] Natural Language Processing Technology: OpenAI's GPT-3

[1725] Text-to-speech software: Amazon Polly

[1726] Example of a prompt

[1727] "Based on what a user says to the system designed for elderly people living alone, please use an emotion engine to generate the most appropriate response. For example, please tell us what kind of response would be appropriate to a conversation like, 'The weather is nice today.'"

[1728] As described above, users can perform various operations using voice commands and receive emotional support. This system is expected to improve the quality of life for elderly people living alone and reduce feelings of loneliness.

[1729] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1730] Basic Information Registration

[1731] Step 1:

[1732] The user enters their basic information (name, age, address, health information, etc.) into the device. Specifically, the user enters the necessary information into the input form using the device's keyboard or touchscreen.

[1733] Input: User's basic information (name, age, address, health information)

[1734] Output: Input basic information data

[1735] Step 2:

[1736] The terminal sends the entered basic information to the server. Specifically, the terminal's application uses an HTTP POST request to send data to the server's information registration API.

[1737] Input: Entered basic information data

[1738] Output: Status of successful transmission to server

[1739] Step 3:

[1740] The server saves the basic information it receives to the database. Specifically, the server opens a database connection and saves the information using an SQL INSERT statement.

[1741] Input: Basic information data sent to the server

[1742] Output: Status of saving to database

[1743] Step 4:

[1744] The server sends a notification to the device when saving is complete. Specifically, the server returns a save completion message in an HTTP response.

[1745] Input: Database save complete status

[1746] Output: Notification that saving to the device is complete.

[1747] Step 5:

[1748] The device receives a notification that the save is complete and displays "Registration complete." Specifically, it updates the GUI and displays a pop-up message.

[1749] Input: Server save completion notification

[1750] Output: Display of the message "Registration complete".

[1751] Daily conversation support

[1752] Step 1:

[1753] The user speaks to the device, saying, "The weather's nice today." Specifically, this is done using the device's microphone function for voice input.

[1754] Input: User voice input ("The weather is nice today.")

[1755] Output: Audio data

[1756] Step 2:

[1757] The device records audio data. Specifically, it captures audio input using the microphone as digital data.

[1758] Input: User's voice

[1759] Output: Recorded audio data

[1760] Step 3:

[1761] The device sends audio data to the server. Specifically, it sends the audio data, which has been Base64 encoded, via an HTTP POST request.

[1762] Input: Recorded audio data

[1763] Output: Status of successful transmission to server

[1764] Step 4:

[1765] The server converts the audio data to text. Specifically, it uses the Google Speech-to-Text API to convert the audio data to text.

[1766] Input: Audio data sent to the server

[1767] Output: Text data ("The weather is nice today.")

[1768] Step 5:

[1769] The server recognizes emotions from text data. Specifically, it uses IBM Watson Natural Language Understanding to extract emotions from the text.

[1770] Input: Text data

[1771] Output: Recognized emotion (positive emotion)

[1772] Step 6:

[1773] The server generates responses based on emotions. Specifically, it uses OpenAI's GPT-3 model to generate appropriate reply text.

[1774] Input: Recognized emotions and text data

[1775] Output: Response text ("What lovely weather we have today!")

[1776] Step 7:

[1777] The server sends the response text to the terminal. Specifically, it returns the response text data as an HTTP response.

[1778] Input: Response text

[1779] Output: Sending status to terminal and response text

[1780] Step 8:

[1781] The device converts the response text into speech and plays it back to the user. Specifically, it uses Amazon Polly Text-to-Speech software to generate the speech and plays it through the device's speaker.

[1782] Input: Response text

[1783] Output: Voice response ("What beautiful weather we have today!")

[1784] Sending and receiving messages

[1785] Step 1:

[1786] The user speaks to the device, saying, "Send a message to my son saying, 'I'm doing well.'" Specifically, they use the voice input function to instruct the device to send the message.

[1787] Input: User's voice command to send a message

[1788] Output: Audio data

[1789] Step 2:

[1790] The device records audio. Specifically, it uses the microphone to capture audio as digital data.

[1791] Input: User's voice command to send a message

[1792] Output: Recording data

[1793] Step 3:

[1794] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[1795] Input: Recording data

[1796] Output: Status of successful transmission to server

[1797] Step 4:

[1798] The server converts the speech to text. Specifically, it uses the Google Speech-to-Text API to convert the speech data to text.

[1799] Input: Audio data

[1800] Output: Text data ("Send a message to your son saying 'I'm doing well.'")

[1801] Step 5:

[1802] The server recognizes emotions and generates messages in an appropriate tone. Specifically, it uses IBM Watson Natural Language Understanding to recognize emotions and OpenAI GPT-3 to generate messages.

[1803] Input: Text data, user sentiment

[1804] Output: Response text ("I'm fine. I'm a little lonely though.")

[1805] Step 6:

[1806] The server sends the message to the recipient. Specifically, it uses the SMS API or email sending API to send the message.

[1807] Input: Response text

[1808] Output: Message sent successfully status

[1809] Step 7:

[1810] The server notifies the terminal that the transmission is complete. Specifically, it sends a transmission completion notification as an HTTP response.

[1811] Input: Message sent status

[1812] Output: Notification of successful transmission to the terminal

[1813] Step 8:

[1814] The terminal notifies the user that the message has been sent. Specifically, it updates the GUI and displays "Message sent."

[1815] Input: Server transmission completion notification

[1816] Output: Displays "Message sent".

[1817] Online medical consultation booking

[1818] Step 1:

[1819] The user speaks into the terminal, saying, "Please schedule an appointment for me this Friday at 3 PM." Specifically, they use the voice input function to enter their desired appointment date and time.

[1820] Input: User voice commands

[1821] Output: Audio data

[1822] Step 2:

[1823] The device records audio. Specifically, it uses the microphone to capture the user's voice as digital data.

[1824] Input: User voice commands

[1825] Output: Recording data

[1826] Step 3:

[1827] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[1828] Input: Recording data

[1829] Output: Status of successful transmission to server

[1830] Step 4:

[1831] The server analyzes the audio and converts the desired date and time into text. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text data and extract the date and time information.

[1832] Input: Audio data

[1833] Output: Text data of the desired date and time ("This Friday at 3 PM")

[1834] Step 5:

[1835] The server accesses the online medical appointment system and attempts to book an appointment for the desired date and time. Specifically, it sends an appointment request using the medical appointment API.

[1836] Input: Text data of desired date and time

[1837] Output: Booking complete status

[1838] Step 6:

[1839] The server notifies the terminal that the reservation is complete. Specifically, it sends a reservation completion notification as an HTTP response.

[1840] Input: Booking complete status

[1841] Output: Reservation completion notification to the device

[1842] Step 7:

[1843] The device notifies the user that the reservation is complete. Specifically, it updates the GUI and displays "Reservation complete."

[1844] Input: Reservation completion notification from the server

[1845] Output: "Reservation complete" is displayed.

[1846] Implementation of requested items

[1847] Step 1:

[1848] The user speaks to the device, saying, "Add milk to my shopping list." Specifically, they use the voice input function to enter their request.

[1849] Input: User voice commands

[1850] Output: Audio data

[1851] Step 2:

[1852] The device records audio. Specifically, it uses the microphone to capture the user's voice as digital data.

[1853] Input: User voice commands

[1854] Output: Recording data

[1855] Step 3:

[1856] The device sends audio data to the server. Specifically, it sends Base64 encoded audio data via HTTP POST.

[1857] Input: Recording data

[1858] Output: Status of successful transmission to server

[1859] Step 4:

[1860] The server analyzes the audio data and converts the request into text. Specifically, it uses the Google Speech-to-Text API to convert the audio data into text data and extract the request.

[1861] Input: Audio data

[1862] Output: Text data of the request ("Add milk to the shopping list")

[1863] Step 5:

[1864] The server analyzes the request, recognizes the emotion using an emotion engine, and performs the appropriate action. Specifically, it uses a generative AI model (GPT-3) to generate an emotion-based response and perform the corresponding action. For example, adding an item to a shopping list.

[1865] Input: Text data of the request and the recognized emotion.

[1866] Output: Operation complete status ("Milk has been added to shopping list")

[1867] Step 6:

[1868] The server notifies the terminal that the operation is complete. Specifically, it sends an operation completion notification as an HTTP response.

[1869] Input: Operation completion status

[1870] Output: Operation completion notification to the terminal

[1871] Step 7:

[1872] The device notifies the user that the operation is complete. Specifically, it updates the GUI and displays a message such as "Milk has been added to your shopping list." Alternatively, it provides an audio notification.

[1873] Input: Operation completion notification from the server

[1874] Output: A message or notification saying "Milk has been added to your shopping list."

[1875] The above outlines the specific processing steps of the system's program. In this way, even users unfamiliar with digital technology can perform various daily tasks while receiving emotional support through natural conversation.

[1876] (Application Example 2)

[1877] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1878] When elderly people living alone or older seniors look for products in physical stores or have questions, it can be time-consuming to ask store staff for assistance. Furthermore, appropriate responses tailored to their emotional state are required, but not all store staff necessarily possess the ability or time to do so. In this situation, there is a need to improve the convenience and sense of security for elderly people living alone or older seniors when they are doing their daily shopping.

[1879] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for the user to input basic information, means for transmitting the basic information to the server, and means for storing the basic information in a database. This ensures that the user's basic information is securely stored on the server, providing a foundation for daily support. The system also includes means for voice input of daily conversations with the user, means for transmitting the voice data to the server, and means for converting the voice data into text. This allows the system to recognize and transcribe questions even when the user asks them by voice.

[1880] Furthermore, the system includes means for recognizing the user's emotions using an emotion engine, means for generating a response based on the text and adjusting it based on the user's emotions, means for outputting the response to the user as audio, and means for providing the optimal response to the user's questions in a physical store. This makes it possible to generate a response that matches the user's emotional state, enabling a more human-like interaction.

[1881] As a result, elderly people living alone and older adults can shop with peace of mind, reduce stress in stores, and be provided with an environment where they can easily find products and ask questions.

[1882] "Means for inputting basic user information" refers to devices or interfaces that allow users to input information about themselves (such as name, age, and address) into a terminal.

[1883] "Means of sending basic information to a server" refers to functions or protocols that transfer basic information entered by a user to a server via the internet or a local network.

[1884] A "server that stores basic information in a database" is a server equipped with a digital storage system that efficiently and securely stores the basic information of users that has been transmitted.

[1885] "Means of notifying that saving is complete" refers to functions or mechanisms that inform the user that basic information has been successfully saved to the database.

[1886] "Means for inputting everyday conversations by voice" refers to microphones or voice recognition devices that allow users to input everyday conversations as voice into a system.

[1887] "Means for sending audio data to a server" refers to the functions and protocols for transferring acquired audio data to a server via the internet or a local network.

[1888] A "server that converts audio data to text" is a server equipped with software and hardware to receive audio data and convert it into text format.

[1889] A "response-generating server" is a server equipped with algorithms and software that generate responses to users based on text data.

[1890] "Means of outputting a response to the user in audio" refers to speakers or headphones that play back the generated response as audio to the user using speech synthesis technology.

[1891] "Means for inputting the recipient and message content as voice" refers to a microphone or voice recognition device used by a user to input the content of a message they intend to send to a specific recipient as voice.

[1892] A "server that sends audio data to a server and sends a message to a recipient" is a server equipped with the function of converting audio data into text and sending a message to a designated recipient.

[1893] "A means of inputting desired date and time for medical appointments by voice" refers to a microphone or voice recognition device that allows the user to input their desired date and time for medical appointments into the system by voice.

[1894] "Means for sending desired date and time to the server" refers to a function or protocol that transmits the user's voice-inputted desired date and time for a medical consultation to a server via the internet or a local network.

[1895] A "server that makes medical appointments based on requested dates and times" is a server equipped with the function to confirm appointments in conjunction with a medical appointment system based on the received requested dates and times.

[1896] An "emotion engine" is an algorithm or software that estimates emotions from a user's words and actions and adjusts its response accordingly.

[1897] "Means of adjusting responses based on user emotions" refers to a function that appropriately modifies the tone and content of generated responses based on the user's emotional state recognized by the emotion engine.

[1898] "Means of providing optimal responses to user questions within a physical store" refers to interfaces and systems that provide appropriate information and instructions via voice in response to user questions and requests within a physical store.

[1899] This invention's system is designed to make it easier and safer for elderly people living alone or those in general to navigate situations such as shopping at physical stores. The system begins with the input of the user's basic information and then supports daily conversations, sending and receiving messages, scheduling online medical appointments, and fulfilling requests. In addition, it features an emotion engine that recognizes the user's emotions and adjusts its responses accordingly.

[1900] 1. Management of user basic information

[1901] The user first enters their basic information into their smartphone or smart glasses, and this information is sent to the server. The server stores the information in a database and sends a notification to the device when the saving is complete.

[1902] 2. Support for everyday conversation

[1903] Users can speak to a terminal in a physical store. For example, they might ask, "Where is this product?" The voice device records this, and the voice data is sent to a server. The server converts the voice data into text and uses an emotion engine to recognize emotions. Based on this information, the most appropriate response is generated and sent to the terminal. The terminal converts this response back into voice and provides the user with a response such as, "That product is in the second column from the right."

[1904] 3. Sending and receiving messages

[1905] For example, if a user says, "Send my son a message saying 'I'm doing well'," the voice is recorded and sent to the server. The server converts the voice data into text and uses an emotion engine to recognize the emotion. It then generates a message in an appropriate tone and sends it to the recipient. For example, it might provide a tailored message like, "I'm doing well. I'm a little lonely, though."

[1906] 4. Booking an online medical consultation

[1907] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the voice data is sent to the server. The server analyzes this data and enters the desired date and time into the online booking system. Once the booking is complete, the information is sent to the terminal, and the terminal notifies the user that "Your booking is complete."

[1908] 5. Implementation of requested items

[1909] User requests are entered via voice input and sent to the server. For example, if a user says, "Add milk to my shopping list," that voice data is sent to the server and the corresponding action is performed. An emotion engine is also included, and if the user is feeling stressed, a response in a more appropriate tone will be provided, such as, "Milk has been added to the list. Is there anything else I can help you with?"

[1910] Hardware / software to use

[1911] The primary hardware used will be smartphones and smart glasses, utilizing their built-in microphones and speakers. The following technologies and tools will be used in the software:

[1912] Speech recognition: Use the speech_recognition library to convert speech data into text.

[1913] Emotion Recognition: Implement an emotion recognition model using the transformers library.

[1914] Speech synthesis: Generate speech using gTTS or other speech synthesis libraries, and play it back using the playsound library.

[1915] Examples of specific cases and prompt statements

[1916] Specific example

[1917] Assume the user will be using smart glasses in the store:

[1918] User: "Where can I find this product?"

[1919] System: "That item is in the second row from the right. Please let us know if you need further information."

[1920] Example of a prompt

[1921] If the user wants to have a casual conversation:

[1922] "It's a beautiful day today."

[1923] System response: "It truly is a beautiful day!"

[1924] This system will improve the in-store shopping experience and allow elderly people living alone and seniors to shop with peace of mind.

[1925] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1926] Step 1:

[1927] The user enters basic information. The user uses the input interface of their smartphone or smart glasses to enter basic information such as their name, age, and address. This entered information is stored in a database, creating a baseline of user information.

[1928] Step 2:

[1929] The terminal sends the entered basic information to the server. The server receives this information and processes it for storage in the database. As output, a save completion notification is generated and sent to the terminal.

[1930] Step 3:

[1931] The system uses voice input for everyday conversations. When a user speaks into the device, the voice device (microphone) records it. The input voice data is sent to a server, which stores it in a database and processes it for impact analysis and recognition.

[1932] Step 4:

[1933] The device sends the recorded audio data to the server. This data is stored on the server and used in the next step.

[1934] Step 5:

[1935] The process converts audio data to text. The server uses speech recognition software (e.g., the speech_recognition library) to convert the transmitted audio data into text. The input is audio data, and the output is text data. During this conversion process, data calculations are performed to analyze the characteristics of the audio data and convert it into text format.

[1936] Step 6:

[1937] The system recognizes emotions and generates responses. The server uses a generative AI model to recognize user emotions from text data (e.g., the emotion recognition model in the transformers library). Input is text data, and output is response text with emotion labels. Based on this, the system generates responses in an appropriate emotional tone.

[1938] Step 7:

[1939] The response is output as audio. After the server recognizes the emotion, it uses speech synthesis software (e.g., gTTS library) to convert the response text into audio data. The terminal receives this audio data and plays it back to the user through speakers or headphones. The input is the response text, and the output is the audio data.

[1940] Step 8:

[1941] Sending a message. When the user issues a message sending command, the device sends it as audio data to the server. The server converts the audio data to text, performs emotion recognition, generates a message text in an appropriate tone, and sends it to the designated recipient. The input is the audio data, and the output is the message to be sent.

[1942] Step 9:

[1943] The process involves making a medical appointment. The user inputs their desired appointment date and time via voice, and the terminal sends this to the server. The server converts the desired date and time into text, accesses the online booking system, and confirms the appointment. The input is the voice data of the desired date and time, and the output is a booking confirmation notification.

[1944] Step 10:

[1945] The system executes the requested actions. The user inputs their specific request via voice, and the terminal sends this to the server. The server analyzes the request, recognizes the emotion, generates a response in an appropriate tone, and performs the corresponding action. The input is the voice data of the request, and the output is the appropriate response and the execution result.

[1946] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1947] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1948] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1949] [Fourth Embodiment]

[1950] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1951] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1952] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1953] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1954] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1955] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1956] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1957] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1958] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1959] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1960] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1961] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1962] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1963] The communication tool for elderly people living alone according to the present invention is designed to support daily life and improve the convenience of communication. This system includes a user-friendly terminal and a server that processes data and provides necessary services.

[1964] Specifically, the user enters basic information through their device and sends that information to the server. The server then stores the information in a database and notifies the device that registration is complete. Through this process, the system can manage the user's basic information.

[1965] Furthermore, users can enjoy everyday conversations by speaking directly to the device. For example, if a user says, "The weather's nice today," the device records the voice and sends it to the server. The server converts the voice data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[1966] Furthermore, sending and receiving messages is easy. When a user says to the device, "Send a message to my son saying 'I'm doing well'," the device records the voice and sends it to the server. The server retrieves the recipient's (in this case, the son's) contact information from its database and sends the specified message to that contact. Once the server has finished processing, it notifies the device, and the device then informs the user by voice, "The message has been sent."

[1967] Making appointments is also easy for users. When a user says, "Please schedule an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data and makes an online appointment for the specified date and time. Once the appointment is complete, the server sends the information back to the device, which then notifies the user that "the appointment is complete."

[1968] Furthermore, it can also respond to user requests. For example, if a user says, "Add milk to my shopping list," the device sends voice data to the server, and the server adds the appropriate item to the shopping list. Once the server has finished processing, it notifies the device of the result, and the device informs the user of the result.

[1969] Through these processes, the system of the present invention provides elderly people living alone with the communication and support they need in their daily lives, significantly improving their convenience. Furthermore, since all operations are voice-based, it is designed to be easy to use even for users unfamiliar with digital technology.

[1970] The following describes the processing flow.

[1971] User Registration

[1972] Step 1:

[1973] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[1974] Step 2:

[1975] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[1976] Step 3:

[1977] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[1978] Step 4:

[1979] The server notifies the terminal that saving the user information is complete.

[1980] Step 5:

[1981] The device receives a notification and displays "Registration Complete".

[1982] Support for everyday conversation

[1983] Step 1:

[1984] The user speaks to the device. For example, they might say, "The weather's nice today."

[1985] Step 2:

[1986] The device records the user's voice and sends the audio data to the server.

[1987] Step 3:

[1988] The server uses a speech recognition engine to convert the audio data into text.

[1989] Step 4:

[1990] The server generates an appropriate response based on the converted text. Natural language processing is used in this process.

[1991] Step 5:

[1992] The server sends a response text to the terminal.

[1993] Step 6:

[1994] The device converts the received response into speech and plays it back to the user. For example, it might say, "What beautiful weather we have today."

[1995] Send message

[1996] Step 1:

[1997] The user speaks to the device, specifying the recipient and content of the message. For example, they might say, "Send my son a message saying 'I'm doing well.'"

[1998] Step 2:

[1999] The device sends audio data to the server. The data is encoded in the appropriate format.

[2000] Step 3:

[2001] The server converts the audio data into text and analyzes the message content for the recipient.

[2002] Step 4:

[2003] The server retrieves the recipient's contact information from the database and sends the specified message content to that contact.

[2004] Step 5:

[2005] The server confirms that the message has been successfully sent and notifies the terminal of the result.

[2006] Step 6:

[2007] The device informs the user of notifications it has received. For example, it might say, "A message has been sent."

[2008] Online medical consultation booking

[2009] Step 1:

[2010] The user speaks into the device, specifying the date and time they wish to make an appointment. For example, they might say, "Please schedule an appointment for me this Friday at 3 PM."

[2011] Step 2:

[2012] The device sends voice data to the server.

[2013] Step 3:

[2014] The server analyzes the voice data and converts the desired reservation date and time into text.

[2015] Step 4:

[2016] The server accesses the online medical appointment system and attempts to make a reservation for that date and time.

[2017] Step 5:

[2018] The server receives the reservation confirmation information and sends the result to the terminal.

[2019] Step 6:

[2020] The device plays a voice message to the user confirming that the reservation has been received. For example, it might say, "Your reservation is complete."

[2021] Implementing requested actions (such as updating the shopping list)

[2022] Step 1:

[2023] The user speaks a request to the device. For example, they might say, "Add milk to my shopping list."

[2024] Step 2:

[2025] The device sends voice data to the server.

[2026] Step 3:

[2027] The server analyzes the audio data and converts the request into text.

[2028] Step 4:

[2029] The server performs the necessary operations based on the request. For example, it adds "milk" to the shopping list in the database.

[2030] Step 5:

[2031] The server confirms that the request has been completed and notifies the terminal.

[2032] Step 6:

[2033] The device will verbally inform the user of notifications it has received. For example, it might say, "Milk has been added to your shopping list."

[2034] (Example 1)

[2035] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2036] The objective of this invention is to address the lack of communication and support that elderly people living alone face in their daily lives. In particular, it aims to provide a voice-based interface that can be easily used even by elderly people who are unfamiliar with digital technology, enabling them to enjoy everyday conversations, obtain necessary information, send and receive messages, and easily make medical appointments.

[2037] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[2038] In this invention, the server includes means for the user to input general information, means for transmitting the general information to an information processing device, and an information processing device for storing the general information in a storage device. This makes it possible for the user to easily register basic information, enjoy everyday conversations, send messages, make medical appointments, and more, all through voice.

[2039] A "user" is an individual who utilizes the system of the present invention.

[2040] "General information" refers to basic information such as name, address, and contact information entered by the user.

[2041] An "information processing device" is a device that processes information transmitted by a user and provides necessary services.

[2042] A "storage device" is a device used by an information processing device to permanently store information.

[2043] "Voice input" is a method of inputting information or instructions using voice.

[2044] "Speech" refers to sound information generated by the user's utterances.

[2045] A "document" is information in the form obtained when an information processing device converts speech into text.

[2046] A "response" is the answer generated by an information processing device in response to a user's inquiry.

[2047] A "recipient" is an individual or organization that receives a message from a user.

[2048] "Message content" refers to the information or notification that the user wants to send.

[2049] "Appointment scheduling" is the process of making a medical appointment for a date and time specified by the user.

[2050] "Desired date and time" refers to the specific date and time the user wishes to schedule an appointment.

[2051] A "request" is a specific demand or instruction that a user wants to perform through the system.

[2052] "Natural language processing" is a technology that enables information processing devices to understand human language and generate appropriate responses.

[2053] "Methods of confirming online" refer to methods of confirming procedures such as reservations via the internet.

[2054] "The relevant operation" refers to the specific actions or procedures that the system performs in response to the user's request.

[2055] This invention relates to a communication tool for elderly people living alone, and is a system that uses a voice-based interface to allow users to easily input basic information, enjoy daily conversations, send messages, make medical appointments, and perform other requests.

[2056] How users enter information

[2057] The user inputs general information by voice through the microphone into the terminal. The terminal has a recording function, and the input voice is saved as digital audio data. The terminal sends this digital audio data to an information processing device (server). Specifically, the data is transmitted securely using the HTTPS protocol.

[2058] The server receives the transmitted information and saves it to a storage device (e.g., a MySQL database). It also notifies the terminal when saving is complete, and the terminal notifies the user that "registration is complete." A speech synthesis engine (e.g., Amazon Polly) is used for this notification.

[2059] Processing everyday conversations

[2060] When a user speaks to the device, for example, saying "The weather's nice today," the device records the audio and sends it to the server. The server uses Google Cloud Speech-to-Text to convert the audio data into text. Then, it uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response and sends it to the device. The device plays the response using a speech synthesis engine, allowing the user to enjoy a natural conversation.

[2061] Send message

[2062] When the user says, "Send my son a message saying 'I'm doing well'," the device records the voice and sends it to the server. The server converts the voice data to text and retrieves the recipient's (the son's) contact information from its database. The server uses Twilio's SMS API to send the message to the recipient. Once the server notifies the user that processing is complete, the device informs the user that "the message has been sent."

[2063] Appointment

[2064] When a user says, "Please schedule an appointment for me this Friday at 3 PM," the device sends the voice data to the server. The server converts the voice data into text and extracts the specified date and time information. The server uses the online medical consultation system's API to schedule the appointment. Once the appointment is complete, the server notifies the device that "the appointment is complete."

[2065] Responding to requests

[2066] When a user says, "Add milk to my shopping list," the device sends the voice data to the server. The server converts the voice data to text and adds the milk to the shopping list. Once the process is complete, the server notifies the device that "Milk has been added to your shopping list."

[2067] Examples of specific cases and prompt statements

[2068] For example, if a user says, "The weather's nice today," the following specific actions will be taken:

[2069] Voice input: "The weather is nice today."

[2070] Server processing: Converts audio data to text and generates responses using a generative AI model.

[2071] Response: "It really is a beautiful day today."

[2072] Example of a prompt:

[2073] "Please tell me what kind of response you should generate if a user says, 'The weather is nice today.'"

[2074] Hardware or software to use

[2075] Voice recording device: Terminal with microphone

[2076] Data transmission: HTTPS protocol

[2077] Speech recognition software: Google Cloud Speech-to-Text

[2078] Response generation model: OpenAI's GPT-3

[2079] Speech synthesis engine: Amazon Polly

[2080] Message sending API: Twilio

[2081] Data storage: MySQL database

[2082] The above describes the specific embodiments of the present invention. This system solves the lack of communication and support faced by elderly people living alone and supports their daily lives.

[2083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[2084] Basic Information Registration

[2085] Step 1:

[2086] The user inputs general information (name, address, contact information, etc.) by voice into the device. The device records this voice data as input.

[2087] Step 2:

[2088] The device converts recorded audio data into text data. This process uses speech recognition software (e.g., Google Cloud Speech-to-Text). Audio data is taken as input, and text data is obtained as output.

[2089] Step 3:

[2090] The terminal sends the converted text data to the server using the HTTPS protocol. Text data is used as input and sent to the server as output.

[2091] Step 4:

[2092] The server saves the received text data to a database. A storage device (e.g., a MySQL database) is used to permanently store the input text data. The output is a status indicating that saving is complete.

[2093] Step 5:

[2094] The server sends a notification to the device when saving is complete. The save completion status is used as input, and the notification is sent to the device as output.

[2095] Step 6:

[2096] The device converts the save completion notification into speech using a speech synthesis engine (e.g., Amazon Polly) and notifies the user that "Registration complete." The notification status is used as input, and the output is a voice message.

[2097] Processing everyday conversations

[2098] Step 1:

[2099] The user speaks into the device saying, "The weather's nice today." The device receives the user's voice data as input. The device then records this voice data.

[2100] Step 2:

[2101] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[2102] Step 3:

[2103] The server uses speech recognition software (e.g., Google Cloud Speech-to-Text) to convert the audio data into text data. Audio data is used as input, and text data is obtained as output.

[2104] Step 4:

[2105] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response based on the text. Text data is used as input, and the response text is obtained as output.

[2106] Step 5:

[2107] The server sends the generated response text to the terminal. The response text is used as input and sent to the terminal as output.

[2108] Step 6:

[2109] The device converts the response text into speech using a text-to-speech engine (e.g., Amazon Polly) and plays it back to the user. The response text is used as input, and the output is a voiced message.

[2110] Send message

[2111] Step 1:

[2112] The user says, "Send a message to my son saying, 'I'm doing well.'" The input is the user's voice data. The device records this voice data.

[2113] Step 2:

[2114] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[2115] Step 3:

[2116] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[2117] Step 4:

[2118] The server retrieves the recipient's contact information from the database. Text data is used as input, and the recipient's contact information is obtained as output.

[2119] Step 5:

[2120] The server uses Twilio's SMS API to send a message to the recipient. The recipient's contact information and message content are used as input, and the message sending status is obtained as output.

[2121] Step 6:

[2122] The server sends a status message to the terminal indicating that the transmission is complete. The message transmission status is used as input and sent to the terminal as output.

[2123] Step 7:

[2124] The device converts the message transmission completion notification into speech using a speech synthesis engine and notifies the user that "the message has been sent." The notification status is used as input, and the voice message is obtained as output.

[2125] Appointment booking

[2126] Step 1:

[2127] The user says, "Please schedule an appointment for me this Friday at 3 PM." The user's voice data is obtained as input. The device records this voice data.

[2128] Step 2:

[2129] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[2130] Step 3:

[2131] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[2132] Step 4:

[2133] The server analyzes the text data and extracts information about the desired date and time. Text data is used as input, and the output is information about the desired date and time.

[2134] Step 5:

[2135] The server uses the online medical consultation system's API to make reservations. The desired date and time are used as input, and the output is a reservation completion status.

[2136] Step 6:

[2137] The server sends a reservation completion notification to the device. The reservation completion status is used as input and sent to the device as output.

[2138] Step 7:

[2139] The device converts the reservation completion notification into speech using a speech synthesis engine and notifies the user that "Your reservation is complete." The notification status is used as input, and the voice message is obtained as output.

[2140] Responding to requests

[2141] Step 1:

[2142] The user says, "Add milk to my shopping list." The device receives the user's voice data as input. The device then records this voice data.

[2143] Step 2:

[2144] The device sends the recorded audio data to the server. The audio data is used as input and sent to the server as output.

[2145] Step 3:

[2146] The server uses speech recognition software to convert audio data into text data. Audio data is used as input, and text data is obtained as output.

[2147] Step 4:

[2148] The server analyzes the text data and identifies the relevant requests. Text data is used as input, and the requests are obtained as output.

[2149] Step 5:

[2150] The server performs the appropriate operation based on the requested information. The requested information is used as input, and the operation execution status is obtained as output.

[2151] Step 6:

[2152] The server sends a notification to the terminal that the operation is complete. The operation execution status is used as input and sent to the terminal as output.

[2153] Step 7:

[2154] The device converts the completion notification into speech using a speech synthesis engine and notifies the user that "Milk has been added to your shopping list." The notification status is used as input, and the voice message is obtained as output.

[2155] The above is the specific processing flow of the system.

[2156] (Application Example 1)

[2157] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2158] Elderly people living alone face the problem of being unable to rely on others for daily life support, such as ordering food delivery or other assistance, because traditional methods are cumbersome to use. Furthermore, users unfamiliar with digital technology require natural communication using voice recognition. Therefore, there is a need for a system that allows elderly people living alone to easily and naturally order food delivery and other daily life support.

[2159] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[2160] In this invention, the server is

[2161] A means for users to input basic information,

[2162] Means for transmitting the aforementioned basic information to a server,

[2163] A server that stores the aforementioned basic information in a database,

[2164] A means for notifying that the aforementioned storage has been completed,

[2165] A means of inputting everyday conversations with the user via voice,

[2166] Means for transmitting the aforementioned audio data to a server,

[2167] A server that converts the aforementioned audio data into text,

[2168] A server that generates a response based on the aforementioned text,

[2169] A means for outputting the aforementioned response to the user in voice,

[2170] A means of inputting the recipient and message content by the user,

[2171] A server that sends the aforementioned audio data to a server and sends a message to a recipient,

[2172] A method for users to input their preferred date and time for medical appointments via voice input,

[2173] A means for sending the aforementioned desired date and time to the server,

[2174] A server that makes medical appointments based on the aforementioned preferred date and time,

[2175] A means of inputting user requests by voice,

[2176] A server that transmits the aforementioned audio data to a server and performs the corresponding operation,

[2177] A means for users to input food delivery orders by voice,

[2178] Means for sending the food delivery order to the server,

[2179] A server that processes orders based on the aforementioned food delivery orders,

[2180] A means for notifying the completion of the aforementioned order processing,

[2181] Includes.

[2182] This will make it easy and natural for elderly people living alone to order food delivery and other forms of assistance for daily life.

[2183] "User" refers to any person who uses this system.

[2184] "Basic information" refers to necessary information about the user, such as their name, address, and contact information.

[2185] A "server" refers to a computer system that processes and stores data, and notifies users.

[2186] A "database" refers to an information system used to systematically store basic information and other data.

[2187] "Means of notification" refers to methods of informing users about the completion of information storage or processing.

[2188] "Voice input" refers to a method of inputting information by voice through a microphone or similar device.

[2189] "Audio data" refers to digital audio information obtained from voice input.

[2190] "Converting to text" refers to the process of converting audio data into written text.

[2191] "Response" refers to the reply or instruction generated by the server based on the user's voice input.

[2192] A "message" refers to information with specified content that a user sends to other recipients.

[2193] "Medical appointment booking" refers to making a reservation for a specific date and time to receive medical treatment at a medical institution or similar facility.

[2194] "Requests" refer to the operations or requests that users want to perform on the system.

[2195] "Food delivery" refers to the process of a user ordering food to be delivered to a specified location.

[2196] "Order processing" refers to the procedure for confirming the details of a food delivery order and handling it appropriately.

[2197] "Means of notifying completion" refers to methods of informing users that an order or operation has been completed.

[2198] This invention is a system that provides the communication and support that elderly people living alone need in their daily lives. This system includes a terminal that is easy for the user to operate and a server that processes data and provides the necessary services.

[2199] The system is configured as follows:

[2200] The user first enters basic information (name, address, contact information, etc.) into the terminal. This basic information is sent from the terminal to the server, which stores it in a database. Once the saving is complete, the server notifies the terminal, and the terminal informs the user that the saving is complete. In this way, the system can manage the user's basic information.

[2201] Regarding support for everyday conversation, when a user says, "The weather's nice today," the device records the audio and sends it to the server. The server converts the audio data into text, generates an appropriate response, and sends it back to the device. The device then plays the response back aloud, allowing the user to enjoy a natural conversation.

[2202] It is also possible to send and receive messages. For example, if a user says, "Send a message to my son saying, 'I'm doing well'," the device will record the voice and send it to the server. The server will retrieve the recipient's contact information from its database and send the specified message. Once the server has finished processing, it will notify the device, and the device will inform the user by voice, "The message has been sent."

[2203] Similarly, when scheduling an appointment, if the user says, "Please schedule an appointment for me this Friday at 3pm," the device sends that voice message to the server. The server analyzes the voice data and schedules an appointment for the specified date and time. Once the appointment is complete, the server notifies the device, and the device informs the user that "the appointment is complete."

[2204] This system also supports ordering food delivery. When a user says, "Order curry for lunch today," the terminal records the voice and sends it to the server. The server converts the voice data into text, selects an appropriate food delivery service, and places the order. Once the order is complete, the server notifies the terminal of the result, and the terminal informs the user via voice, "Your order is complete."

[2205] The specific hardware used to realize this invention can be a smartphone or tablet. For the software, the "speech_recognition" library is used for speech recognition, the "pyttsx3" library for speech synthesis, and the "requests" library for processing API requests. Furthermore, a generative AI model is used for natural language processing (NLP).

[2206] Examples of specific prompt messages include the following:

[2207] "I'd like to order ~ for lunch today."

[2208] "Please schedule an appointment for this Friday at 3 PM."

[2209] This system offers significant convenience by allowing elderly people living alone to easily perform various daily tasks using voice commands. This reduces barriers to digital technology and improves the quality of daily life.

[2210] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[2211] Step 1:

[2212] The user enters basic information. The user enters their name, address, contact information, etc., on the screen of their smartphone or tablet. This constitutes input.

[2213] Step 2:

[2214] The terminal sends the entered basic information to the server. The basic information is converted into a digital format and sent as a request to the server. The server then receives the data.

[2215] Step 3:

[2216] The server saves the received basic information to the database. Data processing includes correctly formatting the input information and writing it to the database. The server then verifies that saving is complete.

[2217] Step 4:

[2218] The server notifies the user when the basic information has been saved. It sends the status of save completion to the terminal. The terminal informs the user of this information via voice or screen display.

[2219] Step 5:

[2220] The user inputs everyday conversations via voice. For example, they might say, "The weather's nice today." This becomes the input.

[2221] Step 6:

[2222] The terminal records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data. A process of converting it to a data format takes place here.

[2223] Step 7:

[2224] The server converts the received audio data into text. A generative AI model is used for the conversion from speech to text. The converted text is obtained as output.

[2225] Step 8:

[2226] The server generates a response based on the converted text. It uses natural language processing (NLP) to analyze the text and create an appropriate response text, which is then output as the response.

[2227] Step 9:

[2228] The server sends the generated response to the terminal. The text data of the response is sent to the terminal, and preparations for speech synthesis are made.

[2229] Step 10:

[2230] The terminal plays back the response received from the server as audio. It uses a speech synthesis tool (pyttsx3) to convert text data into speech and plays it back to the user. This is the output to the user.

[2231] Step 11:

[2232] The user inputs the recipient and message content by voice. For example, they might say, "Send a message to my son saying, 'I'm doing well.'" This is the input.

[2233] Step 12:

[2234] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[2235] Step 13:

[2236] The server sends a message based on the recipient and message content. It retrieves the recipient's contact information from the database and sends the specified message.

[2237] Step 14:

[2238] The server notifies the terminal that the message has been sent. It sends a status message to the terminal indicating that the message has been sent, and the terminal then verbally informs the user that "the message has been sent."

[2239] Step 15:

[2240] Users can place food delivery orders by voice. For example, they might say, "Order curry for lunch today." This constitutes the input.

[2241] Step 16:

[2242] The device records audio data and sends it to the server. The recorded data is digitized and sent to the server as audio data.

[2243] Step 17:

[2244] The server converts voice data into text and processes the necessary food delivery orders. It uses a generative AI model to convert voice to text and then places orders with food delivery services based on that information.

[2245] Step 18:

[2246] The server notifies the terminal that the order processing is complete. It confirms that the order is complete and sends the status to the terminal. The terminal then informs the user by voice, "Your order is complete."

[2247] Through these steps, the system will be able to naturally use voice commands to order food deliveries and provide daily support for elderly people living alone.

[2248] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[2249] The communication tool for elderly people living alone according to the present invention is designed to support the user's daily life and improve the convenience of communication. This system includes means for inputting and saving the user's basic information, support for daily conversation, sending and receiving messages, scheduling online medical appointments, and fulfilling requests, as well as a function that uses an emotion engine to recognize the user's emotions and generate responses.

[2250] First, the user enters their basic information into the device and sends that information to the server. The server saves the received information to a database and notifies the device when the saving is complete. The device displays "Registration Complete," and the management of the basic information is finished.

[2251] Next, the user speaks to the device to enjoy a casual conversation. For example, if the user says, "The weather's nice today," the device records the voice and sends the voice data to the server. The server converts the voice data into text and recognizes the emotion using an emotion engine. Based on the recognized emotion, it generates the most appropriate response and sends the response text to the device. The device then converts the response back into voice and plays it back to the user, saying, "The weather's really nice today." In this way, the user can receive emotional support through a natural conversation.

[2252] Furthermore, sending and receiving messages is easy. When a user speaks to the device, saying, "Send a message to my son saying, 'I'm fine,'" the device records the voice and sends it to the server. The server converts the recipient and message content into text, recognizes the user's emotions using an emotion engine, and generates a message in an appropriate tone. For example, if the user sends a message in a sad tone, the emotion engine recognizes that emotion and generates a message in a tone such as, "I'm fine. I'm a little lonely, though." The server sends the message to the recipient and notifies the device when it is complete. The device then informs the user by voice, "Message sent."

[2253] Furthermore, booking an online consultation is also easy. When a user says, "Please book an appointment for me this Friday at 3pm," the device sends the voice data to the server. The server analyzes the voice data, transcribes the desired date and time into text, and attempts to book the appointment at that time by accessing the online consultation booking system. Once the booking is complete, the server sends that information to the device, which then notifies the user that "the booking is complete."

[2254] Regarding requests, when a user says, "Add milk to my shopping list," the device sends voice data to the server. The server analyzes the request, recognizes the user's emotions using an emotion engine, and performs the appropriate action. If the user is stressed, the system responds in a tone such as, "Milk has been added to the list. Is there anything else I can help with?" The server notifies the device that the action is complete, and the device then informs the user, "Milk has been added to your shopping list."

[2255] Through these processes, the system of the present invention recognizes the emotions of elderly people living alone and generates responses based on them, thereby providing more humane support. Since all operations are voice-based, it can be easily used even by users unfamiliar with digital technology. This system is expected to support the user's daily life and reduce feelings of loneliness.

[2256] The following describes the processing flow.

[2257] User Registration

[2258] Step 1:

[2259] The user opens the registration screen on their device and enters their basic information (name, contact information, emergency contact information, etc.).

[2260] Step 2:

[2261] The terminal sends the entered information to the server. The information is encoded in an appropriate format, such as JSON.

[2262] Step 3:

[2263] The server saves the user information it receives to the database. The information is correctly stored in the corresponding field within the database.

[2264] Step 4:

[2265] The server notifies the terminal that saving the user information is complete.

[2266] Step 5:

[2267] The device r...

Claims

1. A means for users to input basic information, Means for transmitting the aforementioned basic information to a server, A server that stores the aforementioned basic information in a database, A means for notifying that the aforementioned storage has been completed, A means of inputting everyday conversations with the user via voice, Means for transmitting the aforementioned audio data to a server, A server that converts the aforementioned audio data into text, A server that generates a response based on the aforementioned text, A means for outputting the aforementioned response to the user in voice, A means of inputting the recipient and message content by the user, A server that sends the aforementioned audio data to a server and sends a message to a recipient, A method for users to input their preferred date and time for medical appointments via voice input, A means for sending the aforementioned desired date and time to the server, A server that makes medical appointments based on the aforementioned preferred date and time, A means of inputting user requests by voice, A server that transmits the aforementioned audio data to a server and performs the corresponding operation, A system that includes this.

2. The system according to claim 1, wherein the response generation means utilizes natural language processing.

3. The system according to claim 1, wherein the voice input means for the desired date and time of the medical appointment includes means for confirming the appointment online.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A