System

The system addresses the lack of personalized conversations in smart speakers by generating and updating customized conversation scripts for elderly users, reducing loneliness and dementia progression through AI-enabled IoT devices.

JP2026019819APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121567
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current smart speakers and chatbots fail to provide personalized conversations tailored to the individual backgrounds and interests of elderly individuals, leading to feelings of loneliness and potential health issues such as dementia progression.

Method used

A system that inputs a target person's profile information, uses a server to generate a customized conversation script based on this data, and continuously updates the script with new information to provide personalized and engaging conversations.

Benefits of technology

The system reduces feelings of loneliness and slows the progression of dementia by offering enjoyable and relevant conversations using AI-enabled IoT devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019819000001_ABST
    Figure 2026019819000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting profile information of a target person; means for transmitting the profile information of the target person to a server; means for collecting relevant information on the Internet based on the transmitted profile information and generating a customized conversation script; means for transmitting the generated customized conversation script to a terminal; and means for recognizing a user's voice and responding based on the customized conversation script.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] As society as a whole continues to experience a declining birthrate and aging population, elderly people are spending more time alone at home, leading to feelings of loneliness and various mental and physical health problems. Daily conversation and communication are important, particularly for slowing the progression of dementia. However, current smart speakers and chatbots have the problem of being unable to provide personalized conversations tailored to the individual backgrounds and interests of the elderly. This creates a demand for AI-enabled devices that can act as "companions" that provide satisfaction to the elderly. [Means for solving the problem]

[0005] To solve this problem, we invented a system that provides the following means. First, we provide a means for inputting the target person's profile information, which allows us to acquire information such as the elderly person's name, age, hobbies, and health status. Next, we provide a means for transmitting the acquired profile information to a server, which then uses that information to collect related information on the Internet and generate a customized conversation script for each user. This generated conversation script is then sent to the device, which recognizes the user's voice and responds based on the customized conversation script. Furthermore, the server includes a means for continuously collecting and analyzing the user's conversation data and periodically updating the customized conversation script, enabling continuous learning and the provision of the latest information. In this way, elderly people can enjoy enjoyable conversations at any time, which is expected to reduce feelings of loneliness and slow the progression of dementia.

[0006] "Target users" refers to users of this system, primarily elderly people.

[0007] "Profile information" is information about a subject, including name, age, hobbies, health status, etc.

[0008] "Means" refers to a device or method for achieving a particular function.

[0009] "Terminal" refers to a hardware device used by a subject, and includes a physical or doll-shaped device.

[0010] "Server" refers to the central system that receives and analyzes profile information and generates customized conversation scripts.

[0011] "Internet" means the global network for collecting and transmitting information.

[0012] A "conversation script" refers to a series of dialogue data generated by the server to manage a conversation with a target person.

[0013] "Speech recognition" refers to the technology of analyzing a person's speech as digital data and understanding its content.

[0014] "Customization" refers to providing content that is individually optimized based on the target person's profile information.

[0015] "Continuous learning" refers to the process by which a system acquires new knowledge and skills and improves its performance based on data collected during use.

[0016] "Updating" refers to the process of replacing existing conversation scripts and data with the latest information. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention provides a system that uses AI-equipped IoT devices to enable elderly people to enjoy everyday conversations without feeling lonely. Specifically, the following process is performed.

[0039] ---

[0040] Enter your profile information

[0041] Device:

[0042] Using the initial setup screen that appears when the device is started, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[0043] Examples:

[0044] The user's family member enters the following information into the terminal: "Taro, 75 years old, hobbies are gardening and Go, suffers from mild dementia."

[0045] ---

[0046] Sending profile information

[0047] Device:

[0048] The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[0049] Examples:

[0050] The device will notify you that "Taro's information has been sent to the server."

[0051] ---

[0052] Profile information analysis and customization

[0053] server:

[0054] The system analyzes the received profile information and gathers relevant information from the Internet (e.g., the latest gardening techniques or Go news). Based on this information, it generates a conversation script customized for each user.

[0055] Examples:

[0056] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro and generated a customized conversation script."

[0057] ---

[0058] Submitting a customization script

[0059] server:

[0060] The generated customized conversation script is sent to the terminal.

[0061] Device:

[0062] Receives customized conversation scripts and saves them in internal memory. Notifies the user or their family that "conversation data has been updated."

[0063] Examples:

[0064] The server sends "Gardening trivia" and the device notifies the user that "Taro's conversation data has been updated."

[0065] ---

[0066] Start a conversation

[0067] User:

[0068] The user speaks to the device (e.g., "How's the gardening going lately?").

[0069] Device:

[0070] The user's voice is analyzed by a voice recognition module, and based on the analysis results, an appropriate response is selected from a customized conversation script and responded to by voice.

[0071] Examples:

[0072] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[0073] ---

[0074] Continuous learning and updates

[0075] server:

[0076] It continuously collects, analyzes, and provides feedback on user conversation data, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[0077] Device:

[0078] Every time a new script is sent, it is received, stored in internal memory, and updated. It notifies you that "Conversation data has been updated based on new information."

[0079] Examples:

[0080] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0081] ---

[0082] In this way, AI-equipped IoT devices can collect the latest information from the internet based on the target person's profile information and provide optimized conversations, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely, which is expected to reduce loneliness and slow the progression of dementia.

[0083] The processing flow will be explained below.

[0084] Step 1:

[0085] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is stored in the device's temporary memory.

[0086] Step 2:

[0087] Device: The entered profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[0088] Step 3:

[0089] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[0090] Step 4:

[0091] Server: Based on the extracted information, the server collects the latest information related to the target person from the Internet. For example, if the target person's hobby is gardening, it collects information on the latest gardening techniques and plants.

[0092] Step 5:

[0093] Server: Generates a customized conversation script based on the collected information, including content that is interesting and relevant to the target audience.

[0094] Step 6:

[0095] Server: Sends the generated customized conversation script to the device.

[0096] Step 7:

[0097] Device: Saves the received conversation script in its internal memory. Notifies the user or their family that "the conversation data has been updated."

[0098] Step 8:

[0099] User: The user speaks to the device (e.g., "How's your gardening going lately?"). The device analyzes the user's voice using a speech recognition module.

[0100] Step 9:

[0101] Terminal: Based on the analysis results, the terminal selects an appropriate response from a customized conversation script and returns the response to the user through a voice output device.

[0102] Step 10:

[0103] Server: Continuously collects user conversation data, analyzes it, and provides feedback. This data serves as the basis for further customization based on the user's interests.

[0104] Step 11:

[0105] Server: Periodically collects new information and topics, updates the customized conversation script, and sends the updated script to the device.

[0106] Step 12:

[0107] Device: Receives each new script sent, saves it in internal memory, and updates it. Once the update is complete, it notifies you that "Conversation data has been updated based on new information."

[0108] Through each step, the AI-enabled IoT device provides personalized conversations, allowing the elderly to enjoy their daily lives without feeling lonely.

[0109] Example 1

[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0111] To enable elderly people to enjoy everyday conversations without feeling lonely, conversation content needs to be customized to each individual. However, conventional systems lack such customization, and often provide topics and information that do not interest elderly people. This leads to a decrease in satisfaction. In addition, there is no mechanism for continuously updating new information, which means that conversation content easily becomes outdated.

[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0113] In this invention, the server includes a means for transmitting the target person's attribute information to the network, a means for collecting related information in a data store based on the transmitted attribute information and generating customized conversation data, and a means for transmitting the generated customized conversation data to the client device. This enables conversations using the latest information specific to the target person, allowing elderly people to participate in conversations with interest without feeling lonely. Furthermore, by continuously collecting new information and updating the conversation content, it is possible to always provide fresh topics of conversation.

[0114] "Subject attribute information" refers to personal information such as the subject's name, age, hobbies, and health status.

[0115] "Data store" refers to a system or database for storing information and data collected from the Internet.

[0116] "Customized conversation data" refers to a conversation script that is generated based on the individual attribute information of the subject and tailored to the subject.

[0117] "Client device" refers to an AI-enabled IoT device or terminal used by the subject.

[0118] "Network" refers to the Internet and other communication means that provide paths for transmitting and receiving data.

[0119] A "conversation script" refers to data that describes the content of utterances and responses that the system uses in dialogue with the user.

[0120] This invention provides a system that uses AI-equipped IoT devices to help elderly people enjoy everyday conversations without feeling lonely. The entire system is realized by collecting profile information of the target person and providing customized conversation data to the user.

[0121] Hardware and software used

[0122] Terminal: An AI-enabled IoT device placed close to the user. This terminal displays the necessary initial setup screens and allows the user or their family members to enter their profile information.

[0123] Server: Refers to the back-end system that analyzes data and generates customized conversation scripts. The server is connected to the Internet and collects various information.

[0124] Speech Recognition Module: For example, using Google Speech-to-Text. Used to convert the user's speech into text.

[0125] Generative AI models, such as GPT-3, are used to generate customized conversation scripts based on collected information.

[0126] A speech synthesis module, e.g., using Google Text-to-Speech, is used to convert the generated text into speech and respond to the user.

[0127] Explanation of program processing

[0128] Device:

[0129] 1. When the device starts up, the initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored on the device.

[0130] 2. The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[0131] server:

[0132] 1. The server analyzes the received profile information and collects relevant information from a data store (e.g., the latest gardening techniques or Go news). Based on this information, a customized conversation script is generated for each user. This is generated using a generative AI model.

[0133] 2. Send the generated customized conversation script to the terminal.

[0134] Device:

[0135] 1. The device saves the received conversation script in its internal memory and displays "Conversation data updated" on the screen.

[0136] 2. When the user speaks to the device, the device analyzes it using the voice recognition module, selects a response based on the customized conversation data, and responds to the user using the voice synthesis module.

[0137] Specific examples

[0138] Enter your profile information:

[0139] The user's family member enters information into the device, such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." This information is temporarily stored in the device.

[0140] Send profile information:

[0141] When the device is connected to a network such as Wi-Fi, it sends the saved profile information to the server as a POST request. If the request is successful, a pop-up message will appear saying, "Taro's information has been sent to the server."

[0142] Profile information analysis and customization:

[0143] The server analyzes the profile information "Taro, 75 years old, gardening, Go, mild dementia," collects relevant latest gardening techniques and recent Go news, and generates a conversation script using a generative AI model.

[0144] Send a customized conversation script:

[0145] The server sends the conversation script data in JSON format, and the device receives this data, parses it, and saves it in its internal memory. The updated content is notified with a message saying "The conversation data has been updated."

[0146] Start a conversation:

[0147] When a user speaks to the device, "How's your gardening going these days?", the device converts the speech into text using a speech recognition module, and then uses a generative AI model to respond, "The latest gardening techniques place great importance on soil preparation," which is then played back as audio using a speech synthesis module.

[0148] Example prompt sentence:

[0149] "Please tell me the latest gardening knowledge for 75-year-old Taro."

[0150] "Tell me the latest news about Go"

[0151] "How's your gardening going lately?"

[0152] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0153] Step 1:

[0154] Enter and save your profile information

[0155] Device:

[0156] When the device starts up, an initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[0157] Input: The user enters profile information as text.

[0158] Output: The entered profile information is saved in the device's temporary memory.

[0159] Specific operation: A user's family member enters information such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." The device saves this information in text format.

[0160] Step 2:

[0161] Sending profile information

[0162] Device:

[0163] The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[0164] Input: Profile information stored on your device.

[0165] Output: The profile information is sent to the server.

[0166] Specific operation: The device sends profile information in the form of a POST request via a network such as Wi-Fi. A message pops up when the transmission is successful.

[0167] Step 3:

[0168] Profile information analysis and customization

[0169] server:

[0170] The server analyzes the received profile information and gathers relevant information from a data store, such as the latest gardening techniques or Go news, and then uses a generative AI model to generate a customized conversation script based on the collected information.

[0171] Input: Profile information stored on the server.

[0172] Output: A customized conversation script.

[0173] Specific operation: The server analyzes the information "Taro, 75 years old, hobbies are gardening and Go, mild dementia," collects related information using web scraping and APIs, and generates a conversation script using a generative AI model (e.g., GPT-3).

[0174] Step 4:

[0175] Sending customized conversation scripts

[0176] server:

[0177] The generated customized conversation script is sent to the terminal.

[0178] Device:

[0179] The terminal stores the received conversation script in its internal memory and displays the message "Conversation data has been updated" on the screen.

[0180] Input: A customized conversation script generated by the server.

[0181] Output: The customized conversation script is saved in the device's memory.

[0182] Specific operation: The server sends script data in JSON format, the device receives it, parses it, saves it in its internal memory, and displays a pop-up message saying "Conversation data has been updated."

[0183] Step 5:

[0184] Start a conversation

[0185] User:

[0186] The user provides the device with a question or topic (e.g., "How's your gardening going lately?").

[0187] Device:

[0188] The device uses a speech recognition module to convert the user's speech into text, generates appropriate responses based on customized conversation data, and responds to the user using a speech synthesis module.

[0189] Input: User's voice input.

[0190] Output: Audio response from the device.

[0191] What it does: When a user says, "How's your gardening going these days?", the device converts the speech to text and selects a response from a customized conversation script: "Modern gardening techniques emphasize soil conditioning," and plays it back as audio.

[0192] Step 6:

[0193] Continuous learning and updates

[0194] server:

[0195] The server continuously collects user conversation data, analyzes it, and provides feedback, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[0196] Device:

[0197] Every time a new conversation script is sent from the server, it is received, saved in internal memory, and updated. When the update is complete, a message will appear saying, "The conversation data has been updated based on the new information."

[0198] Input: User conversation data and new information.

[0199] Output: Updated customized conversation script.

[0200] Specific operation: The server analyzes the user's conversation history, collects new gardening information, the latest news on Go, etc., and sends an updated conversation script to the device. The device receives this and notifies the user that "the conversation data has been updated with the latest information."

[0201] (Application example 1)

[0202] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0203] There is a need for systems that can not only reduce loneliness among the elderly and provide daily communication, but also monitor their health status and respond to emergency situations. There is also a need for systems that can provide customized conversations using information specific to the elderly and the latest topics.

[0204] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0205] In this invention, the server includes means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for collecting sensor data for monitoring health status, and means for making an emergency call if an abnormality occurs. This not only reduces the sense of loneliness of elderly people and provides communication, but also enables health status monitoring and prompt response in emergencies.

[0206] "Subject profile information" is basic information about an individual, such as the subject's name, age, hobbies, and health status.

[0207] The "means for inputting profile information" refers to an interface that allows the subject or their family to input profile information using a terminal.

[0208] The "means for transmitting profile information to a server" refers to a communication means for transmitting profile information from a terminal to a server via the Internet.

[0209] A "customized conversation script" is an individually optimized conversation scenario that is generated by collecting relevant information on the Internet based on the target person's profile information.

[0210] The "means for generating a customized conversation script" is a processing means for analyzing profile information and creating a conversation script based on the latest related information on the Internet.

[0211] The "means for recognizing and responding to the user's voice" is a function for analyzing the voice uttered by the user and selecting and responding to an appropriate response based on a customized conversation script.

[0212] "Sensor data for monitoring health status" refers to information obtained from sensors that collect physiological and activity data of subjects, such as heart rate, number of steps, and fall detection.

[0213] The "means for collecting sensor data" refers to devices and software for collecting physiological data and activity data through sensors, which are necessary for monitoring the health status of a subject.

[0214] The "means for making an emergency report when an abnormality occurs" is a means for making a report to a pre-set emergency contact when an abnormality is detected based on collected sensor data.

[0215] "Means for gathering information from the Internet" refers to the Internet browsing ability to gather relevant and up-to-date information and news on the Internet based on the profile information.

[0216] This invention provides a system that can reduce loneliness among the elderly, monitor their health status, and respond to emergencies. The system provides customized conversations based on the target person's profile information and makes an emergency call when an abnormality in their health status is detected.

[0217] Program processing explanation

[0218] Hardware and Software Configuration

[0219] Hardware:

[0220] Robot device (equipped with microphone, speaker, and sensor)

[0221] Software and Libraries:

[0222] requests library: Used for server communication.

[0223] speech_recognition library: Used for speech recognition.

[0224] pyttsx3 library: Used for speech synthesis.

[0225] Processing explanation

[0226] (1) Enter and submit profile information

[0227] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device and then sent to the server. The server receives the information and sends a confirmation message back to the device.

[0228] Examples:

[0229] The user's family member enters information such as "Mr. Tanaka, 75 years old, hobbies are gardening and Go, mild dementia" into the terminal and sends it to the server.

[0230] (2) Analyzing profile information and generating conversation scripts

[0231] The server analyzes the received profile information and collects relevant information from the Internet (such as the latest gardening techniques or Go news), and generates a conversation script customized for each user based on this information.

[0232] Examples:

[0233] The server "collects the latest gardening trivia and Go game information for Tanaka and generates a customized conversation script."

[0234] (3) Sending customized conversation scripts

[0235] The server sends the generated customized conversation script to the device. The device receives the conversation script and saves it in its internal memory. Once the saving is complete, the device displays a notification to the user that the update is complete.

[0236] Examples:

[0237] The server sends "Gardening trivia" and the device notifies the user that "Tanaka's conversation data has been updated."

[0238] (4) Communication using voice recognition

[0239] When a user speaks to the terminal, the terminal analyzes the user's voice using a voice recognition module, selects an appropriate response from a customized conversation script, and responds verbally.

[0240] Examples:

[0241] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[0242] (5) Health monitoring and emergency notification

[0243] The device's built-in sensors constantly monitor the user's health. If an abnormality is detected (e.g., an abnormal heart rate or a fall), the device immediately sends the abnormal data to the server, which then notifies the user's emergency contacts.

[0244] Examples:

[0245] The device notifies the server that an abnormal heart rate has been detected, and the server then calls a pre-set emergency number.

[0246] (6) Continuous learning and updating

[0247] The server continuously collects and analyzes the user's conversation data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory. The user is notified that the conversation data has been updated with new information.

[0248] Examples:

[0249] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0250] Example prompts for generative AI models

[0251] Generate customized conversation scripts about the latest gardening news and Go topics based on the user's profile information. For example, respond to the question "Tell me about gardening" with "Soil preparation is becoming increasingly important as a modern gardening technique."

[0252] In this way, the embodiments of the invention allow elderly people to enjoy everyday conversations without feeling lonely, while their health condition is monitored, and if an abnormality occurs, a prompt response can be made.

[0253] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0254] Step 1:

[0255] Enter your profile information

[0256] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device. The input is done using the user interface, and the input data is stored in the device's internal storage.

[0257] Input: Profile information such as name, age, hobbies, health status, etc.

[0258] Output: Profile information temporarily saved on the device

[0259] Step 2:

[0260] Sending profile information

[0261] The device will then send the saved profile information to the server via the internet, and once the transfer is complete, a confirmation message will appear saying "Profile information has been sent to the server."

[0262] Input: Profile information temporarily saved on the device

[0263] Output: Profile information sent to server, confirmation message to user

[0264] Step 3:

[0265] Analyzing profile information

[0266] The server analyzes the received profile information and collects relevant information on the Internet (e.g., the latest gardening techniques or Go news). Based on the profile information, the analysis generates relevant search queries and retrieves relevant information.

[0267] Input: Profile information sent to the server

[0268] Output: Analysis results and related collected information

[0269] Step 4:

[0270] Generate customized conversation scripts

[0271] The server generates a customized conversation script for each user based on the collected relevant information. It uses a generative AI model to script a conversation scenario based on the profile information and collected information.

[0272] Input: Analysis results and related collected information

[0273] Output: Customized conversation script

[0274] Step 5:

[0275] Sending customized conversation scripts

[0276] The server sends the generated customized conversation script to the terminal. The terminal receives the conversation script and saves it in its internal memory. Once the saving is complete, a notification that "Conversation data has been updated" is displayed to the user.

[0277] Input: Customized conversation script

[0278] Output: Conversation script saved on the device, update completion notification to the user

[0279] Step 6:

[0280] User voice recognition

[0281] When a user speaks to the device, the device uses a voice recognition module to analyze the user's voice, converting the voice data into text data, and then selects an appropriate response from a customized conversation script based on the text.

[0282] Input: User voice input

[0283] Output: Recognized text data, selected response

[0284] Step 7:

[0285] Voice response

[0286] The device responds to the user vocally based on the selected customized conversation script, which is generated using a speech synthesis module and output audibly through the speaker.

[0287] Input: Selected Response

[0288] Output: Audio response

[0289] Step 8:

[0290] Health monitoring

[0291] The device's built-in sensors constantly monitor the subject's health condition, and the collected sensor data is periodically sent to a server, which immediately notifies the user if any abnormalities are detected.

[0292] Input: Sensor data (heart rate, steps, fall detection, etc.)

[0293] Output: Health data sent to the server, anomaly detection notification

[0294] Step 9:

[0295] emergency call

[0296] If an abnormality is detected, the device or server will make an emergency call, and the server will send an emergency notification to the configured emergency contacts, reporting the specific situation.

[0297] Input: Anomaly detection notification

[0298] Output: Notification to emergency number

[0299] Step 10:

[0300] Continuous learning and updates

[0301] The server continuously collects and analyzes the user's conversation data and health data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory.

[0302] Input: conversation data, health data

[0303] Output: Updated conversation script, new information notification

[0304] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0305] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Detailed embodiments of this system are described below.

[0306] ---

[0307] Enter your profile information

[0308] Device:

[0309] On the initial setup screen of the device, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). This information is temporarily stored in the device.

[0310] Examples:

[0311] The user's family member enters the following information: "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia."

[0312] ---

[0313] Sending profile information

[0314] Device:

[0315] Your profile information will be sent to the server via the Internet. When the sending is complete, the message "Profile information has been sent to the server" will be displayed.

[0316] Examples:

[0317] The device will notify you that "Taro Sato's information has been sent to the server."

[0318] ---

[0319] Profile information analysis and customization

[0320] server:

[0321] The system analyzes the received profile information, gathers relevant information on the Internet, and generates a customized conversation script based on the target person's hobbies and interests.

[0322] Examples:

[0323] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro Sato and generated a customized conversation script."

[0324] ---

[0325] Submitting a customization script

[0326] server:

[0327] The generated customized conversation script is sent to the terminal.

[0328] Device:

[0329] The received conversation script is saved in the internal memory and a message is displayed stating "Conversation data has been updated."

[0330] Examples:

[0331] The server sends "Gardening trivia" and the device notifies the user that "Taro Sato's conversation data has been updated."

[0332] ---

[0333] Speech and Emotion Recognition

[0334] User:

[0335] The user speaks to the device (e.g., "How's the gardening going lately?").

[0336] Device:

[0337] The user's voice is analyzed using a voice recognition module, and the emotion engine is used to recognize the user's emotions from their voice and facial expressions.

[0338] Examples:

[0339] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil preparation is becoming important," while determining whether the user looks happy.

[0340] ---

[0341] Providing customized responses

[0342] Device:

[0343] The system selects appropriate responses based on the user's emotions. For example, if the user is in good spirits, it selects more positive conversation content, and if they are feeling down, it offers encouraging words.

[0344] Examples:

[0345] If the user is feeling unwell, the device will gently ask, "You seem to be feeling unwell lately. Is there anything that is bothering you?"

[0346] ---

[0347] Continuous learning and updates

[0348] server:

[0349] By continuously collecting and analyzing user conversational and emotional data, the system learns the user's interests and emotional patterns and regularly updates the conversation script.

[0350] Device:

[0351] When a new script is received, it is saved in the internal memory and updated. Once the update is complete, a message will appear saying "Conversation data has been updated based on new information."

[0352] Examples:

[0353] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0354] ---

[0355] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[0356] The processing flow will be explained below.

[0357] Step 1:

[0358] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is temporarily stored in the device.

[0359] Step 2:

[0360] Device: The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[0361] Step 3:

[0362] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[0363] Step 4:

[0364] Server: Based on the extracted information, collects the latest information relevant to the target person from the Internet (e.g., the latest gardening techniques or Go news).

[0365] Step 5:

[0366] Server: Based on the collected information, a customized conversation script is generated for each target audience, including content that will interest the target audience and relevant information.

[0367] Step 6:

[0368] Server: Sends the generated customized conversation script to the device.

[0369] Step 7:

[0370] Device: Saves the received conversation script in its internal memory and notifies the user or their family that "the conversation data has been updated."

[0371] Step 8:

[0372] User: The user speaks to the device (e.g., "How's the gardening going lately?").

[0373] Step 9:

[0374] Device: The user's voice is analyzed by the speech recognition module. Then, the emotion engine recognizes emotions from the user's voice and facial expressions.

[0375] Step 10:

[0376] The device: Based on the recognized emotion, it selects an appropriate response from a customized conversation script. If the user is in good spirits, it selects positive responses, and if they are not, it offers encouraging words.

[0377] Step 11:

[0378] Terminal: The selected response is returned to the user through a voice output device. For example, "Soil conditioning is becoming increasingly important in modern gardening techniques."

[0379] Step 12:

[0380] Server: Continuously collects and analyzes user conversations and recognized emotional data, thereby learning the user's emotional patterns and interests.

[0381] Step 13:

[0382] Server: Periodically gathers new information and topics from the Internet and updates the customized conversation script.

[0383] Step 14:

[0384] Server: Sends the updated script to the device.

[0385] Step 15:

[0386] Device: When a new script is received, it is saved in internal memory and a notification is sent to the user or their family member stating, "Conversation data has been updated based on new information."

[0387] Through each step, the AI-powered IoT device can provide individually optimized conversations based on the target person's profile information and emotional data, allowing the elderly to enjoy their daily lives without feeling lonely. Furthermore, by combining it with an emotion engine, the device can understand the user's emotional state and provide appropriate responses, enabling deeper communication.

[0388] Example 2

[0389] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0390] Today's elderly often feel social isolation and loneliness, which can accelerate the progression of dementia. However, increasing opportunities for these elderly people to interact through everyday conversations could contribute to reducing loneliness and maintaining cognitive function. However, current technology is limited in systems that provide personalized conversations based on individual hobbies and interests, and there are particularly few systems that can recognize emotions. Therefore, there is a need for a system that provides conversations tailored to each individual elderly person and enables deep interaction.

[0391] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0392] In this invention, the server includes means for analyzing the profile information of the subject, collecting related information on the Internet, and generating a customized conversation script using a generative AI model, means for transmitting the customized conversation script to the terminal, and means for the terminal to analyze the user's voice using a voice recognition module and recognize the user's emotions using an emotion engine. This allows the user to enjoy conversations based on their hobbies and interests, and furthermore, receives appropriate responses through emotion recognition, which is expected to reduce feelings of loneliness among the elderly and slow the progression of dementia.

[0393] "Profile information" refers to personal information such as the subject's name, age, hobbies, and health status.

[0394] "Server" refers to a computer system that receives and analyzes profile information, gathers relevant information from the Internet, and generates customized conversation scripts.

[0395] "Conversation script" refers to customized conversation content generated using a generative AI model based on the target person's profile information and related information.

[0396] "Terminal" refers to a device used by a user to input profile information, receive customized conversation scripts, recognize voice and emotions, and display responses.

[0397] "Speech recognition module" refers to software or hardware that has the function of analyzing a user's voice and converting it into text.

[0398] "Emotion engine" refers to a system that includes technology to identify and analyze a user's emotions from voice and facial expression data.

[0399] A "generative AI model" refers to an artificial intelligence algorithm or machine learning model that generates customized conversation scripts based on input data.

[0400] "Speech synthesis engine" refers to a system that includes technology for converting text data into natural-sounding speech.

[0401] "Collecting related information on the Internet" refers to the server obtaining information related to the subject's hobbies and interests based on the profile information through methods such as web scraping or API calls.

[0402] "Response" refers to the voice or text response that the terminal returns to the user based on the analysis results of the voice recognition module and emotion engine.

[0403] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Specific embodiments for carrying out this invention are described below.

[0404] The main components of the system are a device, a server, a generative AI model, an emotion engine, and related software modules. The device has an interface for inputting and sending user profile information, which is then sent to the server. The server generates a customized conversation script based on this profile information. The generative AI model is used for generation.

[0405] Device:

[0406] The device is initially set up by the user or their family, and profile information (such as name, age, hobbies, and health status) is entered. The entered information is temporarily stored in the device's memory. Examples of devices include tablets and smart home devices. These devices are equipped with touch panels and keyboards, making it easy to enter information.

[0407] Examples:

[0408] The user's family member enters information such as "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" into the tablet.

[0409] server:

[0410] The server receives profile information sent from the device via the internet. Based on the received information, the server collects the latest information related to the target person's hobbies and interests from the internet. This information is collected using web scraping and API calls. A customized conversation script is then generated using a generative AI model and sent to the device.

[0411] Examples:

[0412] The prompt sentence "75-year-old male, hobbies are gardening and Go, mild dementia" is input into the generative AI model, which then performs processes such as "collecting the latest gardening trivia and Go game information for Taro Sato and generating a customized conversation script."

[0413] Device:

[0414] The device receives the customized conversation script sent from the server and saves it in its internal memory. After saving, it notifies the user that "the conversation data has been updated." The device then analyzes the user's voice using a voice recognition module and uses an emotion engine to recognize the user's emotions from their voice and facial expressions.

[0415] Examples:

[0416] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil adjustment is becoming important," while determining whether the user looks happy.

[0417] Emotion Engine:

[0418] The emotion engine includes technology that analyzes the user's emotions from voice and facial expression data. If the device determines that the user is not in good spirits, the emotion engine selects an appropriate response.

[0419] Examples:

[0420] If the user is not feeling well, gently ask, "You seem to be feeling unwell lately. Is there anything that's bothering you?"

[0421] The server continuously collects and analyzes the user's conversational and emotional data to learn about the user's interests and emotional patterns, allowing it to periodically update the conversation script and provide more personalized conversations.

[0422] Examples:

[0423] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0424] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[0425] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0426] Step 1:

[0427] Device:

[0428] On the initial setup screen, the user or their family member enters the target person's profile information. For example, they might enter "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" on the tablet's touch panel. The entered data is temporarily stored in the device's memory.

[0429] Input: Profile information (name, age, hobbies, health status)

[0430] Data processing: Input via touch panel or keyboard

[0431] Output: Profile information temporarily stored in the device memory

[0432] Step 2:

[0433] Device:

[0434] Your profile information will be sent to the server via the Internet. Once the sending is complete, the message "Profile information sent to server" will be displayed on the screen.

[0435] Input: Profile information stored in the device memory

[0436] Data processing: Sending HTTP requests via the network module

[0437] Output: Notification of completion of sending profile information to the server

[0438] Step 3:

[0439] server:

[0440] The submitted profile information is analyzed. For example, "75 years old, male, hobbies are gardening and Go, mild dementia." Related information is then collected from the internet. Web scraping and APIs are used to collect this information.

[0441] Input: Submitted profile information

[0442] Data processing: Analysis of profile information and collection of related information on the Internet

[0443] Output: Related information

[0444] Step 4:

[0445] server:

[0446] A generative AI model is used to generate a customized conversation script based on the collected relevant information. The prompt sentence, "75-year-old male, hobbies are gardening and Go, mild dementia," is input into the generative AI model.

[0447] Input: Profile information and related information

[0448] Data processing: Generating conversation scripts using generative AI models

[0449] Output: Customized conversation script

[0450] Step 5:

[0451] server:

[0452] The generated customized conversation script is sent to the terminal.

[0453] Input: Customized conversation script

[0454] Data processing: Sending HTTP requests via the network module

[0455] Output: Notification of completion of transmission to the terminal

[0456] Step 6:

[0457] Device:

[0458] The received conversation script is saved to the internal memory. After saving is complete, a message will be displayed saying "Conversation data has been updated."

[0459] Input: Conversation script sent from the server

[0460] Data processing: Saving to internal memory and notifying on the user interface

[0461] Output: "Conversation data updated" notification

[0462] Step 7:

[0463] User:

[0464] The user speaks to the device (e.g., "How's the gardening going lately?").

[0465] Input: User voice input

[0466] Data processing: None

[0467] Output: User's voice input

[0468] Step 8:

[0469] Device:

[0470] The user's voice is analyzed by the voice recognition module. The voice data is sent to the STT (Speech-to-Text) engine to convert it into text, and the emotion engine recognizes emotions from the voice and facial expressions.

[0471] Input: User's voice data

[0472] Data processing: Speech to text and emotion recognition

[0473] Output: User utterance text and recognized sentiment

[0474] Step 9:

[0475] Device:

[0476] Based on the user's emotions, the robot selects an appropriate response from a customized conversation script. For example, if the user is in good spirits, it selects positive conversational content, and if they are feeling down, it responds with encouraging words. The selected response is then converted into voice by a speech synthesis engine and returned to the user.

[0477] Input: User utterance text, recognized emotions, and a customized conversation script

[0478] Data processing: Response selection and speech synthesis based on spoken text and emotional data

[0479] Output: Audio response

[0480] Step 10:

[0481] server:

[0482] It continuously collects and analyzes user conversational and emotional data, and uses this data to periodically update the generative AI model with customized conversation scripts.

[0483] Input: User conversation data and emotion data

[0484] Data processing: Data analysis and updating of conversation scripts using generative AI models

[0485] Output: Updated conversation script

[0486] Step 11:

[0487] Device:

[0488] The updated conversation script is received and saved in the internal memory. After the update is complete, a message will be displayed stating, "The conversation data has been updated based on the latest information."

[0489] Input: Updated conversation script

[0490] Data processing: Saving to internal memory and notifying on the user interface

[0491] Output: Notification "Conversation data has been updated based on the latest information"

[0492] (Application example 2)

[0493] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0494] In modern society, loneliness and mental stress are serious problems for the elderly and for elderly factory workers. In particular, the psychological impact of monotonous work in factories and lack of communication with others cannot be overlooked. There are also concerns about dementia and health conditions. Therefore, there is a need for a system that helps elderly people stay healthy through appropriate communication, without feeling lonely, and allows them to work safely and efficiently.

[0495] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting profile information of the target person, means for transmitting the profile information of the target person to the server, means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for transmitting the generated customized conversation script to the terminal, means for recognizing the user's voice and responding based on the customized conversation script, and means for analyzing the user's emotional state using an emotion engine and providing appropriate conversation content. This not only enables elderly people to enjoy everyday conversations without feeling lonely, but also improves work efficiency and safety.

[0496] "Subject profile information" refers to basic information including the subject's name, age, hobbies, health status, past work experience, etc.

[0497] "Server" refers to a computer system for analyzing profile information, collecting related information, and generating and transmitting customized conversation scripts.

[0498] A "terminal" refers to a device for interacting with a user, which has the function of recognizing the user's voice and responding based on a conversation script.

[0499] "Customized conversation script" refers to personalized dialogue content generated based on the target person's profile information.

[0500] "Speech recognition" refers to the technology that analyzes a user's voice and converts it into text data.

[0501] An "emotion engine" refers to technology that analyzes a user's emotional state from data such as voice and facial expressions, and selects appropriate conversation content.

[0502] "Emotional state" refers to the emotions the user is feeling at the time, such as joy, sadness, interest, indifference, etc.

[0503] MODE FOR CARRYING OUT THE INVENTION

[0504] System configuration

[0505] The system for implementing this invention includes the following main components: a terminal for inputting profile information of a target person; a server for analyzing the profile information, collecting related information, and generating a customized conversation script; and a terminal for recognizing the user's voice and providing appropriate conversation content using an emotion engine.

[0506] Hardware and software used

[0507] Hardware

[0508] Smartphones and tablets (as audio input / output devices)

[0509] Factory robots (for work assistance and interaction)

[0510] software

[0511] OpenAI API (for dialogue generation)

[0512] SpeechRecognition

[0513] gTTS (Text to Speech)

[0514] mpg321 (audio playback)

[0515] Program processing flow and data processing

[0516] 1. Enter your profile information

[0517] Enter the subject's basic information (name, age, hobbies, health status, past work experience, etc.) on the terminal.

[0518] The terminal transmits the input information to the server.

[0519] 2. Analyzing profile information and generating customized conversation scripts

[0520] The server analyzes the received profile information and collects related information on the Internet.

[0521] The server generates a customized conversation script based on the collected information and sends it to the terminal.

[0522] 3. Speech and Emotion Recognition

[0523] When the user speaks to the device, it recognizes the voice and converts it into text data.

[0524] An emotion engine is used to analyze the user's emotional state from their voice and facial expressions.

[0525] 4. Providing customized responses

[0526] The device selects appropriate conversation content based on the analyzed emotional state.

[0527] The selected conversation content is converted from text to speech and provided to the user.

[0528] Specific examples

[0529] For example, suppose an elderly person working in a factory asks the device, "How's the gardening going lately?" The device first recognizes the user's voice and sends the information to the server. The server has already registered the person's profile information (for example, that gardening is a hobby), so it collects the latest information about gardening and generates a customized conversation script. Next, it uses an emotion engine to analyze the user's emotional state from their voice and facial expressions, and selects the most appropriate response to respond. For example, it provides specific advice such as, "A recent gardening technique that places importance on adjusting the soil."

[0530] Prompt Sentence Examples

[0531] The specific prompt for the generative AI model is as follows:

[0532] Generate a script for a conversation with Taro Sato, an elderly worker. User utterance: How's your gardening going lately?

[0533] Based on this prompt, the AI ​​model can generate a response that combines the user's profile information and relevant information.

[0534] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0535] Step 1:

[0536] Enter your profile information

[0537] The terminal (smartphone or tablet) provides an interface where the user or their family members can input the subject's name, age, hobbies, health status, past work experience, etc.

[0538] Input: Subject's profile information (e.g., "Name: Ichiro Takahashi, Age: 70, Hobbies: Gardening and Go, Health Status: Mild Dementia")

[0539] Output: The entered profile information is temporarily stored in the device's local memory.

[0540] Step 2:

[0541] Sending profile information

[0542] The terminal transmits the saved profile information to a server via the Internet.

[0543] Input: Profile information stored in local memory

[0544] Output: The profile information is sent to the server and received at the server side.

[0545] Step 3:

[0546] Analysis of profile information and collection of related information

[0547] The server analyzes the received profile information and collects related information from the Internet based on the subject's hobbies and interests.

[0548] Input: Received profile information

[0549] Output: A customized conversation script (e.g., "Latest gardening news," "New Go strategies") is generated based on the profile information.

[0550] Step 4:

[0551] Sending customized conversation scripts

[0552] The server transmits the generated customized conversation script to the terminal.

[0553] Input: Customized conversation script

[0554] Output: The customized conversation script is sent to the device and stored in the device's internal memory.

[0555] Step 5:

[0556] Voice Recognition

[0557] When a user speaks to the terminal, the terminal uses a speech recognition module (SpeechRecognition) to convert the voice data into text data.

[0558] Input: User speech (e.g., "How's the gardening going lately?")

[0559] Output: The converted text data (e.g., "How's your gardening going lately?")

[0560] Step 6:

[0561] emotion recognition

[0562] The device uses an emotion engine to analyze the user's emotional state from the voice data.

[0563] Input: Converted text data and audio data

[0564] Output: User's emotional state (e.g., "Interested")

[0565] Step 7:

[0566] Providing customized responses

[0567] The device selects appropriate responses based on the analyzed emotional state and a customized conversation script.

[0568] Input: User's emotional state, customized conversation script

[0569] Output: Text response (e.g., "Soil conditioning is becoming increasingly important as a gardening technique these days.")

[0570] Step 8:

[0571] Text-to-speech conversion and output

[0572] The terminal converts the selected response into speech using gTTS and outputs it as speech.

[0573] Input: Text data response

[0574] Output: Provided to the user as audio data (e.g., "Soil preparation is becoming increasingly important as a gardening technique these days.")

[0575] This will provide an environment where elderly people and elderly factory workers can work with peace of mind without feeling lonely.

[0576] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0577] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0578] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0579] [Second embodiment]

[0580] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0581] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0582] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0583] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0584] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0585] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0586] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0587] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0588] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0589] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0590] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0591] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0592] This invention provides a system that uses AI-equipped IoT devices to enable elderly people to enjoy everyday conversations without feeling lonely. Specifically, the following process is performed.

[0593] ---

[0594] Enter your profile information

[0595] Device:

[0596] Using the initial setup screen that appears when the device is started, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[0597] Examples:

[0598] The user's family member enters the following information into the terminal: "Taro, 75 years old, hobbies are gardening and Go, suffers from mild dementia."

[0599] ---

[0600] Sending profile information

[0601] Device:

[0602] The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[0603] Examples:

[0604] The device will notify you that "Taro's information has been sent to the server."

[0605] ---

[0606] Profile information analysis and customization

[0607] server:

[0608] The system analyzes the received profile information and gathers relevant information from the Internet (e.g., the latest gardening techniques or Go news). Based on this information, it generates a conversation script customized for each user.

[0609] Examples:

[0610] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro and generated a customized conversation script."

[0611] ---

[0612] Submitting a customization script

[0613] server:

[0614] The generated customized conversation script is sent to the terminal.

[0615] Device:

[0616] Receives customized conversation scripts and saves them in internal memory. Notifies the user or their family that "conversation data has been updated."

[0617] Examples:

[0618] The server sends "Gardening trivia" and the device notifies the user that "Taro's conversation data has been updated."

[0619] ---

[0620] Start a conversation

[0621] User:

[0622] The user speaks to the device (e.g., "How's the gardening going lately?").

[0623] Device:

[0624] The user's voice is analyzed by a voice recognition module, and based on the analysis results, an appropriate response is selected from a customized conversation script and responded to by voice.

[0625] Examples:

[0626] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[0627] ---

[0628] Continuous learning and updates

[0629] server:

[0630] It continuously collects, analyzes, and provides feedback on user conversation data, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[0631] Device:

[0632] Every time a new script is sent, it is received, stored in internal memory, and updated. It notifies you that "Conversation data has been updated based on new information."

[0633] Examples:

[0634] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0635] ---

[0636] In this way, AI-equipped IoT devices can collect the latest information from the internet based on the target person's profile information and provide optimized conversations, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely, which is expected to reduce loneliness and slow the progression of dementia.

[0637] The processing flow will be explained below.

[0638] Step 1:

[0639] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is stored in the device's temporary memory.

[0640] Step 2:

[0641] Device: The entered profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[0642] Step 3:

[0643] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[0644] Step 4:

[0645] Server: Based on the extracted information, the server collects the latest information related to the target person from the Internet. For example, if the target person's hobby is gardening, it collects information on the latest gardening techniques and plants.

[0646] Step 5:

[0647] Server: Generates a customized conversation script based on the collected information, including content that is interesting and relevant to the target audience.

[0648] Step 6:

[0649] Server: Sends the generated customized conversation script to the device.

[0650] Step 7:

[0651] Device: Saves the received conversation script in its internal memory. Notifies the user or their family that "the conversation data has been updated."

[0652] Step 8:

[0653] User: The user speaks to the device (e.g., "How's your gardening going lately?"). The device analyzes the user's voice using a speech recognition module.

[0654] Step 9:

[0655] Terminal: Based on the analysis results, the terminal selects an appropriate response from a customized conversation script and returns the response to the user through a voice output device.

[0656] Step 10:

[0657] Server: Continuously collects user conversation data, analyzes it, and provides feedback. This data serves as the basis for further customization based on the user's interests.

[0658] Step 11:

[0659] Server: Periodically collects new information and topics, updates the customized conversation script, and sends the updated script to the device.

[0660] Step 12:

[0661] Device: Receives each new script sent, saves it in internal memory, and updates it. Once the update is complete, it notifies you that "Conversation data has been updated based on new information."

[0662] Through each step, the AI-enabled IoT device provides personalized conversations, allowing the elderly to enjoy their daily lives without feeling lonely.

[0663] Example 1

[0664] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0665] To enable elderly people to enjoy everyday conversations without feeling lonely, conversation content needs to be customized to each individual. However, conventional systems lack such customization, and often provide topics and information that do not interest elderly people. This leads to a decrease in satisfaction. In addition, there is no mechanism for continuously updating new information, which means that conversation content easily becomes outdated.

[0666] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0667] In this invention, the server includes a means for transmitting the target person's attribute information to the network, a means for collecting related information in a data store based on the transmitted attribute information and generating customized conversation data, and a means for transmitting the generated customized conversation data to the client device. This enables conversations using the latest information specific to the target person, allowing elderly people to participate in conversations with interest without feeling lonely. Furthermore, by continuously collecting new information and updating the conversation content, it is possible to always provide fresh topics of conversation.

[0668] "Subject attribute information" refers to personal information such as the subject's name, age, hobbies, and health status.

[0669] "Data store" refers to a system or database for storing information and data collected from the Internet.

[0670] "Customized conversation data" refers to a conversation script that is generated based on the individual attribute information of the subject and tailored to the subject.

[0671] "Client device" refers to an AI-enabled IoT device or terminal used by the subject.

[0672] "Network" refers to the Internet and other communication means that provide paths for transmitting and receiving data.

[0673] A "conversation script" refers to data that describes the content of utterances and responses that the system uses in dialogue with the user.

[0674] This invention provides a system that uses AI-equipped IoT devices to help elderly people enjoy everyday conversations without feeling lonely. The entire system is realized by collecting profile information of the target person and providing customized conversation data to the user.

[0675] Hardware and software used

[0676] Terminal: An AI-enabled IoT device placed close to the user. This terminal displays the necessary initial setup screens and allows the user or their family members to enter their profile information.

[0677] Server: Refers to the back-end system that analyzes data and generates customized conversation scripts. The server is connected to the Internet and collects various information.

[0678] Speech Recognition Module: For example, using Google Speech-to-Text. Used to convert the user's speech into text.

[0679] Generative AI models, such as GPT-3, are used to generate customized conversation scripts based on collected information.

[0680] A speech synthesis module, e.g., using Google Text-to-Speech, is used to convert the generated text into speech and respond to the user.

[0681] Explanation of program processing

[0682] Device:

[0683] 1. When the device starts up, the initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored on the device.

[0684] 2. The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[0685] server:

[0686] 1. The server analyzes the received profile information and collects relevant information from a data store (e.g., the latest gardening techniques or Go news). Based on this information, a customized conversation script is generated for each user. This is generated using a generative AI model.

[0687] 2. Send the generated customized conversation script to the terminal.

[0688] Device:

[0689] 1. The device saves the received conversation script in its internal memory and displays "Conversation data updated" on the screen.

[0690] 2. When the user speaks to the device, the device analyzes it using the voice recognition module, selects a response based on the customized conversation data, and responds to the user using the voice synthesis module.

[0691] Specific examples

[0692] Enter your profile information:

[0693] The user's family member enters information into the device, such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." This information is temporarily stored in the device.

[0694] Send profile information:

[0695] When the device is connected to a network such as Wi-Fi, it sends the saved profile information to the server as a POST request. If the request is successful, a pop-up message will appear saying, "Taro's information has been sent to the server."

[0696] Profile information analysis and customization:

[0697] The server analyzes the profile information "Taro, 75 years old, gardening, Go, mild dementia," collects relevant latest gardening techniques and recent Go news, and generates a conversation script using a generative AI model.

[0698] Send a customized conversation script:

[0699] The server sends the conversation script data in JSON format, and the device receives this data, parses it, and saves it in its internal memory. The updated content is notified with a message saying "The conversation data has been updated."

[0700] Start a conversation:

[0701] When a user speaks to the device, "How's your gardening going these days?", the device converts the speech into text using a speech recognition module, and then uses a generative AI model to respond, "The latest gardening techniques place great importance on soil preparation," which is then played back as audio using a speech synthesis module.

[0702] Example prompt sentence:

[0703] "Please tell me the latest gardening knowledge for 75-year-old Taro."

[0704] "Tell me the latest news about Go"

[0705] "How's your gardening going lately?"

[0706] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0707] Step 1:

[0708] Enter and save your profile information

[0709] Device:

[0710] When the device starts up, an initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[0711] Input: The user enters profile information as text.

[0712] Output: The entered profile information is saved in the device's temporary memory.

[0713] Specific operation: A user's family member enters information such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." The device saves this information in text format.

[0714] Step 2:

[0715] Sending profile information

[0716] Device:

[0717] The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[0718] Input: Profile information stored on your device.

[0719] Output: The profile information is sent to the server.

[0720] Specific operation: The device sends profile information in the form of a POST request via a network such as Wi-Fi. A message pops up when the transmission is successful.

[0721] Step 3:

[0722] Profile information analysis and customization

[0723] server:

[0724] The server analyzes the received profile information and gathers relevant information from a data store, such as the latest gardening techniques or Go news, and then uses a generative AI model to generate a customized conversation script based on the collected information.

[0725] Input: Profile information stored on the server.

[0726] Output: A customized conversation script.

[0727] Specific operation: The server analyzes the information "Taro, 75 years old, hobbies are gardening and Go, mild dementia," collects related information using web scraping and APIs, and generates a conversation script using a generative AI model (e.g., GPT-3).

[0728] Step 4:

[0729] Sending customized conversation scripts

[0730] server:

[0731] The generated customized conversation script is sent to the terminal.

[0732] Device:

[0733] The terminal stores the received conversation script in its internal memory and displays the message "Conversation data has been updated" on the screen.

[0734] Input: A customized conversation script generated by the server.

[0735] Output: The customized conversation script is saved in the device's memory.

[0736] Specific operation: The server sends script data in JSON format, the device receives it, parses it, saves it in its internal memory, and displays a pop-up message saying "Conversation data has been updated."

[0737] Step 5:

[0738] Start a conversation

[0739] User:

[0740] The user provides the device with a question or topic (e.g., "How's your gardening going lately?").

[0741] Device:

[0742] The device uses a speech recognition module to convert the user's speech into text, generates appropriate responses based on customized conversation data, and responds to the user using a speech synthesis module.

[0743] Input: User's voice input.

[0744] Output: Audio response from the device.

[0745] What it does: When a user says, "How's your gardening going these days?", the device converts the speech to text and selects a response from a customized conversation script: "Modern gardening techniques emphasize soil conditioning," and plays it back as audio.

[0746] Step 6:

[0747] Continuous learning and updates

[0748] server:

[0749] The server continuously collects user conversation data, analyzes it, and provides feedback, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[0750] Device:

[0751] Every time a new conversation script is sent from the server, it is received, saved in internal memory, and updated. When the update is complete, a message will appear saying, "The conversation data has been updated based on the new information."

[0752] Input: User conversation data and new information.

[0753] Output: Updated customized conversation script.

[0754] Specific operation: The server analyzes the user's conversation history, collects new gardening information, the latest news on Go, etc., and sends an updated conversation script to the device. The device receives this and notifies the user that "the conversation data has been updated with the latest information."

[0755] (Application example 1)

[0756] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0757] There is a need for systems that can not only reduce loneliness among the elderly and provide daily communication, but also monitor their health status and respond to emergency situations. There is also a need for systems that can provide customized conversations using information specific to the elderly and the latest topics.

[0758] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0759] In this invention, the server includes means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for collecting sensor data for monitoring health status, and means for making an emergency call if an abnormality occurs. This not only reduces the sense of loneliness of elderly people and provides communication, but also enables health status monitoring and prompt response in emergencies.

[0760] "Subject profile information" is basic information about an individual, such as the subject's name, age, hobbies, and health status.

[0761] The "means for inputting profile information" refers to an interface that allows the subject or their family to input profile information using a terminal.

[0762] The "means for transmitting profile information to a server" refers to a communication means for transmitting profile information from a terminal to a server via the Internet.

[0763] A "customized conversation script" is an individually optimized conversation scenario that is generated by collecting relevant information on the Internet based on the target person's profile information.

[0764] The "means for generating a customized conversation script" is a processing means for analyzing profile information and creating a conversation script based on the latest related information on the Internet.

[0765] The "means for recognizing and responding to the user's voice" is a function for analyzing the voice uttered by the user and selecting and responding to an appropriate response based on a customized conversation script.

[0766] "Sensor data for monitoring health status" refers to information obtained from sensors that collect physiological and activity data of subjects, such as heart rate, number of steps, and fall detection.

[0767] The "means for collecting sensor data" refers to devices and software for collecting physiological data and activity data through sensors, which are necessary for monitoring the health status of a subject.

[0768] The "means for making an emergency report when an abnormality occurs" is a means for making a report to a pre-set emergency contact when an abnormality is detected based on collected sensor data.

[0769] "Means for gathering information from the Internet" refers to the Internet browsing ability to gather relevant and up-to-date information and news on the Internet based on the profile information.

[0770] This invention provides a system that can reduce loneliness among the elderly, monitor their health status, and respond to emergencies. The system provides customized conversations based on the target person's profile information and makes an emergency call when an abnormality in their health status is detected.

[0771] Program processing explanation

[0772] Hardware and Software Configuration

[0773] Hardware:

[0774] Robot device (equipped with microphone, speaker, and sensor)

[0775] Software and Libraries:

[0776] requests library: Used for server communication.

[0777] speech_recognition library: Used for speech recognition.

[0778] pyttsx3 library: Used for speech synthesis.

[0779] Processing explanation

[0780] (1) Enter and submit profile information

[0781] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device and then sent to the server. The server receives the information and sends a confirmation message back to the device.

[0782] Examples:

[0783] The user's family member enters information such as "Mr. Tanaka, 75 years old, hobbies are gardening and Go, mild dementia" into the terminal and sends it to the server.

[0784] (2) Analyzing profile information and generating conversation scripts

[0785] The server analyzes the received profile information and collects relevant information from the Internet (such as the latest gardening techniques or Go news), and generates a conversation script customized for each user based on this information.

[0786] Examples:

[0787] The server "collects the latest gardening trivia and Go game information for Tanaka and generates a customized conversation script."

[0788] (3) Sending customized conversation scripts

[0789] The server sends the generated customized conversation script to the device. The device receives the conversation script and saves it in its internal memory. Once the saving is complete, the device displays a notification to the user that the update is complete.

[0790] Examples:

[0791] The server sends "Gardening trivia" and the device notifies the user that "Tanaka's conversation data has been updated."

[0792] (4) Communication using voice recognition

[0793] When a user speaks to the terminal, the terminal analyzes the user's voice using a voice recognition module, selects an appropriate response from a customized conversation script, and responds verbally.

[0794] Examples:

[0795] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[0796] (5) Health monitoring and emergency notification

[0797] The device's built-in sensors constantly monitor the user's health. If an abnormality is detected (e.g., an abnormal heart rate or a fall), the device immediately sends the abnormal data to the server, which then notifies the user's emergency contacts.

[0798] Examples:

[0799] The device notifies the server that an abnormal heart rate has been detected, and the server then calls a pre-set emergency number.

[0800] (6) Continuous learning and updating

[0801] The server continuously collects and analyzes the user's conversation data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory. The user is notified that the conversation data has been updated with new information.

[0802] Examples:

[0803] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0804] Example prompts for generative AI models

[0805] Generate customized conversation scripts about the latest gardening news and Go topics based on the user's profile information. For example, respond to the question "Tell me about gardening" with "Soil preparation is becoming increasingly important as a modern gardening technique."

[0806] In this way, the embodiments of the invention allow elderly people to enjoy everyday conversations without feeling lonely, while their health condition is monitored, and if an abnormality occurs, a prompt response can be made.

[0807] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0808] Step 1:

[0809] Enter your profile information

[0810] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device. The input is done using the user interface, and the input data is stored in the device's internal storage.

[0811] Input: Profile information such as name, age, hobbies, health status, etc.

[0812] Output: Profile information temporarily saved on the device

[0813] Step 2:

[0814] Sending profile information

[0815] The device will then send the saved profile information to the server via the internet, and once the transfer is complete, a confirmation message will appear saying "Profile information has been sent to the server."

[0816] Input: Profile information temporarily saved on the device

[0817] Output: Profile information sent to server, confirmation message to user

[0818] Step 3:

[0819] Analyzing profile information

[0820] The server analyzes the received profile information and collects relevant information on the Internet (e.g., the latest gardening techniques or Go news). Based on the profile information, the analysis generates relevant search queries and retrieves relevant information.

[0821] Input: Profile information sent to the server

[0822] Output: Analysis results and related collected information

[0823] Step 4:

[0824] Generate customized conversation scripts

[0825] The server generates a customized conversation script for each user based on the collected relevant information. It uses a generative AI model to script a conversation scenario based on the profile information and collected information.

[0826] Input: Analysis results and related collected information

[0827] Output: Customized conversation script

[0828] Step 5:

[0829] Sending customized conversation scripts

[0830] The server sends the generated customized conversation script to the terminal. The terminal receives the conversation script and saves it in its internal memory. Once the saving is complete, a notification that "Conversation data has been updated" is displayed to the user.

[0831] Input: Customized conversation script

[0832] Output: Conversation script saved on the device, update completion notification to the user

[0833] Step 6:

[0834] User voice recognition

[0835] When a user speaks to the device, the device uses a voice recognition module to analyze the user's voice, converting the voice data into text data, and then selects an appropriate response from a customized conversation script based on the text.

[0836] Input: User voice input

[0837] Output: Recognized text data, selected response

[0838] Step 7:

[0839] Voice response

[0840] The device responds to the user vocally based on the selected customized conversation script, which is generated using a speech synthesis module and output audibly through the speaker.

[0841] Input: Selected Response

[0842] Output: Audio response

[0843] Step 8:

[0844] Health monitoring

[0845] The device's built-in sensors constantly monitor the subject's health condition, and the collected sensor data is periodically sent to a server, which immediately notifies the user if any abnormalities are detected.

[0846] Input: Sensor data (heart rate, steps, fall detection, etc.)

[0847] Output: Health data sent to the server, anomaly detection notification

[0848] Step 9:

[0849] emergency call

[0850] If an abnormality is detected, the device or server will make an emergency call, and the server will send an emergency notification to the configured emergency contacts, reporting the specific situation.

[0851] Input: Anomaly detection notification

[0852] Output: Notification to emergency number

[0853] Step 10:

[0854] Continuous learning and updates

[0855] The server continuously collects and analyzes the user's conversation data and health data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory.

[0856] Input: conversation data, health data

[0857] Output: Updated conversation script, new information notification

[0858] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0859] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Detailed embodiments of this system are described below.

[0860] ---

[0861] Enter your profile information

[0862] Device:

[0863] On the initial setup screen of the device, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). This information is temporarily stored in the device.

[0864] Examples:

[0865] The user's family member enters the following information: "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia."

[0866] ---

[0867] Sending profile information

[0868] Device:

[0869] Your profile information will be sent to the server via the Internet. When the sending is complete, the message "Profile information has been sent to the server" will be displayed.

[0870] Examples:

[0871] The device will notify you that "Taro Sato's information has been sent to the server."

[0872] ---

[0873] Profile information analysis and customization

[0874] server:

[0875] The system analyzes the received profile information, gathers relevant information on the Internet, and generates a customized conversation script based on the target person's hobbies and interests.

[0876] Examples:

[0877] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro Sato and generated a customized conversation script."

[0878] ---

[0879] Submitting a customization script

[0880] server:

[0881] The generated customized conversation script is sent to the terminal.

[0882] Device:

[0883] The received conversation script is saved in the internal memory and a message is displayed stating "Conversation data has been updated."

[0884] Examples:

[0885] The server sends "Gardening trivia" and the device notifies the user that "Taro Sato's conversation data has been updated."

[0886] ---

[0887] Speech and Emotion Recognition

[0888] User:

[0889] The user speaks to the device (e.g., "How's the gardening going lately?").

[0890] Device:

[0891] The user's voice is analyzed using a voice recognition module, and the emotion engine is used to recognize the user's emotions from their voice and facial expressions.

[0892] Examples:

[0893] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil preparation is becoming important," while determining whether the user looks happy.

[0894] ---

[0895] Providing customized responses

[0896] Device:

[0897] The system selects appropriate responses based on the user's emotions. For example, if the user is in good spirits, it selects more positive conversation content, and if they are feeling down, it offers encouraging words.

[0898] Examples:

[0899] If the user is feeling unwell, the device will gently ask, "You seem to be feeling unwell lately. Is there anything that is bothering you?"

[0900] ---

[0901] Continuous learning and updates

[0902] server:

[0903] By continuously collecting and analyzing user conversational and emotional data, the system learns the user's interests and emotional patterns and regularly updates the conversation script.

[0904] Device:

[0905] When a new script is received, it is saved in the internal memory and updated. Once the update is complete, a message will appear saying "Conversation data has been updated based on new information."

[0906] Examples:

[0907] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0908] ---

[0909] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[0910] The processing flow will be explained below.

[0911] Step 1:

[0912] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is temporarily stored in the device.

[0913] Step 2:

[0914] Device: The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[0915] Step 3:

[0916] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[0917] Step 4:

[0918] Server: Based on the extracted information, collects the latest information relevant to the target person from the Internet (e.g., the latest gardening techniques or Go news).

[0919] Step 5:

[0920] Server: Based on the collected information, a customized conversation script is generated for each target audience, including content that will interest the target audience and relevant information.

[0921] Step 6:

[0922] Server: Sends the generated customized conversation script to the device.

[0923] Step 7:

[0924] Device: Saves the received conversation script in its internal memory and notifies the user or their family that "the conversation data has been updated."

[0925] Step 8:

[0926] User: The user speaks to the device (e.g., "How's the gardening going lately?").

[0927] Step 9:

[0928] Device: The user's voice is analyzed by the speech recognition module. Then, the emotion engine recognizes emotions from the user's voice and facial expressions.

[0929] Step 10:

[0930] The device: Based on the recognized emotion, it selects an appropriate response from a customized conversation script. If the user is in good spirits, it selects positive responses, and if they are not, it offers encouraging words.

[0931] Step 11:

[0932] Terminal: The selected response is returned to the user through a voice output device. For example, "Soil conditioning is becoming increasingly important in modern gardening techniques."

[0933] Step 12:

[0934] Server: Continuously collects and analyzes user conversations and recognized emotional data, thereby learning the user's emotional patterns and interests.

[0935] Step 13:

[0936] Server: Periodically gathers new information and topics from the Internet and updates the customized conversation script.

[0937] Step 14:

[0938] Server: Sends the updated script to the device.

[0939] Step 15:

[0940] Device: When a new script is received, it is saved in internal memory and a notification is sent to the user or their family member stating, "Conversation data has been updated based on new information."

[0941] Through each step, the AI-powered IoT device can provide individually optimized conversations based on the target person's profile information and emotional data, allowing the elderly to enjoy their daily lives without feeling lonely. Furthermore, by combining it with an emotion engine, the device can understand the user's emotional state and provide appropriate responses, enabling deeper communication.

[0942] Example 2

[0943] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0944] Today's elderly often feel social isolation and loneliness, which can accelerate the progression of dementia. However, increasing opportunities for these elderly people to interact through everyday conversations could contribute to reducing loneliness and maintaining cognitive function. However, current technology is limited in systems that provide personalized conversations based on individual hobbies and interests, and there are particularly few systems that can recognize emotions. Therefore, there is a need for a system that provides conversations tailored to each individual elderly person and enables deep interaction.

[0945] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0946] In this invention, the server includes means for analyzing the profile information of the subject, collecting related information on the Internet, and generating a customized conversation script using a generative AI model, means for transmitting the customized conversation script to the terminal, and means for the terminal to analyze the user's voice using a voice recognition module and recognize the user's emotions using an emotion engine. This allows the user to enjoy conversations based on their hobbies and interests, and furthermore, receives appropriate responses through emotion recognition, which is expected to reduce feelings of loneliness among the elderly and slow the progression of dementia.

[0947] "Profile information" refers to personal information such as the subject's name, age, hobbies, and health status.

[0948] "Server" refers to a computer system that receives and analyzes profile information, gathers relevant information from the Internet, and generates customized conversation scripts.

[0949] "Conversation script" refers to customized conversation content generated using a generative AI model based on the target person's profile information and related information.

[0950] "Terminal" refers to a device used by a user to input profile information, receive customized conversation scripts, recognize voice and emotions, and display responses.

[0951] "Speech recognition module" refers to software or hardware that has the function of analyzing a user's voice and converting it into text.

[0952] "Emotion engine" refers to a system that includes technology to identify and analyze a user's emotions from voice and facial expression data.

[0953] A "generative AI model" refers to an artificial intelligence algorithm or machine learning model that generates customized conversation scripts based on input data.

[0954] "Speech synthesis engine" refers to a system that includes technology for converting text data into natural-sounding speech.

[0955] "Collecting related information on the Internet" refers to the server obtaining information related to the subject's hobbies and interests based on the profile information through methods such as web scraping or API calls.

[0956] "Response" refers to the voice or text response that the terminal returns to the user based on the analysis results of the voice recognition module and emotion engine.

[0957] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Specific embodiments for carrying out this invention are described below.

[0958] The main components of the system are a device, a server, a generative AI model, an emotion engine, and related software modules. The device has an interface for inputting and sending user profile information, which is then sent to the server. The server generates a customized conversation script based on this profile information. The generative AI model is used for generation.

[0959] Device:

[0960] The device is initially set up by the user or their family, and profile information (such as name, age, hobbies, and health status) is entered. The entered information is temporarily stored in the device's memory. Examples of devices include tablets and smart home devices. These devices are equipped with touch panels and keyboards, making it easy to enter information.

[0961] Examples:

[0962] The user's family member enters information such as "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" into the tablet.

[0963] server:

[0964] The server receives profile information sent from the device via the internet. Based on the received information, the server collects the latest information related to the target person's hobbies and interests from the internet. This information is collected using web scraping and API calls. A customized conversation script is then generated using a generative AI model and sent to the device.

[0965] Examples:

[0966] The prompt sentence "75-year-old male, hobbies are gardening and Go, mild dementia" is input into the generative AI model, which then performs processes such as "collecting the latest gardening trivia and Go game information for Taro Sato and generating a customized conversation script."

[0967] Device:

[0968] The device receives the customized conversation script sent from the server and saves it in its internal memory. After saving, it notifies the user that "the conversation data has been updated." The device then analyzes the user's voice using a voice recognition module and uses an emotion engine to recognize the user's emotions from their voice and facial expressions.

[0969] Examples:

[0970] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil adjustment is becoming important," while determining whether the user looks happy.

[0971] Emotion Engine:

[0972] The emotion engine includes technology that analyzes the user's emotions from voice and facial expression data. If the device determines that the user is not in good spirits, the emotion engine selects an appropriate response.

[0973] Examples:

[0974] If the user is not feeling well, gently ask, "You seem to be feeling unwell lately. Is there anything that's bothering you?"

[0975] The server continuously collects and analyzes the user's conversational and emotional data to learn about the user's interests and emotional patterns, allowing it to periodically update the conversation script and provide more personalized conversations.

[0976] Examples:

[0977] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[0978] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[0979] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0980] Step 1:

[0981] Device:

[0982] On the initial setup screen, the user or their family member enters the target person's profile information. For example, they might enter "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" on the tablet's touch panel. The entered data is temporarily stored in the device's memory.

[0983] Input: Profile information (name, age, hobbies, health status)

[0984] Data processing: Input via touch panel or keyboard

[0985] Output: Profile information temporarily stored in the device memory

[0986] Step 2:

[0987] Device:

[0988] Your profile information will be sent to the server via the Internet. Once the sending is complete, the message "Profile information sent to server" will be displayed on the screen.

[0989] Input: Profile information stored in the device memory

[0990] Data processing: Sending HTTP requests via the network module

[0991] Output: Notification of completion of sending profile information to the server

[0992] Step 3:

[0993] server:

[0994] The submitted profile information is analyzed. For example, "75 years old, male, hobbies are gardening and Go, mild dementia." Related information is then collected from the internet. Web scraping and APIs are used to collect this information.

[0995] Input: Submitted profile information

[0996] Data processing: Analysis of profile information and collection of related information on the Internet

[0997] Output: Related information

[0998] Step 4:

[0999] server:

[1000] A generative AI model is used to generate a customized conversation script based on the collected relevant information. The prompt sentence, "75-year-old male, hobbies are gardening and Go, mild dementia," is input into the generative AI model.

[1001] Input: Profile information and related information

[1002] Data processing: Generating conversation scripts using generative AI models

[1003] Output: Customized conversation script

[1004] Step 5:

[1005] server:

[1006] The generated customized conversation script is sent to the terminal.

[1007] Input: Customized conversation script

[1008] Data processing: Sending HTTP requests via the network module

[1009] Output: Notification of completion of transmission to the terminal

[1010] Step 6:

[1011] Device:

[1012] The received conversation script is saved to the internal memory. After saving is complete, a message will be displayed saying "Conversation data has been updated."

[1013] Input: Conversation script sent from the server

[1014] Data processing: Saving to internal memory and notifying on the user interface

[1015] Output: "Conversation data updated" notification

[1016] Step 7:

[1017] User:

[1018] The user speaks to the device (e.g., "How's the gardening going lately?").

[1019] Input: User voice input

[1020] Data processing: None

[1021] Output: User's voice input

[1022] Step 8:

[1023] Device:

[1024] The user's voice is analyzed by the voice recognition module. The voice data is sent to the STT (Speech-to-Text) engine to convert it into text, and the emotion engine recognizes emotions from the voice and facial expressions.

[1025] Input: User's voice data

[1026] Data processing: Speech to text and emotion recognition

[1027] Output: User utterance text and recognized sentiment

[1028] Step 9:

[1029] Device:

[1030] Based on the user's emotions, the robot selects an appropriate response from a customized conversation script. For example, if the user is in good spirits, it selects positive conversational content, and if they are feeling down, it responds with encouraging words. The selected response is then converted into voice by a speech synthesis engine and returned to the user.

[1031] Input: User utterance text, recognized emotions, and a customized conversation script

[1032] Data processing: Response selection and speech synthesis based on spoken text and emotional data

[1033] Output: Audio response

[1034] Step 10:

[1035] server:

[1036] It continuously collects and analyzes user conversational and emotional data, and uses this data to periodically update the generative AI model with customized conversation scripts.

[1037] Input: User conversation data and emotion data

[1038] Data processing: Data analysis and updating of conversation scripts using generative AI models

[1039] Output: Updated conversation script

[1040] Step 11:

[1041] Device:

[1042] The updated conversation script is received and saved in the internal memory. After the update is complete, a message will be displayed stating, "The conversation data has been updated based on the latest information."

[1043] Input: Updated conversation script

[1044] Data processing: Saving to internal memory and notifying on the user interface

[1045] Output: Notification "Conversation data has been updated based on the latest information"

[1046] (Application example 2)

[1047] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1048] In modern society, loneliness and mental stress are serious problems for the elderly and for elderly factory workers. In particular, the psychological impact of monotonous work in factories and lack of communication with others cannot be overlooked. There are also concerns about dementia and health conditions. Therefore, there is a need for a system that helps elderly people stay healthy through appropriate communication, without feeling lonely, and allows them to work safely and efficiently.

[1049] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting profile information of the target person, means for transmitting the profile information of the target person to the server, means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for transmitting the generated customized conversation script to the terminal, means for recognizing the user's voice and responding based on the customized conversation script, and means for analyzing the user's emotional state using an emotion engine and providing appropriate conversation content. This not only enables elderly people to enjoy everyday conversations without feeling lonely, but also improves work efficiency and safety.

[1050] "Subject profile information" refers to basic information including the subject's name, age, hobbies, health status, past work experience, etc.

[1051] "Server" refers to a computer system for analyzing profile information, collecting related information, and generating and transmitting customized conversation scripts.

[1052] A "terminal" refers to a device for interacting with a user, which has the function of recognizing the user's voice and responding based on a conversation script.

[1053] "Customized conversation script" refers to personalized dialogue content generated based on the target person's profile information.

[1054] "Speech recognition" refers to the technology that analyzes a user's voice and converts it into text data.

[1055] An "emotion engine" refers to technology that analyzes a user's emotional state from data such as voice and facial expressions, and selects appropriate conversation content.

[1056] "Emotional state" refers to the emotions the user is feeling at the time, such as joy, sadness, interest, indifference, etc.

[1057] MODE FOR CARRYING OUT THE INVENTION

[1058] System configuration

[1059] The system for implementing this invention includes the following main components: a terminal for inputting profile information of a target person; a server for analyzing the profile information, collecting related information, and generating a customized conversation script; and a terminal for recognizing the user's voice and providing appropriate conversation content using an emotion engine.

[1060] Hardware and software used

[1061] Hardware

[1062] Smartphones and tablets (as audio input / output devices)

[1063] Factory robots (for work assistance and interaction)

[1064] software

[1065] OpenAI API (for dialogue generation)

[1066] SpeechRecognition

[1067] gTTS (Text to Speech)

[1068] mpg321 (audio playback)

[1069] Program processing flow and data processing

[1070] 1. Enter your profile information

[1071] Enter the subject's basic information (name, age, hobbies, health status, past work experience, etc.) on the terminal.

[1072] The terminal transmits the input information to the server.

[1073] 2. Analyzing profile information and generating customized conversation scripts

[1074] The server analyzes the received profile information and collects related information on the Internet.

[1075] The server generates a customized conversation script based on the collected information and sends it to the terminal.

[1076] 3. Speech and Emotion Recognition

[1077] When the user speaks to the device, it recognizes the voice and converts it into text data.

[1078] An emotion engine is used to analyze the user's emotional state from their voice and facial expressions.

[1079] 4. Providing customized responses

[1080] The device selects appropriate conversation content based on the analyzed emotional state.

[1081] The selected conversation content is converted from text to speech and provided to the user.

[1082] Specific examples

[1083] For example, suppose an elderly person working in a factory asks the device, "How's the gardening going lately?" The device first recognizes the user's voice and sends the information to the server. The server has already registered the person's profile information (for example, that gardening is a hobby), so it collects the latest information about gardening and generates a customized conversation script. Next, it uses an emotion engine to analyze the user's emotional state from their voice and facial expressions, and selects the most appropriate response to respond. For example, it provides specific advice such as, "A recent gardening technique that places importance on adjusting the soil."

[1084] Prompt Sentence Examples

[1085] The specific prompt for the generative AI model is as follows:

[1086] Generate a script for a conversation with Taro Sato, an elderly worker. User utterance: How's your gardening going lately?

[1087] Based on this prompt, the AI ​​model can generate a response that combines the user's profile information and relevant information.

[1088] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1089] Step 1:

[1090] Enter your profile information

[1091] The terminal (smartphone or tablet) provides an interface where the user or their family members can input the subject's name, age, hobbies, health status, past work experience, etc.

[1092] Input: Subject's profile information (e.g., "Name: Ichiro Takahashi, Age: 70, Hobbies: Gardening and Go, Health Status: Mild Dementia")

[1093] Output: The entered profile information is temporarily stored in the device's local memory.

[1094] Step 2:

[1095] Sending profile information

[1096] The terminal transmits the saved profile information to a server via the Internet.

[1097] Input: Profile information stored in local memory

[1098] Output: The profile information is sent to the server and received at the server side.

[1099] Step 3:

[1100] Analysis of profile information and collection of related information

[1101] The server analyzes the received profile information and collects related information from the Internet based on the subject's hobbies and interests.

[1102] Input: Received profile information

[1103] Output: A customized conversation script (e.g., "Latest gardening news," "New Go strategies") is generated based on the profile information.

[1104] Step 4:

[1105] Sending customized conversation scripts

[1106] The server transmits the generated customized conversation script to the terminal.

[1107] Input: Customized conversation script

[1108] Output: The customized conversation script is sent to the device and stored in the device's internal memory.

[1109] Step 5:

[1110] Voice Recognition

[1111] When a user speaks to the terminal, the terminal uses a speech recognition module (SpeechRecognition) to convert the voice data into text data.

[1112] Input: User speech (e.g., "How's the gardening going lately?")

[1113] Output: The converted text data (e.g., "How's your gardening going lately?")

[1114] Step 6:

[1115] emotion recognition

[1116] The device uses an emotion engine to analyze the user's emotional state from the voice data.

[1117] Input: Converted text data and audio data

[1118] Output: User's emotional state (e.g., "Interested")

[1119] Step 7:

[1120] Providing customized responses

[1121] The device selects appropriate responses based on the analyzed emotional state and a customized conversation script.

[1122] Input: User's emotional state, customized conversation script

[1123] Output: Text response (e.g., "Soil conditioning is becoming increasingly important as a gardening technique these days.")

[1124] Step 8:

[1125] Text-to-speech conversion and output

[1126] The terminal converts the selected response into speech using gTTS and outputs it as speech.

[1127] Input: Text data response

[1128] Output: Provided to the user as audio data (e.g., "Soil preparation is becoming increasingly important as a gardening technique these days.")

[1129] This will provide an environment where elderly people and elderly factory workers can work with peace of mind without feeling lonely.

[1130] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1131] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1132] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1133] [Third embodiment]

[1134] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1135] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1136] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1137] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1138] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1140] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1141] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1142] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1143] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1144] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1145] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1146] This invention provides a system that uses AI-equipped IoT devices to enable elderly people to enjoy everyday conversations without feeling lonely. Specifically, the following process is performed.

[1147] ---

[1148] Enter your profile information

[1149] Device:

[1150] Using the initial setup screen that appears when the device is started, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[1151] Examples:

[1152] The user's family member enters the following information into the terminal: "Taro, 75 years old, hobbies are gardening and Go, suffers from mild dementia."

[1153] ---

[1154] Sending profile information

[1155] Device:

[1156] The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[1157] Examples:

[1158] The device will notify you that "Taro's information has been sent to the server."

[1159] ---

[1160] Profile information analysis and customization

[1161] server:

[1162] The system analyzes the received profile information and gathers relevant information from the Internet (e.g., the latest gardening techniques or Go news). Based on this information, it generates a conversation script customized for each user.

[1163] Examples:

[1164] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro and generated a customized conversation script."

[1165] ---

[1166] Submitting a customization script

[1167] server:

[1168] The generated customized conversation script is sent to the terminal.

[1169] Device:

[1170] Receives customized conversation scripts and saves them in internal memory. Notifies the user or their family that "conversation data has been updated."

[1171] Examples:

[1172] The server sends "Gardening trivia" and the device notifies the user that "Taro's conversation data has been updated."

[1173] ---

[1174] Start a conversation

[1175] User:

[1176] The user speaks to the device (e.g., "How's the gardening going lately?").

[1177] Device:

[1178] The user's voice is analyzed by a voice recognition module, and based on the analysis results, an appropriate response is selected from a customized conversation script and responded to by voice.

[1179] Examples:

[1180] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[1181] ---

[1182] Continuous learning and updates

[1183] server:

[1184] It continuously collects, analyzes, and provides feedback on user conversation data, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[1185] Device:

[1186] Every time a new script is sent, it is received, stored in internal memory, and updated. It notifies you that "Conversation data has been updated based on new information."

[1187] Examples:

[1188] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[1189] ---

[1190] In this way, AI-equipped IoT devices can collect the latest information from the internet based on the target person's profile information and provide optimized conversations, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely, which is expected to reduce loneliness and slow the progression of dementia.

[1191] The processing flow will be explained below.

[1192] Step 1:

[1193] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is stored in the device's temporary memory.

[1194] Step 2:

[1195] Device: The entered profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[1196] Step 3:

[1197] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[1198] Step 4:

[1199] Server: Based on the extracted information, the server collects the latest information related to the target person from the Internet. For example, if the target person's hobby is gardening, it collects information on the latest gardening techniques and plants.

[1200] Step 5:

[1201] Server: Generates a customized conversation script based on the collected information, including content that is interesting and relevant to the target audience.

[1202] Step 6:

[1203] Server: Sends the generated customized conversation script to the device.

[1204] Step 7:

[1205] Device: Saves the received conversation script in its internal memory. Notifies the user or their family that "the conversation data has been updated."

[1206] Step 8:

[1207] User: The user speaks to the device (e.g., "How's your gardening going lately?"). The device analyzes the user's voice using a speech recognition module.

[1208] Step 9:

[1209] Terminal: Based on the analysis results, the terminal selects an appropriate response from a customized conversation script and returns the response to the user through a voice output device.

[1210] Step 10:

[1211] Server: Continuously collects user conversation data, analyzes it, and provides feedback. This data serves as the basis for further customization based on the user's interests.

[1212] Step 11:

[1213] Server: Periodically collects new information and topics, updates the customized conversation script, and sends the updated script to the device.

[1214] Step 12:

[1215] Device: Receives each new script sent, saves it in internal memory, and updates it. Once the update is complete, it notifies you that "Conversation data has been updated based on new information."

[1216] Through each step, the AI-enabled IoT device provides personalized conversations, allowing the elderly to enjoy their daily lives without feeling lonely.

[1217] Example 1

[1218] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1219] To enable elderly people to enjoy everyday conversations without feeling lonely, conversation content needs to be customized to each individual. However, conventional systems lack such customization, and often provide topics and information that do not interest elderly people. This leads to a decrease in satisfaction. In addition, there is no mechanism for continuously updating new information, which means that conversation content easily becomes outdated.

[1220] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1221] In this invention, the server includes a means for transmitting the target person's attribute information to the network, a means for collecting related information in a data store based on the transmitted attribute information and generating customized conversation data, and a means for transmitting the generated customized conversation data to the client device. This enables conversations using the latest information specific to the target person, allowing elderly people to participate in conversations with interest without feeling lonely. Furthermore, by continuously collecting new information and updating the conversation content, it is possible to always provide fresh topics of conversation.

[1222] "Subject attribute information" refers to personal information such as the subject's name, age, hobbies, and health status.

[1223] "Data store" refers to a system or database for storing information and data collected from the Internet.

[1224] "Customized conversation data" refers to a conversation script that is generated based on the individual attribute information of the subject and tailored to the subject.

[1225] "Client device" refers to an AI-enabled IoT device or terminal used by the subject.

[1226] "Network" refers to the Internet and other communication means that provide paths for transmitting and receiving data.

[1227] A "conversation script" refers to data that describes the content of utterances and responses that the system uses in dialogue with the user.

[1228] This invention provides a system that uses AI-equipped IoT devices to help elderly people enjoy everyday conversations without feeling lonely. The entire system is realized by collecting profile information of the target person and providing customized conversation data to the user.

[1229] Hardware and software used

[1230] Terminal: An AI-enabled IoT device placed close to the user. This terminal displays the necessary initial setup screens and allows the user or their family members to enter their profile information.

[1231] Server: Refers to the back-end system that analyzes data and generates customized conversation scripts. The server is connected to the Internet and collects various information.

[1232] Speech Recognition Module: For example, using Google Speech-to-Text. Used to convert the user's speech into text.

[1233] Generative AI models, such as GPT-3, are used to generate customized conversation scripts based on collected information.

[1234] A speech synthesis module, e.g., using Google Text-to-Speech, is used to convert the generated text into speech and respond to the user.

[1235] Explanation of program processing

[1236] Device:

[1237] 1. When the device starts up, the initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored on the device.

[1238] 2. The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[1239] server:

[1240] 1. The server analyzes the received profile information and collects relevant information from a data store (e.g., the latest gardening techniques or Go news). Based on this information, a customized conversation script is generated for each user. This is generated using a generative AI model.

[1241] 2. Send the generated customized conversation script to the terminal.

[1242] Device:

[1243] 1. The device saves the received conversation script in its internal memory and displays "Conversation data updated" on the screen.

[1244] 2. When the user speaks to the device, the device analyzes it using the voice recognition module, selects a response based on the customized conversation data, and responds to the user using the voice synthesis module.

[1245] Specific examples

[1246] Enter your profile information:

[1247] The user's family member enters information into the device, such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." This information is temporarily stored in the device.

[1248] Send profile information:

[1249] When the device is connected to a network such as Wi-Fi, it sends the saved profile information to the server as a POST request. If the request is successful, a pop-up message will appear saying, "Taro's information has been sent to the server."

[1250] Profile information analysis and customization:

[1251] The server analyzes the profile information "Taro, 75 years old, gardening, Go, mild dementia," collects relevant latest gardening techniques and recent Go news, and generates a conversation script using a generative AI model.

[1252] Send a customized conversation script:

[1253] The server sends the conversation script data in JSON format, and the device receives this data, parses it, and saves it in its internal memory. The updated content is notified with a message saying "The conversation data has been updated."

[1254] Start a conversation:

[1255] When a user speaks to the device, "How's your gardening going these days?", the device converts the speech into text using a speech recognition module, and then uses a generative AI model to respond, "The latest gardening techniques place great importance on soil preparation," which is then played back as audio using a speech synthesis module.

[1256] Example prompt sentence:

[1257] "Please tell me the latest gardening knowledge for 75-year-old Taro."

[1258] "Tell me the latest news about Go"

[1259] "How's your gardening going lately?"

[1260] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1261] Step 1:

[1262] Enter and save your profile information

[1263] Device:

[1264] When the device starts up, an initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[1265] Input: The user enters profile information as text.

[1266] Output: The entered profile information is saved in the device's temporary memory.

[1267] Specific operation: A user's family member enters information such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." The device saves this information in text format.

[1268] Step 2:

[1269] Sending profile information

[1270] Device:

[1271] The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[1272] Input: Profile information stored on your device.

[1273] Output: The profile information is sent to the server.

[1274] Specific operation: The device sends profile information in the form of a POST request via a network such as Wi-Fi. A message pops up when the transmission is successful.

[1275] Step 3:

[1276] Profile information analysis and customization

[1277] server:

[1278] The server analyzes the received profile information and gathers relevant information from a data store, such as the latest gardening techniques or Go news, and then uses a generative AI model to generate a customized conversation script based on the collected information.

[1279] Input: Profile information stored on the server.

[1280] Output: A customized conversation script.

[1281] Specific operation: The server analyzes the information "Taro, 75 years old, hobbies are gardening and Go, mild dementia," collects related information using web scraping and APIs, and generates a conversation script using a generative AI model (e.g., GPT-3).

[1282] Step 4:

[1283] Sending customized conversation scripts

[1284] server:

[1285] The generated customized conversation script is sent to the terminal.

[1286] Device:

[1287] The terminal stores the received conversation script in its internal memory and displays the message "Conversation data has been updated" on the screen.

[1288] Input: A customized conversation script generated by the server.

[1289] Output: The customized conversation script is saved in the device's memory.

[1290] Specific operation: The server sends script data in JSON format, the device receives it, parses it, saves it in its internal memory, and displays a pop-up message saying "Conversation data has been updated."

[1291] Step 5:

[1292] Start a conversation

[1293] User:

[1294] The user provides the device with a question or topic (e.g., "How's your gardening going lately?").

[1295] Device:

[1296] The device uses a speech recognition module to convert the user's speech into text, generates appropriate responses based on customized conversation data, and responds to the user using a speech synthesis module.

[1297] Input: User's voice input.

[1298] Output: Audio response from the device.

[1299] What it does: When a user says, "How's your gardening going these days?", the device converts the speech to text and selects a response from a customized conversation script: "Modern gardening techniques emphasize soil conditioning," and plays it back as audio.

[1300] Step 6:

[1301] Continuous learning and updates

[1302] server:

[1303] The server continuously collects user conversation data, analyzes it, and provides feedback, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[1304] Device:

[1305] Every time a new conversation script is sent from the server, it is received, saved in internal memory, and updated. When the update is complete, a message will appear saying, "The conversation data has been updated based on the new information."

[1306] Input: User conversation data and new information.

[1307] Output: Updated customized conversation script.

[1308] Specific operation: The server analyzes the user's conversation history, collects new gardening information, the latest news on Go, etc., and sends an updated conversation script to the device. The device receives this and notifies the user that "the conversation data has been updated with the latest information."

[1309] (Application example 1)

[1310] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1311] There is a need for systems that can not only reduce loneliness among the elderly and provide daily communication, but also monitor their health status and respond to emergency situations. There is also a need for systems that can provide customized conversations using information specific to the elderly and the latest topics.

[1312] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1313] In this invention, the server includes means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for collecting sensor data for monitoring health status, and means for making an emergency call if an abnormality occurs. This not only reduces the sense of loneliness of elderly people and provides communication, but also enables health status monitoring and prompt response in emergencies.

[1314] "Subject profile information" is basic information about an individual, such as the subject's name, age, hobbies, and health status.

[1315] The "means for inputting profile information" refers to an interface that allows the subject or their family to input profile information using a terminal.

[1316] The "means for transmitting profile information to a server" refers to a communication means for transmitting profile information from a terminal to a server via the Internet.

[1317] A "customized conversation script" is an individually optimized conversation scenario that is generated by collecting relevant information on the Internet based on the target person's profile information.

[1318] The "means for generating a customized conversation script" is a processing means for analyzing profile information and creating a conversation script based on the latest related information on the Internet.

[1319] The "means for recognizing and responding to the user's voice" is a function for analyzing the voice uttered by the user and selecting and responding to an appropriate response based on a customized conversation script.

[1320] "Sensor data for monitoring health status" refers to information obtained from sensors that collect physiological and activity data of subjects, such as heart rate, number of steps, and fall detection.

[1321] The "means for collecting sensor data" refers to devices and software for collecting physiological data and activity data through sensors, which are necessary for monitoring the health status of a subject.

[1322] The "means for making an emergency report when an abnormality occurs" is a means for making a report to a pre-set emergency contact when an abnormality is detected based on collected sensor data.

[1323] "Means for gathering information from the Internet" refers to the Internet browsing ability to gather relevant and up-to-date information and news on the Internet based on the profile information.

[1324] This invention provides a system that can reduce loneliness among the elderly, monitor their health status, and respond to emergencies. The system provides customized conversations based on the target person's profile information and makes an emergency call when an abnormality in their health status is detected.

[1325] Program processing explanation

[1326] Hardware and Software Configuration

[1327] Hardware:

[1328] Robot device (equipped with microphone, speaker, and sensor)

[1329] Software and Libraries:

[1330] requests library: Used for server communication.

[1331] speech_recognition library: Used for speech recognition.

[1332] pyttsx3 library: Used for speech synthesis.

[1333] Processing explanation

[1334] (1) Enter and submit profile information

[1335] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device and then sent to the server. The server receives the information and sends a confirmation message back to the device.

[1336] Examples:

[1337] The user's family member enters information such as "Mr. Tanaka, 75 years old, hobbies are gardening and Go, mild dementia" into the terminal and sends it to the server.

[1338] (2) Analyzing profile information and generating conversation scripts

[1339] The server analyzes the received profile information and collects relevant information from the Internet (such as the latest gardening techniques or Go news), and generates a conversation script customized for each user based on this information.

[1340] Examples:

[1341] The server "collects the latest gardening trivia and Go game information for Tanaka and generates a customized conversation script."

[1342] (3) Sending customized conversation scripts

[1343] The server sends the generated customized conversation script to the device. The device receives the conversation script and saves it in its internal memory. Once the saving is complete, the device displays a notification to the user that the update is complete.

[1344] Examples:

[1345] The server sends "Gardening trivia" and the device notifies the user that "Tanaka's conversation data has been updated."

[1346] (4) Communication using voice recognition

[1347] When a user speaks to the terminal, the terminal analyzes the user's voice using a voice recognition module, selects an appropriate response from a customized conversation script, and responds verbally.

[1348] Examples:

[1349] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[1350] (5) Health monitoring and emergency notification

[1351] The device's built-in sensors constantly monitor the user's health. If an abnormality is detected (e.g., an abnormal heart rate or a fall), the device immediately sends the abnormal data to the server, which then notifies the user's emergency contacts.

[1352] Examples:

[1353] The device notifies the server that an abnormal heart rate has been detected, and the server then calls a pre-set emergency number.

[1354] (6) Continuous learning and updating

[1355] The server continuously collects and analyzes the user's conversation data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory. The user is notified that the conversation data has been updated with new information.

[1356] Examples:

[1357] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[1358] Example prompts for generative AI models

[1359] Generate customized conversation scripts about the latest gardening news and Go topics based on the user's profile information. For example, respond to the question "Tell me about gardening" with "Soil preparation is becoming increasingly important as a modern gardening technique."

[1360] In this way, the embodiments of the invention allow elderly people to enjoy everyday conversations without feeling lonely, while their health condition is monitored, and if an abnormality occurs, a prompt response can be made.

[1361] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1362] Step 1:

[1363] Enter your profile information

[1364] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device. The input is done using the user interface, and the input data is stored in the device's internal storage.

[1365] Input: Profile information such as name, age, hobbies, health status, etc.

[1366] Output: Profile information temporarily saved on the device

[1367] Step 2:

[1368] Sending profile information

[1369] The device will then send the saved profile information to the server via the internet, and once the transfer is complete, a confirmation message will appear saying "Profile information has been sent to the server."

[1370] Input: Profile information temporarily saved on the device

[1371] Output: Profile information sent to server, confirmation message to user

[1372] Step 3:

[1373] Analyzing profile information

[1374] The server analyzes the received profile information and collects relevant information on the Internet (e.g., the latest gardening techniques or Go news). Based on the profile information, the analysis generates relevant search queries and retrieves relevant information.

[1375] Input: Profile information sent to the server

[1376] Output: Analysis results and related collected information

[1377] Step 4:

[1378] Generate customized conversation scripts

[1379] The server generates a customized conversation script for each user based on the collected relevant information. It uses a generative AI model to script a conversation scenario based on the profile information and collected information.

[1380] Input: Analysis results and related collected information

[1381] Output: Customized conversation script

[1382] Step 5:

[1383] Sending customized conversation scripts

[1384] The server sends the generated customized conversation script to the terminal. The terminal receives the conversation script and saves it in its internal memory. Once the saving is complete, a notification that "Conversation data has been updated" is displayed to the user.

[1385] Input: Customized conversation script

[1386] Output: Conversation script saved on the device, update completion notification to the user

[1387] Step 6:

[1388] User voice recognition

[1389] When a user speaks to the device, the device uses a voice recognition module to analyze the user's voice, converting the voice data into text data, and then selects an appropriate response from a customized conversation script based on the text.

[1390] Input: User voice input

[1391] Output: Recognized text data, selected response

[1392] Step 7:

[1393] Voice response

[1394] The device responds to the user vocally based on the selected customized conversation script, which is generated using a speech synthesis module and output audibly through the speaker.

[1395] Input: Selected Response

[1396] Output: Audio response

[1397] Step 8:

[1398] Health monitoring

[1399] The device's built-in sensors constantly monitor the subject's health condition, and the collected sensor data is periodically sent to a server, which immediately notifies the user if any abnormalities are detected.

[1400] Input: Sensor data (heart rate, steps, fall detection, etc.)

[1401] Output: Health data sent to the server, anomaly detection notification

[1402] Step 9:

[1403] emergency call

[1404] If an abnormality is detected, the device or server will make an emergency call, and the server will send an emergency notification to the configured emergency contacts, reporting the specific situation.

[1405] Input: Anomaly detection notification

[1406] Output: Notification to emergency number

[1407] Step 10:

[1408] Continuous learning and updates

[1409] The server continuously collects and analyzes the user's conversation data and health data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory.

[1410] Input: conversation data, health data

[1411] Output: Updated conversation script, new information notification

[1412] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1413] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Detailed embodiments of this system are described below.

[1414] ---

[1415] Enter your profile information

[1416] Device:

[1417] On the initial setup screen of the device, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). This information is temporarily stored in the device.

[1418] Examples:

[1419] The user's family member enters the following information: "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia."

[1420] ---

[1421] Sending profile information

[1422] Device:

[1423] Your profile information will be sent to the server via the Internet. When the sending is complete, the message "Profile information has been sent to the server" will be displayed.

[1424] Examples:

[1425] The device will notify you that "Taro Sato's information has been sent to the server."

[1426] ---

[1427] Profile information analysis and customization

[1428] server:

[1429] The system analyzes the received profile information, gathers relevant information on the Internet, and generates a customized conversation script based on the target person's hobbies and interests.

[1430] Examples:

[1431] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro Sato and generated a customized conversation script."

[1432] ---

[1433] Submitting a customization script

[1434] server:

[1435] The generated customized conversation script is sent to the terminal.

[1436] Device:

[1437] The received conversation script is saved in the internal memory and a message is displayed stating "Conversation data has been updated."

[1438] Examples:

[1439] The server sends "Gardening trivia" and the device notifies the user that "Taro Sato's conversation data has been updated."

[1440] ---

[1441] Speech and Emotion Recognition

[1442] User:

[1443] The user speaks to the device (e.g., "How's the gardening going lately?").

[1444] Device:

[1445] The user's voice is analyzed using a voice recognition module, and the emotion engine is used to recognize the user's emotions from their voice and facial expressions.

[1446] Examples:

[1447] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil preparation is becoming important," while determining whether the user looks happy.

[1448] ---

[1449] Providing customized responses

[1450] Device:

[1451] The system selects appropriate responses based on the user's emotions. For example, if the user is in good spirits, it selects more positive conversation content, and if they are feeling down, it offers encouraging words.

[1452] Examples:

[1453] If the user is feeling unwell, the device will gently ask, "You seem to be feeling unwell lately. Is there anything that is bothering you?"

[1454] ---

[1455] Continuous learning and updates

[1456] server:

[1457] By continuously collecting and analyzing user conversational and emotional data, the system learns the user's interests and emotional patterns and regularly updates the conversation script.

[1458] Device:

[1459] When a new script is received, it is saved in the internal memory and updated. Once the update is complete, a message will appear saying "Conversation data has been updated based on new information."

[1460] Examples:

[1461] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[1462] ---

[1463] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[1464] The processing flow will be explained below.

[1465] Step 1:

[1466] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is temporarily stored in the device.

[1467] Step 2:

[1468] Device: The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[1469] Step 3:

[1470] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[1471] Step 4:

[1472] Server: Based on the extracted information, collects the latest information relevant to the target person from the Internet (e.g., the latest gardening techniques or Go news).

[1473] Step 5:

[1474] Server: Based on the collected information, a customized conversation script is generated for each target audience, including content that will interest the target audience and relevant information.

[1475] Step 6:

[1476] Server: Sends the generated customized conversation script to the device.

[1477] Step 7:

[1478] Device: Saves the received conversation script in its internal memory and notifies the user or their family that "the conversation data has been updated."

[1479] Step 8:

[1480] User: The user speaks to the device (e.g., "How's the gardening going lately?").

[1481] Step 9:

[1482] Device: The user's voice is analyzed by the speech recognition module. Then, the emotion engine recognizes emotions from the user's voice and facial expressions.

[1483] Step 10:

[1484] The device: Based on the recognized emotion, it selects an appropriate response from a customized conversation script. If the user is in good spirits, it selects positive responses, and if they are not, it offers encouraging words.

[1485] Step 11:

[1486] Terminal: The selected response is returned to the user through a voice output device. For example, "Soil conditioning is becoming increasingly important in modern gardening techniques."

[1487] Step 12:

[1488] Server: Continuously collects and analyzes user conversations and recognized emotional data, thereby learning the user's emotional patterns and interests.

[1489] Step 13:

[1490] Server: Periodically gathers new information and topics from the Internet and updates the customized conversation script.

[1491] Step 14:

[1492] Server: Sends the updated script to the device.

[1493] Step 15:

[1494] Device: When a new script is received, it is saved in internal memory and a notification is sent to the user or their family member stating, "Conversation data has been updated based on new information."

[1495] Through each step, the AI-powered IoT device can provide individually optimized conversations based on the target person's profile information and emotional data, allowing the elderly to enjoy their daily lives without feeling lonely. Furthermore, by combining it with an emotion engine, the device can understand the user's emotional state and provide appropriate responses, enabling deeper communication.

[1496] Example 2

[1497] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1498] Today's elderly often feel social isolation and loneliness, which can accelerate the progression of dementia. However, increasing opportunities for these elderly people to interact through everyday conversations could contribute to reducing loneliness and maintaining cognitive function. However, current technology is limited in systems that provide personalized conversations based on individual hobbies and interests, and there are particularly few systems that can recognize emotions. Therefore, there is a need for a system that provides conversations tailored to each individual elderly person and enables deep interaction.

[1499] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1500] In this invention, the server includes means for analyzing the profile information of the subject, collecting related information on the Internet, and generating a customized conversation script using a generative AI model, means for transmitting the customized conversation script to the terminal, and means for the terminal to analyze the user's voice using a voice recognition module and recognize the user's emotions using an emotion engine. This allows the user to enjoy conversations based on their hobbies and interests, and furthermore, receives appropriate responses through emotion recognition, which is expected to reduce feelings of loneliness among the elderly and slow the progression of dementia.

[1501] "Profile information" refers to personal information such as the subject's name, age, hobbies, and health status.

[1502] "Server" refers to a computer system that receives and analyzes profile information, gathers relevant information from the Internet, and generates customized conversation scripts.

[1503] "Conversation script" refers to customized conversation content generated using a generative AI model based on the target person's profile information and related information.

[1504] "Terminal" refers to a device used by a user to input profile information, receive customized conversation scripts, recognize voice and emotions, and display responses.

[1505] "Speech recognition module" refers to software or hardware that has the function of analyzing a user's voice and converting it into text.

[1506] "Emotion engine" refers to a system that includes technology to identify and analyze a user's emotions from voice and facial expression data.

[1507] A "generative AI model" refers to an artificial intelligence algorithm or machine learning model that generates customized conversation scripts based on input data.

[1508] "Speech synthesis engine" refers to a system that includes technology for converting text data into natural-sounding speech.

[1509] "Collecting related information on the Internet" refers to the server obtaining information related to the subject's hobbies and interests based on the profile information through methods such as web scraping or API calls.

[1510] "Response" refers to the voice or text response that the terminal returns to the user based on the analysis results of the voice recognition module and emotion engine.

[1511] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Specific embodiments for carrying out this invention are described below.

[1512] The main components of the system are a device, a server, a generative AI model, an emotion engine, and related software modules. The device has an interface for inputting and sending user profile information, which is then sent to the server. The server generates a customized conversation script based on this profile information. The generative AI model is used for generation.

[1513] Device:

[1514] The device is initially set up by the user or their family, and profile information (such as name, age, hobbies, and health status) is entered. The entered information is temporarily stored in the device's memory. Examples of devices include tablets and smart home devices. These devices are equipped with touch panels and keyboards, making it easy to enter information.

[1515] Examples:

[1516] The user's family member enters information such as "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" into the tablet.

[1517] server:

[1518] The server receives profile information sent from the device via the internet. Based on the received information, the server collects the latest information related to the target person's hobbies and interests from the internet. This information is collected using web scraping and API calls. A customized conversation script is then generated using a generative AI model and sent to the device.

[1519] Examples:

[1520] The prompt sentence "75-year-old male, hobbies are gardening and Go, mild dementia" is input into the generative AI model, which then performs processes such as "collecting the latest gardening trivia and Go game information for Taro Sato and generating a customized conversation script."

[1521] Device:

[1522] The device receives the customized conversation script sent from the server and saves it in its internal memory. After saving, it notifies the user that "the conversation data has been updated." The device then analyzes the user's voice using a voice recognition module and uses an emotion engine to recognize the user's emotions from their voice and facial expressions.

[1523] Examples:

[1524] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil adjustment is becoming important," while determining whether the user looks happy.

[1525] Emotion Engine:

[1526] The emotion engine includes technology that analyzes the user's emotions from voice and facial expression data. If the device determines that the user is not in good spirits, the emotion engine selects an appropriate response.

[1527] Examples:

[1528] If the user is not feeling well, gently ask, "You seem to be feeling unwell lately. Is there anything that's bothering you?"

[1529] The server continuously collects and analyzes the user's conversational and emotional data to learn about the user's interests and emotional patterns, allowing it to periodically update the conversation script and provide more personalized conversations.

[1530] Examples:

[1531] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[1532] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[1533] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1534] Step 1:

[1535] Device:

[1536] On the initial setup screen, the user or their family member enters the target person's profile information. For example, they might enter "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" on the tablet's touch panel. The entered data is temporarily stored in the device's memory.

[1537] Input: Profile information (name, age, hobbies, health status)

[1538] Data processing: Input via touch panel or keyboard

[1539] Output: Profile information temporarily stored in the device memory

[1540] Step 2:

[1541] Device:

[1542] Your profile information will be sent to the server via the Internet. Once the sending is complete, the message "Profile information sent to server" will be displayed on the screen.

[1543] Input: Profile information stored in the device memory

[1544] Data processing: Sending HTTP requests via the network module

[1545] Output: Notification of completion of sending profile information to the server

[1546] Step 3:

[1547] server:

[1548] The submitted profile information is analyzed. For example, "75 years old, male, hobbies are gardening and Go, mild dementia." Related information is then collected from the internet. Web scraping and APIs are used to collect this information.

[1549] Input: Submitted profile information

[1550] Data processing: Analysis of profile information and collection of related information on the Internet

[1551] Output: Related information

[1552] Step 4:

[1553] server:

[1554] A generative AI model is used to generate a customized conversation script based on the collected relevant information. The prompt sentence, "75-year-old male, hobbies are gardening and Go, mild dementia," is input into the generative AI model.

[1555] Input: Profile information and related information

[1556] Data processing: Generating conversation scripts using generative AI models

[1557] Output: Customized conversation script

[1558] Step 5:

[1559] server:

[1560] The generated customized conversation script is sent to the terminal.

[1561] Input: Customized conversation script

[1562] Data processing: Sending HTTP requests via the network module

[1563] Output: Notification of completion of transmission to the terminal

[1564] Step 6:

[1565] Device:

[1566] The received conversation script is saved to the internal memory. After saving is complete, a message will be displayed saying "Conversation data has been updated."

[1567] Input: Conversation script sent from the server

[1568] Data processing: Saving to internal memory and notifying on the user interface

[1569] Output: "Conversation data updated" notification

[1570] Step 7:

[1571] User:

[1572] The user speaks to the device (e.g., "How's the gardening going lately?").

[1573] Input: User voice input

[1574] Data processing: None

[1575] Output: User's voice input

[1576] Step 8:

[1577] Device:

[1578] The user's voice is analyzed by the voice recognition module. The voice data is sent to the STT (Speech-to-Text) engine to convert it into text, and the emotion engine recognizes emotions from the voice and facial expressions.

[1579] Input: User's voice data

[1580] Data processing: Speech to text and emotion recognition

[1581] Output: User utterance text and recognized sentiment

[1582] Step 9:

[1583] Device:

[1584] Based on the user's emotions, the robot selects an appropriate response from a customized conversation script. For example, if the user is in good spirits, it selects positive conversational content, and if they are feeling down, it responds with encouraging words. The selected response is then converted into voice by a speech synthesis engine and returned to the user.

[1585] Input: User utterance text, recognized emotions, and a customized conversation script

[1586] Data processing: Response selection and speech synthesis based on spoken text and emotional data

[1587] Output: Audio response

[1588] Step 10:

[1589] server:

[1590] It continuously collects and analyzes user conversational and emotional data, and uses this data to periodically update the generative AI model with customized conversation scripts.

[1591] Input: User conversation data and emotion data

[1592] Data processing: Data analysis and updating of conversation scripts using generative AI models

[1593] Output: Updated conversation script

[1594] Step 11:

[1595] Device:

[1596] The updated conversation script is received and saved in the internal memory. After the update is complete, a message will be displayed stating, "The conversation data has been updated based on the latest information."

[1597] Input: Updated conversation script

[1598] Data processing: Saving to internal memory and notifying on the user interface

[1599] Output: Notification "Conversation data has been updated based on the latest information"

[1600] (Application example 2)

[1601] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1602] In modern society, loneliness and mental stress are serious problems for the elderly and for elderly factory workers. In particular, the psychological impact of monotonous work in factories and lack of communication with others cannot be overlooked. There are also concerns about dementia and health conditions. Therefore, there is a need for a system that helps elderly people stay healthy through appropriate communication, without feeling lonely, and allows them to work safely and efficiently.

[1603] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting profile information of the target person, means for transmitting the profile information of the target person to the server, means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for transmitting the generated customized conversation script to the terminal, means for recognizing the user's voice and responding based on the customized conversation script, and means for analyzing the user's emotional state using an emotion engine and providing appropriate conversation content. This not only enables elderly people to enjoy everyday conversations without feeling lonely, but also improves work efficiency and safety.

[1604] "Subject profile information" refers to basic information including the subject's name, age, hobbies, health status, past work experience, etc.

[1605] "Server" refers to a computer system for analyzing profile information, collecting related information, and generating and transmitting customized conversation scripts.

[1606] A "terminal" refers to a device for interacting with a user, which has the function of recognizing the user's voice and responding based on a conversation script.

[1607] "Customized conversation script" refers to personalized dialogue content generated based on the target person's profile information.

[1608] "Speech recognition" refers to the technology that analyzes a user's voice and converts it into text data.

[1609] An "emotion engine" refers to technology that analyzes a user's emotional state from data such as voice and facial expressions, and selects appropriate conversation content.

[1610] "Emotional state" refers to the emotions the user is feeling at the time, such as joy, sadness, interest, indifference, etc.

[1611] MODE FOR CARRYING OUT THE INVENTION

[1612] System configuration

[1613] The system for implementing this invention includes the following main components: a terminal for inputting profile information of a target person; a server for analyzing the profile information, collecting related information, and generating a customized conversation script; and a terminal for recognizing the user's voice and providing appropriate conversation content using an emotion engine.

[1614] Hardware and software used

[1615] Hardware

[1616] Smartphones and tablets (as audio input / output devices)

[1617] Factory robots (for work assistance and interaction)

[1618] software

[1619] OpenAI API (for dialogue generation)

[1620] SpeechRecognition

[1621] gTTS (Text to Speech)

[1622] mpg321 (audio playback)

[1623] Program processing flow and data processing

[1624] 1. Enter your profile information

[1625] Enter the subject's basic information (name, age, hobbies, health status, past work experience, etc.) on the terminal.

[1626] The terminal transmits the input information to the server.

[1627] 2. Analyzing profile information and generating customized conversation scripts

[1628] The server analyzes the received profile information and collects related information on the Internet.

[1629] The server generates a customized conversation script based on the collected information and sends it to the terminal.

[1630] 3. Speech and Emotion Recognition

[1631] When the user speaks to the device, it recognizes the voice and converts it into text data.

[1632] An emotion engine is used to analyze the user's emotional state from their voice and facial expressions.

[1633] 4. Providing customized responses

[1634] The device selects appropriate conversation content based on the analyzed emotional state.

[1635] The selected conversation content is converted from text to speech and provided to the user.

[1636] Specific examples

[1637] For example, suppose an elderly person working in a factory asks the device, "How's the gardening going lately?" The device first recognizes the user's voice and sends the information to the server. The server has already registered the person's profile information (for example, that gardening is a hobby), so it collects the latest information about gardening and generates a customized conversation script. Next, it uses an emotion engine to analyze the user's emotional state from their voice and facial expressions, and selects the most appropriate response to respond. For example, it provides specific advice such as, "A recent gardening technique that places importance on adjusting the soil."

[1638] Prompt Sentence Examples

[1639] The specific prompt for the generative AI model is as follows:

[1640] Generate a script for a conversation with Taro Sato, an elderly worker. User utterance: How's your gardening going lately?

[1641] Based on this prompt, the AI ​​model can generate a response that combines the user's profile information and relevant information.

[1642] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1643] Step 1:

[1644] Enter your profile information

[1645] The terminal (smartphone or tablet) provides an interface where the user or their family members can input the subject's name, age, hobbies, health status, past work experience, etc.

[1646] Input: Subject's profile information (e.g., "Name: Ichiro Takahashi, Age: 70, Hobbies: Gardening and Go, Health Status: Mild Dementia")

[1647] Output: The entered profile information is temporarily stored in the device's local memory.

[1648] Step 2:

[1649] Sending profile information

[1650] The terminal transmits the saved profile information to a server via the Internet.

[1651] Input: Profile information stored in local memory

[1652] Output: The profile information is sent to the server and received at the server side.

[1653] Step 3:

[1654] Analysis of profile information and collection of related information

[1655] The server analyzes the received profile information and collects related information from the Internet based on the subject's hobbies and interests.

[1656] Input: Received profile information

[1657] Output: A customized conversation script (e.g., "Latest gardening news," "New Go strategies") is generated based on the profile information.

[1658] Step 4:

[1659] Sending customized conversation scripts

[1660] The server transmits the generated customized conversation script to the terminal.

[1661] Input: Customized conversation script

[1662] Output: The customized conversation script is sent to the device and stored in the device's internal memory.

[1663] Step 5:

[1664] Voice Recognition

[1665] When a user speaks to the terminal, the terminal uses a speech recognition module (SpeechRecognition) to convert the voice data into text data.

[1666] Input: User speech (e.g., "How's the gardening going lately?")

[1667] Output: The converted text data (e.g., "How's your gardening going lately?")

[1668] Step 6:

[1669] emotion recognition

[1670] The device uses an emotion engine to analyze the user's emotional state from the voice data.

[1671] Input: Converted text data and audio data

[1672] Output: User's emotional state (e.g., "Interested")

[1673] Step 7:

[1674] Providing customized responses

[1675] The device selects appropriate responses based on the analyzed emotional state and a customized conversation script.

[1676] Input: User's emotional state, customized conversation script

[1677] Output: Text response (e.g., "Soil conditioning is becoming increasingly important as a gardening technique these days.")

[1678] Step 8:

[1679] Text-to-speech conversion and output

[1680] The terminal converts the selected response into speech using gTTS and outputs it as speech.

[1681] Input: Text data response

[1682] Output: Provided to the user as audio data (e.g., "Soil preparation is becoming increasingly important as a gardening technique these days.")

[1683] This will provide an environment where elderly people and elderly factory workers can work with peace of mind without feeling lonely.

[1684] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1685] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1686] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1687] [Fourth embodiment]

[1688] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1689] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1690] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1691] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1692] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1693] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1694] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1695] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1696] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1697] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1698] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1699] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1700] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1701] This invention provides a system that uses AI-equipped IoT devices to enable elderly people to enjoy everyday conversations without feeling lonely. Specifically, the following process is performed.

[1702] ---

[1703] Enter your profile information

[1704] Device:

[1705] Using the initial setup screen that appears when the device is started, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[1706] Examples:

[1707] The user's family member enters the following information into the terminal: "Taro, 75 years old, hobbies are gardening and Go, suffers from mild dementia."

[1708] ---

[1709] Sending profile information

[1710] Device:

[1711] The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[1712] Examples:

[1713] The device will notify you that "Taro's information has been sent to the server."

[1714] ---

[1715] Profile information analysis and customization

[1716] server:

[1717] The system analyzes the received profile information and gathers relevant information from the Internet (e.g., the latest gardening techniques or Go news). Based on this information, it generates a conversation script customized for each user.

[1718] Examples:

[1719] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro and generated a customized conversation script."

[1720] ---

[1721] Submitting a customization script

[1722] server:

[1723] The generated customized conversation script is sent to the terminal.

[1724] Device:

[1725] Receives customized conversation scripts and saves them in internal memory. Notifies the user or their family that "conversation data has been updated."

[1726] Examples:

[1727] The server sends "Gardening trivia" and the device notifies the user that "Taro's conversation data has been updated."

[1728] ---

[1729] Start a conversation

[1730] User:

[1731] The user speaks to the device (e.g., "How's the gardening going lately?").

[1732] Device:

[1733] The user's voice is analyzed by a voice recognition module, and based on the analysis results, an appropriate response is selected from a customized conversation script and responded to by voice.

[1734] Examples:

[1735] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[1736] ---

[1737] Continuous learning and updates

[1738] server:

[1739] It continuously collects, analyzes, and provides feedback on user conversation data, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[1740] Device:

[1741] Every time a new script is sent, it is received, stored in internal memory, and updated. It notifies you that "Conversation data has been updated based on new information."

[1742] Examples:

[1743] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[1744] ---

[1745] In this way, AI-equipped IoT devices can collect the latest information from the internet based on the target person's profile information and provide optimized conversations, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely, which is expected to reduce loneliness and slow the progression of dementia.

[1746] The processing flow will be explained below.

[1747] Step 1:

[1748] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is stored in the device's temporary memory.

[1749] Step 2:

[1750] Device: The entered profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[1751] Step 3:

[1752] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[1753] Step 4:

[1754] Server: Based on the extracted information, the server collects the latest information related to the target person from the Internet. For example, if the target person's hobby is gardening, it collects information on the latest gardening techniques and plants.

[1755] Step 5:

[1756] Server: Generates a customized conversation script based on the collected information, including content that is interesting and relevant to the target audience.

[1757] Step 6:

[1758] Server: Sends the generated customized conversation script to the device.

[1759] Step 7:

[1760] Device: Saves the received conversation script in its internal memory. Notifies the user or their family that "the conversation data has been updated."

[1761] Step 8:

[1762] User: The user speaks to the device (e.g., "How's your gardening going lately?"). The device analyzes the user's voice using a speech recognition module.

[1763] Step 9:

[1764] Terminal: Based on the analysis results, the terminal selects an appropriate response from a customized conversation script and returns the response to the user through a voice output device.

[1765] Step 10:

[1766] Server: Continuously collects user conversation data, analyzes it, and provides feedback. This data serves as the basis for further customization based on the user's interests.

[1767] Step 11:

[1768] Server: Periodically collects new information and topics, updates the customized conversation script, and sends the updated script to the device.

[1769] Step 12:

[1770] Device: Receives each new script sent, saves it in internal memory, and updates it. Once the update is complete, it notifies you that "Conversation data has been updated based on new information."

[1771] Through each step, the AI-enabled IoT device provides personalized conversations, allowing the elderly to enjoy their daily lives without feeling lonely.

[1772] Example 1

[1773] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1774] To enable elderly people to enjoy everyday conversations without feeling lonely, conversation content needs to be customized to each individual. However, conventional systems lack such customization, and often provide topics and information that do not interest elderly people. This leads to a decrease in satisfaction. In addition, there is no mechanism for continuously updating new information, which means that conversation content easily becomes outdated.

[1775] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1776] In this invention, the server includes a means for transmitting the target person's attribute information to the network, a means for collecting related information in a data store based on the transmitted attribute information and generating customized conversation data, and a means for transmitting the generated customized conversation data to the client device. This enables conversations using the latest information specific to the target person, allowing elderly people to participate in conversations with interest without feeling lonely. Furthermore, by continuously collecting new information and updating the conversation content, it is possible to always provide fresh topics of conversation.

[1777] "Subject attribute information" refers to personal information such as the subject's name, age, hobbies, and health status.

[1778] "Data store" refers to a system or database for storing information and data collected from the Internet.

[1779] "Customized conversation data" refers to a conversation script that is generated based on the individual attribute information of the subject and tailored to the subject.

[1780] "Client device" refers to an AI-enabled IoT device or terminal used by the subject.

[1781] "Network" refers to the Internet and other communication means that provide paths for transmitting and receiving data.

[1782] A "conversation script" refers to data that describes the content of utterances and responses that the system uses in dialogue with the user.

[1783] This invention provides a system that uses AI-equipped IoT devices to help elderly people enjoy everyday conversations without feeling lonely. The entire system is realized by collecting profile information of the target person and providing customized conversation data to the user.

[1784] Hardware and software used

[1785] Terminal: An AI-enabled IoT device placed close to the user. This terminal displays the necessary initial setup screens and allows the user or their family members to enter their profile information.

[1786] Server: Refers to the back-end system that analyzes data and generates customized conversation scripts. The server is connected to the Internet and collects various information.

[1787] Speech Recognition Module: For example, using Google Speech-to-Text. Used to convert the user's speech into text.

[1788] Generative AI models, such as GPT-3, are used to generate customized conversation scripts based on collected information.

[1789] A speech synthesis module, e.g., using Google Text-to-Speech, is used to convert the generated text into speech and respond to the user.

[1790] Explanation of program processing

[1791] Device:

[1792] 1. When the device starts up, the initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored on the device.

[1793] 2. The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[1794] server:

[1795] 1. The server analyzes the received profile information and collects relevant information from a data store (e.g., the latest gardening techniques or Go news). Based on this information, a customized conversation script is generated for each user. This is generated using a generative AI model.

[1796] 2. Send the generated customized conversation script to the terminal.

[1797] Device:

[1798] 1. The device saves the received conversation script in its internal memory and displays "Conversation data updated" on the screen.

[1799] 2. When the user speaks to the device, the device analyzes it using the voice recognition module, selects a response based on the customized conversation data, and responds to the user using the voice synthesis module.

[1800] Specific examples

[1801] Enter your profile information:

[1802] The user's family member enters information into the device, such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." This information is temporarily stored in the device.

[1803] Send profile information:

[1804] When the device is connected to a network such as Wi-Fi, it sends the saved profile information to the server as a POST request. If the request is successful, a pop-up message will appear saying, "Taro's information has been sent to the server."

[1805] Profile information analysis and customization:

[1806] The server analyzes the profile information "Taro, 75 years old, gardening, Go, mild dementia," collects relevant latest gardening techniques and recent Go news, and generates a conversation script using a generative AI model.

[1807] Send a customized conversation script:

[1808] The server sends the conversation script data in JSON format, and the device receives this data, parses it, and saves it in its internal memory. The updated content is notified with a message saying "The conversation data has been updated."

[1809] Start a conversation:

[1810] When a user speaks to the device, "How's your gardening going these days?", the device converts the speech into text using a speech recognition module, and then uses a generative AI model to respond, "The latest gardening techniques place great importance on soil preparation," which is then played back as audio using a speech synthesis module.

[1811] Example prompt sentence:

[1812] "Please tell me the latest gardening knowledge for 75-year-old Taro."

[1813] "Tell me the latest news about Go"

[1814] "How's your gardening going lately?"

[1815] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1816] Step 1:

[1817] Enter and save your profile information

[1818] Device:

[1819] When the device starts up, an initial setup screen is displayed. On this screen, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). The entered information is temporarily stored in the device.

[1820] Input: The user enters profile information as text.

[1821] Output: The entered profile information is saved in the device's temporary memory.

[1822] Specific operation: A user's family member enters information such as "Taro, 75 years old, hobbies are gardening and Go, mild dementia." The device saves this information in text format.

[1823] Step 2:

[1824] Sending profile information

[1825] Device:

[1826] The saved profile information is sent to the server via the Internet, and once the sending is complete, the message "Profile information sent to server" is displayed on the screen.

[1827] Input: Profile information stored on your device.

[1828] Output: The profile information is sent to the server.

[1829] Specific operation: The device sends profile information in the form of a POST request via a network such as Wi-Fi. A message pops up when the transmission is successful.

[1830] Step 3:

[1831] Profile information analysis and customization

[1832] server:

[1833] The server analyzes the received profile information and gathers relevant information from a data store, such as the latest gardening techniques or Go news, and then uses a generative AI model to generate a customized conversation script based on the collected information.

[1834] Input: Profile information stored on the server.

[1835] Output: A customized conversation script.

[1836] Specific operation: The server analyzes the information "Taro, 75 years old, hobbies are gardening and Go, mild dementia," collects related information using web scraping and APIs, and generates a conversation script using a generative AI model (e.g., GPT-3).

[1837] Step 4:

[1838] Sending customized conversation scripts

[1839] server:

[1840] The generated customized conversation script is sent to the terminal.

[1841] Device:

[1842] The terminal stores the received conversation script in its internal memory and displays the message "Conversation data has been updated" on the screen.

[1843] Input: A customized conversation script generated by the server.

[1844] Output: The customized conversation script is saved in the device's memory.

[1845] Specific operation: The server sends script data in JSON format, the device receives it, parses it, saves it in its internal memory, and displays a pop-up message saying "Conversation data has been updated."

[1846] Step 5:

[1847] Start a conversation

[1848] User:

[1849] The user provides the device with a question or topic (e.g., "How's your gardening going lately?").

[1850] Device:

[1851] The device uses a speech recognition module to convert the user's speech into text, generates appropriate responses based on customized conversation data, and responds to the user using a speech synthesis module.

[1852] Input: User's voice input.

[1853] Output: Audio response from the device.

[1854] What it does: When a user says, "How's your gardening going these days?", the device converts the speech to text and selects a response from a customized conversation script: "Modern gardening techniques emphasize soil conditioning," and plays it back as audio.

[1855] Step 6:

[1856] Continuous learning and updates

[1857] server:

[1858] The server continuously collects user conversation data, analyzes it, and provides feedback, allowing it to learn new information and topics and periodically update its customized conversation scripts.

[1859] Device:

[1860] Every time a new conversation script is sent from the server, it is received, saved in internal memory, and updated. When the update is complete, a message will appear saying, "The conversation data has been updated based on the new information."

[1861] Input: User conversation data and new information.

[1862] Output: Updated customized conversation script.

[1863] Specific operation: The server analyzes the user's conversation history, collects new gardening information, the latest news on Go, etc., and sends an updated conversation script to the device. The device receives this and notifies the user that "the conversation data has been updated with the latest information."

[1864] (Application example 1)

[1865] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1866] There is a need for systems that can not only reduce loneliness among the elderly and provide daily communication, but also monitor their health status and respond to emergency situations. There is also a need for systems that can provide customized conversations using information specific to the elderly and the latest topics.

[1867] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1868] In this invention, the server includes means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for collecting sensor data for monitoring health status, and means for making an emergency call if an abnormality occurs. This not only reduces the sense of loneliness of elderly people and provides communication, but also enables health status monitoring and prompt response in emergencies.

[1869] "Subject profile information" is basic information about an individual, such as the subject's name, age, hobbies, and health status.

[1870] The "means for inputting profile information" refers to an interface that allows the subject or their family to input profile information using a terminal.

[1871] The "means for transmitting profile information to a server" refers to a communication means for transmitting profile information from a terminal to a server via the Internet.

[1872] A "customized conversation script" is an individually optimized conversation scenario that is generated by collecting relevant information on the Internet based on the target person's profile information.

[1873] The "means for generating a customized conversation script" is a processing means for analyzing profile information and creating a conversation script based on the latest related information on the Internet.

[1874] The "means for recognizing and responding to the user's voice" is a function for analyzing the voice uttered by the user and selecting and responding to an appropriate response based on a customized conversation script.

[1875] "Sensor data for monitoring health status" refers to information obtained from sensors that collect physiological and activity data of subjects, such as heart rate, number of steps, and fall detection.

[1876] The "means for collecting sensor data" refers to devices and software for collecting physiological data and activity data through sensors, which are necessary for monitoring the health status of a subject.

[1877] The "means for making an emergency report when an abnormality occurs" is a means for making a report to a pre-set emergency contact when an abnormality is detected based on collected sensor data.

[1878] "Means for gathering information from the Internet" refers to the Internet browsing ability to gather relevant and up-to-date information and news on the Internet based on the profile information.

[1879] This invention provides a system that can reduce loneliness among the elderly, monitor their health status, and respond to emergencies. The system provides customized conversations based on the target person's profile information and makes an emergency call when an abnormality in their health status is detected.

[1880] Program processing explanation

[1881] Hardware and Software Configuration

[1882] Hardware:

[1883] Robot device (equipped with microphone, speaker, and sensor)

[1884] Software and Libraries:

[1885] requests library: Used for server communication.

[1886] speech_recognition library: Used for speech recognition.

[1887] pyttsx3 library: Used for speech synthesis.

[1888] Processing explanation

[1889] (1) Enter and submit profile information

[1890] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device and then sent to the server. The server receives the information and sends a confirmation message back to the device.

[1891] Examples:

[1892] The user's family member enters information such as "Mr. Tanaka, 75 years old, hobbies are gardening and Go, mild dementia" into the terminal and sends it to the server.

[1893] (2) Analyzing profile information and generating conversation scripts

[1894] The server analyzes the received profile information and collects relevant information from the Internet (such as the latest gardening techniques or Go news), and generates a conversation script customized for each user based on this information.

[1895] Examples:

[1896] The server "collects the latest gardening trivia and Go game information for Tanaka and generates a customized conversation script."

[1897] (3) Sending customized conversation scripts

[1898] The server sends the generated customized conversation script to the device. The device receives the conversation script and saves it in its internal memory. Once the saving is complete, the device displays a notification to the user that the update is complete.

[1899] Examples:

[1900] The server sends "Gardening trivia" and the device notifies the user that "Tanaka's conversation data has been updated."

[1901] (4) Communication using voice recognition

[1902] When a user speaks to the terminal, the terminal analyzes the user's voice using a voice recognition module, selects an appropriate response from a customized conversation script, and responds verbally.

[1903] Examples:

[1904] When the user asks, "How's gardening going these days?" the device responds, "The latest gardening technique is that soil adjustment is becoming increasingly important."

[1905] (5) Health monitoring and emergency notification

[1906] The device's built-in sensors constantly monitor the user's health. If an abnormality is detected (e.g., an abnormal heart rate or a fall), the device immediately sends the abnormal data to the server, which then notifies the user's emergency contacts.

[1907] Examples:

[1908] The device notifies the server that an abnormal heart rate has been detected, and the server then calls a pre-set emergency number.

[1909] (6) Continuous learning and updating

[1910] The server continuously collects and analyzes the user's conversation data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory. The user is notified that the conversation data has been updated with new information.

[1911] Examples:

[1912] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[1913] Example prompts for generative AI models

[1914] Generate customized conversation scripts about the latest gardening news and Go topics based on the user's profile information. For example, respond to the question "Tell me about gardening" with "Soil preparation is becoming increasingly important as a modern gardening technique."

[1915] In this way, the embodiments of the invention allow elderly people to enjoy everyday conversations without feeling lonely, while their health condition is monitored, and if an abnormality occurs, a prompt response can be made.

[1916] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1917] Step 1:

[1918] Enter your profile information

[1919] The user or their family member uses the device to enter the target person's profile information (such as name, age, hobbies, and health status). The entered information is temporarily stored on the device. The input is done using the user interface, and the input data is stored in the device's internal storage.

[1920] Input: Profile information such as name, age, hobbies, health status, etc.

[1921] Output: Profile information temporarily saved on the device

[1922] Step 2:

[1923] Sending profile information

[1924] The device will then send the saved profile information to the server via the internet, and once the transfer is complete, a confirmation message will appear saying "Profile information has been sent to the server."

[1925] Input: Profile information temporarily saved on the device

[1926] Output: Profile information sent to server, confirmation message to user

[1927] Step 3:

[1928] Analyzing profile information

[1929] The server analyzes the received profile information and collects relevant information on the Internet (e.g., the latest gardening techniques or Go news). Based on the profile information, the analysis generates relevant search queries and retrieves relevant information.

[1930] Input: Profile information sent to the server

[1931] Output: Analysis results and related collected information

[1932] Step 4:

[1933] Generate customized conversation scripts

[1934] The server generates a customized conversation script for each user based on the collected relevant information. It uses a generative AI model to script a conversation scenario based on the profile information and collected information.

[1935] Input: Analysis results and related collected information

[1936] Output: Customized conversation script

[1937] Step 5:

[1938] Sending customized conversation scripts

[1939] The server sends the generated customized conversation script to the terminal. The terminal receives the conversation script and saves it in its internal memory. Once the saving is complete, a notification that "Conversation data has been updated" is displayed to the user.

[1940] Input: Customized conversation script

[1941] Output: Conversation script saved on the device, update completion notification to the user

[1942] Step 6:

[1943] User voice recognition

[1944] When a user speaks to the device, the device uses a voice recognition module to analyze the user's voice, converting the voice data into text data, and then selects an appropriate response from a customized conversation script based on the text.

[1945] Input: User voice input

[1946] Output: Recognized text data, selected response

[1947] Step 7:

[1948] Voice response

[1949] The device responds to the user vocally based on the selected customized conversation script, which is generated using a speech synthesis module and output audibly through the speaker.

[1950] Input: Selected Response

[1951] Output: Audio response

[1952] Step 8:

[1953] Health monitoring

[1954] The device's built-in sensors constantly monitor the subject's health condition, and the collected sensor data is periodically sent to a server, which immediately notifies the user if any abnormalities are detected.

[1955] Input: Sensor data (heart rate, steps, fall detection, etc.)

[1956] Output: Health data sent to the server, anomaly detection notification

[1957] Step 9:

[1958] emergency call

[1959] If an abnormality is detected, the device or server will make an emergency call, and the server will send an emergency notification to the configured emergency contacts, reporting the specific situation.

[1960] Input: Anomaly detection notification

[1961] Output: Notification to emergency number

[1962] Step 10:

[1963] Continuous learning and updates

[1964] The server continuously collects and analyzes the user's conversation data and health data, learns new information and topics from the Internet, and periodically updates the customized conversation script. Each time a new script is generated, the device receives it and stores it in its internal memory.

[1965] Input: conversation data, health data

[1966] Output: Updated conversation script, new information notification

[1967] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1968] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Detailed embodiments of this system are described below.

[1969] ---

[1970] Enter your profile information

[1971] Device:

[1972] On the initial setup screen of the device, the user or their family member enters the target person's profile information (name, age, hobbies, health status, etc.). This information is temporarily stored in the device.

[1973] Examples:

[1974] The user's family member enters the following information: "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia."

[1975] ---

[1976] Sending profile information

[1977] Device:

[1978] Your profile information will be sent to the server via the Internet. When the sending is complete, the message "Profile information has been sent to the server" will be displayed.

[1979] Examples:

[1980] The device will notify you that "Taro Sato's information has been sent to the server."

[1981] ---

[1982] Profile information analysis and customization

[1983] server:

[1984] The system analyzes the received profile information, gathers relevant information on the Internet, and generates a customized conversation script based on the target person's hobbies and interests.

[1985] Examples:

[1986] The server processes the message as follows: "We have collected the latest gardening trivia and Go game information for Taro Sato and generated a customized conversation script."

[1987] ---

[1988] Submitting a customization script

[1989] server:

[1990] The generated customized conversation script is sent to the terminal.

[1991] Device:

[1992] The received conversation script is saved in the internal memory and a message is displayed stating "Conversation data has been updated."

[1993] Examples:

[1994] The server sends "Gardening trivia" and the device notifies the user that "Taro Sato's conversation data has been updated."

[1995] ---

[1996] Speech and Emotion Recognition

[1997] User:

[1998] The user speaks to the device (e.g., "How's the gardening going lately?").

[1999] Device:

[2000] The user's voice is analyzed using a voice recognition module, and the emotion engine is used to recognize the user's emotions from their voice and facial expressions.

[2001] Examples:

[2002] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil preparation is becoming important," while determining whether the user looks happy.

[2003] ---

[2004] Providing customized responses

[2005] Device:

[2006] The system selects appropriate responses based on the user's emotions. For example, if the user is in good spirits, it selects more positive conversation content, and if they are feeling down, it offers encouraging words.

[2007] Examples:

[2008] If the user is feeling unwell, the device will gently ask, "You seem to be feeling unwell lately. Is there anything that is bothering you?"

[2009] ---

[2010] Continuous learning and updates

[2011] server:

[2012] By continuously collecting and analyzing user conversational and emotional data, the system learns the user's interests and emotional patterns and regularly updates the conversation script.

[2013] Device:

[2014] When a new script is received, it is saved in the internal memory and updated. Once the update is complete, a message will appear saying "Conversation data has been updated based on new information."

[2015] Examples:

[2016] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[2017] ---

[2018] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[2019] The processing flow will be explained below.

[2020] Step 1:

[2021] Device: When the device is started, an initial setup screen appears. The user or their family member enters their profile information (name, age, hobbies, health status, etc.) through this screen. The entered information is temporarily stored in the device.

[2022] Step 2:

[2023] Device: The saved profile information is sent to the server via the Internet. When the sending is complete, the message "Profile information sent to server" is displayed.

[2024] Step 3:

[2025] Server: Analyzes the received profile information, extracting information about the target person's hobbies, interests, and health status.

[2026] Step 4:

[2027] Server: Based on the extracted information, collects the latest information relevant to the target person from the Internet (e.g., the latest gardening techniques or Go news).

[2028] Step 5:

[2029] Server: Based on the collected information, a customized conversation script is generated for each target audience, including content that will interest the target audience and relevant information.

[2030] Step 6:

[2031] Server: Sends the generated customized conversation script to the device.

[2032] Step 7:

[2033] Device: Saves the received conversation script in its internal memory and notifies the user or their family that "the conversation data has been updated."

[2034] Step 8:

[2035] User: The user speaks to the device (e.g., "How's the gardening going lately?").

[2036] Step 9:

[2037] Device: The user's voice is analyzed by the speech recognition module. Then, the emotion engine recognizes emotions from the user's voice and facial expressions.

[2038] Step 10:

[2039] The device: Based on the recognized emotion, it selects an appropriate response from a customized conversation script. If the user is in good spirits, it selects positive responses, and if they are not, it offers encouraging words.

[2040] Step 11:

[2041] Terminal: The selected response is returned to the user through a voice output device. For example, "Soil conditioning is becoming increasingly important in modern gardening techniques."

[2042] Step 12:

[2043] Server: Continuously collects and analyzes user conversations and recognized emotional data, thereby learning the user's emotional patterns and interests.

[2044] Step 13:

[2045] Server: Periodically gathers new information and topics from the Internet and updates the customized conversation script.

[2046] Step 14:

[2047] Server: Sends the updated script to the device.

[2048] Step 15:

[2049] Device: When a new script is received, it is saved in internal memory and a notification is sent to the user or their family member stating, "Conversation data has been updated based on new information."

[2050] Through each step, the AI-powered IoT device can provide individually optimized conversations based on the target person's profile information and emotional data, allowing the elderly to enjoy their daily lives without feeling lonely. Furthermore, by combining it with an emotion engine, the device can understand the user's emotional state and provide appropriate responses, enabling deeper communication.

[2051] Example 2

[2052] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2053] Today's elderly often feel social isolation and loneliness, which can accelerate the progression of dementia. However, increasing opportunities for these elderly people to interact through everyday conversations could contribute to reducing loneliness and maintaining cognitive function. However, current technology is limited in systems that provide personalized conversations based on individual hobbies and interests, and there are particularly few systems that can recognize emotions. Therefore, there is a need for a system that provides conversations tailored to each individual elderly person and enables deep interaction.

[2054] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2055] In this invention, the server includes means for analyzing the profile information of the subject, collecting related information on the Internet, and generating a customized conversation script using a generative AI model, means for transmitting the customized conversation script to the terminal, and means for the terminal to analyze the user's voice using a voice recognition module and recognize the user's emotions using an emotion engine. This allows the user to enjoy conversations based on their hobbies and interests, and furthermore, receives appropriate responses through emotion recognition, which is expected to reduce feelings of loneliness among the elderly and slow the progression of dementia.

[2056] "Profile information" refers to personal information such as the subject's name, age, hobbies, and health status.

[2057] "Server" refers to a computer system that receives and analyzes profile information, gathers relevant information from the Internet, and generates customized conversation scripts.

[2058] "Conversation script" refers to customized conversation content generated using a generative AI model based on the target person's profile information and related information.

[2059] "Terminal" refers to a device used by a user to input profile information, receive customized conversation scripts, recognize voice and emotions, and display responses.

[2060] "Speech recognition module" refers to software or hardware that has the function of analyzing a user's voice and converting it into text.

[2061] "Emotion engine" refers to a system that includes technology to identify and analyze a user's emotions from voice and facial expression data.

[2062] A "generative AI model" refers to an artificial intelligence algorithm or machine learning model that generates customized conversation scripts based on input data.

[2063] "Speech synthesis engine" refers to a system that includes technology for converting text data into natural-sounding speech.

[2064] "Collecting related information on the Internet" refers to the server obtaining information related to the subject's hobbies and interests based on the profile information through methods such as web scraping or API calls.

[2065] "Response" refers to the voice or text response that the terminal returns to the user based on the analysis results of the voice recognition module and emotion engine.

[2066] This invention provides a system that combines an AI-equipped IoT device with an emotion engine to enable elderly people to enjoy everyday conversations without feeling lonely. Specific embodiments for carrying out this invention are described below.

[2067] The main components of the system are a device, a server, a generative AI model, an emotion engine, and related software modules. The device has an interface for inputting and sending user profile information, which is then sent to the server. The server generates a customized conversation script based on this profile information. The generative AI model is used for generation.

[2068] Device:

[2069] The device is initially set up by the user or their family, and profile information (such as name, age, hobbies, and health status) is entered. The entered information is temporarily stored in the device's memory. Examples of devices include tablets and smart home devices. These devices are equipped with touch panels and keyboards, making it easy to enter information.

[2070] Examples:

[2071] The user's family member enters information such as "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" into the tablet.

[2072] server:

[2073] The server receives profile information sent from the device via the internet. Based on the received information, the server collects the latest information related to the target person's hobbies and interests from the internet. This information is collected using web scraping and API calls. A customized conversation script is then generated using a generative AI model and sent to the device.

[2074] Examples:

[2075] The prompt sentence "75-year-old male, hobbies are gardening and Go, mild dementia" is input into the generative AI model, which then performs processes such as "collecting the latest gardening trivia and Go game information for Taro Sato and generating a customized conversation script."

[2076] Device:

[2077] The device receives the customized conversation script sent from the server and saves it in its internal memory. After saving, it notifies the user that "the conversation data has been updated." The device then analyzes the user's voice using a voice recognition module and uses an emotion engine to recognize the user's emotions from their voice and facial expressions.

[2078] Examples:

[2079] When a user asks, "How's your gardening going these days?", the device recognizes the user's interest from their voice and responds, "The latest gardening technique is that soil adjustment is becoming important," while determining whether the user looks happy.

[2080] Emotion Engine:

[2081] The emotion engine includes technology that analyzes the user's emotions from voice and facial expression data. If the device determines that the user is not in good spirits, the emotion engine selects an appropriate response.

[2082] Examples:

[2083] If the user is not feeling well, gently ask, "You seem to be feeling unwell lately. Is there anything that's bothering you?"

[2084] The server continuously collects and analyzes the user's conversational and emotional data to learn about the user's interests and emotional patterns, allowing it to periodically update the conversation script and provide more personalized conversations.

[2085] Examples:

[2086] The server collects new "Autumn Gardening Tips" and sends them to the device. The device then notifies the user that "The latest autumn gardening information has been updated."

[2087] In this way, AI-equipped IoT devices can provide more personalized conversations based on the target person's profile information and emotional data, creating an environment where seniors can enjoy enjoyable conversations without feeling lonely. Furthermore, by combining an emotion engine, it is possible to understand the user's emotional state and provide appropriate responses, enabling deeper communication. This is expected to help reduce feelings of loneliness and slow the progression of dementia.

[2088] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2089] Step 1:

[2090] Device:

[2091] On the initial setup screen, the user or their family member enters the target person's profile information. For example, they might enter "Taro Sato, 75 years old, hobbies are gardening and Go, mild dementia" on the tablet's touch panel. The entered data is temporarily stored in the device's memory.

[2092] Input: Profile information (name, age, hobbies, health status)

[2093] Data processing: Input via touch panel or keyboard

[2094] Output: Profile information temporarily stored in the device memory

[2095] Step 2:

[2096] Device:

[2097] Your profile information will be sent to the server via the Internet. Once the sending is complete, the message "Profile information sent to server" will be displayed on the screen.

[2098] Input: Profile information stored in the device memory

[2099] Data processing: Sending HTTP requests via the network module

[2100] Output: Notification of completion of sending profile information to the server

[2101] Step 3:

[2102] server:

[2103] The submitted profile information is analyzed. For example, "75 years old, male, hobbies are gardening and Go, mild dementia." Related information is then collected from the internet. Web scraping and APIs are used to collect this information.

[2104] Input: Submitted profile information

[2105] Data processing: Analysis of profile information and collection of related information on the Internet

[2106] Output: Related information

[2107] Step 4:

[2108] server:

[2109] A generative AI model is used to generate a customized conversation script based on the collected relevant information. The prompt sentence, "75-year-old male, hobbies are gardening and Go, mild dementia," is input into the generative AI model.

[2110] Input: Profile information and related information

[2111] Data processing: Generating conversation scripts using generative AI models

[2112] Output: Customized conversation script

[2113] Step 5:

[2114] server:

[2115] The generated customized conversation script is sent to the terminal.

[2116] Input: Customized conversation script

[2117] Data processing: Sending HTTP requests via the network module

[2118] Output: Notification of completion of transmission to the terminal

[2119] Step 6:

[2120] Device:

[2121] The received conversation script is saved to the internal memory. After saving is complete, a message will be displayed saying "Conversation data has been updated."

[2122] Input: Conversation script sent from the server

[2123] Data processing: Saving to internal memory and notifying on the user interface

[2124] Output: "Conversation data updated" notification

[2125] Step 7:

[2126] User:

[2127] The user speaks to the device (e.g., "How's the gardening going lately?").

[2128] Input: User voice input

[2129] Data processing: None

[2130] Output: User's voice input

[2131] Step 8:

[2132] Device:

[2133] The user's voice is analyzed by the voice recognition module. The voice data is sent to the STT (Speech-to-Text) engine to convert it into text, and the emotion engine recognizes emotions from the voice and facial expressions.

[2134] Input: User's voice data

[2135] Data processing: Speech to text and emotion recognition

[2136] Output: User utterance text and recognized sentiment

[2137] Step 9:

[2138] Device:

[2139] Based on the user's emotions, the robot selects an appropriate response from a customized conversation script. For example, if the user is in good spirits, it selects positive conversational content, and if they are feeling down, it responds with encouraging words. The selected response is then converted into voice by a speech synthesis engine and returned to the user.

[2140] Input: User utterance text, recognized emotions, and a customized conversation script

[2141] Data processing: Response selection and speech synthesis based on spoken text and emotional data

[2142] Output: Audio response

[2143] Step 10:

[2144] server:

[2145] It continuously collects and analyzes user conversational and emotional data, and uses this data to periodically update the generative AI model with customized conversation scripts.

[2146] Input: User conversation data and emotion data

[2147] Data processing: Data analysis and updating of conversation scripts using generative AI models

[2148] Output: Updated conversation script

[2149] Step 11:

[2150] Device:

[2151] The updated conversation script is received and saved in the internal memory. After the update is complete, a message will be displayed stating, "The conversation data has been updated based on the latest information."

[2152] Input: Updated conversation script

[2153] Data processing: Saving to internal memory and notifying on the user interface

[2154] Output: Notification "Conversation data has been updated based on the latest information"

[2155] (Application example 2)

[2156] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2157] In modern society, loneliness and mental stress are serious problems for the elderly and for elderly factory workers. In particular, the psychological impact of monotonous work in factories and lack of communication with others cannot be overlooked. There are also concerns about dementia and health conditions. Therefore, there is a need for a system that helps elderly people stay healthy through appropriate communication, without feeling lonely, and allows them to work safely and efficiently.

[2158] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting profile information of the target person, means for transmitting the profile information of the target person to the server, means for collecting related information on the Internet based on the transmitted profile information and generating a customized conversation script, means for transmitting the generated customized conversation script to the terminal, means for recognizing the user's voice and responding based on the customized conversation script, and means for analyzing the user's emotional state using an emotion engine and providing appropriate conversation content. This not only enables elderly people to enjoy everyday conversations without feeling lonely, but also improves work efficiency and safety.

[2159] "Subject profile information" refers to basic information including the subject's name, age, hobbies, health status, past work experience, etc.

[2160] "Server" refers to a computer system for analyzing profile information, collecting related information, and generating and transmitting customized conversation scripts.

[2161] A "terminal" refers to a device for interacting with a user, which has the function of recognizing the user's voice and responding based on a conversation script.

[2162] "Customized conversation script" refers to personalized dialogue content generated based on the target person's profile information.

[2163] "Speech recognition" refers to the technology that analyzes a user's voice and converts it into text data.

[2164] An "emotion engine" refers to technology that analyzes a user's emotional state from data such as voice and facial expressions, and selects appropriate conversation content.

[2165] "Emotional state" refers to the emotions the user is feeling at the time, such as joy, sadness, interest, indifference, etc.

[2166] MODE FOR CARRYING OUT THE INVENTION

[2167] System configuration

[2168] The system for implementing this invention includes the following main components: a terminal for inputting profile information of a target person; a server for analyzing the profile information, collecting related information, and generating a customized conversation script; and a terminal for recognizing the user's voice and providing appropriate conversation content using an emotion engine.

[2169] Hardware and software used

[2170] Hardware

[2171] Smartphones and tablets (as audio input / output devices)

[2172] Factory robots (for work assistance and interaction)

[2173] software

[2174] OpenAI API (for dialogue generation)

[2175] SpeechRecognition

[2176] gTTS (Text to Speech)

[2177] mpg321 (audio playback)

[2178] Program processing flow and data processing

[2179] 1. Enter your profile information

[2180] Enter the subject's basic information (name, age, hobbies, health status, past work experience, etc.) on the terminal.

[2181] The terminal transmits the input information to the server.

[2182] 2. Analyzing profile information and generating customized conversation scripts

[2183] The server analyzes the received profile information and collects related information on the Internet.

[2184] The server generates a customized conversation script based on the collected information and sends it to the terminal.

[2185] 3. Speech and Emotion Recognition

[2186] When the user speaks to the device, it recognizes the voice and converts it into text data.

[2187] An emotion engine is used to analyze the user's emotional state from their voice and facial expressions.

[2188] 4. Providing customized responses

[2189] The device selects appropriate conversation content based on the analyzed emotional state.

[2190] The selected conversation content is converted from text to speech and provided to the user.

[2191] Specific examples

[2192] For example, suppose an elderly person working in a factory asks the device, "How's the gardening going lately?" The device first recognizes the user's voice and sends the information to the server. The server has already registered the person's profile information (for example, that gardening is a hobby), so it collects the latest information about gardening and generates a customized conversation script. Next, it uses an emotion engine to analyze the user's emotional state from their voice and facial expressions, and selects the most appropriate response to respond. For example, it provides specific advice such as, "A recent gardening technique that places importance on adjusting the soil."

[2193] Prompt Sentence Examples

[2194] The specific prompt for the generative AI model is as follows:

[2195] Generate a script for a conversation with Taro Sato, an elderly worker. User utterance: How's your gardening going lately?

[2196] Based on this prompt, the AI ​​model can generate a response that combines the user's profile information and relevant information.

[2197] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2198] Step 1:

[2199] Enter your profile information

[2200] The terminal (smartphone or tablet) provides an interface where the user or their family members can input the subject's name, age, hobbies, health status, past work experience, etc.

[2201] Input: Subject's profile information (e.g., "Name: Ichiro Takahashi, Age: 70, Hobbies: Gardening and Go, Health Status: Mild Dementia")

[2202] Output: The entered profile information is temporarily stored in the device's local memory.

[2203] Step 2:

[2204] Sending profile information

[2205] The terminal transmits the saved profile information to a server via the Internet.

[2206] Input: Profile information stored in local memory

[2207] Output: The profile information is sent to the server and received at the server side.

[2208] Step 3:

[2209] Analysis of profile information and collection of related information

[2210] The server analyzes the received profile information and collects related information from the Internet based on the subject's hobbies and interests.

[2211] Input: Received profile information

[2212] Output: A customized conversation script (e.g., "Latest gardening news," "New Go strategies") is generated based on the profile information.

[2213] Step 4:

[2214] Sending customized conversation scripts

[2215] The server transmits the generated customized conversation script to the terminal.

[2216] Input: Customized conversation script

[2217] Output: The customized conversation script is sent to the device and stored in the device's internal memory.

[2218] Step 5:

[2219] Voice Recognition

[2220] When a user speaks to the terminal, the terminal uses a speech recognition module (SpeechRecognition) to convert the voice data into text data.

[2221] Input: User speech (e.g., "How's the gardening going lately?")

[2222] Output: The converted text data (e.g., "How's your gardening going lately?")

[2223] Step 6:

[2224] emotion recognition

[2225] The device uses an emotion engine to analyze the user's emotional state from the voice data.

[2226] Input: Converted text data and audio data

[2227] Output: User's emotional state (e.g., "Interested")

[2228] Step 7:

[2229] Providing customized responses

[2230] The device selects appropriate responses based on the analyzed emotional state and a customized conversation script.

[2231] Input: User's emotional state, customized conversation script

[2232] Output: Text response (e.g., "Soil conditioning is becoming increasingly important as a gardening technique these days.")

[2233] Step 8:

[2234] Text-to-speech conversion and output

[2235] The terminal converts the selected response into speech using gTTS and outputs it as speech.

[2236] Input: Text data response

[2237] Output: Provided to the user as audio data (e.g., "Soil preparation is becoming increasingly important as a gardening technique these days.")

[2238] This will provide an environment where elderly people and elderly factory workers can work with peace of mind without feeling lonely.

[2239] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2240] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2241] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2242] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2243] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2244] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2245] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2246] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2247] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2248] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2249] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2250] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2251] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2252] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2253] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2254] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2255] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2256] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2257] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2258] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2259] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2260] The following is further disclosed regarding the above embodiment.

[2261] (Claim 1)

[2262] a means for inputting subject profile information;

[2263] means for transmitting profile information of the subject to a server;

[2264] a means for collecting relevant information on the Internet based on the submitted profile information and generating a customized conversation script;

[2265] means for transmitting the generated customized conversation script to the terminal;

[2266] means for recognizing the user's voice and responding based on a customized conversation script;

[2267] A system including:

[2268] (Claim 2)

[2269] 10. The system of claim 1, further comprising means for analyzing the submitted profile information and gathering up-to-date information related to the user from the Internet.

[2270] (Claim 3)

[2271] 10. The system of claim 1, further comprising means for continuously collecting and analyzing user conversation data and periodically updating the customized conversation script.

[2272] "Example 1"

[2273] (Claim 1)

[2274] A means for inputting attribute information of a subject;

[2275] means for transmitting attribute information of the subject to the network;

[2276] a means for collecting related information in a data store based on the transmitted attribute information and generating customized conversation data;

[2277] means for transmitting the generated customized conversation data to the client device;

[2278] means for recognizing the user's speech and responding based on the customized speech data;

[2279] A system including:

[2280] (Claim 2)

[2281] 10. The system of claim 1, further comprising means for analyzing the transmitted attribute information and collecting up-to-date information related to the user from a data store.

[2282] (Claim 3)

[2283] 10. The system of claim 1, further comprising means for continuously collecting and analyzing user conversation data and periodically updating the customized conversation data.

[2284] "Application Example 1"

[2285] (Claim 1)

[2286] a means for inputting subject profile information;

[2287] means for transmitting profile information of the subject to a server;

[2288] a means for collecting relevant information on the Internet based on the submitted profile information and generating a customized conversation script;

[2289] means for transmitting the generated customized conversation script to the terminal;

[2290] means for recognizing the user's voice and responding based on a customized conversation script;

[2291] means for collecting sensor data for monitoring health status;

[2292] A means for making an emergency call in the event of an abnormality;

[2293] A system including:

[2294] (Claim 2)

[2295] 10. The system of claim 1, further comprising means for analyzing the submitted profile information and gathering up-to-date information related to the user from the Internet.

[2296] (Claim 3)

[2297] 10. The system of claim 1, further comprising means for continuously collecting and analyzing user conversation data and periodically updating the customized conversation script.

[2298] "Example 2: Combining Emotion Engines"

[2299] (Claim 1)

[2300] a means for inputting subject profile information;

[2301] means for transmitting profile information of the subject to a server;

[2302] a server analyzing the submitted profile information, gathering relevant information on the Internet, and generating a customized conversation script using a generative AI model;

[2303] means for transmitting the generated customized conversation script to the terminal;

[2304] a means for the terminal to analyze the user's voice using a voice recognition module and recognize the user's emotion using an emotion engine;

[2305] a means for selecting an appropriate response from a customized conversation script based on the user's emotions and responding with the speech synthesis engine;

[2306] A system including:

[2307] (Claim 2)

[2308] 10. The system of claim 1, further comprising means for analyzing the profile information and collecting up-to-date information related to the subject from the Internet.

[2309] (Claim 3)

[2310] 10. The system of claim 1, further comprising means for continuously collecting and analyzing user conversational and emotional data and periodically updating the customized conversational script.

[2311] "Application example 2 when combining emotion engines"

[2312] (Claim 1)

[2313] a means for inputting subject profile information;

[2314] means for transmitting profile information of the subject to a server;

[2315] a means for collecting relevant information on the Internet based on the submitted profile information and generating a customized conversation script;

[2316] means for transmitting the generated customized conversation script to the terminal;

[2317] means for recognizing the user's voice and responding based on a customized conversation script;

[2318] means for analyzing the user's emotional state using an emotion engine and providing appropriate conversation content;

[2319] A system including:

[2320] (Claim 2)

[2321] 10. The system of claim 1, further comprising means for analyzing the submitted profile information and gathering up-to-date information related to the user from the Internet.

[2322] (Claim 3)

[2323] 10. The system of claim 1, further comprising means for continuously collecting and analyzing user conversation data and periodically updating the customized conversation script. [Explanation of symbols]

[2324] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for inputting subject profile information; means for transmitting profile information of the subject to a server; a means for collecting relevant information on the Internet based on the submitted profile information and generating a customized conversation script; means for transmitting the generated customized conversation script to the terminal; means for recognizing the user's voice and responding based on a customized conversation script; A system including:

2. 10. The system of claim 1, further comprising means for analyzing the submitted profile information and gathering up-to-date information related to the user from the Internet.

3. 10. The system of claim 1, further comprising means for continuously collecting and analyzing user conversation data and periodically updating the customized conversation script.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A