system

An information processing device with audio prompts, generative model summaries, and biometric monitoring addresses the challenge of remote elderly care by facilitating communication and ensuring timely alerts for health and safety issues.

JP2026071660APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

There is a need for a system that allows family members living away from elderly individuals to easily grasp their health status and living situation, providing continuous remote monitoring and prompt information sharing, especially in situations where direct communication is difficult due to differences in daily life cycles.

Method used

An information processing device that transmits audio signals at fixed times to encourage conversation with the elderly, summarizes conversation points using a generative model, stores this information, and sends alerts or notifications based on detected keywords or abnormal biometric values, enabling remote health and safety monitoring.

Benefits of technology

Facilitates natural communication and comprehensive remote management of the elderly's health and safety, ensuring timely delivery of important information and enabling early action by family members.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071660000001_ABST
    Figure 2026071660000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of prompting conversation by transmitting an audio signal from an information processing device to the user at a fixed time every day, A means for extracting the main points of the conversation using a generative model and generating summary information, Means for storing the summary information in a storage device and making it accessible, A means for detecting keywords included in the summary information based on specific conditions and transmitting a warning signal to the user terminal, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a situation where it is difficult for family members living away from the elderly to communicate directly every day due to differences in daily life cycles, there is a need for a means to easily grasp the health status and living situation of the elderly and receive important information promptly as needed. Also, it aims to provide a system that can continuously monitor the health and safety of the elderly from a remote location.

Means for Solving the Problems

[0005] This invention provides an information processing device that transmits an audio signal to the user at a fixed time every day to encourage conversation with the elderly. It also has a function to summarize the main points of the conversation using a generative model and store them in a storage device. Furthermore, it has a means to detect keywords in the summarized information according to specific conditions and send a warning signal to the user terminal, enabling rapid information sharing. In addition, it has a function to acquire data based on biometric measurements and generate a notification signal when an abnormal value is detected, allowing for remote monitoring of health status and necessary actions to be taken.

[0006] An "information processing device" is a device that transmits audio signals and analyzes conversation content, and has the function of processing information based on user instructions.

[0007] A "user" is someone who uses the system to understand the situation of an elderly person and receive necessary information; this usually refers to a family member who lives separately.

[0008] A "voice signal" is a voice communication signal transmitted from an information processing device, and is used to encourage conversation among elderly people at specific times.

[0009] A "generative model" is an algorithm or machine learning model used for natural language processing, specifically a technique for extracting the main points of a conversation and generating a summary.

[0010] "Summary information" refers to information that summarizes the key points of a conversation extracted by a generative model, and is presented in a format that users can easily understand.

[0011] A "storage device" is a database or memory device that stores information, such as conversation content and summary information, and makes it accessible as needed.

[0012] "Specific conditions" are criteria determined by the system or conditions set by the user, and serve as indicators for extracting important information from summarized information.

[0013] "Keywords" are words or phrases that are considered particularly important within the summary information, and they are elements that trigger warning signals.

[0014] A "warning signal" is alert information generated when specific conditions are met, and it is a signal sent to the user's terminal to draw their attention.

[0015] "Biometric measurement" refers to a method of acquiring physiological data such as heart rate and step count, and is a technology used to monitor the health status of elderly people.

[0016] An "abnormal value" is a value detected by biometric measurement that is different from the normal range and indicates a health condition or situation that requires attention from the user. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention provides a system that facilitates natural communication among the elderly in their daily lives and enables them to share safety and health-related information with family members living remotely. This system primarily consists of an information processing device, a generative model, a storage device, and various communication means.

[0039] First, the device uses a portable communication device such as a smartphone or smartwatch to send an audio signal to the elderly person at a set time. The audio signal prompts the elderly person to start a conversation, and everyday communication is automatically conducted.

[0040] When a conversation begins, the device records the content in real time and saves it as digital audio data. This audio data is transmitted to a server via the internet. Encryption is used during data transfer to ensure privacy protection.

[0041] Next, in the process of processing the received audio data, the server uses a generative model to extract the main points of the conversation and generate summary information as needed. This summary information is then stored in a memory device and made accessible to the user when they need to review it.

[0042] The summary information is organized based on important keywords and context to allow families to easily understand the elderly person's situation. For example, if health-related information is included in a conversation, summaries such as "I felt fine yesterday" or "I have a doctor's appointment today" are generated.

[0043] Furthermore, the server monitors keywords within the summary information based on specific conditions, and if an anomaly or urgent situation is detected, it sends a warning signal to the user's terminal. This ensures that necessary information is delivered immediately, even when the user is in a remote location.

[0044] Furthermore, the system includes a biometric measurement function built into the terminal to continuously acquire health data from elderly individuals, and to immediately send a signal to the server if any abnormal values ​​are detected. This information is provided to the user as an emergency alert, serving as a means to encourage early action.

[0045] In this way, the present invention provides comprehensive remote management of the health and safety of the elderly, offering peace of mind to their families.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] The user downloads the application and completes the initial setup. On the settings screen, they specify settings such as the timing of voice notifications, the voice of the AI ​​avatar, and the conditions for warning signals. Once the setup is complete, the device sends this information to the server, and the account is registered.

[0049] Step 2:

[0050] At the designated time, the device emits an audio signal to prompt the elderly person to begin a conversation. This includes a greeting message played through the speaker. For example, a message such as "Good morning. How are you today?" might be played.

[0051] Step 3:

[0052] When a conversation begins, the device records audio data. The recorded audio is converted into digital data in real time and sent to the server via the communication line. Encryption is applied during data transfer to protect privacy.

[0053] Step 4:

[0054] The server analyzes the received audio data using a generative model. Here, the key points of the conversation are extracted, and summary information is generated. The generative model is based on natural language processing techniques and summarizes important information concisely and clearly.

[0055] Step 5:

[0056] The summary information is stored in the server's storage device. The stored data is made accessible through the user interface so that users can check it at any time.

[0057] Step 6:

[0058] Based on specific criteria, the server searches for important keywords in the summary information. If keywords matching the criteria are found, a warning signal is generated and a notification is sent to the user's terminal. This allows the user to receive information that enables them to take immediate action.

[0059] Step 7:

[0060] The device continuously records biometric data using its built-in sensors and detects abnormal values. If an abnormality is detected, a signal is immediately sent to the server, and an emergency alert is issued to the user.

[0061] Through the process described above, this system can accurately grasp the situation of elderly people remotely and take necessary actions quickly.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] There is a need for a system that allows elderly people to communicate naturally on a daily basis and effectively monitor their health remotely. Furthermore, it is necessary to ensure that abnormal situations are detected quickly while protecting the privacy of the elderly, enabling family members in remote locations to respond immediately.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for transmitting voice signals from an information processing device to a user at regular intervals to encourage conversation, means for extracting the main points of the conversation using a generative model and generating summary information, and means for encrypting and transmitting digital voice data. This makes it possible for family members who live far away to reliably understand the health and safety of the elderly, and to quickly detect abnormal situations while protecting their privacy.

[0067] An "information processing device" is a device that performs data input, processing, storage, and output, and has the function of exchanging information with other devices via a communication network.

[0068] "Users" refers to individuals who use this system, including, but not limited to, the elderly and their families.

[0069] An "audio signal" is a signal that is created by electronically converting sound and transmitting it via a communication means.

[0070] A "generative model" refers to an algorithm or system that uses machine learning to extract specific patterns or key points from input data.

[0071] "Summary information" refers to information that organizes the important points and keywords extracted from conversation content using a generative model.

[0072] "Means of storage" refers to physical or electronic means of storing data and keeping it accessible as needed.

[0073] "Encryption" is a technology that transforms the content of data based on a specific algorithm to prevent unauthorized access or decryption.

[0074] A "portable communication device" refers to a device that is portable and uses wireless technology to communicate.

[0075] "Privacy" is the right or state to prevent an individual's private information from being used or made public without their consent.

[0076] An "abnormal situation" refers to a state in which unusual circumstances or values ​​are detected, and is sometimes used particularly in relation to health.

[0077] The system of this invention aims to enable elderly people to communicate naturally in their daily lives and to effectively provide health information to distant family members. The system includes an information processing device, a communication device, a generative AI model, and a storage means.

[0078] The devices used are portable communication devices such as smartphones and smartwatches. These devices transmit voice signals to elderly people at designated times to encourage everyday conversation. The voice signals include time announcements and messages to start a conversation.

[0079] After an audio signal is transmitted, the terminal records the elderly person's voice and saves it as digital audio data. This audio data is encrypted via an information processing device and transferred to a server over the internet. The AES algorithm is used for encryption, ensuring data privacy.

[0080] The server decodes the received audio data and analyzes it using a generative AI model. During the analysis, natural language processing techniques are used to extract the main points of the conversation and summarize important information. The generated summary information is stored in a memory device for easy access by the user.

[0081] Users access summarized information posted to allow family members to easily understand the condition of elderly individuals, using a web browser or dedicated application. This ensures that important information is clearly organized and displayed. For example, if everyday conversations include information such as "I felt fine yesterday" or "I have an appointment to see a doctor today," this information is summarized and provided.

[0082] Furthermore, the server monitors specific phrases within the summary information, and if an anomaly or urgent case is detected, it sends a warning signal to the user's device in real time. This notification uses a push notification service for smartphones.

[0083] Furthermore, the device has a function to continuously monitor biometric information, and if an abnormal value is recorded, the data is immediately sent to the server. This information is notified to the user as an emergency alert, prompting early action.

[0084] An example of a prompt might be, "Please analyze the recording of a conversation with an elderly person and generate a summary regarding their health and safety." This allows the generative model to analyze the conversation content and organize and provide information according to the purpose.

[0085] By combining the technologies described above, this system provides a means to remotely, comprehensively, and effectively monitor the health and safety of the elderly.

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The device transmits an audio signal to the elderly person at a specified time. The input is the time information set within the device. The output is the audio signal transmitted by the device. This audio signal contains messages designed to facilitate everyday conversation.

[0089] Step 2:

[0090] The device records conversations with elderly individuals in response to voice signals. The input is the voice of the elderly person, which is converted into digital audio data and output. Specifically, this process involves capturing the voice through the device's microphone and converting it into a digital format.

[0091] Step 3:

[0092] The terminal encrypts the generated digital audio data and sends it to the server over the internet. The input is the recorded audio data, and the output is an encrypted data packet. The AES algorithm is used for encryption, ensuring privacy protection.

[0093] Step 4:

[0094] The server decrypts the received encrypted data and obtains digital audio data. The input is encrypted data, and the output is decrypted audio data. Specifically, it decodes the data using a cryptographic decryption algorithm.

[0095] Step 5:

[0096] The server uses a generative AI model to analyze audio data and extract the main points of the conversation. The input is decoded audio data, and the output is summarized information. Natural language processing techniques are used to identify keywords and important context.

[0097] Step 6:

[0098] The server stores the extracted summary information in a storage device and makes it accessible to the user. The input is the summary information, and the output is stored as saved data in the storage device. Specific operations include writing information to a database.

[0099] Step 7:

[0100] The server monitors keywords included in the summary information and sends a warning signal to the user's terminal if an anomaly or emergency condition is detected. The input is the keyword to be monitored, and the output is the warning signal sent to the terminal. The alert is triggered via a push notification service.

[0101] Step 8:

[0102] The device continuously acquires health data of elderly individuals using its built-in biosensors, and if an abnormal value is detected, it sends that information to a server. The input is the data obtained from the biosensors, and the output is the abnormal data sent to the server. Specifically, it monitors the data feed from the sensors in real time and generates an alert if a threshold is exceeded.

[0103] (Application Example 1)

[0104] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0105] There is a need for a system that can monitor the health status of the elderly in real time, facilitate daily communication, and immediately notify families when abnormalities are detected. However, many current systems have limitations in the frequency and content of communication, and are unable to adequately provide real-time notifications of important health information.

[0106] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0107] In this invention, the server includes means for transmitting voice signals from an information processing device to the user at a fixed time each day to facilitate language exchange; means for extracting the key points of the language exchange using a generative model and forming summary information; and means for continuously monitoring biometric information and immediately generating and transmitting a warning when an abnormal value is detected. This enables real-time monitoring of the health status of the elderly and rapid notification in emergencies.

[0108] An "information processing device" is a device that transmits audio signals to facilitate language communication and processes data.

[0109] "User" refers to an individual who receives audio signals via an information processing device and uses the service.

[0110] "Scheduled time" refers to a specific time when audio signals are transmitted or data processing takes place.

[0111] An "audio signal" is a means of transmitting information via sound to a user.

[0112] "Language exchange" refers to communication conducted through voice between the user and the system.

[0113] A "generative model" is a machine learning model that extracts key points from the content of linguistic exchange and forms summarized information.

[0114] "Summary information" refers to information that briefly summarizes the important points of language exchange.

[0115] A "storage medium" is a device or system that stores summary information and makes it accessible as needed.

[0116] A "keyword" is a specific, important word included in the summary information that triggers a warning or notification.

[0117] A "warning signal" is a signal sent to the user's device to alert them when an abnormality is detected.

[0118] "Biometric information" refers to data that indicates the user's health status, including physiological information such as heart rate and activity level.

[0119] An "abnormal value" is a value detected in biological information as an unexpected or exceeding standard.

[0120] A "warning" is a notification that alerts you to an abnormality based on your biometric information, and is primarily sent to family members or medical support staff.

[0121] To implement this invention, a smartwatch or portable device worn by an elderly person is used. The device transmits an audio signal to the elderly person at a set time each day, facilitating communication. The device records the elderly person's conversations, saves them as digital audio data in real time, and transmits them to a server in an encrypted format. This server analyzes the received audio data using a generation AI model, extracts the key points, and generates summary information. The generated summary information is stored on a storage medium and kept accessible to family members to understand the elderly person's situation. In addition, the device constantly monitors the user's biometric information, and if abnormal values ​​such as heart rate or activity level are detected, it immediately generates a warning and sends a notification to the family.

[0122] As a concrete example, the server analyzes the data received from the smartwatch, and if the heart rate exceeds the normal range, it immediately sends a message to the family's smartphone such as, "The user's heart rate is high. Is there anything wrong?" At this stage, the generative AI model detects summary information and reports specific details such as "Activity level increased yesterday" concisely to the family.

[0123] An example of a prompt to be input to the generation AI model is, "Extract the main points of an elderly person's everyday conversation, and extract and summarize keywords that include important medical information." Based on this prompt, the system automatically organizes the important information and immediately issues a warning if any abnormalities occur.

[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0125] Step 1:

[0126] The device, such as a smartwatch or other portable device, transmits voice signals to the elderly person to facilitate everyday conversation. The input for this step is the device's scheduled time information, and the output is the start of a conversation with the elderly person. By sending voice signals, it initiates natural communication.

[0127] Step 2:

[0128] The device records the conversation and digitizes it as audio data. The input for this step is the audio signal from the microphone, and the output is digital audio data. The digital data is stored and prepared for encryption in preparation for subsequent processing.

[0129] Step 3:

[0130] The terminal encrypts digital audio data using encryption technology and sends it to the server in a secure format. The input for this step is unprocessed digital audio data, and the output is encrypted data. Encryption is performed for security purposes.

[0131] Step 4:

[0132] The server decrypts the received encrypted data and converts it into processable audio data. The input for this step is encrypted audio data, and the output is the decrypted audio data. The server then prepares the data for appropriate processing.

[0133] Step 5:

[0134] The server uses a generative AI model to extract key points from audio data and generate summary information. The input for this step is decoded audio data, and the output is the summary information. By extracting key points, it efficiently presents important information.

[0135] Step 6:

[0136] The server stores the summary information on a storage medium and makes it accessible. The input for this step is the generated summary information, and the output is a database of the stored information. The environment is set up so that users can access the information when needed.

[0137] Step 7:

[0138] The server continuously monitors biometric data and, upon detecting an anomaly, immediately generates an alert and sends a notification to the family. The input for this step is real-time biometric data sent from the smartwatch, and the output is an alert message. Processing is performed to ensure rapid notification in emergencies.

[0139] Step 8:

[0140] The user receives summary information and warnings sent by the server to monitor their daily health status and make decisions regarding emergency responses. The input for this step is the notifications sent from the server, and the output is the user's decision. The user then takes appropriate action based on the information.

[0141] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0142] This invention is a system that recognizes the emotions of an elderly person during a conversation and provides that emotional information to their family. The system consists of an information processing device, a generative model, a memory device, an emotion engine, and a communication means.

[0143] First, the device uses a smart device to transmit an audio signal to the elderly person at a set time. This signal initiates a conversation. Once the conversation begins, the device records the audio data and converts it into digital data in real time. This involves transmitting the data to a server via a communication line. The data transfer is encrypted to protect privacy.

[0144] Next, the server analyzes the received audio data using a generative model to extract the main points of the conversation and generate summary information. This summary includes an emotion engine, which also includes emotional states detected during the conversation. The emotion engine identifies emotional data based on facial recognition and speech tone analysis and adds it to the summary information.

[0145] Summary information and sentiment data are stored in the server's storage device and made accessible to the user. The stored data is organized according to specific conditions, and if, for example, a particular emotional state or keyword is detected, an alert signal is sent to the user's device. This allows family members to quickly grasp the necessary information and respond appropriately to the situation.

[0146] Furthermore, the device continuously acquires biometric data and immediately sends a signal to the server if it detects any abnormal values. This information is provided to the user as an emergency alert to encourage prompt action.

[0147] For example, if an elderly person says "I've been feeling irritable a lot lately" during a conversation, the emotion engine will determine that their emotions are unstable. Based on this, the summary information will record "Recently, their emotions have been unstable," and the user will receive a notification immediately. This system allows families to have a detailed understanding of the elderly person's emotions and health condition and to provide necessary care quickly.

[0148] The following describes the processing flow.

[0149] Step 1:

[0150] The user installs the application and performs initial setup on a screen where they can configure the voice notification time, emotion recognition settings, and warning signal conditions. This configuration information is sent from the device to the server and stored there.

[0151] Step 2:

[0152] At a set time, the device emits an audio signal to prompt the elderly person to begin a conversation. The audio signal includes greetings and questions delivered through a speaker, such as "Hello, how was your day?"

[0153] Step 3:

[0154] When a conversation begins, the device records the audio data in real time and saves it as digital data. This data is encrypted and sent to the server via the communication line.

[0155] Step 4:

[0156] The server analyzes the received audio data using a generative model. The generative model utilizes natural language processing techniques to grasp the main points of the conversation and create a summary. At this point, the emotion engine operates, analyzing the voice tone and word choice to extract emotion data.

[0157] Step 5:

[0158] The server adds sentiment data to the generated summary information and stores it in storage. This makes the user able to access the summary, including sentiment information, at any time.

[0159] Step 6:

[0160] If the server detects keywords or emotional states that meet specific criteria from the summary information, it generates a warning signal and sends it to the user's terminal. For example, if the emotion engine detects "unstable emotions," a notification is immediately sent to the user.

[0161] Step 7:

[0162] The device uses biometric measurement functions to measure health data of elderly individuals (e.g., heart rate and activity level). If abnormal values ​​are detected, the device sends the collected data to a server, and an emergency alert is sent to the user.

[0163] This entire process provides a system that accurately assesses the health and emotional state of elderly individuals remotely, enabling families to respond quickly.

[0164] (Example 2)

[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0166] There are challenges such as a lack of communication with the elderly and the inability for families to quickly grasp changes in their health and take appropriate action. Furthermore, there is a lack of means to appropriately extract important information and emotional changes from conversations and promptly notify families.

[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0168] In this invention, the server includes means for periodically transmitting voice signals from an information processing device to a user to facilitate dialogue, means for extracting the main points of the dialogue and generating summary information using a generation algorithm, and means for adding emotion data generated by facial recognition and voice tone analysis included in the summary information. This enables family members to grasp changes in the emotional and health conditions of elderly people in real time and take necessary actions quickly.

[0169] An "information processing device" is a computer system that receives audio data and performs calculations for analysis and generation.

[0170] "Users" refers to individuals or organizations that receive conversational information and emotional data from elderly people through this system.

[0171] A "generative algorithm" is a computer program that extracts key points and features from audio data and performs summarization and sentiment analysis.

[0172] "Emotional data" refers to information that indicates the emotional state of a subject, generated based on changes in voice tone and facial recognition.

[0173] "Facial recognition" is a technology that uses image processing to analyze the movements and expressions of a person's face and determine their emotions.

[0174] "Voice tone analysis" is a process that analyzes the intonation and speed of speech to evaluate the speaker's emotions and state of mind.

[0175] A "storage device" is a storage medium for permanently or temporarily recording digital information.

[0176] A "warning signal" is alert information sent to the user's device when specific conditions are detected.

[0177] This invention is a system that analyzes the emotional and health status of elderly individuals through communication and provides this information to their families. The embodiments are described in detail below.

[0178] The device utilizes smart technology to transmit voice signals to elderly individuals and initiate conversations. These voice signals are transmitted automatically based on time. High-performance microphones and voice recognition software are used to convert the voice into digital data in real time, and encryption technology (e.g., AES) is used to transmit it to the server to maintain privacy.

[0179] The server uses a generative AI model to analyze the received audio data. The analysis extracts key points and important information from the conversation and generates a summary. Furthermore, the server incorporates an emotion engine that generates emotion data through voice tone analysis and facial expression recognition, and adds this to the summary.

[0180] The generated summary information and sentiment data are recorded on the server's storage device and made accessible to the user. The system has a function to immediately send a warning signal to the user's terminal when specific emotional states or keywords are detected. This allows families to obtain necessary information in real time and respond promptly to the elderly person's condition.

[0181] Furthermore, the device continuously monitors the biometric indicators of elderly individuals, and if abnormal data is detected, the server sends an urgent notification to the user. This collaboration enables early detection of health abnormalities and supports prompt response.

[0182] As a concrete example, if an elderly person mentions during a conversation that they have been feeling irritable lately, the emotion engine will determine that their condition is unstable, record it in the summary information, and immediately notify the user. This functionality allows families to appropriately monitor changes in the elderly person's emotions and health and provide necessary support quickly.

[0183] Example of a prompt:

[0184] "Please explain the mechanism for notifying a summary of information if unstable emotions are detected in recent conversations."

[0185] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0186] Step 1:

[0187] The terminal initiates a conversation by sending an audio signal to the elderly person using a smart device based on a set time. The input is the set time information, and the output is the transmission of an audio signal. At this time, the program utilizes the timer function of the smart device to execute the process of sending an audio signal at the specified time.

[0188] Step 2:

[0189] The device records the initiated conversation using a high-precision microphone and converts the audio data into digital data in real time. The input is the audio signal of the conversation, and the output is the digitized audio data. Speech recognition software operates in the background, analyzing the input audio waveform and converting it into text data and speech feature data.

[0190] Step 3:

[0191] The terminal encrypts digital data using AES encryption technology and transmits it to the server via the communication line while protecting privacy. The input is digitized voice data, and the output is encrypted data communication. The encryption module activates and converts the plaintext data based on the encryption key.

[0192] Step 4:

[0193] The server decrypts the received encrypted data and analyzes the audio data using a generative AI model. The input is encrypted audio data, and the output is summary information and sentiment data. The generative AI model simultaneously runs speech-to-text conversion and sentiment analysis algorithms to extract the main points and emotional state of the conversation.

[0194] Step 5:

[0195] The server uses an emotion engine to perform facial recognition and voice tone analysis based on the generated summary information. The input is the summary information, and the output is the summary information with added emotion data. The emotion engine evaluates various voice parameters and facial expression data, classifying emotions as either numerical or categorical.

[0196] Step 6:

[0197] The server stores the final summary information and sentiment data in storage, keeping it accessible to users. The input is the summary information and sentiment data, and the output is the entries in the stored database. Database management software organizes and indexes the data in an automated process.

[0198] Step 7:

[0199] The server immediately sends an alert signal to the user's device when a specific emotional state or keyword is detected. The input consists of stored summary information and triggering conditions, while the output is the alert signal. The alert system continuously monitors the system and activates a notification protocol when the conditions are met.

[0200] Step 8:

[0201] The device continuously acquires biometric data from elderly individuals via biosensors and reports any abnormal measurements to the server. The input is biometric data, and the output is a notification signal containing abnormal values. A vital signs monitoring system analyzes the data and triggers an alert if a threshold is exceeded.

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0204] In situations where elderly people experience anxiety and loneliness on a daily basis, it is difficult for family members and caregivers to quickly recognize these emotional changes and respond appropriately. Therefore, there is a need for a system that monitors the physical and mental state of elderly people in real time and promptly notifies relevant parties if any abnormalities are detected.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes means for transmitting an audio signal from an information processing device to a user at a fixed time each day to encourage conversation; means for extracting the main points of the conversation using a generative model and generating summary information; means for storing the summary information in a storage device and making it accessible; means for analyzing the audio data in real time and evaluating the emotional state; and means for sending a notification to relevant parties when an abnormal emotional state is detected. This makes it possible to quickly grasp the physical and mental state of elderly people and respond promptly as needed.

[0207] An "information processing device" is an electronic device that transmits audio signals to the user to facilitate conversation.

[0208] A "generative model" is an algorithmic method for extracting the main points of a conversation and summarizing information.

[0209] A "storage device" is a digital storage device used to store summary information and make it accessible as needed.

[0210] "Emotion recognition means" refers to technology that analyzes voice data and evaluates the user's emotional state.

[0211] A "portable communication device" is a portable communication device used to transmit voice signals.

[0212] An "abnormal value" is a value in biological data that exceeds the normal range and indicates a condition that requires attention.

[0213] "Means of sending notifications to relevant parties" refers to methods for communicating important information to users and stakeholders in real time.

[0214] The system implementing this invention consists of an information processing device, a generative model, a storage device, an emotion recognition engine, and a communication means. The server has a function to prompt conversation by sending an audio signal to the user at a fixed time every day via the information processing device. When a conversation begins, the terminal converts the audio data into digital data in real time, encrypts it, and sends it to the server. The server converts the audio data into text data using software such as Google® Cloud Speech-to-Text, and based on that, extracts the main points of the conversation using a generative model and generates summary information. In this process, an emotion recognition engine such as IBM Watson® Tone Analyzer is used to perform voice tone analysis and evaluate the emotional state. If an anomaly is detected by emotion recognition, a push notification is sent to the relevant parties.

[0215] For example, if a device records a conversation with an elderly person and detects a phrase like "I haven't been able to sleep lately," the emotion recognition engine will detect "anxiety," and the user will be promptly notified with a message such as "You've been experiencing persistent anxiety lately." An example of a prompt to the generative AI model used in this case would be, "Analyze the following conversation emotionally and detect any significant changes: 'I've been having more trouble sleeping lately...'" In this way, the server and the device work together, allowing users to quickly understand the emotions and health status of elderly people.

[0216] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0217] Step 1:

[0218] The terminal transmits an audio signal to the user (an elderly person) at regular intervals via an information processing device. The input is the scheduled alarm setting data, and the output is an audio signal sent to the user. This action prompts the user to initiate conversation.

[0219] Step 2:

[0220] As soon as a conversation with the user begins, the device converts the audio data into digital data in real time. The input is raw audio data, and the output is a stream of digital data. The audio data is captured on the device via the microphone and converted into text data using Google Cloud Speech-to-Text.

[0221] Step 3:

[0222] The terminal encrypts the converted text data and sends it to the server. The input is text data, and the output is encrypted text data. AES encryption or similar methods are used to protect the data during transmission.

[0223] Step 4:

[0224] The server inputs the received text data into a generative model to extract the main points of the conversation. The input is decrypted text data, and the output is summarized information. Natural language processing techniques are applied to extract the most important parts of the conversation.

[0225] Step 5:

[0226] The server processes the extracted summary information into an emotion recognition engine to analyze the user's emotional state. The input is summary information, and the output is emotional state data. IBM Watson Tone Analyzer is used to analyze the tone of the text and identify emotions.

[0227] Step 6:

[0228] The server detects anomalies from the emotional state data in the analysis results and sends notifications to relevant parties as needed. The input is emotional state data, and the output is a notification message. When an anomaly is detected, a push notification function is used to send an alert to family members or caregivers with specific information.

[0229] Throughout this entire process, users can quickly grasp the emotions and health status of elderly individuals.

[0230] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0231] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0232] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0233] [Second Embodiment]

[0234] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0235] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0236] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0237] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0238] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0239] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0240] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0241] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0242] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0243] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0244] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0245] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0246] This invention provides a system that facilitates natural communication among the elderly in their daily lives and enables them to share safety and health-related information with family members living remotely. This system primarily consists of an information processing device, a generative model, a storage device, and various communication means.

[0247] First, the device uses a portable communication device such as a smartphone or smartwatch to send an audio signal to the elderly person at a set time. The audio signal prompts the elderly person to start a conversation, and everyday communication is automatically conducted.

[0248] When a conversation begins, the device records the content in real time and saves it as digital audio data. This audio data is transmitted to a server via the internet. Encryption is used during data transfer to ensure privacy protection.

[0249] Next, in the process of processing the received audio data, the server uses a generative model to extract the main points of the conversation and generate summary information as needed. This summary information is then stored in a memory device and made accessible to the user when they need to review it.

[0250] The summary information is organized based on important keywords and context to allow families to easily understand the elderly person's situation. For example, if health-related information is included in a conversation, summaries such as "I felt fine yesterday" or "I have a doctor's appointment today" are generated.

[0251] Furthermore, the server monitors keywords within the summary information based on specific conditions, and if an anomaly or urgent situation is detected, it sends a warning signal to the user's terminal. This ensures that necessary information is delivered immediately, even when the user is in a remote location.

[0252] Furthermore, the system includes a biometric measurement function built into the terminal to continuously acquire health data from elderly individuals, and to immediately send a signal to the server if any abnormal values ​​are detected. This information is provided to the user as an emergency alert, serving as a means to encourage early action.

[0253] In this way, the present invention provides comprehensive remote management of the health and safety of the elderly, offering peace of mind to their families.

[0254] The following describes the processing flow.

[0255] Step 1:

[0256] The user downloads the application and completes the initial setup. On the settings screen, they specify settings such as the timing of voice notifications, the voice of the AI ​​avatar, and the conditions for warning signals. Once the setup is complete, the device sends this information to the server, and the account is registered.

[0257] Step 2:

[0258] At the designated time, the device emits an audio signal to prompt the elderly person to begin a conversation. This includes a greeting message played through the speaker. For example, a message such as "Good morning. How are you today?" might be played.

[0259] Step 3:

[0260] When a conversation begins, the device records audio data. The recorded audio is converted into digital data in real time and sent to the server via the communication line. Encryption is applied during data transfer to protect privacy.

[0261] Step 4:

[0262] The server analyzes the received audio data using a generative model. Here, the key points of the conversation are extracted, and summary information is generated. The generative model is based on natural language processing techniques and summarizes important information concisely and clearly.

[0263] Step 5:

[0264] The summary information is stored in the server's storage device. The stored data is made accessible through the user interface so that users can check it at any time.

[0265] Step 6:

[0266] Based on specific criteria, the server searches for important keywords in the summary information. If keywords matching the criteria are found, a warning signal is generated and a notification is sent to the user's terminal. This allows the user to receive information that enables them to take immediate action.

[0267] Step 7:

[0268] The device continuously records biometric data using its built-in sensors and detects abnormal values. If an abnormality is detected, a signal is immediately sent to the server, and an emergency alert is issued to the user.

[0269] Through the process described above, this system can accurately grasp the situation of elderly people remotely and take necessary actions quickly.

[0270] (Example 1)

[0271] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0272] There is a need for a system that allows elderly people to communicate naturally on a daily basis and effectively monitor their health remotely. Furthermore, it is necessary to ensure that abnormal situations are detected quickly while protecting the privacy of the elderly, enabling family members in remote locations to respond immediately.

[0273] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0274] In this invention, the server includes means for transmitting voice signals from an information processing device to a user at regular intervals to encourage conversation, means for extracting the main points of the conversation using a generative model and generating summary information, and means for encrypting and transmitting digital voice data. This makes it possible for family members who live far away to reliably understand the health and safety of the elderly, and to quickly detect abnormal situations while protecting their privacy.

[0275] An "information processing device" is a device for inputting, processing, storing, and outputting data, and has a function of exchanging information with other devices via a communication network.

[0276] A "user" refers to an individual who uses this system, including but not limited to the elderly and their families.

[0277] An "audio signal" is a signal obtained by electronically converting sound and transmitted via a communication means.

[0278] A "generative model" refers to an algorithm or system that uses machine learning to extract specific patterns or key points from input data.

[0279] "Summary information" is information in which important points and keywords extracted from conversation content by a generative model are organized.

[0280] A "storage means" refers to a physical or electronic means for storing data and keeping it accessible as needed.

[0281] "Encryption" is a technology that converts the content of data based on a specific algorithm to prevent unauthorized access and decryption.

[0282] A "portable communication device" refers to a device having a portable shape and performing communication using wireless technology.

[0283] "Privacy" is a right or state for preventing the unauthorized use or disclosure of an individual's private information.

[0284] An "abnormal situation" refers to a state when a situation or value different from normal is detected, and may be particularly used in relation to health.

[0285] The system of this invention aims to enable the elderly to communicate naturally in daily life and effectively provide health information to distant family members. The system includes an information processing device, a communication device, a generative AI model, and storage means.

[0286] The terminal uses portable communication devices such as smartphones and smartwatches. These terminals send voice signals to the elderly at a specified time to encourage daily conversations. The voice signals include time announcements and messages to start conversations.

[0287] After the voice signal is sent, the terminal records the voice of the elderly and records it as digital voice data. This voice data is encrypted via the information processing device and transferred to the server through the Internet. The AES algorithm is used for encryption to protect the privacy of the data.

[0288] The server decrypts the received voice data and analyzes the data using the generative AI model. In the process of analysis, natural language processing technology is used to extract the key points of the conversation and summarize important information. The generated summary information is stored in the storage means so that it can be easily confirmed by the user. ?

[0289] The user accesses the summary information posted for the family to easily grasp the condition of the elderly using a web browser or a dedicated application. Thereby, important information is presented in an easy-to-understand and organized manner. For example, if information such as "I felt well yesterday" or "I have a scheduled medical appointment today" is included in the daily conversation, they are summarized and provided.

[0290] Also, the server monitors specific phrases in the summary information and, when an abnormal or urgent case is detected, sends a warning signal to the user's terminal in real time. A push notification service for smartphones is used for this notification.

[0291] Furthermore, the device has a function to continuously monitor biometric information, and if an abnormal value is recorded, the data is immediately sent to the server. This information is notified to the user as an emergency alert, prompting early action.

[0292] An example of a prompt might be, "Please analyze the recording of a conversation with an elderly person and generate a summary regarding their health and safety." This allows the generative model to analyze the conversation content and organize and provide information according to the purpose.

[0293] By combining the technologies described above, this system provides a means to remotely, comprehensively, and effectively monitor the health and safety of the elderly.

[0294] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0295] Step 1:

[0296] The device transmits an audio signal to the elderly person at a specified time. The input is the time information set within the device. The output is the audio signal transmitted by the device. This audio signal contains messages designed to facilitate everyday conversation.

[0297] Step 2:

[0298] The device records conversations with elderly individuals in response to voice signals. The input is the voice of the elderly person, which is converted into digital audio data and output. Specifically, this process involves capturing the voice through the device's microphone and converting it into a digital format.

[0299] Step 3:

[0300] The terminal encrypts the generated digital audio data and sends it to the server via the Internet. The input is the recorded audio data, and the output is the encrypted data packets. The AES algorithm is used for encryption to ensure the protection of privacy.

[0301] Step 4:

[0302] The server decrypts the received encrypted data to obtain the digital audio data. The input is the encrypted data, and the output is the decrypted audio data. As a specific operation, the data is decoded using a decryption algorithm.

[0303] Step 5:

[0304] The server analyzes the audio data using a generative AI model and extracts the key points of the conversation. The input is the decrypted audio data, and the output is the summary information. Natural language processing technology is utilized to identify keywords and important contexts.

[0305] Step 6:

[0306] The server stores the extracted summary information in a storage means and makes it accessible to the user. The input is the summary information, and the output is stored as data in a storage device. As a specific operation, it includes an information writing operation to a database.

[0307] Step 7:

[0308] The server monitors the keywords included in the summary information and sends a warning signal to the user's terminal if an abnormal or emergency situation is detected. The input is the keyword to be monitored, and the output is the warning signal sent to the terminal. An alert is triggered via a push notification service.

[0309] Step 8:

[0310] The device continuously acquires health data of elderly individuals using its built-in biosensors, and if an abnormal value is detected, it sends that information to a server. The input is the data obtained from the biosensors, and the output is the abnormal data sent to the server. Specifically, it monitors the data feed from the sensors in real time and generates an alert if a threshold is exceeded.

[0311] (Application Example 1)

[0312] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0313] There is a need for a system that can monitor the health status of the elderly in real time, facilitate daily communication, and immediately notify families when abnormalities are detected. However, many current systems have limitations in the frequency and content of communication, and are unable to adequately provide real-time notifications of important health information.

[0314] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0315] In this invention, the server includes means for transmitting voice signals from an information processing device to the user at a fixed time each day to facilitate language exchange; means for extracting the key points of the language exchange using a generative model and forming summary information; and means for continuously monitoring biometric information and immediately generating and transmitting a warning when an abnormal value is detected. This enables real-time monitoring of the health status of the elderly and rapid notification in emergencies.

[0316] An "information processing device" is a device that transmits audio signals to facilitate language communication and processes data.

[0317] "User" refers to an individual who receives audio signals via an information processing device and uses the service.

[0318] "Scheduled time" refers to a specific time when audio signals are transmitted or data processing takes place.

[0319] An "audio signal" is a means of transmitting information via sound to a user.

[0320] "Language exchange" refers to communication conducted through voice between the user and the system.

[0321] A "generative model" is a machine learning model that extracts key points from the content of linguistic exchange and forms summarized information.

[0322] "Summary information" refers to information that briefly summarizes the important points of language exchange.

[0323] A "storage medium" is a device or system that stores summary information and makes it accessible as needed.

[0324] A "keyword" is a specific, important word included in the summary information that triggers a warning or notification.

[0325] A "warning signal" is a signal sent to the user's device to alert them when an abnormality is detected.

[0326] "Biometric information" refers to data that indicates the user's health status, including physiological information such as heart rate and activity level.

[0327] An "abnormal value" is a value detected in biological information as an unexpected or exceeding standard.

[0328] A "warning" is a notification that alerts you to an abnormality based on your biometric information, and is primarily sent to family members or medical support staff.

[0329] To implement this invention, a smartwatch or portable device worn by an elderly person is used. The device transmits an audio signal to the elderly person at a set time each day, facilitating communication. The device records the elderly person's conversations, saves them as digital audio data in real time, and transmits them to a server in an encrypted format. This server analyzes the received audio data using a generation AI model, extracts the key points, and generates summary information. The generated summary information is stored on a storage medium and kept accessible to family members to understand the elderly person's situation. In addition, the device constantly monitors the user's biometric information, and if abnormal values ​​such as heart rate or activity level are detected, it immediately generates a warning and sends a notification to the family.

[0330] As a concrete example, the server analyzes the data received from the smartwatch, and if the heart rate exceeds the normal range, it immediately sends a message to the family's smartphone such as, "The user's heart rate is high. Is there anything wrong?" At this stage, the generative AI model detects summary information and reports specific details such as "Activity level increased yesterday" concisely to the family.

[0331] An example of a prompt to be input into the generation AI model is, "Extract the main points of an elderly person's everyday conversation, and extract and summarize keywords that include important medical information." Based on this prompt, the system automatically organizes the important information and immediately issues a warning if any abnormalities occur.

[0332] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0333] Step 1:

[0334] The device, such as a smartwatch or other portable device, transmits voice signals to the elderly person to facilitate everyday conversation. The input for this step is the device's scheduled time information, and the output is the start of a conversation with the elderly person. By sending voice signals, it initiates natural communication.

[0335] Step 2:

[0336] The device records the conversation and digitizes it as audio data. The input for this step is the audio signal from the microphone, and the output is digital audio data. The digital data is stored and prepared for encryption in preparation for subsequent processing.

[0337] Step 3:

[0338] The terminal encrypts digital audio data using encryption technology and sends it to the server in a secure format. The input for this step is unprocessed digital audio data, and the output is encrypted data. Encryption is performed for security purposes.

[0339] Step 4:

[0340] The server decrypts the received encrypted data and converts it into processable audio data. The input for this step is encrypted audio data, and the output is the decrypted audio data. The server then prepares the data for appropriate processing.

[0341] Step 5:

[0342] The server uses a generative AI model to extract key points from audio data and generate summary information. The input for this step is decoded audio data, and the output is the summary information. By extracting key points, it efficiently presents important information.

[0343] Step 6:

[0344] The server stores the summary information on a storage medium and makes it accessible. The input for this step is the generated summary information, and the output is a database of the stored information. The environment is set up so that users can access the information when needed.

[0345] Step 7:

[0346] The server continuously monitors biometric data and, upon detecting an anomaly, immediately generates an alert and sends a notification to the family. The input for this step is real-time biometric data sent from the smartwatch, and the output is an alert message. Processing is performed to ensure rapid notification in emergencies.

[0347] Step 8:

[0348] The user receives summary information and warnings sent by the server to monitor their daily health status and make decisions regarding emergency responses. The input for this step is the notifications sent from the server, and the output is the user's decision. The user then takes appropriate action based on the information.

[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0350] This invention is a system that recognizes the emotions of an elderly person during a conversation and provides that emotional information to their family. The system consists of an information processing device, a generative model, a memory device, an emotion engine, and a communication means.

[0351] First, the device uses a smart device to transmit an audio signal to the elderly person at a set time. This signal initiates a conversation. Once the conversation begins, the device records the audio data and converts it into digital data in real time. This involves transmitting the data to a server via a communication line. The data transfer is encrypted to protect privacy.

[0352] Next, the server analyzes the received audio data using a generative model to extract the main points of the conversation and generate summary information. This summary includes an emotion engine, which also includes emotional states detected during the conversation. The emotion engine identifies emotional data based on facial recognition and speech tone analysis and adds it to the summary information.

[0353] Summary information and sentiment data are stored in the server's storage device and made accessible to the user. The stored data is organized according to specific conditions, and if, for example, a particular emotional state or keyword is detected, an alert signal is sent to the user's device. This allows family members to quickly grasp the necessary information and respond appropriately to the situation.

[0354] Furthermore, the device continuously acquires biometric data and immediately sends a signal to the server if it detects any abnormal values. This information is provided to the user as an emergency alert to encourage prompt action.

[0355] For example, if an elderly person says "I've been feeling irritable a lot lately" during a conversation, the emotion engine will determine that their emotions are unstable. Based on this, the summary information will record "Recently, their emotions have been unstable," and the user will receive a notification immediately. This system allows families to have a detailed understanding of the elderly person's emotions and health condition and to provide necessary care quickly.

[0356] The following describes the processing flow.

[0357] Step 1:

[0358] The user installs the application and performs initial setup on a screen where they can configure the voice notification time, emotion recognition settings, and warning signal conditions. This configuration information is sent from the device to the server and stored there.

[0359] Step 2:

[0360] At a set time, the device emits an audio signal to prompt the elderly person to begin a conversation. The audio signal includes greetings and questions delivered through a speaker, such as "Hello, how was your day?"

[0361] Step 3:

[0362] When a conversation begins, the device records the audio data in real time and saves it as digital data. This data is encrypted and sent to the server via the communication line.

[0363] Step 4:

[0364] The server analyzes the received audio data using a generative model. The generative model utilizes natural language processing techniques to grasp the main points of the conversation and create a summary. At this point, the emotion engine operates, analyzing the voice tone and word choice to extract emotion data.

[0365] Step 5:

[0366] The server adds sentiment data to the generated summary information and stores it in storage. This makes the user able to access the summary, including sentiment information, at any time.

[0367] Step 6:

[0368] If the server detects keywords or emotional states that meet specific criteria from the summary information, it generates a warning signal and sends it to the user's terminal. For example, if the emotion engine detects "unstable emotions," a notification is immediately sent to the user.

[0369] Step 7:

[0370] The device uses biometric measurement functions to measure health data of elderly individuals (e.g., heart rate and activity level). If abnormal values ​​are detected, the device sends the collected data to a server, and an emergency alert is sent to the user.

[0371] This entire process provides a system that accurately assesses the health and emotional state of elderly individuals remotely, enabling families to respond quickly.

[0372] (Example 2)

[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0374] There are challenges such as a lack of communication with the elderly and the inability for families to quickly grasp changes in their health and take appropriate action. Furthermore, there is a lack of means to appropriately extract important information and emotional changes from conversations and promptly notify families.

[0375] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0376] In this invention, the server includes means for periodically transmitting voice signals from an information processing device to a user to facilitate dialogue, means for extracting the main points of the dialogue and generating summary information using a generation algorithm, and means for adding emotion data generated by facial recognition and voice tone analysis included in the summary information. This enables family members to grasp changes in the emotional and health conditions of elderly people in real time and take necessary actions quickly.

[0377] An "information processing device" is a computer system that receives audio data and performs calculations for analysis and generation.

[0378] "Users" refers to individuals or organizations that receive conversational information and emotional data from elderly people through this system.

[0379] A "generative algorithm" is a computer program that extracts key points and features from audio data and performs summarization and sentiment analysis.

[0380] "Emotional data" refers to information that indicates the emotional state of a subject, generated based on changes in voice tone and facial recognition.

[0381] "Facial recognition" is a technology that uses image processing to analyze the movements and expressions of a person's face and determine their emotions.

[0382] "Voice tone analysis" is a process that analyzes the intonation and speed of speech to evaluate the speaker's emotions and state of mind.

[0383] A "storage device" is a storage medium for permanently or temporarily recording digital information.

[0384] A "warning signal" is alert information sent to the user's device when specific conditions are detected.

[0385] This invention is a system that analyzes the emotional and health status of elderly individuals through communication and provides this information to their families. The embodiments are described in detail below.

[0386] The device utilizes smart technology to transmit voice signals to elderly individuals and initiate conversations. These voice signals are transmitted automatically based on time. High-performance microphones and voice recognition software are used to convert the voice into digital data in real time, and encryption technology (e.g., AES) is used to transmit it to the server to maintain privacy.

[0387] The server uses a generative AI model to analyze the received audio data. The analysis extracts key points and important information from the conversation and generates a summary. Furthermore, the server incorporates an emotion engine that generates emotion data through voice tone analysis and facial expression recognition, and adds this to the summary.

[0388] The generated summary information and sentiment data are recorded on the server's storage device and made accessible to the user. The system has a function to immediately send a warning signal to the user's terminal when specific emotional states or keywords are detected. This allows families to obtain necessary information in real time and respond promptly to the elderly person's condition.

[0389] Furthermore, the device continuously monitors the biometric indicators of elderly individuals, and if abnormal data is detected, the server sends an urgent notification to the user. This collaboration enables early detection of health abnormalities and supports prompt response.

[0390] As a concrete example, if an elderly person mentions during a conversation that they have been feeling irritable lately, the emotion engine will determine that their condition is unstable, record it in the summary information, and immediately notify the user. This functionality allows families to appropriately monitor changes in the elderly person's emotions and health and provide necessary support quickly.

[0391] Example of a prompt:

[0392] "Please explain the mechanism for notifying a summary of information if unstable emotions are detected in recent conversations."

[0393] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0394] Step 1:

[0395] The terminal initiates a conversation by sending an audio signal to the elderly person using a smart device based on a set time. The input is the set time information, and the output is the transmission of an audio signal. At this time, the program utilizes the timer function of the smart device to execute the process of sending an audio signal at the specified time.

[0396] Step 2:

[0397] The device records the initiated conversation using a high-precision microphone and converts the audio data into digital data in real time. The input is the audio signal of the conversation, and the output is the digitized audio data. Speech recognition software operates in the background, analyzing the input audio waveform and converting it into text data and speech feature data.

[0398] Step 3:

[0399] The terminal encrypts digital data using AES encryption technology and transmits it to the server via the communication line while protecting privacy. The input is digitized voice data, and the output is encrypted data communication. The encryption module activates and converts the plaintext data based on the encryption key.

[0400] Step 4:

[0401] The server decrypts the received encrypted data and analyzes the audio data using a generative AI model. The input is encrypted audio data, and the output is summary information and sentiment data. The generative AI model simultaneously runs speech-to-text conversion and sentiment analysis algorithms to extract the main points and emotional state of the conversation.

[0402] Step 5:

[0403] The server uses an emotion engine to perform facial recognition and voice tone analysis based on the generated summary information. The input is the summary information, and the output is the summary information with added emotion data. The emotion engine evaluates various voice parameters and facial expression data, classifying emotions as either numerical or categorical.

[0404] Step 6:

[0405] The server stores the final summary information and sentiment data in storage, keeping it accessible to users. The input is the summary information and sentiment data, and the output is the entries in the stored database. Database management software organizes and indexes the data in an automated process.

[0406] Step 7:

[0407] The server immediately sends an alert signal to the user's device when a specific emotional state or keyword is detected. The input consists of stored summary information and triggering conditions, while the output is the alert signal. The alert system continuously monitors the system and activates a notification protocol when the conditions are met.

[0408] Step 8:

[0409] The device continuously acquires biometric data from elderly individuals via biosensors and reports any abnormal measurements to the server. The input is biometric data, and the output is a notification signal containing abnormal values. A vital signs monitoring system analyzes the data and triggers an alert if a threshold is exceeded.

[0410] (Application Example 2)

[0411] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0412] In situations where elderly people experience anxiety and loneliness on a daily basis, it is difficult for family members and caregivers to quickly recognize these emotional changes and respond appropriately. Therefore, there is a need for a system that monitors the physical and mental state of elderly people in real time and promptly notifies relevant parties if any abnormalities are detected.

[0413] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0414] In this invention, the server includes means for transmitting an audio signal from an information processing device to a user at a fixed time each day to encourage conversation; means for extracting the main points of the conversation using a generative model and generating summary information; means for storing the summary information in a storage device and making it accessible; means for analyzing the audio data in real time and evaluating the emotional state; and means for sending a notification to relevant parties when an abnormal emotional state is detected. This makes it possible to quickly grasp the physical and mental state of elderly people and respond promptly as needed.

[0415] An "information processing device" is an electronic device that transmits audio signals to the user to facilitate conversation.

[0416] A "generative model" is an algorithmic method for extracting the main points of a conversation and summarizing information.

[0417] A "storage device" is a digital storage device used to store summary information and make it accessible as needed.

[0418] "Emotion recognition means" refers to technology that analyzes voice data and evaluates the user's emotional state.

[0419] A "portable communication device" is a portable communication device used to transmit voice signals.

[0420] An "abnormal value" is a value in biological data that exceeds the normal range and indicates a condition that requires attention.

[0421] "Means of sending notifications to relevant parties" refers to methods for communicating important information to users and stakeholders in real time.

[0422] The system implementing this invention consists of an information processing device, a generative model, a storage device, an emotion recognition engine, and a communication means. The server has a function to prompt conversation by sending an audio signal to the user at a fixed time every day via the information processing device. When a conversation begins, the terminal converts the audio data into digital data in real time, encrypts it, and sends it to the server. The server converts the audio data into text data using software such as Google Cloud Speech-to-Text, and based on that, extracts the main points of the conversation using a generative model and generates summary information. In this process, an emotion recognition engine such as IBM Watson Tone Analyzer is used to perform voice tone analysis and evaluate the emotional state. If an anomaly is detected by emotion recognition, a push notification is sent to the relevant parties.

[0423] For example, if a device records a conversation with an elderly person and detects a phrase like "I haven't been able to sleep lately," the emotion recognition engine will detect "anxiety," and the user will be promptly notified with a message such as "You've been experiencing persistent anxiety lately." An example of a prompt to the generative AI model used in this case would be, "Analyze the following conversation emotionally and detect any significant changes: 'I've been having more trouble sleeping lately...'" In this way, the server and the device work together, allowing users to quickly understand the emotions and health status of elderly people.

[0424] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0425] Step 1:

[0426] The terminal transmits an audio signal to the user (an elderly person) at regular intervals via an information processing device. The input is the scheduled alarm setting data, and the output is an audio signal sent to the user. This action prompts the user to initiate conversation.

[0427] Step 2:

[0428] As soon as a conversation with the user begins, the device converts the audio data into digital data in real time. The input is raw audio data, and the output is a stream of digital data. The audio data is captured on the device via the microphone and converted into text data using Google Cloud Speech-to-Text.

[0429] Step 3:

[0430] The terminal encrypts the converted text data and sends it to the server. The input is text data, and the output is encrypted text data. AES encryption or similar methods are used to protect the data during transmission.

[0431] Step 4:

[0432] The server inputs the received text data into a generative model to extract the main points of the conversation. The input is decrypted text data, and the output is summarized information. Natural language processing techniques are applied to extract the most important parts of the conversation.

[0433] Step 5:

[0434] The server processes the extracted summary information into an emotion recognition engine to analyze the user's emotional state. The input is summary information, and the output is emotional state data. IBM Watson Tone Analyzer is used to analyze the tone of the text and identify emotions.

[0435] Step 6:

[0436] The server detects anomalies from the emotional state data in the analysis results and sends notifications to relevant parties as needed. The input is emotional state data, and the output is a notification message. When an anomaly is detected, a push notification function is used to send an alert to family members or caregivers with specific information.

[0437] Throughout this entire process, users can quickly grasp the emotions and health status of elderly individuals.

[0438] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0439] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0440] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0441] [Third Embodiment]

[0442] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0443] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0444] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0445] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0446] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0447] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0448] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0449] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0450] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0451] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0452] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0453] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0454] This invention provides a system that facilitates natural communication among the elderly in their daily lives and enables them to share safety and health-related information with family members living remotely. This system primarily consists of an information processing device, a generative model, a storage device, and various communication means.

[0455] First, the device uses a portable communication device such as a smartphone or smartwatch to send an audio signal to the elderly person at a set time. The audio signal prompts the elderly person to start a conversation, and everyday communication is automatically conducted.

[0456] When a conversation begins, the device records the content in real time and saves it as digital audio data. This audio data is transmitted to a server via the internet. Encryption is used during data transfer to ensure privacy protection.

[0457] Next, in the process of processing the received audio data, the server uses a generative model to extract the main points of the conversation and generate summary information as needed. This summary information is then stored in a memory device and made accessible to the user when they need to review it.

[0458] The summary information is organized based on important keywords and context to allow families to easily understand the elderly person's situation. For example, if health-related information is included in a conversation, summaries such as "I felt fine yesterday" or "I have a doctor's appointment today" are generated.

[0459] Furthermore, the server monitors keywords within the summary information based on specific conditions, and if an anomaly or urgent situation is detected, it sends a warning signal to the user's terminal. This ensures that necessary information is delivered immediately, even when the user is in a remote location.

[0460] Furthermore, the system includes a biometric measurement function built into the terminal to continuously acquire health data from elderly individuals, and to immediately send a signal to the server if any abnormal values ​​are detected. This information is provided to the user as an emergency alert, serving as a means to encourage early action.

[0461] In this way, the present invention provides comprehensive remote management of the health and safety of the elderly, offering peace of mind to their families.

[0462] The following describes the processing flow.

[0463] Step 1:

[0464] The user downloads the application and completes the initial setup. On the settings screen, they specify settings such as the timing of voice notifications, the voice of the AI ​​avatar, and the conditions for warning signals. Once the setup is complete, the device sends this information to the server, and the account is registered.

[0465] Step 2:

[0466] At the designated time, the device emits an audio signal to prompt the elderly person to begin a conversation. This includes a greeting message played through the speaker. For example, a message such as "Good morning. How are you today?" might be played.

[0467] Step 3:

[0468] When a conversation begins, the device records audio data. The recorded audio is converted into digital data in real time and sent to the server via the communication line. Encryption is applied during data transfer to protect privacy.

[0469] Step 4:

[0470] The server analyzes the received audio data using a generative model. Here, the key points of the conversation are extracted, and summary information is generated. The generative model is based on natural language processing techniques and summarizes important information concisely and clearly.

[0471] Step 5:

[0472] The summary information is stored in the server's storage device. The stored data is made accessible through the user interface so that users can check it at any time.

[0473] Step 6:

[0474] Based on specific criteria, the server searches for important keywords in the summary information. If keywords matching the criteria are found, a warning signal is generated and a notification is sent to the user's terminal. This allows the user to receive information that enables them to take immediate action.

[0475] Step 7:

[0476] The device continuously records biometric data using its built-in sensors and detects abnormal values. If an abnormality is detected, a signal is immediately sent to the server, and an emergency alert is issued to the user.

[0477] Through the process described above, this system can accurately grasp the situation of elderly people remotely and take necessary actions quickly.

[0478] (Example 1)

[0479] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0480] There is a need for a system that allows elderly people to communicate naturally on a daily basis and effectively monitor their health remotely. Furthermore, it is necessary to ensure that abnormal situations are detected quickly while protecting the privacy of the elderly, enabling family members in remote locations to respond immediately.

[0481] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0482] In this invention, the server includes means for transmitting voice signals from an information processing device to a user at regular intervals to encourage conversation, means for extracting the main points of the conversation using a generative model and generating summary information, and means for encrypting and transmitting digital voice data. This makes it possible for family members who live far away to reliably understand the health and safety of the elderly, and to quickly detect abnormal situations while protecting their privacy.

[0483] An "information processing device" is a device that performs data input, processing, storage, and output, and has the function of exchanging information with other devices via a communication network.

[0484] "Users" refers to individuals who use this system, including, but not limited to, the elderly and their families.

[0485] An "audio signal" is a signal that is created by electronically converting sound and transmitting it via a communication means.

[0486] A "generative model" refers to an algorithm or system that uses machine learning to extract specific patterns or key points from input data.

[0487] "Summary information" refers to information that organizes the important points and keywords extracted from conversation content using a generative model.

[0488] "Means of storage" refers to physical or electronic means of storing data and keeping it accessible as needed.

[0489] "Encryption" is a technology that transforms the content of data based on a specific algorithm to prevent unauthorized access or decryption.

[0490] A "portable communication device" refers to a device that is portable and uses wireless technology to communicate.

[0491] "Privacy" is the right or state to prevent an individual's private information from being used or made public without their consent.

[0492] An "abnormal situation" refers to a state in which unusual circumstances or values ​​are detected, and is sometimes used particularly in relation to health.

[0493] The system of this invention aims to enable elderly people to communicate naturally in their daily lives and to effectively provide health information to distant family members. The system includes an information processing device, a communication device, a generative AI model, and a storage means.

[0494] The devices used are portable communication devices such as smartphones and smartwatches. These devices transmit voice signals to elderly people at designated times to encourage everyday conversation. The voice signals include time announcements and messages to start a conversation.

[0495] After an audio signal is transmitted, the terminal records the elderly person's voice and saves it as digital audio data. This audio data is encrypted via an information processing device and transferred to a server over the internet. The AES algorithm is used for encryption, ensuring data privacy.

[0496] The server decodes the received audio data and analyzes it using a generative AI model. During the analysis, natural language processing techniques are used to extract the main points of the conversation and summarize important information. The generated summary information is stored in a memory device for easy access by the user.

[0497] Users access summarized information posted to allow family members to easily understand the condition of elderly individuals, using a web browser or dedicated application. This ensures that important information is clearly organized and displayed. For example, if everyday conversations include information such as "I felt fine yesterday" or "I have an appointment to see a doctor today," this information is summarized and provided.

[0498] Furthermore, the server monitors specific phrases within the summary information, and if an anomaly or urgent case is detected, it sends a warning signal to the user's device in real time. This notification uses a push notification service for smartphones.

[0499] Furthermore, the device has a function to continuously monitor biometric information, and if an abnormal value is recorded, the data is immediately sent to the server. This information is notified to the user as an emergency alert, prompting early action.

[0500] An example of a prompt might be, "Please analyze the recording of a conversation with an elderly person and generate a summary regarding their health and safety." This allows the generative model to analyze the conversation content and organize and provide information according to the purpose.

[0501] By combining the technologies described above, this system provides a means to remotely, comprehensively, and effectively monitor the health and safety of the elderly.

[0502] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0503] Step 1:

[0504] The device transmits an audio signal to the elderly person at a specified time. The input is the time information set within the device. The output is the audio signal transmitted by the device. This audio signal contains messages designed to facilitate everyday conversation.

[0505] Step 2:

[0506] The device records conversations with elderly individuals in response to voice signals. The input is the voice of the elderly person, which is converted into digital audio data and output. Specifically, this process involves capturing the voice through the device's microphone and converting it into a digital format.

[0507] Step 3:

[0508] The terminal encrypts the generated digital audio data and sends it to the server over the internet. The input is the recorded audio data, and the output is an encrypted data packet. The AES algorithm is used for encryption, ensuring privacy protection.

[0509] Step 4:

[0510] The server decrypts the received encrypted data and obtains digital audio data. The input is encrypted data, and the output is decrypted audio data. Specifically, it decodes the data using a cryptographic decryption algorithm.

[0511] Step 5:

[0512] The server uses a generative AI model to analyze audio data and extract the main points of the conversation. The input is decoded audio data, and the output is summarized information. Natural language processing techniques are used to identify keywords and important context.

[0513] Step 6:

[0514] The server stores the extracted summary information in a storage device and makes it accessible to the user. The input is the summary information, and the output is stored as saved data in the storage device. Specific operations include writing information to a database.

[0515] Step 7:

[0516] The server monitors keywords included in the summary information and sends a warning signal to the user's terminal if an anomaly or emergency condition is detected. The input is the keyword to be monitored, and the output is the warning signal sent to the terminal. The alert is triggered via a push notification service.

[0517] Step 8:

[0518] The device continuously acquires health data of elderly individuals using its built-in biosensors, and if an abnormal value is detected, it sends that information to a server. The input is the data obtained from the biosensors, and the output is the abnormal data sent to the server. Specifically, it monitors the data feed from the sensors in real time and generates an alert if a threshold is exceeded.

[0519] (Application Example 1)

[0520] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0521] There is a need for a system that can monitor the health status of the elderly in real time, facilitate daily communication, and immediately notify families when abnormalities are detected. However, many current systems have limitations in the frequency and content of communication, and are unable to adequately provide real-time notifications of important health information.

[0522] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0523] In this invention, the server includes means for transmitting voice signals from an information processing device to the user at a fixed time each day to facilitate language exchange; means for extracting the key points of the language exchange using a generative model and forming summary information; and means for continuously monitoring biometric information and immediately generating and transmitting a warning when an abnormal value is detected. This enables real-time monitoring of the health status of the elderly and rapid notification in emergencies.

[0524] An "information processing device" is a device that transmits audio signals to facilitate language communication and processes data.

[0525] "User" refers to an individual who receives audio signals via an information processing device and uses the service.

[0526] "Scheduled time" refers to a specific time when audio signals are transmitted or data processing takes place.

[0527] An "audio signal" is a means of transmitting information via sound to a user.

[0528] "Language exchange" refers to communication conducted through voice between the user and the system.

[0529] A "generative model" is a machine learning model that extracts key points from the content of linguistic exchange and forms summarized information.

[0530] "Summary information" refers to information that briefly summarizes the important points of language exchange.

[0531] A "storage medium" is a device or system that stores summary information and makes it accessible as needed.

[0532] A "keyword" is a specific, important word included in the summary information that triggers a warning or notification.

[0533] A "warning signal" is a signal sent to the user's device to alert them when an abnormality is detected.

[0534] "Biometric information" refers to data that indicates the user's health status, including physiological information such as heart rate and activity level.

[0535] An "abnormal value" is a value detected in biological information as an unexpected or exceeding standard.

[0536] A "warning" is a notification that alerts you to an abnormality based on your biometric information, and is primarily sent to family members or medical support staff.

[0537] To implement this invention, a smartwatch or portable device worn by an elderly person is used. The device transmits an audio signal to the elderly person at a set time each day, facilitating communication. The device records the elderly person's conversations, saves them as digital audio data in real time, and transmits them to a server in an encrypted format. This server analyzes the received audio data using a generation AI model, extracts the key points, and generates summary information. The generated summary information is stored on a storage medium and kept accessible to family members to understand the elderly person's situation. In addition, the device constantly monitors the user's biometric information, and if abnormal values ​​such as heart rate or activity level are detected, it immediately generates a warning and sends a notification to the family.

[0538] As a concrete example, the server analyzes the data received from the smartwatch, and if the heart rate exceeds the normal range, it immediately sends a message to the family's smartphone such as, "The user's heart rate is high. Is there anything wrong?" At this stage, the generative AI model detects summary information and reports specific details such as "Activity level increased yesterday" concisely to the family.

[0539] An example of a prompt to be input to the generation AI model is, "Extract the main points of an elderly person's everyday conversation, and extract and summarize keywords that include important medical information." Based on this prompt, the system automatically organizes the important information and immediately issues a warning if any abnormalities occur.

[0540] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0541] Step 1:

[0542] The device, such as a smartwatch or other portable device, transmits voice signals to the elderly person to facilitate everyday conversation. The input for this step is the device's scheduled time information, and the output is the start of a conversation with the elderly person. By sending voice signals, it initiates natural communication.

[0543] Step 2:

[0544] The device records the conversation and digitizes it as audio data. The input for this step is the audio signal from the microphone, and the output is digital audio data. The digital data is stored and prepared for encryption in preparation for subsequent processing.

[0545] Step 3:

[0546] The terminal encrypts digital audio data using encryption technology and sends it to the server in a secure format. The input for this step is unprocessed digital audio data, and the output is encrypted data. Encryption is performed for security purposes.

[0547] Step 4:

[0548] The server decrypts the received encrypted data and converts it into processable audio data. The input for this step is encrypted audio data, and the output is the decrypted audio data. The server then prepares the data for appropriate processing.

[0549] Step 5:

[0550] The server uses a generative AI model to extract key points from audio data and generate summary information. The input for this step is decoded audio data, and the output is the summary information. By extracting key points, it efficiently presents important information.

[0551] Step 6:

[0552] The server stores the summary information on a storage medium and makes it accessible. The input for this step is the generated summary information, and the output is a database of the stored information. The environment is set up so that users can access the information when needed.

[0553] Step 7:

[0554] The server continuously monitors biometric data and, upon detecting an anomaly, immediately generates an alert and sends a notification to the family. The input for this step is real-time biometric data sent from the smartwatch, and the output is an alert message. Processing is performed to ensure rapid notification in emergencies.

[0555] Step 8:

[0556] The user receives summary information and warnings sent by the server to monitor their daily health status and make decisions regarding emergency responses. The input for this step is the notifications sent from the server, and the output is the user's decision. The user then takes appropriate action based on the information.

[0557] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0558] This invention is a system that recognizes the emotions of an elderly person during a conversation and provides that emotional information to their family. The system consists of an information processing device, a generative model, a memory device, an emotion engine, and a communication means.

[0559] First, the device uses a smart device to transmit an audio signal to the elderly person at a set time. This signal initiates a conversation. Once the conversation begins, the device records the audio data and converts it into digital data in real time. This involves transmitting the data to a server via a communication line. The data transfer is encrypted to protect privacy.

[0560] Next, the server analyzes the received audio data using a generative model to extract the main points of the conversation and generate summary information. This summary includes an emotion engine, which also includes emotional states detected during the conversation. The emotion engine identifies emotional data based on facial recognition and speech tone analysis and adds it to the summary information.

[0561] Summary information and sentiment data are stored in the server's storage device and made accessible to the user. The stored data is organized according to specific conditions, and if, for example, a particular emotional state or keyword is detected, an alert signal is sent to the user's device. This allows family members to quickly grasp the necessary information and respond appropriately to the situation.

[0562] Furthermore, the device continuously acquires biometric data and immediately sends a signal to the server if it detects any abnormal values. This information is provided to the user as an emergency alert to encourage prompt action.

[0563] For example, if an elderly person says "I've been feeling irritable a lot lately" during a conversation, the emotion engine will determine that their emotions are unstable. Based on this, the summary information will record "Recently, their emotions have been unstable," and the user will receive a notification immediately. This system allows families to have a detailed understanding of the elderly person's emotions and health condition and to provide necessary care quickly.

[0564] The following describes the processing flow.

[0565] Step 1:

[0566] The user installs the application and performs initial setup on a screen where they can configure the voice notification time, emotion recognition settings, and warning signal conditions. This configuration information is sent from the device to the server and stored there.

[0567] Step 2:

[0568] At a set time, the device emits an audio signal to prompt the elderly person to begin a conversation. The audio signal includes greetings and questions delivered through a speaker, such as "Hello, how was your day?"

[0569] Step 3:

[0570] When a conversation begins, the device records the audio data in real time and saves it as digital data. This data is encrypted and sent to the server via the communication line.

[0571] Step 4:

[0572] The server analyzes the received audio data using a generative model. The generative model utilizes natural language processing techniques to grasp the main points of the conversation and create a summary. At this point, the emotion engine operates, analyzing the voice tone and word choice to extract emotion data.

[0573] Step 5:

[0574] The server adds sentiment data to the generated summary information and stores it in storage. This makes the user able to access the summary, including sentiment information, at any time.

[0575] Step 6:

[0576] If the server detects keywords or emotional states that meet specific criteria from the summary information, it generates a warning signal and sends it to the user's terminal. For example, if the emotion engine detects "unstable emotions," a notification is immediately sent to the user.

[0577] Step 7:

[0578] The device uses biometric measurement functions to measure health data of elderly individuals (e.g., heart rate and activity level). If abnormal values ​​are detected, the device sends the collected data to a server, and an emergency alert is sent to the user.

[0579] This entire process provides a system that accurately assesses the health and emotional state of elderly individuals remotely, enabling families to respond quickly.

[0580] (Example 2)

[0581] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0582] There are challenges such as a lack of communication with the elderly and the inability for families to quickly grasp changes in their health and take appropriate action. Furthermore, there is a lack of means to appropriately extract important information and emotional changes from conversations and promptly notify families.

[0583] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0584] In this invention, the server includes means for periodically transmitting voice signals from an information processing device to a user to facilitate dialogue, means for extracting the main points of the dialogue and generating summary information using a generation algorithm, and means for adding emotion data generated by facial recognition and voice tone analysis included in the summary information. This enables family members to grasp changes in the emotional and health conditions of elderly people in real time and take necessary actions quickly.

[0585] An "information processing device" is a computer system that receives audio data and performs calculations for analysis and generation.

[0586] "Users" refers to individuals or organizations that receive conversational information and emotional data from elderly people through this system.

[0587] A "generative algorithm" is a computer program that extracts key points and features from audio data and performs summarization and sentiment analysis.

[0588] "Emotional data" refers to information that indicates the emotional state of a subject, generated based on changes in voice tone and facial recognition.

[0589] "Facial recognition" is a technology that uses image processing to analyze the movements and expressions of a person's face and determine their emotions.

[0590] "Voice tone analysis" is a process that analyzes the intonation and speed of speech to evaluate the speaker's emotions and state of mind.

[0591] A "storage device" is a storage medium for permanently or temporarily recording digital information.

[0592] A "warning signal" is alert information sent to the user's device when specific conditions are detected.

[0593] This invention is a system that analyzes the emotional and health status of elderly individuals through communication and provides this information to their families. The embodiments are described in detail below.

[0594] The device utilizes smart technology to transmit voice signals to elderly individuals and initiate conversations. These voice signals are transmitted automatically based on time. High-performance microphones and voice recognition software are used to convert the voice into digital data in real time, and encryption technology (e.g., AES) is used to transmit it to the server to maintain privacy.

[0595] The server uses a generative AI model to analyze the received audio data. The analysis extracts key points and important information from the conversation and generates a summary. Furthermore, the server incorporates an emotion engine that generates emotion data through voice tone analysis and facial expression recognition, and adds this to the summary.

[0596] The generated summary information and sentiment data are recorded on the server's storage device and made accessible to the user. The system has a function to immediately send a warning signal to the user's terminal when specific emotional states or keywords are detected. This allows families to obtain necessary information in real time and respond promptly to the elderly person's condition.

[0597] Furthermore, the device continuously monitors the biometric indicators of elderly individuals, and if abnormal data is detected, the server sends an urgent notification to the user. This collaboration enables early detection of health abnormalities and supports prompt response.

[0598] As a concrete example, if an elderly person mentions during a conversation that they have been feeling irritable lately, the emotion engine will determine that their condition is unstable, record it in the summary information, and immediately notify the user. This functionality allows families to appropriately monitor changes in the elderly person's emotions and health and provide necessary support quickly.

[0599] Example of a prompt:

[0600] "Please explain the mechanism for notifying a summary of information if unstable emotions are detected in recent conversations."

[0601] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0602] Step 1:

[0603] The terminal initiates a conversation by sending an audio signal to the elderly person using a smart device based on a set time. The input is the set time information, and the output is the transmission of an audio signal. At this time, the program utilizes the timer function of the smart device to execute the process of sending an audio signal at the specified time.

[0604] Step 2:

[0605] The device records the initiated conversation using a high-precision microphone and converts the audio data into digital data in real time. The input is the audio signal of the conversation, and the output is the digitized audio data. Speech recognition software operates in the background, analyzing the input audio waveform and converting it into text data and speech feature data.

[0606] Step 3:

[0607] The terminal encrypts digital data using AES encryption technology and transmits it to the server via the communication line while protecting privacy. The input is digitized voice data, and the output is encrypted data communication. The encryption module activates and converts the plaintext data based on the encryption key.

[0608] Step 4:

[0609] The server decrypts the received encrypted data and analyzes the audio data using a generative AI model. The input is encrypted audio data, and the output is summary information and sentiment data. The generative AI model simultaneously runs speech-to-text conversion and sentiment analysis algorithms to extract the main points and emotional state of the conversation.

[0610] Step 5:

[0611] The server uses an emotion engine to perform facial recognition and voice tone analysis based on the generated summary information. The input is the summary information, and the output is the summary information with added emotion data. The emotion engine evaluates various voice parameters and facial expression data, classifying emotions as either numerical or categorical.

[0612] Step 6:

[0613] The server stores the final summary information and sentiment data in storage, keeping it accessible to users. The input is the summary information and sentiment data, and the output is the entries in the stored database. Database management software organizes and indexes the data in an automated process.

[0614] Step 7:

[0615] The server immediately sends an alert signal to the user's device when a specific emotional state or keyword is detected. The input consists of stored summary information and triggering conditions, while the output is the alert signal. The alert system continuously monitors the system and activates a notification protocol when the conditions are met.

[0616] Step 8:

[0617] The device continuously acquires biometric data from elderly individuals via biosensors and reports any abnormal measurements to the server. The input is biometric data, and the output is a notification signal containing abnormal values. A vital signs monitoring system analyzes the data and triggers an alert if a threshold is exceeded.

[0618] (Application Example 2)

[0619] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0620] In situations where elderly people experience anxiety and loneliness on a daily basis, it is difficult for family members and caregivers to quickly recognize these emotional changes and respond appropriately. Therefore, there is a need for a system that monitors the physical and mental state of elderly people in real time and promptly notifies relevant parties if any abnormalities are detected.

[0621] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0622] In this invention, the server includes means for transmitting an audio signal from an information processing device to a user at a fixed time each day to encourage conversation; means for extracting the main points of the conversation using a generative model and generating summary information; means for storing the summary information in a storage device and making it accessible; means for analyzing the audio data in real time and evaluating the emotional state; and means for sending a notification to relevant parties when an abnormal emotional state is detected. This makes it possible to quickly grasp the physical and mental state of elderly people and respond promptly as needed.

[0623] An "information processing device" is an electronic device that transmits audio signals to the user to facilitate conversation.

[0624] A "generative model" is an algorithmic method for extracting the main points of a conversation and summarizing information.

[0625] A "storage device" is a digital storage device used to store summary information and make it accessible as needed.

[0626] "Emotion recognition means" refers to technology that analyzes voice data and evaluates the user's emotional state.

[0627] A "portable communication device" is a portable communication device used to transmit voice signals.

[0628] An "abnormal value" is a value in biological data that exceeds the normal range and indicates a condition that requires attention.

[0629] "Means of sending notifications to relevant parties" refers to methods for communicating important information to users and stakeholders in real time.

[0630] The system implementing this invention consists of an information processing device, a generative model, a storage device, an emotion recognition engine, and a communication means. The server has a function to prompt conversation by sending an audio signal to the user at a fixed time every day via the information processing device. When a conversation begins, the terminal converts the audio data into digital data in real time, encrypts it, and sends it to the server. The server converts the audio data into text data using software such as Google Cloud Speech-to-Text, and based on that, extracts the main points of the conversation using a generative model and generates summary information. In this process, an emotion recognition engine such as IBM Watson Tone Analyzer is used to perform voice tone analysis and evaluate the emotional state. If an anomaly is detected by emotion recognition, a push notification is sent to the relevant parties.

[0631] For example, if a device records a conversation with an elderly person and detects a phrase like "I haven't been able to sleep lately," the emotion recognition engine will detect "anxiety," and the user will be promptly notified with a message such as "You've been experiencing persistent anxiety lately." An example of a prompt to the generative AI model used in this case would be, "Analyze the following conversation emotionally and detect any significant changes: 'I've been having more trouble sleeping lately...'" In this way, the server and the device work together, allowing users to quickly understand the emotions and health status of elderly people.

[0632] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0633] Step 1:

[0634] The terminal transmits an audio signal to the user (an elderly person) at regular intervals via an information processing device. The input is the scheduled alarm setting data, and the output is an audio signal sent to the user. This action prompts the user to initiate conversation.

[0635] Step 2:

[0636] As soon as a conversation with the user begins, the device converts the audio data into digital data in real time. The input is raw audio data, and the output is a stream of digital data. The audio data is captured on the device via the microphone and converted into text data using Google Cloud Speech-to-Text.

[0637] Step 3:

[0638] The terminal encrypts the converted text data and sends it to the server. The input is text data, and the output is encrypted text data. AES encryption or similar methods are used to protect the data during transmission.

[0639] Step 4:

[0640] The server inputs the received text data into a generative model to extract the main points of the conversation. The input is decrypted text data, and the output is summarized information. Natural language processing techniques are applied to extract the most important parts of the conversation.

[0641] Step 5:

[0642] The server processes the extracted summary information into an emotion recognition engine to analyze the user's emotional state. The input is summary information, and the output is emotional state data. IBM Watson Tone Analyzer is used to analyze the tone of the text and identify emotions.

[0643] Step 6:

[0644] The server detects anomalies from the emotional state data in the analysis results and sends notifications to relevant parties as needed. The input is emotional state data, and the output is a notification message. When an anomaly is detected, a push notification function is used to send an alert to family members or caregivers with specific information.

[0645] Throughout this entire process, users can quickly grasp the emotions and health status of elderly individuals.

[0646] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0647] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0648] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0649] [Fourth Embodiment]

[0650] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0651] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0652] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0653] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0654] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0655] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0656] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0657] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0658] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0659] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0660] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0661] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0662] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0663] This invention provides a system that facilitates natural communication among the elderly in their daily lives and enables them to share safety and health-related information with family members living remotely. This system primarily consists of an information processing device, a generative model, a storage device, and various communication means.

[0664] First, the device uses a portable communication device such as a smartphone or smartwatch to send an audio signal to the elderly person at a set time. The audio signal prompts the elderly person to start a conversation, and everyday communication is automatically conducted.

[0665] When a conversation begins, the device records the content in real time and saves it as digital audio data. This audio data is transmitted to a server via the internet. Encryption is used during data transfer to ensure privacy protection.

[0666] Next, in the process of processing the received audio data, the server uses a generative model to extract the main points of the conversation and generate summary information as needed. This summary information is then stored in a memory device and made accessible to the user when they need to review it.

[0667] The summary information is organized based on important keywords and context to allow families to easily understand the elderly person's situation. For example, if health-related information is included in a conversation, summaries such as "I felt fine yesterday" or "I have a doctor's appointment today" are generated.

[0668] Furthermore, the server monitors keywords within the summary information based on specific conditions, and if an anomaly or urgent situation is detected, it sends a warning signal to the user's terminal. This ensures that necessary information is delivered immediately, even when the user is in a remote location.

[0669] Furthermore, the system includes a biometric measurement function built into the terminal to continuously acquire health data from elderly individuals, and to immediately send a signal to the server if any abnormal values ​​are detected. This information is provided to the user as an emergency alert, serving as a means to encourage early action.

[0670] In this way, the present invention provides comprehensive remote management of the health and safety of the elderly, offering peace of mind to their families.

[0671] The following describes the processing flow.

[0672] Step 1:

[0673] The user downloads the application and completes the initial setup. On the settings screen, they specify settings such as the timing of voice notifications, the voice of the AI ​​avatar, and the conditions for warning signals. Once the setup is complete, the device sends this information to the server, and the account is registered.

[0674] Step 2:

[0675] At the designated time, the device emits an audio signal to prompt the elderly person to begin a conversation. This includes a greeting message played through the speaker. For example, a message such as "Good morning. How are you today?" might be played.

[0676] Step 3:

[0677] When a conversation begins, the device records audio data. The recorded audio is converted into digital data in real time and sent to the server via the communication line. Encryption is applied during data transfer to protect privacy.

[0678] Step 4:

[0679] The server analyzes the received audio data using a generative model. Here, the key points of the conversation are extracted, and summary information is generated. The generative model is based on natural language processing techniques and summarizes important information concisely and clearly.

[0680] Step 5:

[0681] The summary information is stored in the server's storage device. The stored data is made accessible through the user interface so that users can check it at any time.

[0682] Step 6:

[0683] Based on specific criteria, the server searches for important keywords in the summary information. If keywords matching the criteria are found, a warning signal is generated and a notification is sent to the user's terminal. This allows the user to receive information that enables them to take immediate action.

[0684] Step 7:

[0685] The device continuously records biometric data using its built-in sensors and detects abnormal values. If an abnormality is detected, a signal is immediately sent to the server, and an emergency alert is issued to the user.

[0686] Through the process described above, this system can accurately grasp the situation of elderly people remotely and take necessary actions quickly.

[0687] (Example 1)

[0688] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0689] There is a need for a system that allows elderly people to communicate naturally on a daily basis and effectively monitor their health remotely. Furthermore, it is necessary to ensure that abnormal situations are detected quickly while protecting the privacy of the elderly, enabling family members in remote locations to respond immediately.

[0690] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0691] In this invention, the server includes means for transmitting voice signals from an information processing device to a user at regular intervals to encourage conversation, means for extracting the main points of the conversation using a generative model and generating summary information, and means for encrypting and transmitting digital voice data. This makes it possible for family members who live far away to reliably understand the health and safety of the elderly, and to quickly detect abnormal situations while protecting their privacy.

[0692] An "information processing device" is a device that performs data input, processing, storage, and output, and has the function of exchanging information with other devices via a communication network.

[0693] "Users" refers to individuals who use this system, including, but not limited to, the elderly and their families.

[0694] An "audio signal" is a signal that is created by electronically converting sound and transmitting it via a communication means.

[0695] A "generative model" refers to an algorithm or system that uses machine learning to extract specific patterns or key points from input data.

[0696] "Summary information" refers to information that organizes the important points and keywords extracted from conversation content using a generative model.

[0697] "Means of storage" refers to physical or electronic means of storing data and keeping it accessible as needed.

[0698] "Encryption" is a technology that transforms the content of data based on a specific algorithm to prevent unauthorized access or decryption.

[0699] A "portable communication device" refers to a device that is portable and uses wireless technology to communicate.

[0700] "Privacy" is the right or state to prevent an individual's private information from being used or made public without their consent.

[0701] An "abnormal situation" refers to a state in which unusual circumstances or values ​​are detected, and is sometimes used particularly in relation to health.

[0702] The system of this invention aims to enable elderly people to communicate naturally in their daily lives and to effectively provide health information to distant family members. The system includes an information processing device, a communication device, a generative AI model, and a storage means.

[0703] The devices used are portable communication devices such as smartphones and smartwatches. These devices transmit voice signals to elderly people at designated times to encourage everyday conversation. The voice signals include time announcements and messages to start a conversation.

[0704] After an audio signal is transmitted, the terminal records the elderly person's voice and saves it as digital audio data. This audio data is encrypted via an information processing device and transferred to a server over the internet. The AES algorithm is used for encryption, ensuring data privacy.

[0705] The server decodes the received audio data and analyzes it using a generative AI model. During the analysis, natural language processing techniques are used to extract the main points of the conversation and summarize important information. The generated summary information is stored in a memory device for easy access by the user.

[0706] Users access summarized information posted to allow family members to easily understand the condition of elderly individuals, using a web browser or dedicated application. This ensures that important information is clearly organized and displayed. For example, if everyday conversations include information such as "I felt fine yesterday" or "I have an appointment to see a doctor today," this information is summarized and provided.

[0707] Furthermore, the server monitors specific phrases within the summary information, and if an anomaly or urgent case is detected, it sends a warning signal to the user's device in real time. This notification uses a push notification service for smartphones.

[0708] Furthermore, the device has a function to continuously monitor biometric information, and if an abnormal value is recorded, the data is immediately sent to the server. This information is notified to the user as an emergency alert, prompting early action.

[0709] An example of a prompt might be, "Please analyze the recording of a conversation with an elderly person and generate a summary regarding their health and safety." This allows the generative model to analyze the conversation content and organize and provide information according to the purpose.

[0710] By combining the technologies described above, this system provides a means to remotely, comprehensively, and effectively monitor the health and safety of the elderly.

[0711] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0712] Step 1:

[0713] The device transmits an audio signal to the elderly person at a specified time. The input is the time information set within the device. The output is the audio signal transmitted by the device. This audio signal contains messages designed to facilitate everyday conversation.

[0714] Step 2:

[0715] The device records conversations with elderly individuals in response to voice signals. The input is the voice of the elderly person, which is converted into digital audio data and output. Specifically, this process involves capturing the voice through the device's microphone and converting it into a digital format.

[0716] Step 3:

[0717] The terminal encrypts the generated digital audio data and sends it to the server over the internet. The input is the recorded audio data, and the output is an encrypted data packet. The AES algorithm is used for encryption, ensuring privacy protection.

[0718] Step 4:

[0719] The server decrypts the received encrypted data and obtains digital audio data. The input is encrypted data, and the output is decrypted audio data. Specifically, it decodes the data using a cryptographic decryption algorithm.

[0720] Step 5:

[0721] The server uses a generative AI model to analyze audio data and extract the main points of the conversation. The input is decoded audio data, and the output is summarized information. Natural language processing techniques are used to identify keywords and important context.

[0722] Step 6:

[0723] The server stores the extracted summary information in a storage device and makes it accessible to the user. The input is the summary information, and the output is stored as saved data in the storage device. Specific operations include writing information to a database.

[0724] Step 7:

[0725] The server monitors keywords included in the summary information and sends a warning signal to the user's terminal if an anomaly or emergency condition is detected. The input is the keyword to be monitored, and the output is the warning signal sent to the terminal. The alert is triggered via a push notification service.

[0726] Step 8:

[0727] The device continuously acquires health data of elderly individuals using its built-in biosensors, and if an abnormal value is detected, it sends that information to a server. The input is the data obtained from the biosensors, and the output is the abnormal data sent to the server. Specifically, it monitors the data feed from the sensors in real time and generates an alert if a threshold is exceeded.

[0728] (Application Example 1)

[0729] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0730] There is a need for a system that can monitor the health status of the elderly in real time, facilitate daily communication, and immediately notify families when abnormalities are detected. However, many current systems have limitations in the frequency and content of communication, and are unable to adequately provide real-time notifications of important health information.

[0731] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0732] In this invention, the server includes means for transmitting voice signals from an information processing device to the user at a fixed time each day to facilitate language exchange; means for extracting the key points of the language exchange using a generative model and forming summary information; and means for continuously monitoring biometric information and immediately generating and transmitting a warning when an abnormal value is detected. This enables real-time monitoring of the health status of the elderly and rapid notification in emergencies.

[0733] An "information processing device" is a device that transmits audio signals to facilitate language communication and processes data.

[0734] "User" refers to an individual who receives audio signals via an information processing device and uses the service.

[0735] "Scheduled time" refers to a specific time when audio signals are transmitted or data processing takes place.

[0736] An "audio signal" is a means of transmitting information via sound to a user.

[0737] "Language exchange" refers to communication conducted through voice between the user and the system.

[0738] A "generative model" is a machine learning model that extracts key points from the content of linguistic exchange and forms summarized information.

[0739] "Summary information" refers to information that briefly summarizes the important points of language exchange.

[0740] A "storage medium" is a device or system that stores summary information and makes it accessible as needed.

[0741] A "keyword" is a specific, important word included in the summary information that triggers a warning or notification.

[0742] A "warning signal" is a signal sent to the user's device to alert them when an abnormality is detected.

[0743] "Biometric information" refers to data that indicates the user's health status, including physiological information such as heart rate and activity level.

[0744] An "abnormal value" is a value detected in biological information as an unexpected or exceeding standard.

[0745] A "warning" is a notification that alerts you to an abnormality based on your biometric information, and is primarily sent to family members or medical support staff.

[0746] To implement this invention, a smartwatch or portable device worn by an elderly person is used. The device transmits an audio signal to the elderly person at a set time each day, facilitating communication. The device records the elderly person's conversations, saves them as digital audio data in real time, and transmits them to a server in an encrypted format. This server analyzes the received audio data using a generation AI model, extracts the key points, and generates summary information. The generated summary information is stored on a storage medium and kept accessible to family members to understand the elderly person's situation. In addition, the device constantly monitors the user's biometric information, and if abnormal values ​​such as heart rate or activity level are detected, it immediately generates a warning and sends a notification to the family.

[0747] As a concrete example, the server analyzes the data received from the smartwatch, and if the heart rate exceeds the normal range, it immediately sends a message to the family's smartphone such as, "The user's heart rate is high. Is there anything wrong?" At this stage, the generative AI model detects summary information and reports specific details such as "Activity level increased yesterday" concisely to the family.

[0748] An example of a prompt to be input into the generation AI model is, "Extract the main points of an elderly person's everyday conversation, and extract and summarize keywords that include important medical information." Based on this prompt, the system automatically organizes the important information and immediately issues a warning if any abnormalities occur.

[0749] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0750] Step 1:

[0751] The device, such as a smartwatch or other portable device, transmits voice signals to the elderly person to facilitate everyday conversation. The input for this step is the device's scheduled time information, and the output is the start of a conversation with the elderly person. By sending voice signals, it initiates natural communication.

[0752] Step 2:

[0753] The device records the conversation and digitizes it as audio data. The input for this step is the audio signal from the microphone, and the output is digital audio data. The digital data is stored and prepared for encryption in preparation for subsequent processing.

[0754] Step 3:

[0755] The terminal encrypts digital audio data using encryption technology and sends it to the server in a secure format. The input for this step is unprocessed digital audio data, and the output is encrypted data. Encryption is performed for security purposes.

[0756] Step 4:

[0757] The server decrypts the received encrypted data and converts it into processable audio data. The input for this step is encrypted audio data, and the output is the decrypted audio data. The server then prepares the data for appropriate processing.

[0758] Step 5:

[0759] The server uses a generative AI model to extract key points from audio data and generate summary information. The input for this step is decoded audio data, and the output is the summary information. By extracting key points, it efficiently presents important information.

[0760] Step 6:

[0761] The server stores the summary information on a storage medium and makes it accessible. The input for this step is the generated summary information, and the output is a database of the stored information. The environment is set up so that users can access the information when needed.

[0762] Step 7:

[0763] The server continuously monitors biometric data and, upon detecting an anomaly, immediately generates an alert and sends a notification to the family. The input for this step is real-time biometric data sent from the smartwatch, and the output is an alert message. Processing is performed to ensure rapid notification in emergencies.

[0764] Step 8:

[0765] The user receives summary information and warnings sent by the server to monitor their daily health status and make decisions regarding emergency responses. The input for this step is the notifications sent from the server, and the output is the user's decision. The user then takes appropriate action based on the information.

[0766] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0767] This invention is a system that recognizes the emotions of an elderly person during a conversation and provides that emotional information to their family. The system consists of an information processing device, a generative model, a memory device, an emotion engine, and a communication means.

[0768] First, the device uses a smart device to transmit an audio signal to the elderly person at a set time. This signal initiates a conversation. Once the conversation begins, the device records the audio data and converts it into digital data in real time. This involves transmitting the data to a server via a communication line. The data transfer is encrypted to protect privacy.

[0769] Next, the server analyzes the received audio data using a generative model to extract the main points of the conversation and generate summary information. This summary includes an emotion engine, which also includes emotional states detected during the conversation. The emotion engine identifies emotional data based on facial recognition and speech tone analysis and adds it to the summary information.

[0770] Summary information and sentiment data are stored in the server's storage device and made accessible to the user. The stored data is organized according to specific conditions, and if, for example, a particular emotional state or keyword is detected, an alert signal is sent to the user's device. This allows family members to quickly grasp the necessary information and respond appropriately to the situation.

[0771] Furthermore, the device continuously acquires biometric data and immediately sends a signal to the server if it detects any abnormal values. This information is provided to the user as an emergency alert to encourage prompt action.

[0772] For example, if an elderly person says "I've been feeling irritable a lot lately" during a conversation, the emotion engine will determine that their emotions are unstable. Based on this, the summary information will record "Recently, their emotions have been unstable," and the user will receive a notification immediately. This system allows families to have a detailed understanding of the elderly person's emotions and health condition and to provide necessary care quickly.

[0773] The following describes the processing flow.

[0774] Step 1:

[0775] The user installs the application and performs initial setup on a screen where they can configure the voice notification time, emotion recognition settings, and warning signal conditions. This configuration information is sent from the device to the server and stored there.

[0776] Step 2:

[0777] At a set time, the device emits an audio signal to prompt the elderly person to begin a conversation. The audio signal includes greetings and questions delivered through a speaker, such as "Hello, how was your day?"

[0778] Step 3:

[0779] When a conversation begins, the device records the audio data in real time and saves it as digital data. This data is encrypted and sent to the server via the communication line.

[0780] Step 4:

[0781] The server analyzes the received audio data using a generative model. The generative model utilizes natural language processing techniques to grasp the main points of the conversation and create a summary. At this point, the emotion engine operates, analyzing the voice tone and word choice to extract emotion data.

[0782] Step 5:

[0783] The server adds sentiment data to the generated summary information and stores it in storage. This makes the user able to access the summary, including sentiment information, at any time.

[0784] Step 6:

[0785] If the server detects keywords or emotional states that meet specific criteria from the summary information, it generates a warning signal and sends it to the user's terminal. For example, if the emotion engine detects "unstable emotions," a notification is immediately sent to the user.

[0786] Step 7:

[0787] The device uses biometric measurement functions to measure health data of elderly individuals (e.g., heart rate and activity level). If abnormal values ​​are detected, the device sends the collected data to a server, and an emergency alert is sent to the user.

[0788] This entire process provides a system that accurately assesses the health and emotional state of elderly individuals remotely, enabling families to respond quickly.

[0789] (Example 2)

[0790] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0791] There are challenges such as a lack of communication with the elderly and the inability for families to quickly grasp changes in their health and take appropriate action. Furthermore, there is a lack of means to appropriately extract important information and emotional changes from conversations and promptly notify families.

[0792] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0793] In this invention, the server includes means for periodically transmitting voice signals from an information processing device to a user to facilitate dialogue, means for extracting the main points of the dialogue and generating summary information using a generation algorithm, and means for adding emotion data generated by facial recognition and voice tone analysis included in the summary information. This enables family members to grasp changes in the emotional and health conditions of elderly people in real time and take necessary actions quickly.

[0794] An "information processing device" is a computer system that receives audio data and performs calculations for analysis and generation.

[0795] "Users" refers to individuals or organizations that receive conversational information and emotional data from elderly people through this system.

[0796] A "generative algorithm" is a computer program that extracts key points and features from audio data and performs summarization and sentiment analysis.

[0797] "Emotional data" refers to information that indicates the emotional state of a subject, generated based on changes in voice tone and facial recognition.

[0798] "Facial recognition" is a technology that uses image processing to analyze the movements and expressions of a person's face and determine their emotions.

[0799] "Voice tone analysis" is a process that analyzes the intonation and speed of speech to evaluate the speaker's emotions and state of mind.

[0800] A "storage device" is a storage medium for permanently or temporarily recording digital information.

[0801] A "warning signal" is alert information sent to the user's device when specific conditions are detected.

[0802] This invention is a system that analyzes the emotional and health status of elderly individuals through communication and provides this information to their families. The embodiments are described in detail below.

[0803] The device utilizes smart technology to transmit voice signals to elderly individuals and initiate conversations. These voice signals are transmitted automatically based on time. High-performance microphones and voice recognition software are used to convert the voice into digital data in real time, and encryption technology (e.g., AES) is used to transmit it to the server to maintain privacy.

[0804] The server uses a generative AI model to analyze the received audio data. The analysis extracts key points and important information from the conversation and generates a summary. Furthermore, the server incorporates an emotion engine that generates emotion data through voice tone analysis and facial expression recognition, and adds this to the summary.

[0805] The generated summary information and sentiment data are recorded on the server's storage device and made accessible to the user. The system has a function to immediately send a warning signal to the user's terminal when specific emotional states or keywords are detected. This allows families to obtain necessary information in real time and respond promptly to the elderly person's condition.

[0806] Furthermore, the device continuously monitors the biometric indicators of elderly individuals, and if abnormal data is detected, the server sends an urgent notification to the user. This collaboration enables early detection of health abnormalities and supports prompt response.

[0807] As a concrete example, if an elderly person mentions during a conversation that they have been feeling irritable lately, the emotion engine will determine that their condition is unstable, record it in the summary information, and immediately notify the user. This functionality allows families to appropriately monitor changes in the elderly person's emotions and health and provide necessary support quickly.

[0808] Example of a prompt:

[0809] "Please explain the mechanism for notifying a summary of information if unstable emotions are detected in recent conversations."

[0810] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0811] Step 1:

[0812] The terminal initiates a conversation by sending an audio signal to the elderly person using a smart device based on a set time. The input is the set time information, and the output is the transmission of an audio signal. At this time, the program utilizes the timer function of the smart device to execute the process of sending an audio signal at the specified time.

[0813] Step 2:

[0814] The device records the initiated conversation using a high-precision microphone and converts the audio data into digital data in real time. The input is the audio signal of the conversation, and the output is the digitized audio data. Speech recognition software operates in the background, analyzing the input audio waveform and converting it into text data and speech feature data.

[0815] Step 3:

[0816] The terminal encrypts digital data using AES encryption technology and transmits it to the server via the communication line while protecting privacy. The input is digitized voice data, and the output is encrypted data communication. The encryption module activates and converts the plaintext data based on the encryption key.

[0817] Step 4:

[0818] The server decrypts the received encrypted data and analyzes the audio data using a generative AI model. The input is encrypted audio data, and the output is summary information and sentiment data. The generative AI model simultaneously runs speech-to-text conversion and sentiment analysis algorithms to extract the main points and emotional state of the conversation.

[0819] Step 5:

[0820] The server uses an emotion engine to perform facial recognition and voice tone analysis based on the generated summary information. The input is the summary information, and the output is the summary information with added emotion data. The emotion engine evaluates various voice parameters and facial expression data, classifying emotions as either numerical or categorical.

[0821] Step 6:

[0822] The server stores the final summary information and sentiment data in storage, keeping it accessible to users. The input is the summary information and sentiment data, and the output is the entries in the stored database. Database management software organizes and indexes the data in an automated process.

[0823] Step 7:

[0824] The server immediately sends an alert signal to the user's device when a specific emotional state or keyword is detected. The input consists of stored summary information and triggering conditions, while the output is the alert signal. The alert system continuously monitors the system and activates a notification protocol when the conditions are met.

[0825] Step 8:

[0826] The device continuously acquires biometric data from elderly individuals via biosensors and reports any abnormal measurements to the server. The input is biometric data, and the output is a notification signal containing abnormal values. A vital signs monitoring system analyzes the data and triggers an alert if a threshold is exceeded.

[0827] (Application Example 2)

[0828] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0829] In situations where elderly people experience anxiety and loneliness on a daily basis, it is difficult for family members and caregivers to quickly recognize these emotional changes and respond appropriately. Therefore, there is a need for a system that monitors the physical and mental state of elderly people in real time and promptly notifies relevant parties if any abnormalities are detected.

[0830] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0831] In this invention, the server includes means for transmitting an audio signal from an information processing device to a user at a fixed time each day to encourage conversation; means for extracting the main points of the conversation using a generative model and generating summary information; means for storing the summary information in a storage device and making it accessible; means for analyzing the audio data in real time and evaluating the emotional state; and means for sending a notification to relevant parties when an abnormal emotional state is detected. This makes it possible to quickly grasp the physical and mental state of elderly people and respond promptly as needed.

[0832] An "information processing device" is an electronic device that transmits audio signals to the user to facilitate conversation.

[0833] A "generative model" is an algorithmic method for extracting the main points of a conversation and summarizing information.

[0834] A "storage device" is a digital storage device used to store summary information and make it accessible as needed.

[0835] "Emotion recognition means" refers to technology that analyzes voice data and evaluates the user's emotional state.

[0836] A "portable communication device" is a portable communication device used to transmit voice signals.

[0837] An "abnormal value" is a value in biological data that exceeds the normal range and indicates a condition that requires attention.

[0838] "Means of sending notifications to relevant parties" refers to methods for communicating important information to users and stakeholders in real time.

[0839] The system implementing this invention consists of an information processing device, a generative model, a storage device, an emotion recognition engine, and a communication means. The server has a function to prompt conversation by sending an audio signal to the user at a fixed time every day via the information processing device. When a conversation begins, the terminal converts the audio data into digital data in real time, encrypts it, and sends it to the server. The server converts the audio data into text data using software such as Google Cloud Speech-to-Text, and based on that, extracts the main points of the conversation using a generative model and generates summary information. In this process, an emotion recognition engine such as IBM Watson Tone Analyzer is used to perform voice tone analysis and evaluate the emotional state. If an anomaly is detected by emotion recognition, a push notification is sent to the relevant parties.

[0840] For example, if a device records a conversation with an elderly person and detects a phrase like "I haven't been able to sleep lately," the emotion recognition engine will detect "anxiety," and the user will be promptly notified with a message such as "You've been experiencing persistent anxiety lately." An example of a prompt to the generative AI model used in this case would be, "Analyze the following conversation emotionally and detect any significant changes: 'I've been having more trouble sleeping lately...'" In this way, the server and the device work together, allowing users to quickly understand the emotions and health status of elderly people.

[0841] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0842] Step 1:

[0843] The terminal transmits an audio signal to the user (an elderly person) at regular intervals via an information processing device. The input is the scheduled alarm setting data, and the output is an audio signal sent to the user. This action prompts the user to initiate conversation.

[0844] Step 2:

[0845] As soon as a conversation with the user begins, the device converts the audio data into digital data in real time. The input is raw audio data, and the output is a stream of digital data. The audio data is captured on the device via the microphone and converted into text data using Google Cloud Speech-to-Text.

[0846] Step 3:

[0847] The terminal encrypts the converted text data and sends it to the server. The input is text data, and the output is encrypted text data. AES encryption or similar methods are used to protect the data during transmission.

[0848] Step 4:

[0849] The server inputs the received text data into a generative model to extract the main points of the conversation. The input is decrypted text data, and the output is summarized information. Natural language processing techniques are applied to extract the most important parts of the conversation.

[0850] Step 5:

[0851] The server processes the extracted summary information into an emotion recognition engine to analyze the user's emotional state. The input is summary information, and the output is emotional state data. IBM Watson Tone Analyzer is used to analyze the tone of the text and identify emotions.

[0852] Step 6:

[0853] The server detects anomalies from the emotional state data in the analysis results and sends notifications to relevant parties as needed. The input is emotional state data, and the output is a notification message. When an anomaly is detected, a push notification function is used to send an alert to family members or caregivers with specific information.

[0854] Throughout this entire process, users can quickly grasp the emotions and health status of elderly individuals.

[0855] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0856] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0857] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0858] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0859] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0860] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0861] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0862] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0863] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0864] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0865] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0866] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0867] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0868] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0869] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0870] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0871] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0872] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0873] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0874] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0875] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0876] The following is further disclosed regarding the embodiments described above.

[0877] (Claim 1)

[0878] A means of prompting conversation by transmitting an audio signal from an information processing device to the user at a fixed time every day,

[0879] A means for extracting the main points of the conversation using a generative model and generating summary information,

[0880] Means for storing the summary information in a storage device and making it accessible,

[0881] A means for detecting keywords included in the summary information based on specific conditions and transmitting a warning signal to the user terminal,

[0882] A system that includes this.

[0883] (Claim 2)

[0884] The system according to claim 1, wherein the information processing device continuously acquires data based on biological measurements and generates a notification signal when an abnormal value is detected.

[0885] (Claim 3)

[0886] The system according to claim 1, wherein the voice signal transmission means is performed by a portable communication device.

[0887] "Example 1"

[0888] (Claim 1)

[0889] A means of prompting conversation by sending voice signals from an information processing device to the user at regular intervals,

[0890] A means for extracting the main points of the conversation using a generative model and generating summary information,

[0891] Means for storing the summary information in a storage means and making it accessible,

[0892] A means for detecting words or phrases contained in the summary information based on specific conditions and transmitting a warning signal to the user terminal,

[0893] A means of encrypting and transmitting digital audio data,

[0894] A means for monitoring the aforementioned summary information and notifying of abnormal situations,

[0895] A system that includes this.

[0896] (Claim 2)

[0897] The system according to claim 1, wherein the information processing device has means for continuously acquiring data based on biological information and for generating a notification signal when an abnormal value is detected.

[0898] (Claim 3)

[0899] The system according to claim 1, wherein the voice signal transmission means is performed by a portable communication device.

[0900] "Application Example 1"

[0901] (Claim 1)

[0902] A means of facilitating language exchange by transmitting voice signals from an information processing device to users at a fixed time every day,

[0903] A means for extracting the key points of the aforementioned language exchange using a generative model and forming summary information,

[0904] Means for storing the aforementioned summary information on a storage medium and making it accessible,

[0905] A means for detecting keywords included in the summary information based on specific conditions and transmitting a warning signal to the user device,

[0906] A means for continuously monitoring biometric information and immediately generating and transmitting a warning when an abnormal value is detected,

[0907] A system that includes this.

[0908] (Claim 2)

[0909] The system according to claim 1, wherein the information processing device provides the generated summary information to family members in a remote location and promptly notifies them in the event of an emergency.

[0910] (Claim 3)

[0911] The system according to claim 1, wherein the voice signal transmission means is performed by a portable communication device.

[0912] "Example 2 of combining an emotion engine"

[0913] (Claim 1)

[0914] A means of facilitating dialogue by periodically transmitting voice signals from an information processing device to the user,

[0915] A means for extracting the key points of the dialogue using a generation algorithm and generating summary information,

[0916] Means for adding emotional data generated by facial recognition and voice tone analysis included in the summary information,

[0917] Means for storing the aforementioned summary information and emotional data in a storage device and making them accessible,

[0918] A means of sending a warning signal to the user's terminal based on a specific emotional state or keyword,

[0919] A system that includes this.

[0920] (Claim 2)

[0921] The system according to claim 1, wherein the information processing device continuously acquires data based on biological measurements and generates a notification signal when an abnormal value is detected.

[0922] (Claim 3)

[0923] The system according to claim 1, wherein the voice signal transmission means is performed by a portable communication device.

[0924] "Application example 2 when combining with an emotional engine"

[0925] (Claim 1)

[0926] A means of prompting conversation by transmitting an audio signal from an information processing device to the user at a fixed time every day,

[0927] A means for extracting the main points of the conversation using a generative model and generating summary information,

[0928] Means for storing the summary information in a storage device and making it accessible,

[0929] A means for detecting keywords included in the summary information based on specific conditions and transmitting a warning signal to the user terminal,

[0930] An emotion recognition method that analyzes audio data in real time and evaluates emotional states,

[0931] A means of sending a notification to relevant parties when an abnormal emotional state is detected,

[0932] A system that includes this.

[0933] (Claim 2)

[0934] The system according to claim 1, wherein the information processing device continuously acquires data based on biological measurements and generates a notification signal when an abnormal value is detected.

[0935] (Claim 3)

[0936] The system according to claim 1, wherein the voice signal transmission means is performed by a mobile communication device. [Explanation of symbols]

[0937] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of prompting conversation by transmitting an audio signal from an information processing device to the user at a fixed time every day, A means for extracting the main points of the conversation using a generative model and generating summary information, Means for storing the summary information in a storage device and making it accessible, A means for detecting keywords included in the summary information based on specific conditions and transmitting a warning signal to the user terminal, A system that includes this.

2. The system according to claim 1, wherein the information processing device continuously acquires data based on biological measurements and generates a notification signal when an abnormal value is detected.

3. The system according to claim 1, wherein the voice signal transmission means is performed by a portable communication device.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A