system

A system that records and analyzes speech and physical data from elderly individuals to detect anomalies, providing real-time health monitoring and privacy-conscious support.

JP2026073446APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In an aging society, it is challenging to monitor and detect early deterioration in the physical or mental condition of the elderly, particularly in isolation, without relying on family or community support, and existing technologies fail to comprehensively manage health and safety issues in real-time.

Method used

A system that records speech information, converts it to text using speech recognition, learns speech patterns, and combines this with physical information measurement to detect abnormalities, sending notifications when necessary, thereby enabling real-time health monitoring and privacy-conscious protection.

Benefits of technology

Enables early detection and notification of health anomalies in elderly individuals, supporting their safety and health through comprehensive monitoring of both speech and physical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073446000001_ABST
    Figure 2026073446000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A recording means for recording speech information of elderly people, A conversion means for converting recorded speech information into text data, A learning method that analyzes text data and learns speech patterns, A detection means for detecting anomalies based on learned patterns, A notification means that generates a notification based on the detected anomaly, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In an aging society, isolation of the elderly and overlooking of health problems have become social issues. In particular, it is difficult to appropriately monitor and judge daily conversations and minor changes in health conditions to detect early deterioration in physical or mental condition of the elderly. In such cases, there is a need for an efficient and privacy-conscious monitoring system that does not rely on support from family or local communities.

Means for Solving the Problems

[0005] This invention provides a system for detecting abnormalities in elderly individuals by introducing means for recording speech information, means for converting speech to text using speech recognition, and means for learning speech patterns through analysis. Furthermore, by combining abnormality detection with the measurement of physical information and its data, it enables a rapid response in the event of an abnormality. In this way, by using both speech and health data, it aims to achieve both real-time health monitoring of elderly individuals and protection of their privacy.

[0006] "Recording means" refers to devices and technologies that capture speech information from elderly people and save it as digital data.

[0007] "Conversion means" refers to speech recognition technology or a program for analyzing recorded audio data and converting it into text data.

[0008] "Learning methods" refer to techniques or systems for extracting speech patterns from text data of elderly individuals and analyzing and memorizing their characteristics using machine learning.

[0009] A "detection method" refers to a technology or mechanism that compares learned speech patterns with data obtained in real time to identify abnormalities in the speech and behavior of elderly individuals.

[0010] A "notification means" is a method for generating necessary information based on detected anomalies and basic data, and for informing relevant parties.

[0011] "Measurement means" refers to a device or technology that acquires physical information of elderly people using sensors, etc., and quantifies and records it.

[0012] "Verification means" refers to the process or function of comparing and verifying current physical data with past historical data based on detected anomalies.

[0013] A "history management system" refers to a system or method for systematically saving and managing past speech and physical information, making it accessible when needed. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a monitoring system that uses speech and physical information of elderly people, and by combining multiple functions, it enables real-time monitoring of the safety and health status of elderly people.

[0036] The device consists of a bracelet-type device worn on the wrist of an elderly person. This device is equipped with a microphone for recording voice and has the function of recording everyday conversations. It also has built-in sensors for measuring physical information such as heart rate and body temperature. This data is transmitted to a server via wireless communication.

[0037] When the server receives voice data transmitted from the terminal, it first converts it into text data using a speech recognition engine. The converted text data is then analyzed by a generative AI to learn the speech patterns of elderly people. The generative AI focuses on the characteristics of the speech content, the frequency of words used, and the flow of context to detect abnormal patterns. If an abnormality is detected, a notification is sent to pre-configured contacts. The server also simultaneously receives physical information transmitted from the terminal and stores the measurement results in a database.

[0038] Users receive and view information transmitted from the server via their smartphones. The application allows users to check logs of the elderly person's health status and receive abnormality notifications. It also includes features to allow users to investigate the nature of the abnormality in more detail or to consult with healthcare professionals, if necessary.

[0039] As a concrete example of this system, suppose an elderly person living alone starts making more frequent remarks about forgetting things during everyday conversations. The server detects this change in speech pattern as an anomaly and sends a notification to the family member (the user). Upon receiving this notification, the family can take appropriate action, such as seeking early medical advice. In this way, the present invention supports the health and safety of the elderly.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The device records the elderly person's voice in real time. This is done using a built-in microphone, and the voice data is saved to a buffer at regular intervals.

[0043] Step 2:

[0044] The terminal transmits voice data to the server via wireless communication. The transmission is triggered by pre-configured conditions, enabling efficient data transmission.

[0045] Step 3:

[0046] The server analyzes the received audio data and converts it into text data using a speech recognition engine. During this process, pre-processing such as noise reduction is performed to improve recognition accuracy.

[0047] Step 4:

[0048] The server passes text data through a generative AI and compares it to past speech patterns. The generative AI uses a pre-trained model to detect changes in the text data, identifying new patterns and unusually frequent utterances.

[0049] Step 5:

[0050] When the server detects an anomaly, it sends an alert to the configured notification recipient. The alert will include a description of the anomaly and recommended actions.

[0051] Step 6:

[0052] The device measures the physical information of elderly individuals. This information includes body temperature and heart rate, and this data is sent to a server.

[0053] Step 7:

[0054] The server analyzes the physical information and compares it with past data to check for any abnormalities. If no abnormalities are found, it makes a comprehensive judgment based on this information along with the voice data.

[0055] Step 8:

[0056] Users receive notifications and data from the server via a smartphone app. Based on the received information, the system monitors the elderly person's condition and considers necessary actions.

[0057] Step 9:

[0058] Users can check their history of abnormalities and health data within the app. This allows for timely medical consultations and appropriate actions. This information is continuously logged.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] Effectively monitoring the safety and health of the elderly is a critical challenge in a growing elderly population. With conventional technologies, it was difficult for those around the elderly to understand their condition unless they themselves were aware of an abnormality. Furthermore, detecting abnormalities based not only on physical information but also on speech patterns was challenging, potentially leading to delays in necessary responses.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes voice collection means, voice recognition means, generative model learning means, anomaly detection means, notification generation means, physical condition measurement means, and physical information transmission means. This enables comprehensive state monitoring using both the speech and physical information of elderly individuals, allowing for early detection and notification of anomalies.

[0064] "Voice collection means" refers to a device or function that records the everyday conversations of elderly people as audio.

[0065] "Speech recognition means" refers to a device or function that converts recorded speech data into text data.

[0066] "Generative model learning means" refers to a device or program that analyzes converted text data and has the function of learning the speech patterns of elderly people.

[0067] An "anomaly detection means" refers to a device or function that detects speech or behavior that is different from the norm, based on learned speech patterns.

[0068] "Notification generation means" refers to a device or function that generates and sends messages or alerts to designated recipients when an anomaly is detected.

[0069] "Means of measuring physical condition" refers to sensors and devices that measure physical information such as heart rate and body temperature of elderly people.

[0070] "Body information transmission means" refers to devices or protocols that have the function of recording measured body information and transmitting it to remote devices such as servers.

[0071] This invention relates to a system for monitoring the health status of elderly individuals using their speech and physical information. This system is implemented using a server, a terminal, and a user component.

[0072] Data collection by devices:

[0073] The device consists of a bracelet-type device worn on the wrist of an elderly person. As a means of voice collection, the device is equipped with a microphone to record the elderly person's daily conversations as audio data. It also has a built-in heart rate monitor and body temperature sensor as means of measuring physical condition, and this physical information is measured periodically. The collected data is transmitted to a server via Bluetooth or Wi-Fi communication using a physical information transmission device.

[0074] Server-based data analysis:

[0075] The server converts speech data into text data using speech recognition means. This transforms the speech data into an analyzable format. Using generative model learning means, the speech patterns of elderly people are analyzed and learned from the text data. Anomaly detection means can detect anomalies based on the learned patterns. If an anomaly is detected, a notification is sent to a pre-configured contact (e.g., family) using notification generation means. These notifications may include historical data.

[0076] User verification of information:

[0077] Users receive health status data and abnormality notifications from a server via communication devices such as smartphones. Through a dedicated application, they can review past history and investigate the details of abnormalities. The application also provides information to consult with healthcare professionals as needed.

[0078] Specific example:

[0079] As a concrete example, if an elderly person living alone frequently says "I've been forgetting things a lot lately" in conversation, the server will detect this change in speech pattern as an anomaly. Subsequently, a notification is sent to the family member (the user), who can then consider seeking medical advice early.

[0080] Example of a prompt:

[0081] "Detect abnormalities in the speech patterns of elderly individuals. As a specific example, analyze changes in frequently used phrases."

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The device records the daily conversations of elderly people as audio data. A microphone is used as the means of audio collection. The input is spoken voice, and the output is an audio data file. Specifically, the device captures audio at regular time intervals and saves it as a data file.

[0085] Step 2:

[0086] The device measures bodily information such as heart rate and body temperature. It acquires data from heart rate sensors and body temperature sensors using a bodily state measurement device. The input is the body's state, and the output is heart rate and body temperature data. Specifically, the sensors read data from the body and record it as numerical data in storage.

[0087] Step 3:

[0088] The device transmits collected voice data and physical information to a server via Bluetooth or Wi-Fi. Here, the recorded data file is the input, and the data packets sent to the server are the output. Specifically, the device periodically transmits data using a communication module.

[0089] Step 4:

[0090] The server converts received audio data into text data using speech recognition technology. The input is an audio data file, and the output is the spoken content in text format. Specifically, the speech recognition algorithm analyzes the audio and converts it into text.

[0091] Step 5:

[0092] The server analyzes text data using a generative model learning method to learn the speech patterns of elderly people. The input is text data, and the output is the learned speech patterns. Specifically, the generative AI analyzes the structure and content of the text and records the frequency of words and the flow of context.

[0093] Step 6:

[0094] The server uses anomaly detection means to detect anomalies based on learned speech patterns. The input is the learned speech pattern, and the output is information about the detected anomaly. Specifically, it compares existing patterns with new data and identifies changes that deviate from the standard.

[0095] Step 7:

[0096] The server uses a notification generation mechanism to create a notification when an anomaly is detected and sends it to a predetermined contact. The input is information about the detected anomaly, and the output is the anomaly notification. Specifically, it generates a message containing details of the anomaly and sends the notification via the communication network.

[0097] Step 8:

[0098] The user receives and displays notifications from the server using a smartphone app. The input is an anomaly notification from the server, and the output is the notification information displayed on the screen. Specifically, the user can open the app, check the notification content, and look up further details if necessary.

[0099] (Application Example 1)

[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] In technologies that monitor the safety and health of the elderly using speech and biometric information, conventional systems only focus on detecting anomalies in speech and biometric information, making it difficult to detect external safety threats from ambient noise. Furthermore, there was a lack of means to comprehensively manage and monitor this information in real time. Therefore, there is a need to further improve the safety and convenience of life for the elderly.

[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0103] In this invention, the server includes a device for recording speech information of elderly people, a conversion device for converting the recorded speech information into string data, a learning device for analyzing the string data and learning speech patterns, a detection device for detecting abnormalities based on the learned patterns, a notification device for generating notifications based on the detected abnormalities, and an environmental analysis device for detecting safety abnormalities by analyzing ambient sounds. This makes it possible to detect not only speech and biological abnormalities of elderly people, but also safety threats in the surrounding environment, thereby realizing more comprehensive monitoring.

[0104] The term "elderly person" refers to an individual for whom health management and safety become particularly important as they age.

[0105] "Speech information" refers to the linguistic data contained in the voices of elderly people, including the content of conversations and monologues.

[0106] "String data" refers to text-formatted data converted from audio data.

[0107] "Biometric information" refers to physical data, including heart rate, body temperature, and activity level, for elderly individuals.

[0108] The term "device" refers to an instrument or device with a specific function, used for various data collection and processing within a system.

[0109] A "conversion device" refers to a device that has the function of converting audio data into text data.

[0110] A "learning device" refers to a device that has the function of continuously analyzing data and identifying and memorizing features and patterns.

[0111] A "detection device" refers to a device that has the function of identifying and reporting anomalies based on learned data.

[0112] A "notification device" refers to a device that has the function of informing relevant parties about detected anomalies.

[0113] An "environmental analysis device" refers to a device that monitors ambient sounds and environmental data and has the function of detecting abnormalities.

[0114] The system that realizes this invention is designed to monitor the daily safety and health of elderly people. The system mainly consists of family members or caregivers as users, terminals worn by the elderly, and a server responsible for data processing and notifications. Its specific functions and operation are described below.

[0115] The device is worn on the wrist of elderly individuals and contains a microphone and multiple sensors (for measuring heart rate and body temperature). This allows for the continuous collection of speech and biometric information from the elderly person. This data is transmitted to a server via wireless communication.

[0116] The server first converts the audio data sent from the terminal into text data using speech recognition software. Specifically, Microsoft's Azure® Speech Services performs this role. This text data is then analyzed by a generative AI model. For example, it learns past speech patterns and detects abnormal speech or unstable contextual changes. By giving the generative AI model prompts such as, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that have decreased, and evaluate the likelihood of abnormalities based on that," the model performs an appropriate analysis.

[0117] Simultaneously, biometric information such as heart rate and body temperature is analyzed by the server. The IBM Watson® Health platform is used for real-time data analysis. In addition, ambient sounds are analyzed using Google® Cloud's voice analysis service to detect abnormal sounds occurring around the elderly person. In this way, the system provides comprehensive monitoring.

[0118] Users receive information sent from the server through a smartphone application. If an anomaly is detected, a notification is sent to the user via Firebase Cloud Messaging. Users can check the details of the anomaly within the app and obtain information to consult with a medical professional if necessary. For example, a notification such as "Your mother may be frequently mistaking place names" may be sent, allowing the user to take early action. In this way, the entire system works together to build a mechanism that supports the lives of the elderly.

[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0120] Step 1:

[0121] The device collects speech and biometric information from elderly individuals. It records voice data using a microphone and measures heart rate and body temperature using built-in sensors. The collected data is transmitted wirelessly to a server. The input is analog signal data recorded by the microphone and sensors, which is then converted to digital data and sent to the server.

[0122] Step 2:

[0123] The server processes the audio data sent from the terminal using speech recognition software and converts it into text data. Azure Speech Services is used for this process. The input is audio data, and the output is text-based string data. The audio signal is analyzed, and a language model is used to generate the correct word sequence.

[0124] Step 3:

[0125] The server analyzes the converted string data using a generative AI model. Specifically, it learns speech patterns using prompt sentences and performs anomaly detection. The generative AI utilizes the prompt, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of anomalies based on these." The input is string data, and the output is the analysis result regarding the presence or absence of anomalies.

[0126] Step 4:

[0127] The server also analyzes biometric information. Using IBM Watson Health, it analyzes heart rate and body temperature and detects abnormalities by comparing them to common health indicators. The input is biometric data collected by sensors, and the output is a judgment of whether it is normal or abnormal.

[0128] Step 5:

[0129] Simultaneously, the server uses Google Cloud's speech analysis service to analyze ambient sounds and detect anomalies. The input is ambient sound data, and the output is the result of safety-related anomaly detection. It identifies specific acoustic patterns and generates a warning if they are deemed anomaly.

[0130] Step 6:

[0131] The server generates and sends a notification to the user based on the detected anomaly. The notification is sent to smartphones and other devices via Firebase Cloud Messaging. The input is the anomaly detection result, and the output is an informational notification to the user. It includes specific details, allowing the user to receive it and take appropriate action.

[0132] Step 7:

[0133] Users view received notifications through a smartphone application. These notifications include detailed information about the anomaly and recommended next actions. Within the app, users can review past history and an overview of their overall health status, and contact healthcare professionals if necessary. The input is the notification message, and the output is the user's confirmation of the situation and subsequent actions.

[0134] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0135] This invention is a system that combines the acquisition of speech and physical information from elderly individuals with an emotion engine that analyzes emotions from their speech. This system comprehensively monitors the physical and emotional health status of elderly individuals, enabling early detection of abnormalities and appropriate responses.

[0136] The device is a bracelet-type device worn by elderly individuals and is equipped with a microphone for voice recognition and sensors to measure heart rate and body temperature. The device records voice in real time and transmits the acquired data to a server sequentially.

[0137] The server receives voice data transmitted from the terminal and performs a series of processes to convert it into text using a speech recognition engine. The converted text data is analyzed by a generative AI and sent to an emotion engine that recognizes emotions from the speech content along with the speech patterns. The emotion engine has a self-learning function and analyzes and stores the emotional states of elderly people from their daily speech to detect abnormal emotional changes.

[0138] Furthermore, the server integrates speech pattern learning and emotion analysis results from artificial intelligence, and performs anomaly detection by comparing this with health data. When an anomaly is detected, an alert is generated for the relevant parties. The alert includes details about the specific nature of the anomaly and what kind of emotional changes are occurring.

[0139] Users can receive this notification through a smartphone app and check the status of the elderly person. The historical data included in the notification serves as a guide for users to understand past health conditions and emotional tendencies. It also provides useful information when considering appropriate measures and medical support.

[0140] As a concrete example, consider a scenario where an elderly person's daily conversations contain more negative words than usual. The emotion engine recognizes this change as a shift from joy to anxiety and reports it as an anomaly to the server. If the server detects an increase in heart rate, it also considers it as a physical health problem and prompts the user to take prompt action. Thus, the present invention is a system that comprehensively supports the physical and mental health of the elderly through complex data analysis.

[0141] The following describes the processing flow.

[0142] Step 1:

[0143] The device records the elderly person's voice in real time. It acquires voice data using a built-in microphone and stores it in a buffer in an appropriate format.

[0144] Step 2:

[0145] The device transmits recorded audio data to the server at regular intervals. Wireless communication technology is used to ensure efficient data transmission.

[0146] Step 3:

[0147] The server passes the received audio data to the speech recognition engine, which converts it into text data. During the conversion process, noise filtering and other techniques are performed to improve accuracy.

[0148] Step 4:

[0149] The server inputs the generated text data into the generating AI, which then compares it to past speech patterns. During this process, the trained algorithm analyzes any changes in the speech content.

[0150] Step 5:

[0151] The server sends text data to the emotion engine, which analyzes the emotional state from the utterance. The emotion engine extracts specific emotional elements and determines the emotional tendency.

[0152] Step 6:

[0153] The server integrates the emotion analysis results and the speech pattern analysis results. From the integrated data, it detects abnormal behavior and emotional changes and prepares to generate alerts.

[0154] Step 7:

[0155] The device measures physical information such as the heart rate and body temperature of elderly individuals and transmits this data to a server. Measurements are performed regularly to maintain real-time accuracy.

[0156] Step 8:

[0157] The server comprehensively compares physical information with speech and emotion analysis data. This allows for multifaceted verification of anomalies and reduces the risk of misrecognition.

[0158] Step 9:

[0159] Users receive alert notifications via their smartphones. The notifications detail the specific nature of the anomaly, changes in emotional state, and possible underlying physical problems.

[0160] Step 10:

[0161] Users can review the information received through the app and take appropriate action regarding the elderly person's situation as needed. Past historical data can also be referenced to support more accurate decision-making.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] In today's aging society, there is a need for a system that can accurately monitor the physical and mental health of the elderly and detect abnormalities early. Conventional technologies have the problem that health monitoring of the elderly is one-sided, and it is difficult to make a comprehensive judgment of their physical and mental state, making it highly likely that abnormalities will be overlooked.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes voice acquisition means, emotion analysis means, and anomaly detection means. This enables comprehensive monitoring of the health status of elderly people based on their speech content and physiological information, as well as rapid anomaly detection.

[0167] "Voice acquisition means" refers to a device for collecting speech information from elderly people in real time, and mainly includes voice input devices such as microphones.

[0168] "Speech recognition means" refers to technology for converting acquired speech data into text data, and this conversion is performed using a speech recognition engine or AI technology.

[0169] "Emotional analysis methods" refer to analytical techniques for recognizing emotional states from text data, and utilize generative AI models to grasp the emotions contained in spoken language.

[0170] An "anomaly detection means" is a device or program that integrates acquired emotional data and health data to detect anomalies.

[0171] A "notification generation means" is a device or system that generates warnings or notifications based on detected abnormal information and transmits them to the user.

[0172] A "biometric measurement device" is a device used to measure physiological information such as heart rate and body temperature in elderly individuals, and utilizes various sensors.

[0173] "Data transmission means" refers to communication means intended for transmitting measured physiological information and analysis results to a server or other device.

[0174] A "data matching means" is a device or program that, when an abnormality in speech or physiological information is detected, matches that information and performs further analysis.

[0175] A "history management system" is a system for recording and managing past speech and physiological information.

[0176] This invention is a system that comprehensively monitors the physical and emotional health status of elderly individuals and enables early detection of abnormalities. The system primarily consists of terminals, a server, and users.

[0177] First, the device is a bracelet-type device worn by elderly individuals. The device incorporates a voice acquisition mechanism for real-time voice recording, as well as biometric measurement devices for measuring heart rate and body temperature. The acquired data is transmitted to a server using a secure communication protocol. This involves data transmission via Bluetooth and an internet connection.

[0178] Next, the server implements a commonly available speech recognition engine (e.g., a speech recognition API) as a means of speech recognition, which converts the speech data sent from the terminal into text data. The converted text data is analyzed by an emotion analysis means to recognize the emotional state from the elderly person's speech. A generative AI model is used for this, which learns the changes in the elderly person's emotions and is capable of self-improvement.

[0179] The server then integrates emotional states and biometric data using anomaly detection mechanisms to monitor the health status of elderly individuals. If an anomaly is detected, a notification generation mechanism generates a warning, which is sent to the user. The warning includes specific details of the anomaly, allowing the user to take prompt action.

[0180] Finally, users can receive alerts via their smartphones and check on the elderly person's condition. The smartphone app includes a history management system that tracks past speech and physiological data, allowing users to precisely track changes in their health status.

[0181] For example, if an elderly person says, "I'm feeling down today and don't want to do anything," the emotion analysis system will recognize this as anxiety or sadness and report it to the server as an abnormality. Furthermore, if a higher-than-normal heart rate is detected, a notification will be generated immediately. An example of a prompt message would be, "Please identify which utterance caused the emotional abnormality and decide which medical institution you should consult."

[0182] In this way, the system provides comprehensive support for the physical and emotional health of the elderly and offers advanced data analysis capabilities for rapid response in emergencies.

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] The device acquires the voice of the elderly person wearing it in real time and simultaneously measures biometric information such as heart rate and body temperature. The input obtained in this process consists of raw voice data and physiological data. This data is collected via a voice acquisition device and transmitted to a server via a data transmission device. The resulting output is voice data and physiological data that are ready for further processing on the server side.

[0186] Step 2:

[0187] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. The input is raw audio data, and the output is the corresponding text data. Specifically, speech-to-text conversion is performed using a speech recognition API or similar method. Through this process, the spoken content is obtained in text format and sent to the next analysis step.

[0188] Step 3:

[0189] The server inputs the converted text data into an emotion analysis device, which then uses a generative AI model to analyze the emotions. The input is text data, and the output is the analyzed emotion data. Specifically, the generative AI model extracts emotional nuances from the text content and evaluates the emotional state of the elderly person. This information prepares the device for anomaly detection.

[0190] Step 4:

[0191] The server integrates emotional and physiological data using anomaly detection mechanisms to monitor the health status of elderly individuals. Inputs are emotional and physiological data, and output is information on whether or not an anomaly occurred. Specifically, the system compares the data to dynamically set baseline values ​​and patterns; an anomaly is detected if the data exceeds the normal range.

[0192] Step 5:

[0193] The server generates a warning to the user using a notification generation mechanism when an anomaly is detected. The input is the anomaly detection result, and the output is a warning message to the user. This message includes detailed information about the health status of the elderly person, and the user can receive the notification through the app.

[0194] Step 6:

[0195] Users receive notifications via a smartphone app and check the health status of elderly individuals through the displayed notifications. The input is the notification message, and the output is information to help the user understand the situation. Based on this information, users can contact medical institutions or perform additional monitoring as needed.

[0196] Through these steps, this system provides comprehensive support for the physical and emotional health of older adults, enabling prompt and appropriate responses.

[0197] (Application Example 2)

[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0199] In today's living environments for the elderly, not only physical health issues but also emotional health issues are becoming increasingly important. However, conventional technology has struggled to effectively detect emotional changes and integrate them with physical information for comprehensive health management. Therefore, there is a need for a system that enables rapid and appropriate responses when elderly individuals exhibit emotional or physical abnormalities.

[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0201] In this invention, the server includes information acquisition means for recording speech information to detect abnormalities in emotional state, conversion means for converting it into text data, and emotion analysis means for analyzing emotions. This makes it possible to monitor the emotional and physical health of elderly people and to promptly notify them if an abnormality is detected.

[0202] "Information acquisition means" refers to a device or system for recording speech information from elderly people.

[0203] "Conversion means" refers to a device or system that has the function of converting recorded speech information into text data.

[0204] "Analysis means" refers to a device or system that has the function of analyzing text data to learn speech patterns.

[0205] An "anomaly detection means" is a device or system that has the function of detecting abnormalities in speech or emotion based on learned patterns.

[0206] "Notification generation means" refers to a device or system that has the function of generating notifications to relevant parties based on detected anomalies.

[0207] "Emotion analysis means" refers to a device or system that has the function of analyzing emotions from speech information.

[0208] "Integration means" refers to a device or system that has the function of integrating the results of emotion analysis means and anomaly detection means to detect abnormalities in emotional states.

[0209] "Information measurement means" refers to a device or system for measuring bodily information.

[0210] "Information transmission means" refers to a device or system for recording and transmitting measured physical information.

[0211] "Data matching means" refers to a device or system that has the function of matching physical information measured when an anomaly is detected with the actual data.

[0212] "History management means" refers to a device or system that has the function of recording and managing past history based on recorded speech information and physical information.

[0213] This invention comprises a system for comprehensively monitoring the emotional and physical health status of elderly individuals. The main components are information acquisition means, conversion means, analysis means, anomaly detection means, notification generation means, emotion analysis means, integration means, and information measurement means.

[0214] Each elderly person will be provided with a multi-functional bracelet device that records voice and physical information. This device is equipped with a microphone, heart rate and body temperature sensors, and digitally acquires daily speech information and physical condition data.

[0215] The server receives acquired speech information in real time and converts it into text data using speech recognition software. The Python `speech_recognition` library is used for this process. This text data undergoes pattern learning using an analysis tool, and anomalies in emotion and speech patterns are detected by an anomaly detection tool and an emotion analysis tool. A machine learning model from scikit-learn is used for emotion analysis. The results are then comprehensively evaluated by an integration tool to determine the presence or absence of anomalies.

[0216] Users receive relevant notifications via their smartphones or smart glasses. These notifications include information about the specific nature of detected anomalies and changes in emotions. This allows users to quickly and appropriately manage the safety and health of the elderly.

[0217] For example, if an elderly person makes many negative statements while walking and their heart rate is simultaneously elevated, the emotion analysis system will recognize this as a stressful state. The server can immediately generate an alert and send a notification to the user stating, "The elderly person is experiencing stress. Their heart rate is also elevated," prompting quick intervention.

[0218] The generative AI model used includes an example of supplying the following prompt: "An elderly person showed emotional changes during a routine activity. Analyze the audio and use the emotion engine to detect anomalies."

[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0220] Step 1:

[0221] The device acquires speech and physical information from elderly individuals using a multi-functional bracelet device. Inputs include voice data from a microphone, heart rate, and data from a body temperature sensor. Outputs include sending the acquired voice and physical data in digital format to a server.

[0222] Step 2:

[0223] The server converts the received audio data into text data using the Python speech_recognition library. At this stage, the audio data is the input, and the speech recognition engine analyzes it to generate text information as output.

[0224] Step 3:

[0225] The server analyzes the converted text data and learns speech patterns. Text data is the input, and important speech patterns are extracted using the analysis method. The learned speech patterns are generated as output.

[0226] Step 4:

[0227] The server analyzes emotions from text data using emotion analysis tools. Here, a machine learning model from scikit-learn is applied to the text data as input, and data indicating the emotional state is obtained as output.

[0228] Step 5:

[0229] The server evaluates speech patterns and sentiment analysis results using an integration mechanism to detect anomalies. In this step, speech patterns and sentiment data are integrated as input, a determination is made as to whether there are any unnatural changes or anomalies, and the anomaly detection result is obtained as output.

[0230] Step 6:

[0231] When the server detects an anomaly, it uses a notification generation mechanism to create an alarm. Here, the anomaly detection result is taken as input, and an alarm message for the user is generated as output.

[0232] Step 7:

[0233] Users receive notifications sent from a server via their smartphones or smart glasses. These notifications include an alert message as input and output that allows the user to understand the situation and prepare appropriate action.

[0234] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0235] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0237] [Second Embodiment]

[0238] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0239] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0241] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0245] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0246] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0247] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0248] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0249] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0250] This invention is a monitoring system that uses speech and physical information of elderly people, and by combining multiple functions, it enables real-time monitoring of the safety and health status of elderly people.

[0251] The device consists of a bracelet-type device worn on the wrist of an elderly person. This device is equipped with a microphone for recording voice and has the function of recording everyday conversations. It also has built-in sensors for measuring physical information such as heart rate and body temperature. This data is transmitted to a server via wireless communication.

[0252] When the server receives voice data transmitted from the terminal, it first converts it into text data using a speech recognition engine. The converted text data is then analyzed by a generative AI to learn the speech patterns of elderly people. The generative AI focuses on the characteristics of the speech content, the frequency of words used, and the flow of context to detect abnormal patterns. If an abnormality is detected, a notification is sent to pre-configured contacts. The server also simultaneously receives physical information transmitted from the terminal and stores the measurement results in a database.

[0253] Users receive and view information transmitted from the server via their smartphones. The application allows users to check logs of the elderly person's health status and receive abnormality notifications. It also includes features to allow users to investigate the nature of the abnormality in more detail or to consult with healthcare professionals, if necessary.

[0254] As a concrete example of this system, suppose an elderly person living alone starts making more frequent remarks about forgetting things during everyday conversations. The server detects this change in speech pattern as an anomaly and sends a notification to the family member (the user). Upon receiving this notification, the family can take appropriate action, such as seeking early medical advice. In this way, the present invention supports the health and safety of the elderly.

[0255] The following describes the processing flow.

[0256] Step 1:

[0257] The device records the elderly person's voice in real time. This is done using a built-in microphone, and the voice data is saved to a buffer at regular intervals.

[0258] Step 2:

[0259] The terminal transmits voice data to the server via wireless communication. The transmission is triggered by pre-configured conditions, enabling efficient data transmission.

[0260] Step 3:

[0261] The server analyzes the received audio data and converts it into text data using a speech recognition engine. During this process, pre-processing such as noise reduction is performed to improve recognition accuracy.

[0262] Step 4:

[0263] The server passes text data through a generative AI and compares it to past speech patterns. The generative AI uses a pre-trained model to detect changes in the text data, identifying new patterns and unusually frequent utterances.

[0264] Step 5:

[0265] When the server detects an anomaly, it sends an alert to the configured notification recipient. The alert will include a description of the anomaly and recommended actions.

[0266] Step 6:

[0267] The device measures the physical information of elderly individuals. This information includes body temperature and heart rate, and this data is sent to a server.

[0268] Step 7:

[0269] The server analyzes the physical information and compares it with past data to check for any abnormalities. If no abnormalities are found, it makes a comprehensive judgment based on this information along with the voice data.

[0270] Step 8:

[0271] Users receive notifications and data from the server via a smartphone app. Based on the received information, the system monitors the elderly person's condition and considers necessary actions.

[0272] Step 9:

[0273] Users can check their history of abnormalities and health data within the app. This allows for timely medical consultations and appropriate actions. This information is continuously logged.

[0274] (Example 1)

[0275] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0276] Effectively monitoring the safety and health of the elderly is a critical challenge in a growing elderly population. With conventional technologies, it was difficult for those around the elderly to understand their condition unless they themselves were aware of an abnormality. Furthermore, detecting abnormalities based not only on physical information but also on speech patterns was challenging, potentially leading to delays in necessary responses.

[0277] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0278] In this invention, the server includes voice collection means, voice recognition means, generative model learning means, anomaly detection means, notification generation means, physical condition measurement means, and physical information transmission means. This enables comprehensive state monitoring using both the speech and physical information of elderly individuals, allowing for early detection and notification of anomalies.

[0279] "Voice collection means" refers to a device or function that records the everyday conversations of elderly people as audio.

[0280] "Speech recognition means" refers to a device or function that converts recorded speech data into text data.

[0281] "Generative model learning means" refers to a device or program that analyzes converted text data and has the function of learning the speech patterns of elderly people.

[0282] The "abnormality detection means" refers to a device or function for detecting speech or actions that are different from normal based on the learned speech patterns.

[0283] The "notification generation means" refers to a device or function that generates and transmits a message or alert to notify a designated recipient when an abnormality is detected.

[0284] The "physical condition measurement means" refers to sensors or devices that measure physical information such as the heart rate and body temperature of the elderly.

[0285] The "physical information transmission means" refers to a device or protocol that has the function of recording the measured physical information and transmitting it to a remote device such as a server.

[0286] This invention relates to a system for monitoring the health status using the speech information and physical information of the elderly. This system is implemented by each component of the server, terminal, and user.

[0287] Data collection by the terminal:

[0288] The terminal is composed of a bracelet-type device worn on the wrist of the elderly. As a voice collection means, a microphone is installed in the terminal to record the daily conversations of the elderly as voice data. In addition, a heart rate monitor and a body temperature sensor are built in as physical condition measurement means to regularly measure this physical information. The collected data is transferred to the server via Bluetooth or Wi-Fi communication using the physical information transmission means.

[0289] Data analysis by the server:

[0290] The server converts speech data into text data using speech recognition means. This transforms the speech data into an analyzable format. Using generative model learning means, the speech patterns of elderly people are analyzed and learned from the text data. Anomaly detection means can detect anomalies based on the learned patterns. If an anomaly is detected, a notification is sent to a pre-configured contact (e.g., family) using notification generation means. These notifications may include historical data.

[0291] User verification of information:

[0292] Users receive health status data and abnormality notifications from a server via communication devices such as smartphones. Through a dedicated application, they can review past history and investigate the details of abnormalities. The application also provides information to consult with healthcare professionals as needed.

[0293] Specific example:

[0294] As a concrete example, if an elderly person living alone frequently says "I've been forgetting things a lot lately" in conversation, the server will detect this change in speech pattern as an anomaly. Subsequently, a notification is sent to the family member (the user), who can then consider seeking medical advice early.

[0295] Example of a prompt:

[0296] "Detect abnormalities in the speech patterns of elderly individuals. As a specific example, analyze changes in frequently used phrases."

[0297] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0298] Step 1:

[0299] The terminal records the daily conversations of the elderly as voice data. A microphone is used as the voice collection means. The input is the spoken voice, and the output is a voice data file. As a specific operation, the terminal captures voices at regular time intervals and saves them as data files.

[0300] Step 2:

[0301] The terminal measures physical information such as heart rate and body temperature. Data is obtained from a heart rate sensor and a body temperature sensor using physical state measurement means. The input is the physical state, and the output is heart rate and body temperature data. As a specific operation, the sensor reads data from the body and records it in storage as numerical data.

[0302] Step 3:

[0303] The terminal transmits the collected voice data and physical information to the server via Bluetooth or Wi-Fi. Here, the recorded data file is the input, and the data packet sent to the server is the output. As a specific operation, the terminal periodically performs data transmission using the communication module.

[0304] Step 4:

[0305] The server converts the received voice data into text data using voice recognition means. The input is the voice data file, and the output is the spoken content in text format. As a specific operation, the voice recognition algorithm analyzes the voice and performs the process of converting it into text.

[0306] Step 5:

[0307] The server analyzes the text data using generative model learning means and learns the speech patterns of the elderly. The input is the text data, and the output is the learned speech pattern. As a specific operation, the generative AI analyzes the structure and content of the text and records the frequency of words and the flow of context.

[0308] Step 6:

[0309] The server uses anomaly detection means to detect anomalies based on learned speech patterns. The input is the learned speech pattern, and the output is information about the detected anomaly. Specifically, it compares existing patterns with new data and identifies changes that deviate from the standard.

[0310] Step 7:

[0311] The server uses a notification generation mechanism to create a notification when an anomaly is detected and sends it to a predetermined contact. The input is information about the detected anomaly, and the output is the anomaly notification. Specifically, it generates a message containing details of the anomaly and sends the notification via the communication network.

[0312] Step 8:

[0313] The user receives and displays notifications from the server using a smartphone app. The input is an anomaly notification from the server, and the output is the notification information displayed on the screen. Specifically, the user can open the app, check the notification content, and look up further details if necessary.

[0314] (Application Example 1)

[0315] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] In technologies that monitor the safety and health of the elderly using speech and biometric information, conventional systems only focus on detecting anomalies in speech and biometric information, making it difficult to detect external safety threats from ambient noise. Furthermore, there was a lack of means to comprehensively manage and monitor this information in real time. Therefore, there is a need to further improve the safety and convenience of life for the elderly.

[0317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0318] In this invention, the server includes a device for recording speech information of elderly people, a conversion device for converting the recorded speech information into string data, a learning device for analyzing the string data and learning speech patterns, a detection device for detecting abnormalities based on the learned patterns, a notification device for generating notifications based on the detected abnormalities, and an environmental analysis device for detecting safety abnormalities by analyzing ambient sounds. This makes it possible to detect not only speech and biological abnormalities of elderly people, but also safety threats in the surrounding environment, thereby realizing more comprehensive monitoring.

[0319] The term "elderly person" refers to an individual for whom health management and safety become particularly important as they age.

[0320] "Speech information" refers to the linguistic data contained in the voices of elderly people, including the content of conversations and monologues.

[0321] "String data" refers to text-formatted data converted from audio data.

[0322] "Biometric information" refers to physical data, including heart rate, body temperature, and activity level, for elderly individuals.

[0323] The term "device" refers to an instrument or device with a specific function, used for various data collection and processing within a system.

[0324] A "conversion device" refers to a device that has the function of converting audio data into text data.

[0325] A "learning device" refers to a device that has the function of continuously analyzing data and identifying and memorizing features and patterns.

[0326] A "detection device" refers to a device that has the function of identifying and reporting anomalies based on learned data.

[0327] A "notification device" refers to a device that has the function of informing relevant parties about detected anomalies.

[0328] An "environmental analysis device" refers to a device that monitors ambient sounds and environmental data and has the function of detecting abnormalities.

[0329] The system that realizes this invention is designed to monitor the daily safety and health of elderly people. The system mainly consists of family members or caregivers as users, terminals worn by the elderly, and a server responsible for data processing and notifications. Its specific functions and operation are described below.

[0330] The device is worn on the wrist of elderly individuals and contains a microphone and multiple sensors (for measuring heart rate and body temperature). This allows for the continuous collection of speech and biometric information from the elderly person. This data is transmitted to a server via wireless communication.

[0331] The server first converts the audio data sent from the terminal into text data using speech recognition software. Specifically, Microsoft Azure Speech Services performs this role. This text data is then analyzed by a generative AI model. For example, it learns past speech patterns and detects abnormal speech or unstable contextual changes. By giving the generative AI model prompts such as, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of abnormalities based on that," the model performs the appropriate analysis.

[0332] Simultaneously, biometric information such as heart rate and body temperature is analyzed by the server. The IBM Watson Health platform is used for real-time data analysis. In addition, ambient sounds are analyzed using Google Cloud's voice analysis service to detect abnormal sounds occurring around the elderly person. In this way, the system provides comprehensive monitoring.

[0333] Users receive information sent from the server through a smartphone application. If an anomaly is detected, a notification is sent to the user via Firebase Cloud Messaging. Users can check the details of the anomaly within the app and obtain information to consult with a medical professional if necessary. For example, a notification such as "Your mother may be frequently mistaking place names" may be sent, allowing the user to take early action. In this way, the entire system works together to build a mechanism that supports the lives of the elderly.

[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0335] Step 1:

[0336] The device collects speech and biometric information from elderly individuals. It records voice data using a microphone and measures heart rate and body temperature using built-in sensors. The collected data is transmitted wirelessly to a server. The input is analog signal data recorded by the microphone and sensors, which is then converted to digital data and sent to the server.

[0337] Step 2:

[0338] The server processes the audio data sent from the terminal using speech recognition software and converts it into text data. Azure Speech Services is used for this process. The input is audio data, and the output is text-based string data. The audio signal is analyzed, and a language model is used to generate the correct word sequence.

[0339] Step 3:

[0340] The server analyzes the converted string data using a generative AI model. Specifically, it learns speech patterns using prompt sentences and performs anomaly detection. The generative AI utilizes the prompt, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of anomalies based on these." The input is string data, and the output is the analysis result regarding the presence or absence of anomalies.

[0341] Step 4:

[0342] The server also analyzes biometric information. Using IBM Watson Health, it analyzes heart rate and body temperature and detects abnormalities by comparing them to common health indicators. The input is biometric data collected by sensors, and the output is a judgment of whether it is normal or abnormal.

[0343] Step 5:

[0344] Simultaneously, the server uses Google Cloud's speech analysis service to analyze ambient sounds and detect anomalies. The input is ambient sound data, and the output is the result of safety-related anomaly detection. It identifies specific acoustic patterns and generates a warning if they are deemed anomaly.

[0345] Step 6:

[0346] The server generates and sends a notification to the user based on the detected anomaly. The notification is sent to smartphones and other devices via Firebase Cloud Messaging. The input is the anomaly detection result, and the output is an informational notification to the user. It includes specific details, allowing the user to receive it and take appropriate action.

[0347] Step 7:

[0348] Users view received notifications through a smartphone application. These notifications include detailed information about the anomaly and recommended next actions. Within the app, users can review past history and an overview of their overall health status, and contact healthcare professionals if necessary. The input is the notification message, and the output is the user's confirmation of the situation and subsequent actions.

[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0350] This invention is a system that combines the acquisition of speech and physical information from elderly individuals with an emotion engine that analyzes emotions from their speech. This system comprehensively monitors the physical and emotional health status of elderly individuals, enabling early detection of abnormalities and appropriate responses.

[0351] The device is a bracelet-type device worn by elderly individuals and is equipped with a microphone for voice recognition and sensors to measure heart rate and body temperature. The device records voice in real time and transmits the acquired data to a server sequentially.

[0352] The server receives voice data transmitted from the terminal and performs a series of processes to convert it into text using a speech recognition engine. The converted text data is analyzed by a generative AI and sent to an emotion engine that recognizes emotions from the speech content along with the speech patterns. The emotion engine has a self-learning function and analyzes and stores the emotional states of elderly people from their daily speech to detect abnormal emotional changes.

[0353] Furthermore, the server integrates speech pattern learning and emotion analysis results from artificial intelligence, and performs anomaly detection by comparing this with health data. When an anomaly is detected, an alert is generated for the relevant parties. The alert includes details about the specific nature of the anomaly and what kind of emotional changes are occurring.

[0354] Users can receive this notification through a smartphone app and check the status of the elderly person. The historical data included in the notification serves as a guide for users to understand past health conditions and emotional tendencies. It also provides useful information when considering appropriate measures and medical support.

[0355] As a concrete example, consider a scenario where an elderly person's daily conversations contain more negative words than usual. The emotion engine recognizes this change as a shift from joy to anxiety and reports it as an anomaly to the server. If the server detects an increase in heart rate, it also considers it as a physical health problem and prompts the user to take prompt action. Thus, the present invention is a system that comprehensively supports the physical and mental health of the elderly through complex data analysis.

[0356] The following describes the processing flow.

[0357] Step 1:

[0358] The device records the elderly person's voice in real time. It acquires voice data using a built-in microphone and stores it in a buffer in an appropriate format.

[0359] Step 2:

[0360] The device transmits recorded audio data to the server at regular intervals. Wireless communication technology is used to ensure efficient data transmission.

[0361] Step 3:

[0362] The server passes the received audio data to the speech recognition engine, which converts it into text data. During the conversion process, noise filtering and other techniques are performed to improve accuracy.

[0363] Step 4:

[0364] The server inputs the generated text data into the generating AI, which then compares it to past speech patterns. During this process, the trained algorithm analyzes any changes in the speech content.

[0365] Step 5:

[0366] The server sends text data to the emotion engine, which analyzes the emotional state from the utterance. The emotion engine extracts specific emotional elements and determines the emotional tendency.

[0367] Step 6:

[0368] The server integrates the emotion analysis results and the speech pattern analysis results. From the integrated data, it detects abnormal behavior and emotional changes and prepares to generate alerts.

[0369] Step 7:

[0370] The device measures physical information such as the heart rate and body temperature of elderly individuals and transmits this data to a server. Measurements are performed regularly to maintain real-time accuracy.

[0371] Step 8:

[0372] The server comprehensively compares physical information with speech and emotion analysis data. This allows for multifaceted verification of anomalies and reduces the risk of misrecognition.

[0373] Step 9:

[0374] Users receive alert notifications via their smartphones. The notifications detail the specific nature of the anomaly, changes in emotional state, and possible underlying physical problems.

[0375] Step 10:

[0376] Users can review the information received through the app and take appropriate action regarding the elderly person's situation as needed. Past historical data can also be referenced to support more accurate decision-making.

[0377] (Example 2)

[0378] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0379] In today's aging society, there is a need for a system that can accurately monitor the physical and mental health of the elderly and detect abnormalities early. Conventional technologies have the problem that health monitoring of the elderly is one-sided, and it is difficult to make a comprehensive judgment of their physical and mental state, making it highly likely that abnormalities will be overlooked.

[0380] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0381] In this invention, the server includes voice acquisition means, emotion analysis means, and anomaly detection means. This enables comprehensive monitoring of the health status of elderly people based on their speech content and physiological information, as well as rapid anomaly detection.

[0382] "Voice acquisition means" refers to a device for collecting speech information from elderly people in real time, and mainly includes voice input devices such as microphones.

[0383] "Speech recognition means" refers to technology for converting acquired speech data into text data, and this conversion is performed using a speech recognition engine or AI technology.

[0384] "Emotional analysis methods" refer to analytical techniques for recognizing emotional states from text data, and utilize generative AI models to grasp the emotions contained in spoken language.

[0385] An "anomaly detection means" is a device or program that integrates acquired emotional data and health data to detect anomalies.

[0386] A "notification generation means" is a device or system that generates warnings or notifications based on detected abnormal information and transmits them to the user.

[0387] A "biometric measurement device" is a device used to measure physiological information such as heart rate and body temperature in elderly individuals, and utilizes various sensors.

[0388] "Data transmission means" refers to communication means intended for transmitting measured physiological information and analysis results to a server or other device.

[0389] A "data matching means" is a device or program that, when an abnormality in speech or physiological information is detected, matches that information and performs further analysis.

[0390] A "history management system" is a system for recording and managing past speech and physiological information.

[0391] This invention is a system that comprehensively monitors the physical and emotional health status of elderly individuals and enables early detection of abnormalities. The system primarily consists of terminals, a server, and users.

[0392] First, the device is a bracelet-type device worn by elderly individuals. The device incorporates a voice acquisition mechanism for real-time voice recording, as well as biometric measurement devices for measuring heart rate and body temperature. The acquired data is transmitted to a server using a secure communication protocol. This involves data transmission via Bluetooth and an internet connection.

[0393] Next, the server implements a commonly available speech recognition engine (e.g., a speech recognition API) as a means of speech recognition, which converts the speech data sent from the terminal into text data. The converted text data is analyzed by an emotion analysis means to recognize the emotional state from the elderly person's speech. A generative AI model is used for this, which learns the changes in the elderly person's emotions and is capable of self-improvement.

[0394] The server then integrates emotional states and biometric data using anomaly detection mechanisms to monitor the health status of elderly individuals. If an anomaly is detected, a notification generation mechanism generates a warning, which is sent to the user. The warning includes specific details of the anomaly, allowing the user to take prompt action.

[0395] Finally, users can receive alerts via their smartphones and check on the elderly person's condition. The smartphone app includes a history management system that tracks past speech and physiological data, allowing users to precisely track changes in their health status.

[0396] For example, if an elderly person says, "I'm feeling down today and don't want to do anything," the emotion analysis system will recognize this as anxiety or sadness and report it to the server as an abnormality. Furthermore, if a higher-than-normal heart rate is detected, a notification will be generated immediately. An example of a prompt message would be, "Please identify which utterance caused the emotional abnormality and decide which medical institution you should consult."

[0397] In this way, the system provides comprehensive support for the physical and emotional health of the elderly and offers advanced data analysis capabilities for rapid response in emergencies.

[0398] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0399] Step 1:

[0400] The device acquires the voice of the elderly person wearing it in real time and simultaneously measures biometric information such as heart rate and body temperature. The input obtained in this process consists of raw voice data and physiological data. This data is collected via a voice acquisition device and transmitted to a server via a data transmission device. The resulting output is voice data and physiological data that are ready for further processing on the server side.

[0401] Step 2:

[0402] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. The input is raw audio data, and the output is the corresponding text data. Specifically, speech-to-text conversion is performed using a speech recognition API or similar method. Through this process, the spoken content is obtained in text format and sent to the next analysis step.

[0403] Step 3:

[0404] The server inputs the converted text data into an emotion analysis device, which then uses a generative AI model to analyze the emotions. The input is text data, and the output is the analyzed emotion data. Specifically, the generative AI model extracts emotional nuances from the text content and evaluates the emotional state of the elderly person. This information prepares the device for anomaly detection.

[0405] Step 4:

[0406] The server integrates emotional and physiological data using anomaly detection mechanisms to monitor the health status of elderly individuals. Inputs are emotional and physiological data, and output is information on whether or not an anomaly occurred. Specifically, the system compares the data to dynamically set baseline values ​​and patterns; an anomaly is detected if the data exceeds the normal range.

[0407] Step 5:

[0408] The server generates a warning to the user using a notification generation mechanism when an anomaly is detected. The input is the anomaly detection result, and the output is a warning message to the user. This message includes detailed information about the health status of the elderly person, and the user can receive the notification through the app.

[0409] Step 6:

[0410] Users receive notifications via a smartphone app and check the health status of elderly individuals through the displayed notifications. The input is the notification message, and the output is information to help the user understand the situation. Based on this information, users can contact medical institutions or perform additional monitoring as needed.

[0411] Through these steps, this system provides comprehensive support for the physical and emotional health of older adults, enabling prompt and appropriate responses.

[0412] (Application Example 2)

[0413] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0414] In today's living environments for the elderly, not only physical health issues but also emotional health issues are becoming increasingly important. However, conventional technology has struggled to effectively detect emotional changes and integrate them with physical information for comprehensive health management. Therefore, there is a need for a system that enables rapid and appropriate responses when elderly individuals exhibit emotional or physical abnormalities.

[0415] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0416] In this invention, the server includes information acquisition means for recording speech information to detect abnormalities in emotional state, conversion means for converting it into text data, and emotion analysis means for analyzing emotions. This makes it possible to monitor the emotional and physical health of elderly people and to promptly notify them if an abnormality is detected.

[0417] "Information acquisition means" refers to a device or system for recording speech information from elderly people.

[0418] "Conversion means" refers to a device or system that has the function of converting recorded speech information into text data.

[0419] "Analysis means" refers to a device or system that has the function of analyzing text data to learn speech patterns.

[0420] An "anomaly detection means" is a device or system that has the function of detecting abnormalities in speech or emotion based on learned patterns.

[0421] "Notification generation means" refers to a device or system that has the function of generating notifications to relevant parties based on detected anomalies.

[0422] "Emotion analysis means" refers to a device or system that has the function of analyzing emotions from speech information.

[0423] "Integration means" refers to a device or system that has the function of integrating the results of emotion analysis means and anomaly detection means to detect abnormalities in emotional states.

[0424] "Information measurement means" refers to a device or system for measuring bodily information.

[0425] "Information transmission means" refers to a device or system for recording and transmitting measured physical information.

[0426] "Data matching means" refers to a device or system that has the function of matching physical information measured when an anomaly is detected with the actual data.

[0427] "History management means" refers to a device or system that has the function of recording and managing past history based on recorded speech information and physical information.

[0428] This invention comprises a system for comprehensively monitoring the emotional and physical health status of elderly individuals. The main components are information acquisition means, conversion means, analysis means, anomaly detection means, notification generation means, emotion analysis means, integration means, and information measurement means.

[0429] Each elderly person will be provided with a multi-functional bracelet device that records voice and physical information. This device is equipped with a microphone, heart rate and body temperature sensors, and digitally acquires daily speech information and physical condition data.

[0430] The server receives acquired speech information in real time and converts it into text data using speech recognition software. The Python `speech_recognition` library is used for this process. This text data undergoes pattern learning using an analysis tool, and anomalies in emotion and speech patterns are detected by an anomaly detection tool and an emotion analysis tool. A machine learning model from scikit-learn is used for emotion analysis. The results are then comprehensively evaluated by an integration tool to determine the presence or absence of anomalies.

[0431] Users receive relevant notifications via their smartphones or smart glasses. These notifications include information about the specific nature of detected anomalies and changes in emotions. This allows users to quickly and appropriately manage the safety and health of the elderly.

[0432] For example, if an elderly person makes many negative statements while walking and their heart rate is simultaneously elevated, the emotion analysis system will recognize this as a stressful state. The server can immediately generate an alert and send a notification to the user stating, "The elderly person is experiencing stress. Their heart rate is also elevated," prompting quick intervention.

[0433] The generative AI model used includes an example of supplying the following prompt: "An elderly person showed emotional changes during a routine activity. Analyze the audio and use the emotion engine to detect anomalies."

[0434] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0435] Step 1:

[0436] The device acquires speech and physical information from elderly individuals using a multi-functional bracelet device. Inputs include voice data from a microphone, heart rate, and data from a body temperature sensor. Outputs include sending the acquired voice and physical data in digital format to a server.

[0437] Step 2:

[0438] The server converts the received audio data into text data using the Python speech_recognition library. At this stage, the audio data is the input, and the speech recognition engine analyzes it to generate text information as output.

[0439] Step 3:

[0440] The server analyzes the converted text data and learns speech patterns. Text data is the input, and important speech patterns are extracted using the analysis method. The learned speech patterns are generated as output.

[0441] Step 4:

[0442] The server analyzes emotions from text data using emotion analysis tools. Here, a machine learning model from scikit-learn is applied to the text data as input, and data indicating the emotional state is obtained as output.

[0443] Step 5:

[0444] The server evaluates speech patterns and sentiment analysis results using an integration mechanism to detect anomalies. In this step, speech patterns and sentiment data are integrated as input, a determination is made as to whether there are any unnatural changes or anomalies, and the anomaly detection result is obtained as output.

[0445] Step 6:

[0446] When the server detects an anomaly, it uses a notification generation mechanism to create an alarm. Here, the anomaly detection result is taken as input, and an alarm message for the user is generated as output.

[0447] Step 7:

[0448] Users receive notifications sent from a server via their smartphones or smart glasses. These notifications include an alert message as input and output that allows the user to understand the situation and prepare appropriate action.

[0449] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0450] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0451] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0452] [Third Embodiment]

[0453] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0454] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0455] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0456] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0457] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0458] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0459] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0460] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0461] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0462] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0463] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0464] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0465] This invention is a monitoring system that uses speech and physical information of elderly people, and by combining multiple functions, it enables real-time monitoring of the safety and health status of elderly people.

[0466] The device consists of a bracelet-type device worn on the wrist of an elderly person. This device is equipped with a microphone for recording voice and has the function of recording everyday conversations. It also has built-in sensors for measuring physical information such as heart rate and body temperature. This data is transmitted to a server via wireless communication.

[0467] When the server receives voice data transmitted from the terminal, it first converts it into text data using a speech recognition engine. The converted text data is then analyzed by a generative AI to learn the speech patterns of elderly people. The generative AI focuses on the characteristics of the speech content, the frequency of words used, and the flow of context to detect abnormal patterns. If an abnormality is detected, a notification is sent to pre-configured contacts. The server also simultaneously receives physical information transmitted from the terminal and stores the measurement results in a database.

[0468] Users receive and view information transmitted from the server via their smartphones. The application allows users to check logs of the elderly person's health status and receive abnormality notifications. It also includes features to allow users to investigate the nature of the abnormality in more detail or to consult with healthcare professionals, if necessary.

[0469] As a concrete example of this system, suppose an elderly person living alone starts making more frequent remarks about forgetting things during everyday conversations. The server detects this change in speech pattern as an anomaly and sends a notification to the family member (the user). Upon receiving this notification, the family can take appropriate action, such as seeking early medical advice. In this way, the present invention supports the health and safety of the elderly.

[0470] The following describes the processing flow.

[0471] Step 1:

[0472] The device records the elderly person's voice in real time. This is done using a built-in microphone, and the voice data is saved to a buffer at regular intervals.

[0473] Step 2:

[0474] The terminal transmits voice data to the server via wireless communication. The transmission is triggered by pre-configured conditions, enabling efficient data transmission.

[0475] Step 3:

[0476] The server analyzes the received audio data and converts it into text data using a speech recognition engine. During this process, pre-processing such as noise reduction is performed to improve recognition accuracy.

[0477] Step 4:

[0478] The server passes text data through a generative AI and compares it to past speech patterns. The generative AI uses a pre-trained model to detect changes in the text data, identifying new patterns and unusually frequent utterances.

[0479] Step 5:

[0480] When the server detects an anomaly, it sends an alert to the configured notification recipient. The alert will include a description of the anomaly and recommended actions.

[0481] Step 6:

[0482] The device measures the physical information of elderly individuals. This information includes body temperature and heart rate, and this data is sent to a server.

[0483] Step 7:

[0484] The server analyzes the physical information and compares it with past data to check for any abnormalities. If no abnormalities are found, it makes a comprehensive judgment based on this information along with the voice data.

[0485] Step 8:

[0486] Users receive notifications and data from the server via a smartphone app. Based on the received information, the system monitors the elderly person's condition and considers necessary actions.

[0487] Step 9:

[0488] Users can check their history of abnormalities and health data within the app. This allows for timely medical consultations and appropriate actions. This information is continuously logged.

[0489] (Example 1)

[0490] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0491] Effectively monitoring the safety and health of the elderly is a critical challenge in a growing elderly population. With conventional technologies, it was difficult for those around the elderly to understand their condition unless they themselves were aware of an abnormality. Furthermore, detecting abnormalities based not only on physical information but also on speech patterns was challenging, potentially leading to delays in necessary responses.

[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0493] In this invention, the server includes voice collection means, voice recognition means, generative model learning means, anomaly detection means, notification generation means, physical condition measurement means, and physical information transmission means. This enables comprehensive state monitoring using both the speech and physical information of elderly individuals, allowing for early detection and notification of anomalies.

[0494] "Voice collection means" refers to a device or function that records the everyday conversations of elderly people as audio.

[0495] "Speech recognition means" refers to a device or function that converts recorded speech data into text data.

[0496] "Generative model learning means" refers to a device or program that analyzes converted text data and has the function of learning the speech patterns of elderly people.

[0497] An "anomaly detection means" refers to a device or function that detects speech or behavior that is different from the norm, based on learned speech patterns.

[0498] "Notification generation means" refers to a device or function that generates and sends messages or alerts to designated recipients when an anomaly is detected.

[0499] "Means of measuring physical condition" refers to sensors and devices that measure physical information such as heart rate and body temperature of elderly people.

[0500] "Body information transmission means" refers to devices or protocols that have the function of recording measured body information and transmitting it to remote devices such as servers.

[0501] This invention relates to a system for monitoring the health status of elderly individuals using their speech and physical information. This system is implemented using a server, a terminal, and a user component.

[0502] Data collection by devices:

[0503] The device consists of a bracelet-type device worn on the wrist of an elderly person. As a means of voice collection, the device is equipped with a microphone to record the elderly person's daily conversations as audio data. It also has a built-in heart rate monitor and body temperature sensor as means of measuring physical condition, and this physical information is measured periodically. The collected data is transmitted to a server via Bluetooth or Wi-Fi communication using a physical information transmission device.

[0504] Server-based data analysis:

[0505] The server converts speech data into text data using speech recognition means. This transforms the speech data into an analyzable format. Using generative model learning means, the speech patterns of elderly people are analyzed and learned from the text data. Anomaly detection means can detect anomalies based on the learned patterns. If an anomaly is detected, a notification is sent to a pre-configured contact (e.g., family) using notification generation means. These notifications may include historical data.

[0506] User verification of information:

[0507] Users receive health status data and abnormality notifications from a server via communication devices such as smartphones. Through a dedicated application, they can review past history and investigate the details of abnormalities. The application also provides information to consult with healthcare professionals as needed.

[0508] Specific example:

[0509] As a concrete example, if an elderly person living alone frequently says "I've been forgetting things a lot lately" in conversation, the server will detect this change in speech pattern as an anomaly. Subsequently, a notification is sent to the family member (the user), who can then consider seeking medical advice early.

[0510] Example of a prompt:

[0511] "Detect abnormalities in the speech patterns of elderly individuals. As a specific example, analyze changes in frequently used phrases."

[0512] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0513] Step 1:

[0514] The device records the daily conversations of elderly people as audio data. A microphone is used as the means of audio collection. The input is spoken voice, and the output is an audio data file. Specifically, the device captures audio at regular time intervals and saves it as a data file.

[0515] Step 2:

[0516] The device measures bodily information such as heart rate and body temperature. It acquires data from heart rate sensors and body temperature sensors using a bodily state measurement device. The input is the body's state, and the output is heart rate and body temperature data. Specifically, the sensors read data from the body and record it as numerical data in storage.

[0517] Step 3:

[0518] The device transmits collected voice data and physical information to a server via Bluetooth or Wi-Fi. Here, the recorded data file is the input, and the data packets sent to the server are the output. Specifically, the device periodically transmits data using a communication module.

[0519] Step 4:

[0520] The server converts received audio data into text data using speech recognition technology. The input is an audio data file, and the output is the spoken content in text format. Specifically, the speech recognition algorithm analyzes the audio and converts it into text.

[0521] Step 5:

[0522] The server analyzes text data using a generative model learning method to learn the speech patterns of elderly people. The input is text data, and the output is the learned speech patterns. Specifically, the generative AI analyzes the structure and content of the text and records the frequency of words and the flow of context.

[0523] Step 6:

[0524] The server uses anomaly detection means to detect anomalies based on learned speech patterns. The input is the learned speech pattern, and the output is information about the detected anomaly. Specifically, it compares existing patterns with new data and identifies changes that deviate from the standard.

[0525] Step 7:

[0526] The server uses a notification generation mechanism to create a notification when an anomaly is detected and sends it to a predetermined contact. The input is information about the detected anomaly, and the output is the anomaly notification. Specifically, it generates a message containing details of the anomaly and sends the notification via the communication network.

[0527] Step 8:

[0528] The user receives and displays notifications from the server using a smartphone app. The input is an anomaly notification from the server, and the output is the notification information displayed on the screen. Specifically, the user can open the app, check the notification content, and look up further details if necessary.

[0529] (Application Example 1)

[0530] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0531] In technologies that monitor the safety and health of the elderly using speech and biometric information, conventional systems only focus on detecting anomalies in speech and biometric information, making it difficult to detect external safety threats from ambient noise. Furthermore, there was a lack of means to comprehensively manage and monitor this information in real time. Therefore, there is a need to further improve the safety and convenience of life for the elderly.

[0532] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0533] In this invention, the server includes a device for recording speech information of elderly people, a conversion device for converting the recorded speech information into string data, a learning device for analyzing the string data and learning speech patterns, a detection device for detecting abnormalities based on the learned patterns, a notification device for generating notifications based on the detected abnormalities, and an environmental analysis device for detecting safety abnormalities by analyzing ambient sounds. This makes it possible to detect not only speech and biological abnormalities of elderly people, but also safety threats in the surrounding environment, thereby realizing more comprehensive monitoring.

[0534] The term "elderly person" refers to an individual for whom health management and safety become particularly important as they age.

[0535] "Speech information" refers to the linguistic data contained in the voices of elderly people, including the content of conversations and monologues.

[0536] "String data" refers to text-formatted data converted from audio data.

[0537] "Biometric information" refers to physical data, including heart rate, body temperature, and activity level, for elderly individuals.

[0538] The term "device" refers to an instrument or device with a specific function, used for various data collection and processing within a system.

[0539] A "conversion device" refers to a device that has the function of converting audio data into text data.

[0540] A "learning device" refers to a device that has the function of continuously analyzing data and identifying and memorizing features and patterns.

[0541] A "detection device" refers to a device that has the function of identifying and reporting anomalies based on learned data.

[0542] A "notification device" refers to a device that has the function of informing relevant parties about detected anomalies.

[0543] An "environmental analysis device" refers to a device that monitors ambient sounds and environmental data and has the function of detecting abnormalities.

[0544] The system that realizes this invention is designed to monitor the daily safety and health of elderly people. The system mainly consists of family members or caregivers as users, terminals worn by the elderly, and a server responsible for data processing and notifications. Its specific functions and operation are described below.

[0545] The device is worn on the wrist of elderly individuals and contains a microphone and multiple sensors (for measuring heart rate and body temperature). This allows for the continuous collection of speech and biometric information from the elderly person. This data is transmitted to a server via wireless communication.

[0546] The server first converts the audio data sent from the terminal into text data using speech recognition software. Specifically, Microsoft Azure Speech Services performs this role. This text data is then analyzed by a generative AI model. For example, it learns past speech patterns and detects abnormal speech or unstable contextual changes. By giving the generative AI model prompts such as, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of abnormalities based on that," the model performs the appropriate analysis.

[0547] Simultaneously, biometric information such as heart rate and body temperature is analyzed by the server. The IBM Watson Health platform is used for real-time data analysis. In addition, ambient sounds are analyzed using Google Cloud's voice analysis service to detect abnormal sounds occurring around the elderly person. In this way, the system provides comprehensive monitoring.

[0548] Users receive information sent from the server through a smartphone application. If an anomaly is detected, a notification is sent to the user via Firebase Cloud Messaging. Users can check the details of the anomaly within the app and obtain information to consult with a medical professional if necessary. For example, a notification such as "Your mother may be frequently mistaking place names" may be sent, allowing the user to take early action. In this way, the entire system works together to build a mechanism that supports the lives of the elderly.

[0549] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0550] Step 1:

[0551] The device collects speech and biometric information from elderly individuals. It records voice data using a microphone and measures heart rate and body temperature using built-in sensors. The collected data is transmitted wirelessly to a server. The input is analog signal data recorded by the microphone and sensors, which is then converted to digital data and sent to the server.

[0552] Step 2:

[0553] The server processes the audio data sent from the terminal using speech recognition software and converts it into text data. Azure Speech Services is used for this process. The input is audio data, and the output is text-based string data. The audio signal is analyzed, and a language model is used to generate the correct word sequence.

[0554] Step 3:

[0555] The server analyzes the converted string data using a generative AI model. Specifically, it learns speech patterns using prompt sentences and performs anomaly detection. The generative AI utilizes the prompt, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of anomalies based on these." The input is string data, and the output is the analysis result regarding the presence or absence of anomalies.

[0556] Step 4:

[0557] The server also analyzes biometric information. Using IBM Watson Health, it analyzes heart rate and body temperature and detects abnormalities by comparing them to common health indicators. The input is biometric data collected by sensors, and the output is a judgment of whether it is normal or abnormal.

[0558] Step 5:

[0559] Simultaneously, the server uses Google Cloud's speech analysis service to analyze ambient sounds and detect anomalies. The input is ambient sound data, and the output is the result of safety-related anomaly detection. It identifies specific acoustic patterns and generates a warning if they are deemed anomaly.

[0560] Step 6:

[0561] The server generates and sends a notification to the user based on the detected anomaly. The notification is sent to smartphones and other devices via Firebase Cloud Messaging. The input is the anomaly detection result, and the output is an informational notification to the user. It includes specific details, allowing the user to receive it and take appropriate action.

[0562] Step 7:

[0563] Users view received notifications through a smartphone application. These notifications include detailed information about the anomaly and recommended next actions. Within the app, users can review past history and an overview of their overall health status, and contact healthcare professionals if necessary. The input is the notification message, and the output is the user's confirmation of the situation and subsequent actions.

[0564] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0565] This invention is a system that combines the acquisition of speech and physical information from elderly individuals with an emotion engine that analyzes emotions from their speech. This system comprehensively monitors the physical and emotional health status of elderly individuals, enabling early detection of abnormalities and appropriate responses.

[0566] The device is a bracelet-type device worn by elderly individuals and is equipped with a microphone for voice recognition and sensors to measure heart rate and body temperature. The device records voice in real time and transmits the acquired data to a server sequentially.

[0567] The server receives voice data transmitted from the terminal and performs a series of processes to convert it into text using a speech recognition engine. The converted text data is analyzed by a generative AI and sent to an emotion engine that recognizes emotions from the speech content along with the speech patterns. The emotion engine has a self-learning function and analyzes and stores the emotional states of elderly people from their daily speech to detect abnormal emotional changes.

[0568] Furthermore, the server integrates speech pattern learning and emotion analysis results from artificial intelligence, and performs anomaly detection by comparing this with health data. When an anomaly is detected, an alert is generated for the relevant parties. The alert includes details about the specific nature of the anomaly and what kind of emotional changes are occurring.

[0569] Users can receive this notification through a smartphone app and check the status of the elderly person. The historical data included in the notification serves as a guide for users to understand past health conditions and emotional tendencies. It also provides useful information when considering appropriate measures and medical support.

[0570] As a concrete example, consider a scenario where an elderly person's daily conversations contain more negative words than usual. The emotion engine recognizes this change as a shift from joy to anxiety and reports it as an anomaly to the server. If the server detects an increase in heart rate, it also considers it as a physical health problem and prompts the user to take prompt action. Thus, the present invention is a system that comprehensively supports the physical and mental health of the elderly through complex data analysis.

[0571] The following describes the processing flow.

[0572] Step 1:

[0573] The device records the elderly person's voice in real time. It acquires voice data using a built-in microphone and stores it in a buffer in an appropriate format.

[0574] Step 2:

[0575] The device transmits recorded audio data to the server at regular intervals. Wireless communication technology is used to ensure efficient data transmission.

[0576] Step 3:

[0577] The server passes the received audio data to the speech recognition engine, which converts it into text data. During the conversion process, noise filtering and other techniques are performed to improve accuracy.

[0578] Step 4:

[0579] The server inputs the generated text data into the generating AI, which then compares it to past speech patterns. During this process, the trained algorithm analyzes any changes in the speech content.

[0580] Step 5:

[0581] The server sends text data to the emotion engine, which analyzes the emotional state from the utterance. The emotion engine extracts specific emotional elements and determines the emotional tendency.

[0582] Step 6:

[0583] The server integrates the emotion analysis results and the speech pattern analysis results. From the integrated data, it detects abnormal behavior and emotional changes and prepares to generate alerts.

[0584] Step 7:

[0585] The device measures physical information such as the heart rate and body temperature of elderly individuals and transmits this data to a server. Measurements are performed regularly to maintain real-time accuracy.

[0586] Step 8:

[0587] The server comprehensively compares physical information with speech and emotion analysis data. This allows for multifaceted verification of anomalies and reduces the risk of misrecognition.

[0588] Step 9:

[0589] Users receive alert notifications via their smartphones. The notifications detail the specific nature of the anomaly, changes in emotional state, and possible underlying physical problems.

[0590] Step 10:

[0591] Users can review the information received through the app and take appropriate action regarding the elderly person's situation as needed. Past historical data can also be referenced to support more accurate decision-making.

[0592] (Example 2)

[0593] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0594] In today's aging society, there is a need for a system that can accurately monitor the physical and mental health of the elderly and detect abnormalities early. Conventional technologies have the problem that health monitoring of the elderly is one-sided, and it is difficult to make a comprehensive judgment of their physical and mental state, making it highly likely that abnormalities will be overlooked.

[0595] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0596] In this invention, the server includes voice acquisition means, emotion analysis means, and anomaly detection means. This enables comprehensive monitoring of the health status of elderly people based on their speech content and physiological information, as well as rapid anomaly detection.

[0597] "Voice acquisition means" refers to a device for collecting speech information from elderly people in real time, and mainly includes voice input devices such as microphones.

[0598] "Speech recognition means" refers to technology for converting acquired speech data into text data, and this conversion is performed using a speech recognition engine or AI technology.

[0599] "Emotional analysis methods" refer to analytical techniques for recognizing emotional states from text data, and utilize generative AI models to grasp the emotions contained in spoken language.

[0600] An "anomaly detection means" is a device or program that integrates acquired emotional data and health data to detect anomalies.

[0601] A "notification generation means" is a device or system that generates warnings or notifications based on detected abnormal information and transmits them to the user.

[0602] A "biometric measurement device" is a device used to measure physiological information such as heart rate and body temperature in elderly individuals, and utilizes various sensors.

[0603] "Data transmission means" refers to communication means intended for transmitting measured physiological information and analysis results to a server or other device.

[0604] A "data matching means" is a device or program that, when an abnormality in speech or physiological information is detected, matches that information and performs further analysis.

[0605] A "history management system" is a system for recording and managing past speech and physiological information.

[0606] This invention is a system that comprehensively monitors the physical and emotional health status of elderly individuals and enables early detection of abnormalities. The system primarily consists of terminals, a server, and users.

[0607] First, the device is a bracelet-type device worn by elderly individuals. The device incorporates a voice acquisition mechanism for real-time voice recording, as well as biometric measurement devices for measuring heart rate and body temperature. The acquired data is transmitted to a server using a secure communication protocol. This involves data transmission via Bluetooth and an internet connection.

[0608] Next, the server implements a commonly available speech recognition engine (e.g., a speech recognition API) as a means of speech recognition, which converts the speech data sent from the terminal into text data. The converted text data is analyzed by an emotion analysis means to recognize the emotional state from the elderly person's speech. A generative AI model is used for this, which learns the changes in the elderly person's emotions and is capable of self-improvement.

[0609] The server then integrates emotional states and biometric data using anomaly detection mechanisms to monitor the health status of elderly individuals. If an anomaly is detected, a notification generation mechanism generates a warning, which is sent to the user. The warning includes specific details of the anomaly, allowing the user to take prompt action.

[0610] Finally, users can receive alerts via their smartphones and check on the elderly person's condition. The smartphone app includes a history management system that tracks past speech and physiological data, allowing users to precisely track changes in their health status.

[0611] For example, if an elderly person says, "I'm feeling down today and don't want to do anything," the emotion analysis system will recognize this as anxiety or sadness and report it to the server as an abnormality. Furthermore, if a higher-than-normal heart rate is detected, a notification will be generated immediately. An example of a prompt message would be, "Please identify which utterance caused the emotional abnormality and decide which medical institution you should consult."

[0612] In this way, the system provides comprehensive support for the physical and emotional health of the elderly and offers advanced data analysis capabilities for rapid response in emergencies.

[0613] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0614] Step 1:

[0615] The device acquires the voice of the elderly person wearing it in real time and simultaneously measures biometric information such as heart rate and body temperature. The input obtained in this process consists of raw voice data and physiological data. This data is collected via a voice acquisition device and transmitted to a server via a data transmission device. The resulting output is voice data and physiological data that are ready for further processing on the server side.

[0616] Step 2:

[0617] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. The input is raw audio data, and the output is the corresponding text data. Specifically, speech-to-text conversion is performed using a speech recognition API or similar method. Through this process, the spoken content is obtained in text format and sent to the next analysis step.

[0618] Step 3:

[0619] The server inputs the converted text data into an emotion analysis device, which then uses a generative AI model to analyze the emotions. The input is text data, and the output is the analyzed emotion data. Specifically, the generative AI model extracts emotional nuances from the text content and evaluates the emotional state of the elderly person. This information prepares the device for anomaly detection.

[0620] Step 4:

[0621] The server integrates emotional and physiological data using anomaly detection mechanisms to monitor the health status of elderly individuals. Inputs are emotional and physiological data, and output is information on whether or not an anomaly occurred. Specifically, the system compares the data to dynamically set baseline values ​​and patterns; an anomaly is detected if the data exceeds the normal range.

[0622] Step 5:

[0623] The server generates a warning to the user using a notification generation mechanism when an anomaly is detected. The input is the anomaly detection result, and the output is a warning message to the user. This message includes detailed information about the health status of the elderly person, and the user can receive the notification through the app.

[0624] Step 6:

[0625] Users receive notifications via a smartphone app and check the health status of elderly individuals through the displayed notifications. The input is the notification message, and the output is information to help the user understand the situation. Based on this information, users can contact medical institutions or perform additional monitoring as needed.

[0626] Through these steps, this system provides comprehensive support for the physical and emotional health of older adults, enabling prompt and appropriate responses.

[0627] (Application Example 2)

[0628] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0629] In today's living environments for the elderly, not only physical health issues but also emotional health issues are becoming increasingly important. However, conventional technology has struggled to effectively detect emotional changes and integrate them with physical information for comprehensive health management. Therefore, there is a need for a system that enables rapid and appropriate responses when elderly individuals exhibit emotional or physical abnormalities.

[0630] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0631] In this invention, the server includes information acquisition means for recording speech information to detect abnormalities in emotional state, conversion means for converting it into text data, and emotion analysis means for analyzing emotions. This makes it possible to monitor the emotional and physical health of elderly people and to promptly notify them if an abnormality is detected.

[0632] "Information acquisition means" refers to a device or system for recording speech information from elderly people.

[0633] "Conversion means" refers to a device or system that has the function of converting recorded speech information into text data.

[0634] "Analysis means" refers to a device or system that has the function of analyzing text data to learn speech patterns.

[0635] An "anomaly detection means" is a device or system that has the function of detecting abnormalities in speech or emotion based on learned patterns.

[0636] "Notification generation means" refers to a device or system that has the function of generating notifications to relevant parties based on detected anomalies.

[0637] "Emotion analysis means" refers to a device or system that has the function of analyzing emotions from speech information.

[0638] "Integration means" refers to a device or system that has the function of integrating the results of emotion analysis means and anomaly detection means to detect abnormalities in emotional states.

[0639] "Information measurement means" refers to a device or system for measuring bodily information.

[0640] "Information transmission means" refers to a device or system for recording and transmitting measured physical information.

[0641] "Data matching means" refers to a device or system that has the function of matching physical information measured when an anomaly is detected with the actual data.

[0642] "History management means" refers to a device or system that has the function of recording and managing past history based on recorded speech information and physical information.

[0643] This invention comprises a system for comprehensively monitoring the emotional and physical health status of elderly individuals. The main components are information acquisition means, conversion means, analysis means, anomaly detection means, notification generation means, emotion analysis means, integration means, and information measurement means.

[0644] Each elderly person will be provided with a multi-functional bracelet device that records voice and physical information. This device is equipped with a microphone, heart rate and body temperature sensors, and digitally acquires daily speech information and physical condition data.

[0645] The server receives acquired speech information in real time and converts it into text data using speech recognition software. The Python `speech_recognition` library is used for this process. This text data undergoes pattern learning using an analysis tool, and anomalies in emotion and speech patterns are detected by an anomaly detection tool and an emotion analysis tool. A machine learning model from scikit-learn is used for emotion analysis. The results are then comprehensively evaluated by an integration tool to determine the presence or absence of anomalies.

[0646] Users receive relevant notifications via their smartphones or smart glasses. These notifications include information about the specific nature of detected anomalies and changes in emotions. This allows users to quickly and appropriately manage the safety and health of the elderly.

[0647] For example, if an elderly person makes many negative statements while walking and their heart rate is simultaneously elevated, the emotion analysis system will recognize this as a stressful state. The server can immediately generate an alert and send a notification to the user stating, "The elderly person is experiencing stress. Their heart rate is also elevated," prompting quick intervention.

[0648] The generative AI model used includes an example of supplying the following prompt: "An elderly person showed emotional changes during a routine activity. Analyze the audio and use the emotion engine to detect anomalies."

[0649] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0650] Step 1:

[0651] The device acquires speech and physical information from elderly individuals using a multi-functional bracelet device. Inputs include voice data from a microphone, heart rate, and data from a body temperature sensor. Outputs include sending the acquired voice and physical data in digital format to a server.

[0652] Step 2:

[0653] The server converts the received audio data into text data using the Python speech_recognition library. At this stage, the audio data is the input, and the speech recognition engine analyzes it to generate text information as output.

[0654] Step 3:

[0655] The server analyzes the converted text data and learns speech patterns. Text data is the input, and important speech patterns are extracted using the analysis method. The learned speech patterns are generated as output.

[0656] Step 4:

[0657] The server analyzes emotions from text data using emotion analysis tools. Here, a machine learning model from scikit-learn is applied to the text data as input, and data indicating the emotional state is obtained as output.

[0658] Step 5:

[0659] The server evaluates speech patterns and sentiment analysis results using an integration mechanism to detect anomalies. In this step, speech patterns and sentiment data are integrated as input, a determination is made as to whether there are any unnatural changes or anomalies, and the anomaly detection result is obtained as output.

[0660] Step 6:

[0661] When the server detects an anomaly, it uses a notification generation mechanism to create an alarm. Here, the anomaly detection result is taken as input, and an alarm message for the user is generated as output.

[0662] Step 7:

[0663] Users receive notifications sent from a server via their smartphones or smart glasses. These notifications include an alert message as input and output that allows the user to understand the situation and prepare appropriate action.

[0664] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0665] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0666] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0667] [Fourth Embodiment]

[0668] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0669] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0670] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0671] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0672] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0673] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0674] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0675] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0676] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0677] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0678] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0679] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0680] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0681] This invention is a monitoring system that uses speech and physical information of elderly people, and by combining multiple functions, it enables real-time monitoring of the safety and health status of elderly people.

[0682] The device consists of a bracelet-type device worn on the wrist of an elderly person. This device is equipped with a microphone for recording voice and has the function of recording everyday conversations. It also has built-in sensors for measuring physical information such as heart rate and body temperature. This data is transmitted to a server via wireless communication.

[0683] When the server receives voice data transmitted from the terminal, it first converts it into text data using a speech recognition engine. The converted text data is then analyzed by a generative AI to learn the speech patterns of elderly people. The generative AI focuses on the characteristics of the speech content, the frequency of words used, and the flow of context to detect abnormal patterns. If an abnormality is detected, a notification is sent to pre-configured contacts. The server also simultaneously receives physical information transmitted from the terminal and stores the measurement results in a database.

[0684] Users receive and view information transmitted from the server via their smartphones. The application allows users to check logs of the elderly person's health status and receive abnormality notifications. It also includes features to allow users to investigate the nature of the abnormality in more detail or to consult with healthcare professionals, if necessary.

[0685] As a concrete example of this system, suppose an elderly person living alone starts making more frequent remarks about forgetting things during everyday conversations. The server detects this change in speech pattern as an anomaly and sends a notification to the family member (the user). Upon receiving this notification, the family can take appropriate action, such as seeking early medical advice. In this way, the present invention supports the health and safety of the elderly.

[0686] The following describes the processing flow.

[0687] Step 1:

[0688] The device records the elderly person's voice in real time. This is done using a built-in microphone, and the voice data is saved to a buffer at regular intervals.

[0689] Step 2:

[0690] The terminal transmits voice data to the server via wireless communication. The transmission is triggered by pre-configured conditions, enabling efficient data transmission.

[0691] Step 3:

[0692] The server analyzes the received audio data and converts it into text data using a speech recognition engine. During this process, pre-processing such as noise reduction is performed to improve recognition accuracy.

[0693] Step 4:

[0694] The server passes text data through a generative AI and compares it to past speech patterns. The generative AI uses a pre-trained model to detect changes in the text data, identifying new patterns and unusually frequent utterances.

[0695] Step 5:

[0696] When the server detects an anomaly, it sends an alert to the configured notification recipient. The alert will include a description of the anomaly and recommended actions.

[0697] Step 6:

[0698] The device measures the physical information of elderly individuals. This information includes body temperature and heart rate, and this data is sent to a server.

[0699] Step 7:

[0700] The server analyzes the physical information and compares it with past data to check for any abnormalities. If no abnormalities are found, it makes a comprehensive judgment based on this information along with the voice data.

[0701] Step 8:

[0702] Users receive notifications and data from the server via a smartphone app. Based on the received information, the system monitors the elderly person's condition and considers necessary actions.

[0703] Step 9:

[0704] Users can check their history of abnormalities and health data within the app. This allows for timely medical consultations and appropriate actions. This information is continuously logged.

[0705] (Example 1)

[0706] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0707] Effectively monitoring the safety and health of the elderly is a critical challenge in a growing elderly population. With conventional technologies, it was difficult for those around the elderly to understand their condition unless they themselves were aware of an abnormality. Furthermore, detecting abnormalities based not only on physical information but also on speech patterns was challenging, potentially leading to delays in necessary responses.

[0708] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0709] In this invention, the server includes voice collection means, voice recognition means, generative model learning means, anomaly detection means, notification generation means, physical condition measurement means, and physical information transmission means. This enables comprehensive state monitoring using both the speech and physical information of elderly individuals, allowing for early detection and notification of anomalies.

[0710] "Voice collection means" refers to a device or function that records the everyday conversations of elderly people as audio.

[0711] "Speech recognition means" refers to a device or function that converts recorded speech data into text data.

[0712] "Generative model learning means" refers to a device or program that analyzes converted text data and has the function of learning the speech patterns of elderly people.

[0713] An "anomaly detection means" refers to a device or function that detects speech or behavior that is different from the norm, based on learned speech patterns.

[0714] "Notification generation means" refers to a device or function that generates and sends messages or alerts to designated recipients when an anomaly is detected.

[0715] "Means of measuring physical condition" refers to sensors and devices that measure physical information such as heart rate and body temperature of elderly people.

[0716] "Body information transmission means" refers to devices or protocols that have the function of recording measured body information and transmitting it to remote devices such as servers.

[0717] This invention relates to a system for monitoring the health status of elderly individuals using their speech and physical information. This system is implemented using a server, a terminal, and a user component.

[0718] Data collection by devices:

[0719] The device consists of a bracelet-type device worn on the wrist of an elderly person. As a means of voice collection, the device is equipped with a microphone to record the elderly person's daily conversations as audio data. It also has a built-in heart rate monitor and body temperature sensor as means of measuring physical condition, and this physical information is measured periodically. The collected data is transmitted to a server via Bluetooth or Wi-Fi communication using a physical information transmission device.

[0720] Server-based data analysis:

[0721] The server converts speech data into text data using speech recognition means. This transforms the speech data into an analyzable format. Using generative model learning means, the speech patterns of elderly people are analyzed and learned from the text data. Anomaly detection means can detect anomalies based on the learned patterns. If an anomaly is detected, a notification is sent to a pre-configured contact (e.g., family) using notification generation means. These notifications may include historical data.

[0722] User verification of information:

[0723] Users receive health status data and abnormality notifications from a server via communication devices such as smartphones. Through a dedicated application, they can review past history and investigate the details of abnormalities. The application also provides information to consult with healthcare professionals as needed.

[0724] Specific example:

[0725] As a concrete example, if an elderly person living alone frequently says "I've been forgetting things a lot lately" in conversation, the server will detect this change in speech pattern as an anomaly. Subsequently, a notification is sent to the family member (the user), who can then consider seeking medical advice early.

[0726] Example of a prompt:

[0727] "Detect abnormalities in the speech patterns of elderly individuals. As a specific example, analyze changes in frequently used phrases."

[0728] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0729] Step 1:

[0730] The device records the daily conversations of elderly people as audio data. A microphone is used as the means of audio collection. The input is spoken voice, and the output is an audio data file. Specifically, the device captures audio at regular time intervals and saves it as a data file.

[0731] Step 2:

[0732] The device measures bodily information such as heart rate and body temperature. It acquires data from heart rate sensors and body temperature sensors using a bodily state measurement device. The input is the body's state, and the output is heart rate and body temperature data. Specifically, the sensors read data from the body and record it as numerical data in storage.

[0733] Step 3:

[0734] The device transmits collected voice data and physical information to a server via Bluetooth or Wi-Fi. Here, the recorded data file is the input, and the data packets sent to the server are the output. Specifically, the device periodically transmits data using a communication module.

[0735] Step 4:

[0736] The server converts received audio data into text data using speech recognition technology. The input is an audio data file, and the output is the spoken content in text format. Specifically, the speech recognition algorithm analyzes the audio and converts it into text.

[0737] Step 5:

[0738] The server analyzes text data using a generative model learning method to learn the speech patterns of elderly people. The input is text data, and the output is the learned speech patterns. Specifically, the generative AI analyzes the structure and content of the text and records the frequency of words and the flow of context.

[0739] Step 6:

[0740] The server uses anomaly detection means to detect anomalies based on learned speech patterns. The input is the learned speech pattern, and the output is information about the detected anomaly. Specifically, it compares existing patterns with new data and identifies changes that deviate from the standard.

[0741] Step 7:

[0742] The server uses a notification generation mechanism to create a notification when an anomaly is detected and sends it to a predetermined contact. The input is information about the detected anomaly, and the output is the anomaly notification. Specifically, it generates a message containing details of the anomaly and sends the notification via the communication network.

[0743] Step 8:

[0744] The user receives and displays notifications from the server using a smartphone app. The input is an anomaly notification from the server, and the output is the notification information displayed on the screen. Specifically, the user can open the app, check the notification content, and look up further details if necessary.

[0745] (Application Example 1)

[0746] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0747] In technologies that monitor the safety and health of the elderly using speech and biometric information, conventional systems only focus on detecting anomalies in speech and biometric information, making it difficult to detect external safety threats from ambient noise. Furthermore, there was a lack of means to comprehensively manage and monitor this information in real time. Therefore, there is a need to further improve the safety and convenience of life for the elderly.

[0748] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0749] In this invention, the server includes a device for recording speech information of elderly people, a conversion device for converting the recorded speech information into string data, a learning device for analyzing the string data and learning speech patterns, a detection device for detecting abnormalities based on the learned patterns, a notification device for generating notifications based on the detected abnormalities, and an environmental analysis device for detecting safety abnormalities by analyzing ambient sounds. This makes it possible to detect not only speech and biological abnormalities of elderly people, but also safety threats in the surrounding environment, thereby realizing more comprehensive monitoring.

[0750] The term "elderly person" refers to an individual for whom health management and safety become particularly important as they age.

[0751] "Speech information" refers to the linguistic data contained in the voices of elderly people, including the content of conversations and monologues.

[0752] "String data" refers to text-formatted data converted from audio data.

[0753] "Biometric information" refers to physical data, including heart rate, body temperature, and activity level, for elderly individuals.

[0754] The term "device" refers to an instrument or device with a specific function, used for various data collection and processing within a system.

[0755] A "conversion device" refers to a device that has the function of converting audio data into text data.

[0756] A "learning device" refers to a device that has the function of continuously analyzing data and identifying and memorizing features and patterns.

[0757] A "detection device" refers to a device that has the function of identifying and reporting anomalies based on learned data.

[0758] A "notification device" refers to a device that has the function of informing relevant parties about detected anomalies.

[0759] An "environmental analysis device" refers to a device that monitors ambient sounds and environmental data and has the function of detecting abnormalities.

[0760] The system that realizes this invention is designed to monitor the daily safety and health of elderly people. The system mainly consists of family members or caregivers as users, terminals worn by the elderly, and a server responsible for data processing and notifications. Its specific functions and operation are described below.

[0761] The device is worn on the wrist of elderly individuals and contains a microphone and multiple sensors (for measuring heart rate and body temperature). This allows for the continuous collection of speech and biometric information from the elderly person. This data is transmitted to a server via wireless communication.

[0762] The server first converts the audio data sent from the terminal into text data using speech recognition software. Specifically, Microsoft Azure Speech Services performs this role. This text data is then analyzed by a generative AI model. For example, it learns past speech patterns and detects abnormal speech or unstable contextual changes. By giving the generative AI model prompts such as, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of abnormalities based on that," the model performs the appropriate analysis.

[0763] Simultaneously, biometric information such as heart rate and body temperature is analyzed by the server. The IBM Watson Health platform is used for real-time data analysis. In addition, ambient sounds are analyzed using Google Cloud's voice analysis service to detect abnormal sounds occurring around the elderly person. In this way, the system provides comprehensive monitoring.

[0764] Users receive information sent from the server through a smartphone application. If an anomaly is detected, a notification is sent to the user via Firebase Cloud Messaging. Users can check the details of the anomaly within the app and obtain information to consult with a medical professional if necessary. For example, a notification such as "Your mother may be frequently mistaking place names" may be sent, allowing the user to take early action. In this way, the entire system works together to build a mechanism that supports the lives of the elderly.

[0765] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0766] Step 1:

[0767] The device collects speech and biometric information from elderly individuals. It records voice data using a microphone and measures heart rate and body temperature using built-in sensors. The collected data is transmitted wirelessly to a server. The input is analog signal data recorded by the microphone and sensors, which is then converted to digital data and sent to the server.

[0768] Step 2:

[0769] The server processes the audio data sent from the terminal using speech recognition software and converts it into text data. Azure Speech Services is used for this process. The input is audio data, and the output is text-based string data. The audio signal is analyzed, and a language model is used to generate the correct word sequence.

[0770] Step 3:

[0771] The server analyzes the converted string data using a generative AI model. Specifically, it learns speech patterns using prompt sentences and performs anomaly detection. The generative AI utilizes the prompt, "Analyze health abnormalities based on the speech recording text of an elderly person. List new words that appear and words that disappear, and evaluate the likelihood of anomalies based on these." The input is string data, and the output is the analysis result regarding the presence or absence of anomalies.

[0772] Step 4:

[0773] The server also analyzes biometric information. Using IBM Watson Health, it analyzes heart rate and body temperature and detects abnormalities by comparing them to common health indicators. The input is biometric data collected by sensors, and the output is a judgment of whether it is normal or abnormal.

[0774] Step 5:

[0775] Simultaneously, the server uses Google Cloud's speech analysis service to analyze ambient sounds and detect anomalies. The input is ambient sound data, and the output is the result of safety-related anomaly detection. It identifies specific acoustic patterns and generates a warning if they are deemed anomaly.

[0776] Step 6:

[0777] The server generates and sends a notification to the user based on the detected anomaly. The notification is sent to smartphones and other devices via Firebase Cloud Messaging. The input is the anomaly detection result, and the output is an informational notification to the user. It includes specific details, allowing the user to receive it and take appropriate action.

[0778] Step 7:

[0779] Users view received notifications through a smartphone application. These notifications include detailed information about the anomaly and recommended next actions. Within the app, users can review past history and an overview of their overall health status, and contact healthcare professionals if necessary. The input is the notification message, and the output is the user's confirmation of the situation and subsequent actions.

[0780] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0781] This invention is a system that combines the acquisition of speech and physical information from elderly individuals with an emotion engine that analyzes emotions from their speech. This system comprehensively monitors the physical and emotional health status of elderly individuals, enabling early detection of abnormalities and appropriate responses.

[0782] The device is a bracelet-type device worn by elderly individuals and is equipped with a microphone for voice recognition and sensors to measure heart rate and body temperature. The device records voice in real time and transmits the acquired data to a server sequentially.

[0783] The server receives voice data transmitted from the terminal and performs a series of processes to convert it into text using a speech recognition engine. The converted text data is analyzed by a generative AI and sent to an emotion engine that recognizes emotions from the speech content along with the speech patterns. The emotion engine has a self-learning function and analyzes and stores the emotional states of elderly people from their daily speech to detect abnormal emotional changes.

[0784] Furthermore, the server integrates speech pattern learning and emotion analysis results from artificial intelligence, and performs anomaly detection by comparing this with health data. When an anomaly is detected, an alert is generated for the relevant parties. The alert includes details about the specific nature of the anomaly and what kind of emotional changes are occurring.

[0785] Users can receive this notification through a smartphone app and check the status of the elderly person. The historical data included in the notification serves as a guide for users to understand past health conditions and emotional tendencies. It also provides useful information when considering appropriate measures and medical support.

[0786] As a concrete example, consider a scenario where an elderly person's daily conversations contain more negative words than usual. The emotion engine recognizes this change as a shift from joy to anxiety and reports it as an anomaly to the server. If the server detects an increase in heart rate, it also considers it as a physical health problem and prompts the user to take prompt action. Thus, the present invention is a system that comprehensively supports the physical and mental health of the elderly through complex data analysis.

[0787] The following describes the processing flow.

[0788] Step 1:

[0789] The device records the elderly person's voice in real time. It acquires voice data using a built-in microphone and stores it in a buffer in an appropriate format.

[0790] Step 2:

[0791] The device transmits recorded audio data to the server at regular intervals. Wireless communication technology is used to ensure efficient data transmission.

[0792] Step 3:

[0793] The server passes the received audio data to the speech recognition engine, which converts it into text data. During the conversion process, noise filtering and other techniques are performed to improve accuracy.

[0794] Step 4:

[0795] The server inputs the generated text data into the generating AI, which then compares it to past speech patterns. During this process, the trained algorithm analyzes any changes in the speech content.

[0796] Step 5:

[0797] The server sends text data to the emotion engine, which analyzes the emotional state from the utterance. The emotion engine extracts specific emotional elements and determines the emotional tendency.

[0798] Step 6:

[0799] The server integrates the emotion analysis results and the speech pattern analysis results. From the integrated data, it detects abnormal behavior and emotional changes and prepares to generate alerts.

[0800] Step 7:

[0801] The device measures physical information such as the heart rate and body temperature of elderly individuals and transmits this data to a server. Measurements are performed regularly to maintain real-time accuracy.

[0802] Step 8:

[0803] The server comprehensively compares physical information with speech and emotion analysis data. This allows for multifaceted verification of anomalies and reduces the risk of misrecognition.

[0804] Step 9:

[0805] Users receive alert notifications via their smartphones. The notifications detail the specific nature of the anomaly, changes in emotional state, and possible underlying physical problems.

[0806] Step 10:

[0807] Users can review the information received through the app and take appropriate action regarding the elderly person's situation as needed. Past historical data can also be referenced to support more accurate decision-making.

[0808] (Example 2)

[0809] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0810] In today's aging society, there is a need for a system that can accurately monitor the physical and mental health of the elderly and detect abnormalities early. Conventional technologies have the problem that health monitoring of the elderly is one-sided, and it is difficult to make a comprehensive judgment of their physical and mental state, making it highly likely that abnormalities will be overlooked.

[0811] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0812] In this invention, the server includes voice acquisition means, emotion analysis means, and anomaly detection means. This enables comprehensive monitoring of the health status of elderly people based on their speech content and physiological information, as well as rapid anomaly detection.

[0813] "Voice acquisition means" refers to a device for collecting speech information from elderly people in real time, and mainly includes voice input devices such as microphones.

[0814] "Speech recognition means" refers to technology for converting acquired speech data into text data, and this conversion is performed using a speech recognition engine or AI technology.

[0815] "Emotional analysis methods" refer to analytical techniques for recognizing emotional states from text data, and utilize generative AI models to grasp the emotions contained in spoken language.

[0816] An "anomaly detection means" is a device or program that integrates acquired emotional data and health data to detect anomalies.

[0817] A "notification generation means" is a device or system that generates warnings or notifications based on detected abnormal information and transmits them to the user.

[0818] A "biometric measurement device" is a device used to measure physiological information such as heart rate and body temperature in elderly individuals, and utilizes various sensors.

[0819] "Data transmission means" refers to communication means intended for transmitting measured physiological information and analysis results to a server or other device.

[0820] A "data matching means" is a device or program that, when an abnormality in speech or physiological information is detected, matches that information and performs further analysis.

[0821] A "history management system" is a system for recording and managing past speech and physiological information.

[0822] This invention is a system that comprehensively monitors the physical and emotional health status of elderly individuals and enables early detection of abnormalities. The system primarily consists of terminals, a server, and users.

[0823] First, the device is a bracelet-type device worn by elderly individuals. The device incorporates a voice acquisition mechanism for real-time voice recording, as well as biometric measurement devices for measuring heart rate and body temperature. The acquired data is transmitted to a server using a secure communication protocol. This involves data transmission via Bluetooth and an internet connection.

[0824] Next, the server implements a commonly available speech recognition engine (e.g., a speech recognition API) as a means of speech recognition, which converts the speech data sent from the terminal into text data. The converted text data is analyzed by an emotion analysis means to recognize the emotional state from the elderly person's speech. A generative AI model is used for this, which learns the changes in the elderly person's emotions and is capable of self-improvement.

[0825] The server then integrates emotional states and biometric data using anomaly detection mechanisms to monitor the health status of elderly individuals. If an anomaly is detected, a notification generation mechanism generates a warning, which is sent to the user. The warning includes specific details of the anomaly, allowing the user to take prompt action.

[0826] Finally, users can receive alerts via their smartphones and check on the elderly person's condition. The smartphone app includes a history management system that tracks past speech and physiological data, allowing users to precisely track changes in their health status.

[0827] For example, if an elderly person says, "I'm feeling down today and don't want to do anything," the emotion analysis system will recognize this as anxiety or sadness and report it to the server as an abnormality. Furthermore, if a higher-than-normal heart rate is detected, a notification will be generated immediately. An example of a prompt message would be, "Please identify which utterance caused the emotional abnormality and decide which medical institution you should consult."

[0828] In this way, the system provides comprehensive support for the physical and emotional health of the elderly and offers advanced data analysis capabilities for rapid response in emergencies.

[0829] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0830] Step 1:

[0831] The device acquires the voice of the elderly person wearing it in real time and simultaneously measures biometric information such as heart rate and body temperature. The input obtained in this process consists of raw voice data and physiological data. This data is collected via a voice acquisition device and transmitted to a server via a data transmission device. The resulting output is voice data and physiological data that are ready for further processing on the server side.

[0832] Step 2:

[0833] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. The input is raw audio data, and the output is the corresponding text data. Specifically, speech-to-text conversion is performed using a speech recognition API or similar method. Through this process, the spoken content is obtained in text format and sent to the next analysis step.

[0834] Step 3:

[0835] The server inputs the converted text data into an emotion analysis device, which then uses a generative AI model to analyze the emotions. The input is text data, and the output is the analyzed emotion data. Specifically, the generative AI model extracts emotional nuances from the text content and evaluates the emotional state of the elderly person. This information prepares the device for anomaly detection.

[0836] Step 4:

[0837] The server integrates emotional and physiological data using anomaly detection mechanisms to monitor the health status of elderly individuals. Inputs are emotional and physiological data, and output is information on whether or not an anomaly occurred. Specifically, the system compares the data to dynamically set baseline values ​​and patterns; an anomaly is detected if the data exceeds the normal range.

[0838] Step 5:

[0839] The server generates a warning to the user using a notification generation mechanism when an anomaly is detected. The input is the anomaly detection result, and the output is a warning message to the user. This message includes detailed information about the health status of the elderly person, and the user can receive the notification through the app.

[0840] Step 6:

[0841] Users receive notifications via a smartphone app and check the health status of elderly individuals through the displayed notifications. The input is the notification message, and the output is information to help the user understand the situation. Based on this information, users can contact medical institutions or perform additional monitoring as needed.

[0842] Through these steps, this system provides comprehensive support for the physical and emotional health of older adults, enabling prompt and appropriate responses.

[0843] (Application Example 2)

[0844] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0845] In today's living environments for the elderly, not only physical health issues but also emotional health issues are becoming increasingly important. However, conventional technology has struggled to effectively detect emotional changes and integrate them with physical information for comprehensive health management. Therefore, there is a need for a system that enables rapid and appropriate responses when elderly individuals exhibit emotional or physical abnormalities.

[0846] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0847] In this invention, the server includes information acquisition means for recording speech information to detect abnormalities in emotional state, conversion means for converting it into text data, and emotion analysis means for analyzing emotions. This makes it possible to monitor the emotional and physical health of elderly people and to promptly notify them if an abnormality is detected.

[0848] "Information acquisition means" refers to a device or system for recording speech information from elderly people.

[0849] "Conversion means" refers to a device or system that has the function of converting recorded speech information into text data.

[0850] "Analysis means" refers to a device or system that has the function of analyzing text data to learn speech patterns.

[0851] An "anomaly detection means" is a device or system that has the function of detecting abnormalities in speech or emotion based on learned patterns.

[0852] "Notification generation means" refers to a device or system that has the function of generating notifications to relevant parties based on detected anomalies.

[0853] "Emotion analysis means" refers to a device or system that has the function of analyzing emotions from speech information.

[0854] "Integration means" refers to a device or system that has the function of integrating the results of emotion analysis means and anomaly detection means to detect abnormalities in emotional states.

[0855] "Information measurement means" refers to a device or system for measuring bodily information.

[0856] "Information transmission means" refers to a device or system for recording and transmitting measured physical information.

[0857] "Data matching means" refers to a device or system that has the function of matching physical information measured when an anomaly is detected with the actual data.

[0858] "History management means" refers to a device or system that has the function of recording and managing past history based on recorded speech information and physical information.

[0859] This invention comprises a system for comprehensively monitoring the emotional and physical health status of elderly individuals. The main components are information acquisition means, conversion means, analysis means, anomaly detection means, notification generation means, emotion analysis means, integration means, and information measurement means.

[0860] Each elderly person will be provided with a multi-functional bracelet device that records voice and physical information. This device is equipped with a microphone, heart rate and body temperature sensors, and digitally acquires daily speech information and physical condition data.

[0861] The server receives acquired speech information in real time and converts it into text data using speech recognition software. The Python `speech_recognition` library is used for this process. This text data undergoes pattern learning using an analysis tool, and anomalies in emotion and speech patterns are detected by an anomaly detection tool and an emotion analysis tool. A machine learning model from scikit-learn is used for emotion analysis. The results are then comprehensively evaluated by an integration tool to determine the presence or absence of anomalies.

[0862] Users receive relevant notifications via their smartphones or smart glasses. These notifications include information about the specific nature of detected anomalies and changes in emotions. This allows users to quickly and appropriately manage the safety and health of the elderly.

[0863] For example, if an elderly person makes many negative statements while walking and their heart rate is simultaneously elevated, the emotion analysis system will recognize this as a stressful state. The server can immediately generate an alert and send a notification to the user stating, "The elderly person is experiencing stress. Their heart rate is also elevated," prompting quick intervention.

[0864] The generative AI model used includes an example of supplying the following prompt: "An elderly person showed emotional changes during a routine activity. Analyze the audio and use the emotion engine to detect anomalies."

[0865] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0866] Step 1:

[0867] The device acquires speech and physical information from elderly individuals using a multi-functional bracelet device. Inputs include voice data from a microphone, heart rate, and data from a body temperature sensor. Outputs include sending the acquired voice and physical data in digital format to a server.

[0868] Step 2:

[0869] The server converts the received audio data into text data using the Python speech_recognition library. At this stage, the audio data is the input, and the speech recognition engine analyzes it to generate text information as output.

[0870] Step 3:

[0871] The server analyzes the converted text data and learns speech patterns. Text data is the input, and important speech patterns are extracted using the analysis method. The learned speech patterns are generated as output.

[0872] Step 4:

[0873] The server analyzes emotions from text data using emotion analysis tools. Here, a machine learning model from scikit-learn is applied to the text data as input, and data indicating the emotional state is obtained as output.

[0874] Step 5:

[0875] The server evaluates speech patterns and sentiment analysis results using an integration mechanism to detect anomalies. In this step, speech patterns and sentiment data are integrated as input, a determination is made as to whether there are any unnatural changes or anomalies, and the anomaly detection result is obtained as output.

[0876] Step 6:

[0877] When the server detects an anomaly, it uses a notification generation mechanism to create an alarm. Here, the anomaly detection result is taken as input, and an alarm message for the user is generated as output.

[0878] Step 7:

[0879] Users receive notifications sent from a server via their smartphones or smart glasses. These notifications include an alert message as input and output that allows the user to understand the situation and prepare appropriate action.

[0880] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0881] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0882] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0883] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0884] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0885] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0886] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0887] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0888] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0889] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0890] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0891] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0892] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0893] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0894] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0895] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0896] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0897] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0898] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0899] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0900] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0901] The following is further disclosed regarding the embodiments described above.

[0902] (Claim 1)

[0903] A recording means for recording speech information of elderly people,

[0904] A conversion means for converting recorded speech information into text data,

[0905] A learning method that analyzes text data and learns speech patterns,

[0906] A detection means for detecting anomalies based on learned patterns,

[0907] A notification means that generates a notification based on the detected anomaly,

[0908] A system that includes this.

[0909] (Claim 2)

[0910] Measurement methods for measuring the physical information of elderly people,

[0911] A means for recording and transmitting measured physical information,

[0912] When an anomaly is detected by the aforementioned detection means, a matching means is provided to compare it with the measured information.

[0913] The system according to claim 1, including the following:

[0914] (Claim 3)

[0915] It is equipped with a history management means that records and manages past history based on recorded speech information and physical information.

[0916] The system according to claim 1, which has a function to include historical data in notifications by the notification means.

[0917] "Example 1"

[0918] (Claim 1)

[0919] A voice collection method for recording speech information of elderly people,

[0920] A speech recognition means that converts information recorded by a speech collection means into text data,

[0921] A generative model learning method that analyzes text data and learns the speech patterns of elderly people,

[0922] An anomaly detection means that detects anomalies based on speech patterns learned by a generative model learning means,

[0923] A notification generation means generates a notification based on an anomaly detected by an anomaly detection means and transmits it to a predetermined communication destination,

[0924] A means for measuring the physical condition of elderly people,

[0925] A means for transmitting body information that records and transmits measured body information,

[0926] A system that includes this.

[0927] (Claim 2)

[0928] The system according to claim 1, which includes a matching means for comparing information measured by a physical condition measuring means with information detected by an anomaly detection means.

[0929] (Claim 3)

[0930] The system according to claim 1, comprising a history management means for recording and managing past history based on recorded speech information and physical information, and having a function to include history data in notifications from a notification generation means.

[0931] "Application Example 1"

[0932] (Claim 1)

[0933] A device for recording speech information from elderly people,

[0934] A conversion device that converts recorded speech information into string data,

[0935] A learning device that analyzes string data and learns speech patterns,

[0936] A detection device that detects anomalies based on learned patterns,

[0937] A notification device that generates notifications based on detected anomalies,

[0938] A system including an environmental analysis device that analyzes ambient sounds to detect safety anomalies.

[0939] (Claim 2)

[0940] A measuring device for measuring the vital signs of elderly people,

[0941] A device that records and transmits measured biological information,

[0942] When an anomaly is detected by the aforementioned detection device, a verification device is used to confirm consistency with the measured information.

[0943] The system according to claim 1, including the following:

[0944] (Claim 3)

[0945] Equipped with a history management device that records and manages past history based on recorded speech information and biometric information,

[0946] The system according to claim 1, which has a function to include historical data in notifications from the notification device.

[0947] "Example 2 of combining an emotion engine"

[0948] (Claim 1)

[0949] A voice acquisition method for recording speech information of elderly people,

[0950] A speech recognition means that converts recorded speech information into text data,

[0951] A means of sentiment analysis that analyzes text data and recognizes emotional states,

[0952] An anomaly detection method that integrates generated emotional data and health data to detect abnormalities,

[0953] A notification generation means that generates a warning based on the detected anomaly,

[0954] A system that includes this.

[0955] (Claim 2)

[0956] A biometric measurement device for measuring physiological information in elderly people,

[0957] A data transmission means for recording and transmitting measured physiological information,

[0958] When an anomaly is detected by the anomaly detection means, a data matching means is provided to compare it with the measured physiological information.

[0959] The system according to claim 1, including the following:

[0960] (Claim 3)

[0961] It is equipped with a history management means that manages past history based on recorded speech information and physiological information.

[0962] The system according to claim 1, which has a function to include historical data in the warnings issued by the notification generation means.

[0963] "Application example 2 when combining with an emotional engine"

[0964] (Claim 1)

[0965] A means of acquiring information to record speech information of elderly people,

[0966] A conversion means for converting recorded speech information into text data,

[0967] An analysis method that analyzes text data and learns speech patterns,

[0968] An anomaly detection means that detects anomalies based on learned patterns,

[0969] A notification generation means that generates a notification based on the detected anomaly,

[0970] A means of analyzing emotions from speech information,

[0971] To detect abnormalities in emotional states, an integration means is provided for integrating the results of the emotion analysis means and the abnormality detection means,

[0972] A system that includes this.

[0973] (Claim 2)

[0974] Information measurement means for measuring the physical information of the elderly,

[0975] Information transmission means for recording and transmitting measured physical information,

[0976] When an anomaly is detected by the anomaly detection means, a data matching means is provided to compare it with the measured physical information.

[0977] It features a notification method that sends notifications via smart devices.

[0978] The system according to claim 1.

[0979] (Claim 3)

[0980] It is equipped with a history management means that records and manages past history based on recorded speech information and physical information,

[0981] The notification generation means has a function to include historical information in the notification,

[0982] Includes means of issuing warnings to relevant parties in real time.

[0983] The system according to claim 1. [Explanation of Symbols]

[0984] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A recording means for recording speech information of elderly people, A conversion means for converting recorded speech information into text data, A learning method that analyzes text data and learns speech patterns, A detection means for detecting anomalies based on learned patterns, A notification means that generates a notification based on the detected anomaly, A system that includes this.

2. Measurement methods for measuring the physical information of elderly people, A means for recording and transmitting measured physical information, When an anomaly is detected by the aforementioned detection means, a matching means is provided to compare it with the measured information. The system according to claim 1, including the following:

3. It is equipped with a history management means that records and manages past history based on recorded speech information and physical information, The system according to claim 1, which has a function to include historical data in notifications by the notification means.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A