system

A system for remote health monitoring of elderly individuals uses voice recognition and anomaly detection to analyze speech patterns, addressing the challenge of early symptom detection in remote areas, ensuring timely interventions and peace of mind.

JP2026101234APending Publication Date: 2026-06-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-10
Publication Date
2026-06-22

AI Technical Summary

Technical Problem

There is a challenge in routinely monitoring the health status of elderly individuals living in remote areas, particularly in detecting early symptoms like cerebral infarction through changes in language ability, which can lead to serious health risks if overlooked.

Method used

A system that includes voice receiving, voice recognition, natural language processing, and anomaly detection means to analyze speech patterns, comparing them with past evaluations, and sending warnings to caregivers if abnormalities are detected.

Benefits of technology

Enables accurate remote monitoring of elderly individuals' health status by identifying anomalies through everyday conversations, allowing for prompt interventions and providing peace of mind for both the elderly and their families.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026101234000001_ABST
    Figure 2026101234000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A voice input method that receives the user's voice in real time, A speech recognition means for converting the audio into text format, A natural language processing means for analyzing the aforementioned text data and evaluating the health status, A notification means that detects an anomaly based on the evaluation result and sends a warning, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, it is difficult to routinely monitor the health status of the elderly living in remote areas. Also, symptoms such as cerebral infarction can be detected early due to changes in language ability, but there is a problem that overlooking this can lead to serious health risks. There is a need for a system that can solve these problems and simply and lightly monitor the health status of the elderly in a home environment.

Means for Solving the Problems

[0005] To address this challenge, the system includes a voice receiving means for receiving voice signals transmitted by a user in a remote location, and a voice recognition means for converting the voice signals into text information. Furthermore, it includes a natural language processing means for analyzing this text information and evaluating the fluency of speech, and an anomaly detection means for detecting and notifying abnormalities based on the evaluation results. This system has a function to compare speech patterns with past evaluation results using the anomaly detection means, and to send a warning to the user's family or caregiver if an abnormality is detected. This makes it possible to accurately monitor the health status of elderly people even from a distance.

[0006] An "audio signal" is an electrical signal obtained by converting the voice spoken by the user into an electrical signal.

[0007] "Voice receiving means" refers to a device or system equipped with the function of receiving voice signals transmitted from a user located in a remote location.

[0008] "Speech recognition means" refers to a technology or device that converts received speech signals into string information.

[0009] "Text information" refers to string data converted by speech recognition technology, which is a written representation of the content of the speech.

[0010] "Natural language processing means" refers to technologies or devices that analyze text information and evaluate the fluency of speech, contextual consistency, and other aspects.

[0011] "Speech fluency" is an indicator that evaluates the smoothness and uninterruptedness of the words spoken by the user.

[0012] An "anomaly detection means" is a device or system that has the function of identifying linguistic anomalies based on the analysis results of a natural language processing means.

[0013] A "notification means" is a device or system equipped with the function of transmitting warnings or important information to relevant parties when an anomaly is detected.

[0014] "Speech patterns" are data that represent the characteristics of a user's voice communication and are used for comparison with past speech. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] Embodiments for carrying out the present invention will now be described. This system is designed to monitor the health status of elderly people in remote locations, and users can receive health checks through everyday conversations.

[0037] System Configuration

[0038] This system consists of a smartphone used by an elderly person (hereinafter referred to as the terminal) and a cloud-based analysis platform (hereinafter referred to as the server). The terminal is equipped with voice receiving and voice recognition means and inputs the voice spoken by the user. The server is equipped with natural language processing means and anomaly detection means and analyzes the voice data to evaluate the user's health status.

[0039] Details of the program's processing

[0040] 1. Voice capture and text conversion

[0041] The device captures the user's voice in real time. When the user speaks to the device, saying "Good morning," it receives this as an audio signal.

[0042] The speech recognition system instantly converts this audio signal into text information. For example, the phrase "Good morning" is converted directly into text.

[0043] 2. Data transmission and analysis

[0044] The terminal sends the converted text to the server. The server uses the received text to perform natural language processing.

[0045] The server analyzes the received text data and evaluates the fluency and contextual appropriateness of the speech. For example, it checks for hesitations and pronunciation errors.

[0046] 3. Anomaly detection and notification

[0047] The server detects abnormal patterns by comparing them with past data from a healthy state. If an abnormality is detected, it records the details and sends a warning to the user's family or caregiver via a notification system.

[0048] For example, you can send a message to your family saying, "We noticed an unusual pattern in your conversation this morning. We recommend you consult a medical professional as a precaution."

[0049] This system enables rapid anomaly detection by routinely collecting and analyzing data on the user's health status. Furthermore, because it is conducted through natural communication with elderly users, health monitoring can be seamlessly integrated into their daily lives. This approach allows elderly individuals to live their daily lives with peace of mind, and also provides peace of mind for their families.

[0050] The following describes the processing flow.

[0051] Step 1:

[0052] The device captures the voice spoken by the user. The user starts a normal conversation into the smartphone, and the device's microphone picks up the voice signal.

[0053] Step 2:

[0054] The device's voice recognition system converts the captured audio signal into text information. The voice recognition engine performs real-time audio analysis and converts it into text such as "Good morning."

[0055] Step 3:

[0056] The device sends the converted text information to the server. The text data is sent to a cloud-based server via the network.

[0057] Step 4:

[0058] The server analyzes the received text data using natural language processing techniques. The analysis checks for speech fluency, grammatical consistency, and conversational coherence.

[0059] Step 5:

[0060] The server detects anomalies by comparing the analysis results with past data. It analyzes the differences from past normal speech patterns and determines whether an anomaly exists.

[0061] Step 6:

[0062] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. For example, it might send a message such as, "An anomaly was detected in this morning's conversation. We recommend consulting a medical professional."

[0063] Step 7:

[0064] The device notifies the user of the analysis results as feedback. If there are no problems, it displays a message such as "Have a great day!"

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] There is a challenge in remotely monitoring the health status of elderly people in real time. In particular, when elderly people stumble over their words or exhibit unusual voice patterns, prompt action is required, but there is a lack of efficient and natural methods to achieve this.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes voice input means, voice analysis means, language processing means, anomaly detection means, and notification means. This makes it possible to naturally monitor the health status of a user in a remote location through everyday conversation and to quickly notify them if an anomaly is detected.

[0070] "Voice input means" refers to a device or technology for receiving voice data from a user located in a remote location.

[0071] "Speech analysis means" refers to a device or technique that converts received audio material into textual information.

[0072] "Language processing means" refers to a device or method for analyzing textual information and evaluating the fluency and contextual appropriateness of speech.

[0073] An "anomaly detection method" is a device or technique for detecting anomalies based on evaluation results and comparing them with past health data.

[0074] "Notification means" refers to a device or method for sending a warning to the necessary recipients when an anomaly is detected.

[0075] This invention relates to a system for monitoring the health status of a user located in a remote location via voice. This system consists of the user's smartphone or other voice-enabled device (hereinafter referred to as the terminal) and an analysis platform on the cloud (hereinafter referred to as the server).

[0076] The device captures the user's voice using a microphone as a voice input method. A specific example is the microphone built into a smartphone. The device is equipped with a voice analysis system that converts the voice signal into text information. This uses a speech recognition engine (e.g., Google® Cloud Speech-to-Text API).

[0077] The server analyzes the received text information using language processing tools to evaluate the fluency and contextual appropriateness of the utterance. This evaluation utilizes a generative AI model (e.g., GPT-3®). This analysis triggers an anomaly detection mechanism to detect deviations from normal speech patterns. If an anomaly is detected, the server sends a warning to the user's family or care staff via a notification mechanism. Communication takes place via email or a dedicated application.

[0078] For example, if a user casually says, "The weather was nice today, so I went for a walk," the content of that statement is analyzed to check for any abnormalities. In particular, the fluency of the speech and the consistency of the content are evaluated.

[0079] Examples of prompt statements include the following:

[0080] "How does this system support the daily lives of the elderly?"

[0081] "What are the advantages of health checks conducted through everyday conversation?"

[0082] This system allows elderly individuals to have their health monitored naturally through everyday conversations, and enables a rapid response if any abnormalities are detected.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] Audio Capture

[0086] The device collects voice data from the user in real time. The input is the voice the user speaks into the device's microphone. The device captures this voice using a voice input device and transmits it as an audio signal to an internal computer. This operation usually occurs without the user pressing any specific trigger.

[0087] Step 2:

[0088] Speech recognition and text conversion

[0089] The device converts the audio signal into text data. The input for this conversion is the audio signal acquired in step 1. A speech recognition engine (e.g., Google Cloud Speech-to-Text API) is used as the speech analysis tool to convert the audio signal into text information. This conversion results in the audio "Good morning" being output as the text string "Good morning".

[0090] Step 3:

[0091] Send text

[0092] The terminal sends the character information converted in step 2 to the server. The input is the character data obtained by speech recognition, and the output is the transmission of data to the server. The terminal uploads the character data to the server via an internet connection and applies encrypted communication to ensure the security of the data.

[0093] Step 4:

[0094] Natural Language Processing and Analysis

[0095] The server applies natural language processing to the received text information to evaluate the fluency and contextual appropriateness of the utterance. The input is the text information received from step 3. A generative AI model (e.g., GPT-3) is used to analyze the text and evaluate the naturalness and semantic content of the utterance. The output is the evaluation result, which determines whether there are any abnormalities in the utterance.

[0096] Step 5:

[0097] Anomaly detection and notification

[0098] The server determines whether or not there is an abnormality based on the evaluation results obtained in step 4. The input is the evaluation results after analysis, and the abnormality detection means compares it with past health data to identify abnormal patterns. If an abnormality is detected, the server records the information and generates a warning message as output. Using the notification means, it sends warnings via email or app to the user's family or care staff as needed.

[0099] (Application Example 1)

[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] There is a need to effectively monitor the health status of elderly people living in remote locations and to respond quickly when abnormalities occur. However, conventional methods often require the wearing of devices or special operation, which can be burdensome for the elderly, making them resistant to such methods. Therefore, there is a need to develop a system that monitors health status through natural everyday conversation and promptly notifies users when abnormalities occur.

[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0103] In this invention, the server includes voice input means for receiving the user's voice in real time, voice recognition means for converting the voice into text format, natural language processing means for analyzing the text data and evaluating the health status, and notification means for detecting abnormalities based on the evaluation results and sending warnings. This makes it possible to monitor the health status of the elderly while reducing their burden, and to effectively detect and notify abnormalities.

[0104] "Voice input means" refers to a device or process for receiving voice signals from a user in real time.

[0105] "Speech recognition means" refers to a device or software that has the function of converting received speech into digital text.

[0106] A "natural language processing tool" is a program or process that analyzes text data and evaluates the user's health status based on its content.

[0107] A "notification system" is a mechanism that issues an alarm when an abnormality is detected in the user's health condition and communicates the warning to the relevant parties.

[0108] The system implementing this invention utilizes voice analysis technology that allows users to monitor their health status through everyday conversations. This system mainly consists of three sections.

[0109] First, the device receives the user's everyday conversation in real time through a "voice input method." The smartphone's microphone is used as the voice input method. The device converts the recognized voice into digital text using a "speech recognition method." Cloud-based services such as the Google Cloud Speech-to-Text API can be used for speech recognition.

[0110] Next, the server uses "natural language processing tools" to process the converted text data and evaluate the user's health status. The analysis employs a natural language processing (NLP) model built using the Hugging Face Transformers library. This evaluates the user's utterances and speech fluency, and detects abnormal patterns.

[0111] Finally, if the server detects an anomaly, it will use a "notification system" to send an alert to the relevant parties. This notification utilizes the Twilio API to send SMS or emails to family members and caregivers. For example, if a user says, "I'm not feeling very well today," this is considered an abnormal pattern, and a message is sent to the family saying, "Your parent is showing an unusual health pattern. Please check on them."

[0112] Such a system allows for monitoring of users' health status through natural, everyday conversations, providing an environment where elderly people can continue to live with peace of mind. Furthermore, it enables prompt and appropriate notification to relevant parties, allowing for early intervention.

[0113] An example of a prompt that utilizes a generative AI model is: "Please describe the steps to analyze the user's voice and check for any abnormalities in their health condition."

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The device receives the user's voice in real time through a "voice input means." This input is voice data captured using the smartphone's microphone. The collected voice data is then sent directly to the next processing step.

[0117] Step 2:

[0118] The device converts the received audio data into digital text using a "speech recognition tool." The Google Cloud Speech-to-Text API is used for the speech recognition process. In this process, the audio data is converted into natural language text data, and the converted text information is transferred to the server.

[0119] Step 3:

[0120] The server analyzes the received text information using "natural language processing tools." It uses the Hugging Face Transformers library to analyze the meaning of the text and assess the health status. This analysis identifies abnormal patterns based on keywords and context within the text, obtaining health status assessment data.

[0121] Step 4:

[0122] The server uses the analysis results to determine if there is an abnormality in the health status. This determination is made by comparing it with past health status data, and if an abnormality is detected, the server determines the specific nature of the abnormality.

[0123] Step 5:

[0124] The server immediately issues an alarm if an anomaly is detected by the "notification mechanism." Specifically, it uses the Twilio API to send warning messages via SMS or email to the user's family or caregivers. In this way, it supports relevant parties in responding quickly.

[0125] This enables a system that monitors users' health status through natural, everyday conversations and allows for rapid response in case of abnormalities.

[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0127] This invention relates to a system for monitoring the health and emotional state of elderly individuals in remote locations. The system aims to support the user's peace of mind and safety by detecting abnormalities in health and emotional state through the user's daily conversations and providing appropriate notifications.

[0128] System Configuration

[0129] This system consists of a terminal used by the user and a server that processes and analyzes the data. The terminal is equipped with voice receiving and voice recognition capabilities, capturing the user's voice and converting it into text information in real time. It also incorporates an emotion engine that analyzes the user's emotions from their voice tone and speech content. The converted text information and emotion data are sent to the server.

[0130] Details of the program processing

[0131] 1. Voice capture and text conversion

[0132] The device receives the user's voice via its microphone. The user says to the device, "I'm a little tired today."

[0133] The speech recognition system converts this speech signal into text information and generates the text, "I'm a little tired today."

[0134] 2. Analysis of emotions

[0135] The device's emotion engine analyzes the tone of voice (for example, if the way of speaking is heavier or slower than usual) and evaluates the emotional state. In this case, it determines that the user is experiencing fatigue or stress.

[0136] 3. Data transmission and analysis

[0137] The device sends text information and sentiment data to the server.

[0138] The server uses natural language processing to analyze text information and evaluate the fluency and abnormalities of speech.

[0139] 4. Anomaly detection and notification

[0140] The server determines anomalies based on the analysis results and emotional assessment. Here, the emotional engine detects an anomaly based on the assessment that the user is "more tired than usual."

[0141] If an anomaly is detected, a warning message will be sent to the user's family or caregiver using the notification system. For example, a message such as, "You seemed tired during our conversation this morning, so we recommend checking in," might be sent.

[0142] This system allows for comprehensive monitoring of users' health and emotions through everyday conversations, enabling prompt responses to any abnormalities. This allows users and their families to live their daily lives with greater peace of mind.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] The device captures the user's voice. The user speaks into the smartphone, saying everyday greetings or conversations, such as "I'm a little tired today." The device's microphone receives this voice signal.

[0146] Step 2:

[0147] The device's voice recognition system converts the received audio signal into text information. The voice recognition engine analyzes this audio in real time and generates the sentence "I'm a little tired today" as text.

[0148] Step 3:

[0149] The device's emotion engine begins analyzing the voice to evaluate its characteristics. It analyzes the tone, speed, and intonation of the voice, and in this case, evaluates emotions such as "feeling tired."

[0150] Step 4:

[0151] The device sends text information and sentiment data to the server. The collected data is transferred to a cloud-based server using secure data communication methods.

[0152] Step 5:

[0153] The server analyzes the received text information using natural language processing techniques. The analysis evaluates whether the utterance is fluent and consistent in content.

[0154] Step 6:

[0155] The server analyzes the current emotional data and speech content by comparing it with past data. By comparing it to past normal states, it determines that "a particular level of fatigue is being recognized this time."

[0156] Step 7:

[0157] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. The notification message may include something like, "You appeared more fatigued than usual during today's conversation. We recommend checking on you."

[0158] Step 8:

[0159] The device provides the user with feedback based on the analysis results. For example, if there is a problem, it will display a message such as "Please rest well today." This allows the user to become more aware of their own health status.

[0160] (Example 2)

[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0162] This system addresses the technical challenges of providing peace of mind to users and their families by closely monitoring the health and emotional state of users in remote locations through everyday conversations and promptly notifying them of any abnormalities. In particular, accurately detecting emotional abnormalities based on voice intonation and speech content is crucial.

[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0164] In this invention, the server includes terminal means for receiving acoustic information transmitted by a user, recognition means for converting the acoustic information into linguistic information, and analysis means for analyzing the linguistic information and the intonation of the voice to evaluate the emotional state. This makes it possible to effectively monitor the user's emotions and health status and to quickly notify them if an abnormality is detected.

[0165] "Terminal means" refers to a device for receiving acoustic information transmitted by a user. This device may include microphones, sensors, and other similar devices.

[0166] "Recognition means" refers to technologies and algorithms for converting acoustic information into text. This may include speech recognition software and hardware.

[0167] "Analysis means" refers to methods for evaluating a user's emotional state by analyzing text information and speech intonation. This includes natural language processing and sentiment analysis algorithms.

[0168] "Determination means" refers to a mechanism for determining an anomaly based on the analyzed evaluation results and for issuing notifications based on those determination results. This may include an anomaly detection algorithm and a notification system.

[0169] A "notification system" refers to a system used to send warnings to the user's associates or supporters when an anomaly is detected. Email or messaging services are commonly used for this purpose.

[0170] This invention relates to a system that monitors the health and emotional state of users located remotely and notifies them of abnormalities as needed. This system mainly consists of terminal means and a server, and a detailed embodiment thereof is shown below.

[0171] Embodiment of terminal means

[0172] A terminal is a device equipped with a microphone and various sensors to receive speech from the user. The user's voice is first converted into text information by the terminal's speech recognition system. This conversion often utilizes general-purpose speech recognition software. Specifically, it employs a technology that combines a speech signal processing module and a conversion algorithm.

[0173] Server Embodiment

[0174] The server receives text information and speech intonation data transmitted from the terminal and uses analysis tools to evaluate the user's emotional state based on this information. This analysis utilizes natural language processing technology and sentiment analysis algorithms. Based on the evaluation results, a judgment tool identifies any abnormalities and, if necessary, sends a warning message to the user's associates or supporters using a notification tool.

[0175] Specific example

[0176] For example, when a user says to their device, "I'm a little tired today," this voice is instantly converted into text. Based on this text data and the tone of the voice, the device's emotion analysis engine detects "fatigue." If the server detects an anomaly, a message is sent saying, "We sensed fatigue in your conversation this morning; we recommend you check it."

[0177] Example of a prompt

[0178] "If the user is elderly, what kind of voice tone or text would indicate that they are experiencing fatigue?"

[0179] In this way, the entire system works together to realize advanced monitoring functions that provide peace of mind and security to users and their stakeholders. This configuration supports the technical scope defined in the claims and provides concrete, implementable details.

[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0181] Step 1:

[0182] The device receives the audio signal emitted by the user via a microphone. The input is the user's raw voice, which is converted into digital audio data. Specifically, sound waves are converted into electrical signals, and then processed into a format that can be processed through digital signal processing.

[0183] Step 2:

[0184] The terminal's speech recognition means converts digital audio data into text information. The input is the digital audio data obtained in step 1, and the output is the converted text information. Specifically, this involves a process of sequentially converting speech into text using an acoustic model and a language model.

[0185] Step 3:

[0186] The device's emotion analysis engine analyzes text information and acoustic data to evaluate the emotional state. The input is the text information and acoustic data obtained in step 2, and the evaluation result as an emotional state is output. Specifically, an emotion score is calculated based on voice tone, speaking speed, and keywords in the text.

[0187] Step 4:

[0188] The terminal sends text information and emotional state data to the server. In this process, the input is the emotional analysis results and text data, and the server receives the data as output. Specifically, encrypted data packets are sent via network communication.

[0189] Step 5:

[0190] The server analyzes the received data. The input consists of text information and sentiment data received in step 4. The server uses natural language processing to evaluate the fluency and abnormalities of the speech and obtains the analysis results as output. This process involves keyword extraction and contextual analysis to detect abnormal patterns.

[0191] Step 6:

[0192] The server's judgment mechanism determines whether an anomaly exists based on the analysis results and decides whether notification is necessary. The input is the analysis results obtained in step 5, and the output is the judgment result. Specifically, it evaluates whether an anomaly exists by comparing it with predefined criteria, and if an anomaly is found, it activates the notification process.

[0193] Step 7:

[0194] If an anomaly is detected, the server sends a warning to the user's stakeholders or supporters using a notification mechanism. The input is the determination result obtained in step 6, and the output is the notification message to the stakeholders. Specifically, the message is sent via email or a messaging service.

[0195] (Application Example 2)

[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0197] There is a need for a system that can effectively monitor the health and emotional state of elderly people remotely. However, conventional technology has struggled to accurately detect changes in emotions and to quickly notify family members or caregivers in the event of an anomaly. Therefore, the challenge is to provide a more advanced monitoring system that ensures the safety and security of users and enables a rapid response.

[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0199] In this invention, the server includes means for receiving voice information transmitted by a user in a remote location, means for converting the voice information into text information, and means for analyzing the text information and the tone of the voice to evaluate the fluency of speech and emotional state. This enables high-precision monitoring of the user's health and emotional state and accurate notification in the event of an abnormality.

[0200] "Remote location" refers to a place that is physically distant from the user, and the concept includes environments where observation and interaction take place via communication technology.

[0201] "User" refers to an individual whose health and emotional state are monitored through this system.

[0202] "Audio information" refers to digital or analog data that includes all words and sounds emitted by the user.

[0203] "Voice receiving means" refers to hardware or software used to acquire voice information emitted by a user and incorporate it into the system.

[0204] "Speech recognition means" refers to a technology or process for analyzing received speech information and converting it into corresponding text information.

[0205] "Text information" refers to data in string format converted by speech recognition technology, representing the content of the user's speech.

[0206] "Natural language processing means" refers to technologies or processes for analyzing text information and determining its meaning and context.

[0207] "Speech fluency" is an index that evaluates how natural and smooth a user's speech is, based on the analysis of text information.

[0208] "Emotional state" refers to the psychological or emotional state that can be inferred from the tone of the user's voice and the content of their speech.

[0209] "Anomaly detection means" refers to a technology or process used to determine whether there are abnormalities in a user's health or emotional state based on analyzed data.

[0210] A "notification method" is a technology or process for sending warnings or messages to pre-designated parties when an anomaly is detected.

[0211] The system for carrying out the present invention mainly consists of a terminal and a server. The terminal is equipped with a microphone and plays the role of capturing the voice of an elderly person as a means of receiving voice. For example, suppose the terminal receives a voice message from the user saying, "I haven't been sleeping well lately."

[0212] The speech recognition system within the device converts received speech information into text information. This conversion can utilize speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text information represents the user's spoken content in digital form.

[0213] Next, the device performs emotion analysis. This involves analyzing the tone of the voice using emotion analysis libraries such as IBM Watson® Tone Analyzer and evaluating the emotional state. Based on this evaluation, emotional states such as "anxiety" and "stress" are inferred.

[0214] Text information and emotional state data are transmitted to the server in real time. The server performs natural language processing on this data to analyze speech fluency and emotional state. Furthermore, anomaly detection measures determine whether there are any abnormalities in the user's state based on the results of the data analysis. For example, if the emotional tone is lower than usual compared to past data, it is detected as an anomaly.

[0215] When an anomaly is detected, the server uses a notification system such as Firebase Cloud Messaging to send alert messages to pre-registered stakeholders. These messages may include specific notifications such as, "It is recommended that you check in based on recent conversations."

[0216] As an example, a user's everyday statement, such as "I'm feeling a little down today," is analyzed, and if the emotional tone is determined to be different from normal, a notification requesting confirmation is sent to family members or caregivers.

[0217] An example of a prompt for a generative AI model is, "Explain how to analyze the health and emotional state of elderly people from their everyday conversations and notify their families if any abnormalities are found." This prompt provides an accurate explanation of the system's processing and operation.

[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0219] Step 1:

[0220] The device receives the user's voice as input. Specifically, the device's microphone captures the user's voice, and this voice information is taken into the device as a digital signal. The output is the digitized voice data.

[0221] Step 2:

[0222] The device's speech recognition system takes the audio data obtained in step 1 as input and converts it into text information. This process utilizes the Google Cloud Speech-to-Text API to perform data calculations that convert the audio signal into string-formatted data. The output is text information representing the user's speech.

[0223] Step 3:

[0224] The terminal performs sentiment analysis using the text information and audio data generated in step 2 as input. Specifically, IBM Watson Tone Analyzer is used to process the data and evaluate emotions based on the tone and content of the audio. The output is an evaluation result regarding the user's emotional state.

[0225] Step 4:

[0226] The device sends the text information and sentiment data obtained in step 3 to the server. This involves transferring data over the network. The server receives this data and prepares it for analysis.

[0227] Step 5:

[0228] The server performs natural language processing using the text information and sentiment data received in step 4 as input. The server analyzes the fluency and sentiment changes of the data and performs data calculations to determine whether or not there are any anomalies. The analysis results are obtained as output.

[0229] Step 6:

[0230] The server uses the analysis results from step 5 as input to perform anomaly detection. A comparison operation is performed to determine whether the data is anomaly by comparing it with past data. As a result, data in which anomalies have been detected is output.

[0231] Step 7:

[0232] If an anomaly is detected in step 6, the server sends an alert message to registered stakeholders using the notification system. Here, Firebase Cloud Messaging is used to ensure that the message is delivered quickly, resulting in a notification to the stakeholders.

[0233] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0234] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0235] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0236] [Second Embodiment]

[0237] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0238] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0239] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0240] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0241] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0243] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0244] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0245] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0246] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0247] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0248] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0249] Embodiments for carrying out the present invention will now be described. This system is designed to monitor the health status of elderly people in remote locations, and users can receive health checks through everyday conversations.

[0250] System Configuration

[0251] This system consists of a smartphone used by an elderly person (hereinafter referred to as the terminal) and a cloud-based analysis platform (hereinafter referred to as the server). The terminal is equipped with voice receiving and voice recognition means and inputs the voice spoken by the user. The server is equipped with natural language processing means and anomaly detection means and analyzes the voice data to evaluate the user's health status.

[0252] Details of the program's processing

[0253] 1. Voice capture and text conversion

[0254] The device captures the user's voice in real time. When the user speaks to the device, saying "Good morning," it receives this as an audio signal.

[0255] The speech recognition system instantly converts this audio signal into text information. For example, the phrase "Good morning" is converted directly into text.

[0256] 2. Data transmission and analysis

[0257] The terminal sends the converted text to the server. The server uses the received text to perform natural language processing.

[0258] The server analyzes the received text data and evaluates the fluency and contextual appropriateness of the speech. For example, it checks for hesitations and pronunciation errors.

[0259] 3. Anomaly detection and notification

[0260] The server detects abnormal patterns by comparing them with past data from a healthy state. If an abnormality is detected, it records the details and sends a warning to the user's family or caregiver via a notification system.

[0261] For example, you can send a message to your family saying, "We noticed an unusual pattern in your conversation this morning. We recommend you consult a medical professional as a precaution."

[0262] This system enables rapid anomaly detection by routinely collecting and analyzing data on the user's health status. Furthermore, because it is conducted through natural communication with elderly users, health monitoring can be seamlessly integrated into their daily lives. This approach allows elderly individuals to live their daily lives with peace of mind, and also provides peace of mind for their families.

[0263] The following describes the processing flow.

[0264] Step 1:

[0265] The device captures the voice spoken by the user. The user starts a normal conversation into the smartphone, and the device's microphone picks up the voice signal.

[0266] Step 2:

[0267] The device's voice recognition system converts the captured audio signal into text information. The voice recognition engine performs real-time audio analysis and converts it into text such as "Good morning."

[0268] Step 3:

[0269] The device sends the converted text information to the server. The text data is sent to a cloud-based server via the network.

[0270] Step 4:

[0271] The server analyzes the received text data using natural language processing techniques. The analysis checks for speech fluency, grammatical consistency, and conversational coherence.

[0272] Step 5:

[0273] The server detects anomalies by comparing the analysis results with past data. It analyzes the differences from past normal speech patterns and determines whether an anomaly exists.

[0274] Step 6:

[0275] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. For example, it might send a message such as, "An anomaly was detected in this morning's conversation. We recommend consulting a medical professional."

[0276] Step 7:

[0277] The device notifies the user of the analysis results as feedback. If there are no problems, it displays a message such as "Have a great day!"

[0278] (Example 1)

[0279] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0280] There is a challenge in remotely monitoring the health status of elderly people in real time. In particular, when elderly people stumble over their words or exhibit unusual voice patterns, prompt action is required, but there is a lack of efficient and natural methods to achieve this.

[0281] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0282] In this invention, the server includes voice input means, voice analysis means, language processing means, anomaly detection means, and notification means. This makes it possible to naturally monitor the health status of a user in a remote location through everyday conversation and to quickly notify them if an anomaly is detected.

[0283] The "voice input means" is a device or technology for receiving voice materials of a user located remotely.

[0284] The "voice analysis means" is a device or technique for converting the received voice materials into character information.

[0285] The "language processing means" is a device or method for analyzing character information and evaluating the fluency of speech and the appropriateness of context.

[0286] The "abnormality determination means" is a device or technique for detecting an abnormality based on the evaluation result and comparing it with past health data.

[0287] The "notification means" is a device or method for sending a warning to a necessary destination when an abnormality is detected.

[0288] This invention relates to a system for monitoring the health status of a user located remotely via voice. This system is composed of a user's smartphone or other voice-responsive devices (hereinafter referred to as terminals) and an analysis platform on the cloud (hereinafter referred to as servers).

[0289] The terminal uses a microphone as the voice input means to capture the user's voice. As a specific example, the microphone built into the smartphone is used. The terminal is equipped with voice analysis means for converting the voice signal into character information. For this, a voice recognition engine (e.g., Google Cloud Speech-to-Text API) is used.

[0290] The server analyzes the received character information by the language processing means and evaluates the fluency of speech and the appropriateness of context. For this evaluation, a generative AI model is utilized (e.g., GPT-3). Through this analysis, the abnormality determination means for detecting an abnormality from the normal speech pattern works. When an abnormality is detected, the server sends a warning to the user's family and care staff by the notification means. The communication is carried out through email or a dedicated application.

[0291] For example, if a user casually says, "The weather was nice today, so I went for a walk," the content of that statement is analyzed to check for any abnormalities. In particular, the fluency of the speech and the consistency of the content are evaluated.

[0292] Examples of prompt statements include the following:

[0293] "How does this system support the daily lives of the elderly?"

[0294] "What are the advantages of health checks conducted through everyday conversation?"

[0295] This system allows elderly individuals to have their health monitored naturally through everyday conversations, and enables a rapid response if any abnormalities are detected.

[0296] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0297] Step 1:

[0298] Audio Capture

[0299] The device collects voice data from the user in real time. The input is the voice the user speaks into the device's microphone. The device captures this voice using a voice input device and transmits it as an audio signal to an internal computer. This operation usually occurs without the user pressing any specific trigger.

[0300] Step 2:

[0301] Speech recognition and text conversion

[0302] The terminal converts the voice signal into character data. At this time, the input is the voice signal obtained in step 1. Using a speech recognition engine (e.g., Google Cloud Speech-to-Text API) as the speech analysis means, the voice signal is converted into character information. Through this conversion, the voice "おはようございます" is output as the character string "おはようございます" as it is.

[0303] Step 3:

[0304] Text transmission

[0305] The terminal sends the character information converted in step 2 to the server. The input is the character data obtained by speech recognition, and the output is the data transmission to the server. The terminal uploads the character data to the server through the Internet connection and applies encrypted communication to ensure data security.

[0306] Step 4:

[0307] Natural language processing and analysis

[0308] The server applies natural language processing to the received character information and evaluates the fluency and context appropriateness of the speech. The input is the character information received from step 3. Utilizing a generative AI model (e.g., GPT-3) to analyze the text and evaluate the naturalness and semantic content of the speech. As the output, an evaluation result is obtained, and it is determined whether there is an abnormality in the speech.

[0309] Step 5:

[0310] Abnormality detection and notification

[0311] The server determines the presence or absence of an abnormality based on the evaluation result obtained in step 4. The input is the evaluation result after analysis, and it is compared with past health data by the abnormality determination means to identify abnormal patterns. When an abnormality is detected, the server records the information and generates a warning message as the output. Using the notification means, warnings are sent to the user's family and care staff by email or app as necessary.

[0312] (Application Example 1)

[0313] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0314] There is a need to effectively monitor the health status of elderly people living in remote locations and to respond quickly when abnormalities occur. However, conventional methods often require the wearing of devices or special operation, which can be burdensome for the elderly, making them resistant to such methods. Therefore, there is a need to develop a system that monitors health status through natural everyday conversation and promptly notifies users when abnormalities occur.

[0315] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0316] In this invention, the server includes voice input means for receiving the user's voice in real time, voice recognition means for converting the voice into text format, natural language processing means for analyzing the text data and evaluating the health status, and notification means for detecting abnormalities based on the evaluation results and sending warnings. This makes it possible to monitor the health status of the elderly while reducing their burden, and to effectively detect and notify abnormalities.

[0317] "Voice input means" refers to a device or process for receiving voice signals from a user in real time.

[0318] "Speech recognition means" refers to a device or software that has the function of converting received speech into digital text.

[0319] A "natural language processing tool" is a program or process that analyzes text data and evaluates the user's health status based on its content.

[0320] A "notification system" is a mechanism that issues an alarm when an abnormality is detected in the user's health condition and communicates the warning to the relevant parties.

[0321] The system implementing this invention utilizes voice analysis technology that allows users to monitor their health status through everyday conversations. This system mainly consists of three sections.

[0322] First, the device receives the user's everyday conversation in real time through a "voice input method." The smartphone's microphone is used as the voice input method. The device converts the recognized voice into digital text using a "speech recognition method." Cloud-based services such as the Google Cloud Speech-to-Text API can be used for speech recognition.

[0323] Next, the server uses "natural language processing tools" to process the converted text data and evaluate the user's health status. The analysis employs a natural language processing (NLP) model built using the Hugging Face Transformers library. This evaluates the user's utterances and speech fluency, and detects abnormal patterns.

[0324] Finally, if the server detects an anomaly, it will use a "notification system" to send an alert to the relevant parties. This notification utilizes the Twilio API to send SMS or emails to family members and caregivers. For example, if a user says, "I'm not feeling very well today," this is considered an abnormal pattern, and a message is sent to the family saying, "Your parent is showing an unusual health pattern. Please check on them."

[0325] Such a system allows for monitoring of users' health status through natural, everyday conversations, providing an environment where elderly people can continue to live with peace of mind. Furthermore, it enables prompt and appropriate notification to relevant parties, allowing for early intervention.

[0326] An example of a prompt that utilizes a generative AI model is: "Please describe the steps to analyze the user's voice and check for any abnormalities in their health condition."

[0327] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0328] Step 1:

[0329] The device receives the user's voice in real time through a "voice input means." This input is voice data captured using the smartphone's microphone. The collected voice data is then sent directly to the next processing step.

[0330] Step 2:

[0331] The device converts the received audio data into digital text using a "speech recognition tool." The Google Cloud Speech-to-Text API is used for the speech recognition process. In this process, the audio data is converted into natural language text data, and the converted text information is transferred to the server.

[0332] Step 3:

[0333] The server analyzes the received text information using "natural language processing tools." It uses the Hugging Face Transformers library to analyze the meaning of the text and assess the health status. This analysis identifies abnormal patterns based on keywords and context within the text, obtaining health status assessment data.

[0334] Step 4:

[0335] The server uses the analysis results to determine if there is an abnormality in the health status. This determination is made by comparing it with past health status data, and if an abnormality is detected, the server determines the specific nature of the abnormality.

[0336] Step 5:

[0337] The server immediately issues an alarm if an anomaly is detected by the "notification mechanism." Specifically, it uses the Twilio API to send warning messages via SMS or email to the user's family or caregivers. In this way, it supports relevant parties in responding quickly.

[0338] This enables a system that monitors users' health status through natural, everyday conversations and allows for rapid response in case of abnormalities.

[0339] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0340] This invention relates to a system for monitoring the health and emotional state of elderly individuals in remote locations. The system aims to support the user's peace of mind and safety by detecting abnormalities in health and emotional state through the user's daily conversations and providing appropriate notifications.

[0341] System Configuration

[0342] This system consists of a terminal used by the user and a server that processes and analyzes the data. The terminal is equipped with voice receiving and voice recognition capabilities, capturing the user's voice and converting it into text information in real time. It also incorporates an emotion engine that analyzes the user's emotions from their voice tone and speech content. The converted text information and emotion data are sent to the server.

[0343] Details of the program processing

[0344] 1. Voice capture and text conversion

[0345] The device receives the user's voice via its microphone. The user says to the device, "I'm a little tired today."

[0346] The speech recognition system converts this speech signal into text information and generates the text, "I'm a little tired today."

[0347] 2. Analysis of emotions

[0348] The device's emotion engine analyzes the tone of voice (for example, if the way of speaking is heavier or slower than usual) and evaluates the emotional state. In this case, it determines that the user is experiencing fatigue or stress.

[0349] 3. Data transmission and analysis

[0350] The device sends text information and sentiment data to the server.

[0351] The server uses natural language processing to analyze text information and evaluate the fluency and abnormalities of speech.

[0352] 4. Anomaly detection and notification

[0353] The server determines anomalies based on the analysis results and emotional assessment. Here, the emotional engine detects an anomaly based on the assessment that the user is "more tired than usual."

[0354] If an anomaly is detected, a warning message will be sent to the user's family or caregiver using the notification system. For example, a message such as, "You seemed tired during our conversation this morning, so we recommend checking in," might be sent.

[0355] This system allows for comprehensive monitoring of users' health and emotions through everyday conversations, enabling prompt responses to any abnormalities. This allows users and their families to live their daily lives with greater peace of mind.

[0356] The following describes the processing flow.

[0357] Step 1:

[0358] The device captures the user's voice. The user speaks into the smartphone, saying everyday greetings or conversations, such as "I'm a little tired today." The device's microphone receives this voice signal.

[0359] Step 2:

[0360] The device's voice recognition system converts the received audio signal into text information. The voice recognition engine analyzes this audio in real time and generates the sentence "I'm a little tired today" as text.

[0361] Step 3:

[0362] The device's emotion engine begins analyzing the voice to evaluate its characteristics. It analyzes the tone, speed, and intonation of the voice, and in this case, evaluates emotions such as "feeling tired."

[0363] Step 4:

[0364] The device sends text information and sentiment data to the server. The collected data is transferred to a cloud-based server using secure data communication methods.

[0365] Step 5:

[0366] The server analyzes the received text information using natural language processing techniques. The analysis evaluates whether the utterance is fluent and consistent in content.

[0367] Step 6:

[0368] The server analyzes the current emotional data and speech content by comparing it with past data. By comparing it to past normal states, it determines that "a particular level of fatigue is being recognized this time."

[0369] Step 7:

[0370] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. The notification message may include something like, "You appeared more fatigued than usual during today's conversation. We recommend checking on you."

[0371] Step 8:

[0372] The device provides the user with feedback based on the analysis results. For example, if there is a problem, it will display a message such as "Please rest well today." This allows the user to become more aware of their own health status.

[0373] (Example 2)

[0374] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0375] This system addresses the technical challenges of providing peace of mind to users and their families by closely monitoring the health and emotional state of users in remote locations through everyday conversations and promptly notifying them of any abnormalities. In particular, accurately detecting emotional abnormalities based on voice intonation and speech content is crucial.

[0376] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0377] In this invention, the server includes terminal means for receiving acoustic information transmitted by a user, recognition means for converting the acoustic information into linguistic information, and analysis means for analyzing the linguistic information and the intonation of the voice to evaluate the emotional state. This makes it possible to effectively monitor the user's emotions and health status and to quickly notify them if an abnormality is detected.

[0378] "Terminal means" refers to a device for receiving acoustic information transmitted by a user. This device may include microphones, sensors, and other similar devices.

[0379] "Recognition means" refers to technologies and algorithms for converting acoustic information into text. This may include speech recognition software and hardware.

[0380] "Analysis means" refers to methods for evaluating a user's emotional state by analyzing text information and speech intonation. This includes natural language processing and sentiment analysis algorithms.

[0381] "Determination means" refers to a mechanism for determining an anomaly based on the analyzed evaluation results and for issuing notifications based on those determination results. This may include an anomaly detection algorithm and a notification system.

[0382] A "notification system" refers to a system used to send warnings to the user's associates or supporters when an anomaly is detected. Email or messaging services are commonly used for this purpose.

[0383] This invention relates to a system that monitors the health and emotional state of users located remotely and notifies them of abnormalities as needed. This system mainly consists of terminal means and a server, and a detailed embodiment thereof is shown below.

[0384] Embodiment of terminal means

[0385] A terminal is a device equipped with a microphone and various sensors to receive speech from the user. The user's voice is first converted into text information by the terminal's speech recognition system. This conversion often utilizes general-purpose speech recognition software. Specifically, it employs a technology that combines a speech signal processing module and a conversion algorithm.

[0386] Server Embodiment

[0387] The server receives text information and speech intonation data transmitted from the terminal and uses analysis tools to evaluate the user's emotional state based on this information. This analysis utilizes natural language processing technology and sentiment analysis algorithms. Based on the evaluation results, a judgment tool identifies any abnormalities and, if necessary, sends a warning message to the user's associates or supporters using a notification tool.

[0388] Specific example

[0389] For example, when a user says to their device, "I'm a little tired today," this voice is instantly converted into text. Based on this text data and the tone of the voice, the device's emotion analysis engine detects "fatigue." If the server detects an anomaly, a message is sent saying, "We sensed fatigue in your conversation this morning; we recommend you check it."

[0390] Example of a prompt

[0391] "If the user is elderly, what kind of voice tone or text would indicate that they are experiencing fatigue?"

[0392] In this way, the entire system works together to realize advanced monitoring functions that provide peace of mind and security to users and their stakeholders. This configuration supports the technical scope defined in the claims and provides concrete, implementable details.

[0393] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0394] Step 1:

[0395] The device receives the audio signal emitted by the user via a microphone. The input is the user's raw voice, which is converted into digital audio data. Specifically, sound waves are converted into electrical signals, and then processed into a format that can be processed through digital signal processing.

[0396] Step 2:

[0397] The terminal's speech recognition means converts digital audio data into text information. The input is the digital audio data obtained in step 1, and the output is the converted text information. Specifically, this involves a process of sequentially converting speech into text using an acoustic model and a language model.

[0398] Step 3:

[0399] The device's emotion analysis engine analyzes text information and acoustic data to evaluate the emotional state. The input is the text information and acoustic data obtained in step 2, and the evaluation result as an emotional state is output. Specifically, an emotion score is calculated based on voice tone, speaking speed, and keywords in the text.

[0400] Step 4:

[0401] The terminal sends text information and emotional state data to the server. In this process, the input is the emotional analysis results and text data, and the server receives the data as output. Specifically, encrypted data packets are sent via network communication.

[0402] Step 5:

[0403] The server analyzes the received data. The input consists of text information and sentiment data received in step 4. The server uses natural language processing to evaluate the fluency and abnormalities of the speech and obtains the analysis results as output. This process involves keyword extraction and contextual analysis to detect abnormal patterns.

[0404] Step 6:

[0405] The server's judgment mechanism determines whether an anomaly exists based on the analysis results and decides whether notification is necessary. The input is the analysis results obtained in step 5, and the output is the judgment result. Specifically, it evaluates whether an anomaly exists by comparing it with predefined criteria, and if an anomaly is found, it activates the notification process.

[0406] Step 7:

[0407] If an anomaly is detected, the server sends a warning to the user's stakeholders or supporters using a notification mechanism. The input is the determination result obtained in step 6, and the output is the notification message to the stakeholders. Specifically, the message is sent via email or a messaging service.

[0408] (Application Example 2)

[0409] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0410] There is a need for a system that can effectively monitor the health and emotional state of elderly people remotely. However, conventional technology has struggled to accurately detect changes in emotions and to quickly notify family members or caregivers in the event of an anomaly. Therefore, the challenge is to provide a more advanced monitoring system that ensures the safety and security of users and enables a rapid response.

[0411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0412] In this invention, the server includes means for receiving voice information transmitted by a user in a remote location, means for converting the voice information into text information, and means for analyzing the text information and the tone of the voice to evaluate the fluency of speech and emotional state. This enables high-precision monitoring of the user's health and emotional state and accurate notification in the event of an abnormality.

[0413] "Remote location" refers to a place that is physically distant from the user, and the concept includes environments where observation and interaction take place via communication technology.

[0414] "User" refers to an individual whose health and emotional state are monitored through this system.

[0415] "Audio information" refers to digital or analog data that includes all words and sounds emitted by the user.

[0416] "Voice receiving means" refers to hardware or software used to acquire voice information emitted by a user and incorporate it into the system.

[0417] "Speech recognition means" refers to a technology or process for analyzing received speech information and converting it into corresponding text information.

[0418] "Text information" refers to data in string format converted by speech recognition technology, representing the content of the user's speech.

[0419] "Natural language processing means" refers to technologies or processes for analyzing text information and determining its meaning and context.

[0420] "Speech fluency" is an index that evaluates how natural and smooth a user's speech is, based on the analysis of text information.

[0421] "Emotional state" refers to the psychological or emotional state that can be inferred from the tone of the user's voice and the content of their speech.

[0422] "Anomaly detection means" refers to a technology or process used to determine whether there are abnormalities in a user's health or emotional state based on analyzed data.

[0423] A "notification method" is a technology or process for sending warnings or messages to pre-designated parties when an anomaly is detected.

[0424] The system for carrying out the present invention mainly consists of a terminal and a server. The terminal is equipped with a microphone and plays the role of capturing the voice of an elderly person as a means of receiving voice. For example, suppose the terminal receives a voice message from the user saying, "I haven't been sleeping well lately."

[0425] The speech recognition system within the device converts received speech information into text information. This conversion can utilize speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text information represents the user's spoken content in digital form.

[0426] Next, the device performs emotion analysis. This involves analyzing the tone of the voice using emotion analysis libraries such as IBM Watson Tone Analyzer and evaluating the emotional state. Based on this evaluation, emotional states such as "anxiety" and "stress" are inferred.

[0427] Text information and emotional state data are transmitted to the server in real time. The server performs natural language processing on this data to analyze speech fluency and emotional state. Furthermore, anomaly detection measures determine whether there are any abnormalities in the user's state based on the results of the data analysis. For example, if the emotional tone is lower than usual compared to past data, it is detected as an anomaly.

[0428] When an anomaly is detected, the server uses a notification system such as Firebase Cloud Messaging to send alert messages to pre-registered stakeholders. These messages may include specific notifications such as, "It is recommended that you check in based on recent conversations."

[0429] As an example, a user's everyday statement, such as "I'm feeling a little down today," is analyzed, and if the emotional tone is determined to be different from normal, a notification requesting confirmation is sent to family members or caregivers.

[0430] An example of a prompt for a generative AI model is, "Explain how to analyze the health and emotional state of elderly people from their everyday conversations and notify their families if any abnormalities are found." This prompt provides an accurate explanation of the system's processing and operation.

[0431] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0432] Step 1:

[0433] The device receives the user's voice as input. Specifically, the device's microphone captures the user's voice, and this voice information is taken into the device as a digital signal. The output is the digitized voice data.

[0434] Step 2:

[0435] The device's speech recognition system takes the audio data obtained in step 1 as input and converts it into text information. This process utilizes the Google Cloud Speech-to-Text API to perform data calculations that convert the audio signal into string-formatted data. The output is text information representing the user's speech.

[0436] Step 3:

[0437] The terminal performs sentiment analysis using the text information and audio data generated in step 2 as input. Specifically, IBM Watson Tone Analyzer is used to process the data and evaluate emotions based on the tone and content of the audio. The output is an evaluation result regarding the user's emotional state.

[0438] Step 4:

[0439] The device sends the text information and sentiment data obtained in step 3 to the server. This involves transferring data over the network. The server receives this data and prepares it for analysis.

[0440] Step 5:

[0441] The server performs natural language processing using the text information and sentiment data received in step 4 as input. The server analyzes the fluency and sentiment changes of the data and performs data calculations to determine whether or not there are any anomalies. The analysis results are obtained as output.

[0442] Step 6:

[0443] The server uses the analysis results from step 5 as input to perform anomaly detection. A comparison operation is performed to determine whether the data is anomaly by comparing it with past data. As a result, data in which anomalies have been detected is output.

[0444] Step 7:

[0445] If an anomaly is detected in step 6, the server sends an alert message to registered stakeholders using the notification system. Here, Firebase Cloud Messaging is used to ensure that the message is delivered quickly, resulting in a notification to the stakeholders.

[0446] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0447] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0448] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0449] [Third Embodiment]

[0450] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0451] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0452] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0453] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0454] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0456] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0457] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0458] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0459] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0460] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0461] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0462] Embodiments for carrying out the present invention will now be described. This system is designed to monitor the health status of elderly people in remote locations, and users can receive health checks through everyday conversations.

[0463] System Configuration

[0464] This system consists of a smartphone used by an elderly person (hereinafter referred to as the terminal) and a cloud-based analysis platform (hereinafter referred to as the server). The terminal is equipped with voice receiving and voice recognition means and inputs the voice spoken by the user. The server is equipped with natural language processing means and anomaly detection means and analyzes the voice data to evaluate the user's health status.

[0465] Details of the program's processing

[0466] 1. Voice capture and text conversion

[0467] The device captures the user's voice in real time. When the user speaks to the device, saying "Good morning," it receives this as an audio signal.

[0468] The speech recognition system instantly converts this audio signal into text information. For example, the phrase "Good morning" is converted directly into text.

[0469] 2. Data transmission and analysis

[0470] The terminal sends the converted text to the server. The server uses the received text to perform natural language processing.

[0471] The server analyzes the received text data and evaluates the fluency and contextual appropriateness of the speech. For example, it checks for hesitations and pronunciation errors.

[0472] 3. Anomaly detection and notification

[0473] The server detects abnormal patterns by comparing them with past data from a healthy state. If an abnormality is detected, it records the details and sends a warning to the user's family or caregiver via a notification system.

[0474] For example, you can send a message to your family saying, "We noticed an unusual pattern in your conversation this morning. We recommend you consult a medical professional as a precaution."

[0475] This system enables rapid anomaly detection by routinely collecting and analyzing data on the user's health status. Furthermore, because it is conducted through natural communication with elderly users, health monitoring can be seamlessly integrated into their daily lives. This approach allows elderly individuals to live their daily lives with peace of mind, and also provides peace of mind for their families.

[0476] The following describes the processing flow.

[0477] Step 1:

[0478] The device captures the voice spoken by the user. The user starts a normal conversation into the smartphone, and the device's microphone picks up the voice signal.

[0479] Step 2:

[0480] The device's voice recognition system converts the captured audio signal into text information. The voice recognition engine performs real-time audio analysis and converts it into text such as "Good morning."

[0481] Step 3:

[0482] The device sends the converted text information to the server. The text data is sent to a cloud-based server via the network.

[0483] Step 4:

[0484] The server analyzes the received text data using natural language processing techniques. The analysis checks for speech fluency, grammatical consistency, and conversational coherence.

[0485] Step 5:

[0486] The server detects anomalies by comparing the analysis results with past data. It analyzes the differences from past normal speech patterns and determines whether an anomaly exists.

[0487] Step 6:

[0488] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. For example, it might send a message such as, "An anomaly was detected in this morning's conversation. We recommend consulting a medical professional."

[0489] Step 7:

[0490] The device notifies the user of the analysis results as feedback. If there are no problems, it displays a message such as "Have a great day!"

[0491] (Example 1)

[0492] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0493] There is a challenge in remotely monitoring the health status of elderly people in real time. In particular, when elderly people stumble over their words or exhibit unusual voice patterns, prompt action is required, but there is a lack of efficient and natural methods to achieve this.

[0494] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0495] In this invention, the server includes voice input means, voice analysis means, language processing means, anomaly detection means, and notification means. This makes it possible to naturally monitor the health status of a user in a remote location through everyday conversation and to quickly notify them if an anomaly is detected.

[0496] "Voice input means" refers to a device or technology for receiving voice data from a user located in a remote location.

[0497] "Speech analysis means" refers to a device or technique that converts received audio material into textual information.

[0498] "Language processing means" refers to a device or method for analyzing textual information and evaluating the fluency and contextual appropriateness of speech.

[0499] An "anomaly detection method" is a device or technique for detecting anomalies based on evaluation results and comparing them with past health data.

[0500] "Notification means" refers to a device or method for sending a warning to the necessary recipients when an anomaly is detected.

[0501] This invention relates to a system for monitoring the health status of a user located in a remote location via voice. This system consists of the user's smartphone or other voice-enabled device (hereinafter referred to as the terminal) and an analysis platform on the cloud (hereinafter referred to as the server).

[0502] The device captures the user's voice using a microphone as a voice input method. A specific example is the microphone built into a smartphone. The device is equipped with a voice analysis system that converts the voice signal into text information. This is done using a speech recognition engine (e.g., Google Cloud Speech-to-Text API).

[0503] The server analyzes the received text information using language processing tools to evaluate the fluency and contextual appropriateness of the utterance. This evaluation utilizes a generative AI model (e.g., GPT-3). This analysis triggers an anomaly detection mechanism that detects deviations from normal speech patterns. If an anomaly is detected, the server sends a warning to the user's family or care staff via a notification mechanism. Communication takes place via email or a dedicated application.

[0504] For example, if a user casually says, "The weather was nice today, so I went for a walk," the content of that statement is analyzed to check for any abnormalities. In particular, the fluency of the speech and the consistency of the content are evaluated.

[0505] Examples of prompt statements include the following:

[0506] "How does this system support the daily lives of the elderly?"

[0507] "What are the advantages of health checks conducted through everyday conversation?"

[0508] This system allows elderly individuals to have their health monitored naturally through everyday conversations, and enables a rapid response if any abnormalities are detected.

[0509] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0510] Step 1:

[0511] Audio Capture

[0512] The device collects voice data from the user in real time. The input is the voice the user speaks into the device's microphone. The device captures this voice using a voice input device and transmits it as an audio signal to an internal computer. This operation usually occurs without the user pressing any specific trigger.

[0513] Step 2:

[0514] Speech recognition and text conversion

[0515] The device converts the audio signal into text data. The input for this conversion is the audio signal acquired in step 1. A speech recognition engine (e.g., Google Cloud Speech-to-Text API) is used as the speech analysis tool to convert the audio signal into text information. This conversion results in the audio "Good morning" being output as the text string "Good morning".

[0516] Step 3:

[0517] Send text

[0518] The terminal sends the character information converted in step 2 to the server. The input is the character data obtained by speech recognition, and the output is the transmission of data to the server. The terminal uploads the character data to the server via an internet connection and applies encrypted communication to ensure the security of the data.

[0519] Step 4:

[0520] Natural Language Processing and Analysis

[0521] The server applies natural language processing to the received text information to evaluate the fluency and contextual appropriateness of the utterance. The input is the text information received from step 3. A generative AI model (e.g., GPT-3) is used to analyze the text and evaluate the naturalness and semantic content of the utterance. The output is the evaluation result, which determines whether there are any abnormalities in the utterance.

[0522] Step 5:

[0523] Anomaly detection and notification

[0524] The server determines whether or not there is an abnormality based on the evaluation results obtained in step 4. The input is the evaluation results after analysis, and the abnormality detection means compares it with past health data to identify abnormal patterns. If an abnormality is detected, the server records the information and generates a warning message as output. Using the notification means, it sends warnings via email or app to the user's family or care staff as needed.

[0525] (Application Example 1)

[0526] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0527] There is a need to effectively monitor the health status of elderly people living in remote locations and to respond quickly when abnormalities occur. However, conventional methods often require the wearing of devices or special operation, which can be burdensome for the elderly, making them resistant to such methods. Therefore, there is a need to develop a system that monitors health status through natural everyday conversation and promptly notifies users when abnormalities occur.

[0528] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0529] In this invention, the server includes voice input means for receiving the user's voice in real time, voice recognition means for converting the voice into text format, natural language processing means for analyzing the text data and evaluating the health status, and notification means for detecting abnormalities based on the evaluation results and sending warnings. This makes it possible to monitor the health status of the elderly while reducing their burden, and to effectively detect and notify abnormalities.

[0530] "Voice input means" refers to a device or process for receiving voice signals from a user in real time.

[0531] "Speech recognition means" refers to a device or software that has the function of converting received speech into digital text.

[0532] A "natural language processing tool" is a program or process that analyzes text data and evaluates the user's health status based on its content.

[0533] A "notification system" is a mechanism that issues an alarm when an abnormality is detected in the user's health condition and communicates the warning to the relevant parties.

[0534] The system implementing this invention utilizes voice analysis technology that allows users to monitor their health status through everyday conversations. This system mainly consists of three sections.

[0535] First, the device receives the user's everyday conversation in real time through a "voice input method." The smartphone's microphone is used as the voice input method. The device converts the recognized voice into digital text using a "speech recognition method." Cloud-based services such as the Google Cloud Speech-to-Text API can be used for speech recognition.

[0536] Next, the server uses "natural language processing tools" to process the converted text data and evaluate the user's health status. The analysis employs a natural language processing (NLP) model built using the Hugging Face Transformers library. This evaluates the user's utterances and speech fluency, and detects abnormal patterns.

[0537] Finally, if the server detects an anomaly, it will use a "notification system" to send an alert to the relevant parties. This notification utilizes the Twilio API to send SMS or emails to family members and caregivers. For example, if a user says, "I'm not feeling very well today," this is considered an abnormal pattern, and a message is sent to the family saying, "Your parent is showing an unusual health pattern. Please check on them."

[0538] Such a system allows for monitoring of users' health status through natural, everyday conversations, providing an environment where elderly people can continue to live with peace of mind. Furthermore, it enables prompt and appropriate notification to relevant parties, allowing for early intervention.

[0539] An example of a prompt that utilizes a generative AI model is: "Please describe the steps to analyze the user's voice and check for any abnormalities in their health condition."

[0540] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0541] Step 1:

[0542] The device receives the user's voice in real time through a "voice input means." This input is voice data captured using the smartphone's microphone. The collected voice data is then sent directly to the next processing step.

[0543] Step 2:

[0544] The device converts the received audio data into digital text using a "speech recognition tool." The Google Cloud Speech-to-Text API is used for the speech recognition process. In this process, the audio data is converted into natural language text data, and the converted text information is transferred to the server.

[0545] Step 3:

[0546] The server analyzes the received text information using "natural language processing tools." It uses the Hugging Face Transformers library to analyze the meaning of the text and assess the health status. This analysis identifies abnormal patterns based on keywords and context within the text, obtaining health status assessment data.

[0547] Step 4:

[0548] The server uses the analysis results to determine if there is an abnormality in the health status. This determination is made by comparing it with past health status data, and if an abnormality is detected, the server determines the specific nature of the abnormality.

[0549] Step 5:

[0550] The server immediately issues an alarm if an anomaly is detected by the "notification mechanism." Specifically, it uses the Twilio API to send warning messages via SMS or email to the user's family or caregivers. In this way, it supports relevant parties in responding quickly.

[0551] This enables a system that monitors users' health status through natural, everyday conversations and allows for rapid response in case of abnormalities.

[0552] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0553] This invention relates to a system for monitoring the health and emotional state of elderly individuals in remote locations. The system aims to support the user's peace of mind and safety by detecting abnormalities in health and emotional state through the user's daily conversations and providing appropriate notifications.

[0554] System Configuration

[0555] This system consists of a terminal used by the user and a server that processes and analyzes the data. The terminal is equipped with voice receiving and voice recognition capabilities, capturing the user's voice and converting it into text information in real time. It also incorporates an emotion engine that analyzes the user's emotions from their voice tone and speech content. The converted text information and emotion data are sent to the server.

[0556] Details of the program processing

[0557] 1. Voice capture and text conversion

[0558] The device receives the user's voice via its microphone. The user says to the device, "I'm a little tired today."

[0559] The speech recognition system converts this speech signal into text information and generates the text, "I'm a little tired today."

[0560] 2. Analysis of emotions

[0561] The device's emotion engine analyzes the tone of voice (for example, if the way of speaking is heavier or slower than usual) and evaluates the emotional state. In this case, it determines that the user is experiencing fatigue or stress.

[0562] 3. Data transmission and analysis

[0563] The device sends text information and sentiment data to the server.

[0564] The server uses natural language processing to analyze text information and evaluate the fluency and abnormalities of speech.

[0565] 4. Anomaly detection and notification

[0566] The server determines anomalies based on the analysis results and emotional assessment. Here, the emotional engine detects an anomaly based on the assessment that the user is "more tired than usual."

[0567] If an anomaly is detected, a warning message will be sent to the user's family or caregiver using the notification system. For example, a message such as, "You seemed tired during our conversation this morning, so we recommend checking in," might be sent.

[0568] This system allows for comprehensive monitoring of users' health and emotions through everyday conversations, enabling prompt responses to any abnormalities. This allows users and their families to live their daily lives with greater peace of mind.

[0569] The following describes the processing flow.

[0570] Step 1:

[0571] The device captures the user's voice. The user speaks into the smartphone, saying everyday greetings or conversations, such as "I'm a little tired today." The device's microphone receives this voice signal.

[0572] Step 2:

[0573] The device's voice recognition system converts the received audio signal into text information. The voice recognition engine analyzes this audio in real time and generates the sentence "I'm a little tired today" as text.

[0574] Step 3:

[0575] The device's emotion engine begins analyzing the voice to evaluate its characteristics. It analyzes the tone, speed, and intonation of the voice, and in this case, evaluates emotions such as "feeling tired."

[0576] Step 4:

[0577] The device sends text information and sentiment data to the server. The collected data is transferred to a cloud-based server using secure data communication methods.

[0578] Step 5:

[0579] The server analyzes the received text information using natural language processing techniques. The analysis evaluates whether the utterance is fluent and consistent in content.

[0580] Step 6:

[0581] The server analyzes the current emotional data and speech content by comparing it with past data. By comparing it to past normal states, it determines that "a particular level of fatigue is being recognized this time."

[0582] Step 7:

[0583] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. The notification message may include something like, "You appeared more fatigued than usual during today's conversation. We recommend checking on you."

[0584] Step 8:

[0585] The device provides the user with feedback based on the analysis results. For example, if there is a problem, it will display a message such as "Please rest well today." This allows the user to become more aware of their own health status.

[0586] (Example 2)

[0587] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] This system addresses the technical challenges of providing peace of mind to users and their families by closely monitoring the health and emotional state of users in remote locations through everyday conversations and promptly notifying them of any abnormalities. In particular, accurately detecting emotional abnormalities based on voice intonation and speech content is crucial.

[0589] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0590] In this invention, the server includes terminal means for receiving acoustic information transmitted by a user, recognition means for converting the acoustic information into linguistic information, and analysis means for analyzing the linguistic information and the intonation of the voice to evaluate the emotional state. This makes it possible to effectively monitor the user's emotions and health status and to quickly notify them if an abnormality is detected.

[0591] "Terminal means" refers to a device for receiving acoustic information transmitted by a user. This device may include microphones, sensors, and other similar devices.

[0592] "Recognition means" refers to technologies and algorithms for converting acoustic information into text. This may include speech recognition software and hardware.

[0593] "Analysis means" refers to methods for evaluating a user's emotional state by analyzing text information and speech intonation. This includes natural language processing and sentiment analysis algorithms.

[0594] "Determination means" refers to a mechanism for determining an anomaly based on the analyzed evaluation results and for issuing notifications based on those determination results. This may include an anomaly detection algorithm and a notification system.

[0595] A "notification system" refers to a system used to send warnings to the user's associates or supporters when an anomaly is detected. Email or messaging services are commonly used for this purpose.

[0596] This invention relates to a system that monitors the health and emotional state of users located remotely and notifies them of abnormalities as needed. This system mainly consists of terminal means and a server, and a detailed embodiment thereof is shown below.

[0597] Embodiment of terminal means

[0598] A terminal is a device equipped with a microphone and various sensors to receive speech from the user. The user's voice is first converted into text information by the terminal's speech recognition system. This conversion often utilizes general-purpose speech recognition software. Specifically, it employs a technology that combines a speech signal processing module and a conversion algorithm.

[0599] Server Embodiment

[0600] The server receives text information and speech intonation data transmitted from the terminal and uses analysis tools to evaluate the user's emotional state based on this information. This analysis utilizes natural language processing technology and sentiment analysis algorithms. Based on the evaluation results, a judgment tool identifies any abnormalities and, if necessary, sends a warning message to the user's associates or supporters using a notification tool.

[0601] Specific example

[0602] For example, when a user says to their device, "I'm a little tired today," this voice is instantly converted into text. Based on this text data and the tone of the voice, the device's emotion analysis engine detects "fatigue." If the server detects an anomaly, a message is sent saying, "We sensed fatigue in your conversation this morning; we recommend you check it."

[0603] Example of a prompt

[0604] "If the user is elderly, what kind of voice tone or text would indicate that they are experiencing fatigue?"

[0605] In this way, the entire system works together to realize advanced monitoring functions that provide peace of mind and security to users and their stakeholders. This configuration supports the technical scope defined in the claims and provides concrete, implementable details.

[0606] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0607] Step 1:

[0608] The device receives the audio signal emitted by the user via a microphone. The input is the user's raw voice, which is converted into digital audio data. Specifically, sound waves are converted into electrical signals, and then processed into a format that can be processed through digital signal processing.

[0609] Step 2:

[0610] The terminal's speech recognition means converts digital audio data into text information. The input is the digital audio data obtained in step 1, and the output is the converted text information. Specifically, this involves a process of sequentially converting speech into text using an acoustic model and a language model.

[0611] Step 3:

[0612] The device's emotion analysis engine analyzes text information and acoustic data to evaluate the emotional state. The input is the text information and acoustic data obtained in step 2, and the evaluation result as an emotional state is output. Specifically, an emotion score is calculated based on voice tone, speaking speed, and keywords in the text.

[0613] Step 4:

[0614] The terminal sends text information and emotional state data to the server. In this process, the input is the emotional analysis results and text data, and the server receives the data as output. Specifically, encrypted data packets are sent via network communication.

[0615] Step 5:

[0616] The server analyzes the received data. The input consists of text information and sentiment data received in step 4. The server uses natural language processing to evaluate the fluency and abnormalities of the speech and obtains the analysis results as output. This process involves keyword extraction and contextual analysis to detect abnormal patterns.

[0617] Step 6:

[0618] The server's judgment mechanism determines whether an anomaly exists based on the analysis results and decides whether notification is necessary. The input is the analysis results obtained in step 5, and the output is the judgment result. Specifically, it evaluates whether an anomaly exists by comparing it with predefined criteria, and if an anomaly is found, it activates the notification process.

[0619] Step 7:

[0620] If an anomaly is detected, the server sends a warning to the user's stakeholders or supporters using a notification mechanism. The input is the determination result obtained in step 6, and the output is the notification message to the stakeholders. Specifically, the message is sent via email or a messaging service.

[0621] (Application Example 2)

[0622] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0623] There is a need for a system that can effectively monitor the health and emotional state of elderly people remotely. However, conventional technology has struggled to accurately detect changes in emotions and to quickly notify family members or caregivers in the event of an anomaly. Therefore, the challenge is to provide a more advanced monitoring system that ensures the safety and security of users and enables a rapid response.

[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0625] In this invention, the server includes means for receiving voice information transmitted by a user in a remote location, means for converting the voice information into text information, and means for analyzing the text information and the tone of the voice to evaluate the fluency of speech and emotional state. This enables high-precision monitoring of the user's health and emotional state and accurate notification in the event of an abnormality.

[0626] "Remote location" refers to a place that is physically distant from the user, and the concept includes environments where observation and interaction take place via communication technology.

[0627] "User" refers to an individual whose health and emotional state are monitored through this system.

[0628] "Audio information" refers to digital or analog data that includes all words and sounds emitted by the user.

[0629] "Voice receiving means" refers to hardware or software used to acquire voice information emitted by a user and incorporate it into the system.

[0630] "Speech recognition means" refers to a technology or process for analyzing received speech information and converting it into corresponding text information.

[0631] "Text information" refers to data in string format converted by speech recognition technology, representing the content of the user's speech.

[0632] "Natural language processing means" refers to technologies or processes for analyzing text information and determining its meaning and context.

[0633] "Speech fluency" is an index that evaluates how natural and smooth a user's speech is, based on the analysis of text information.

[0634] "Emotional state" refers to the psychological or emotional state that can be inferred from the tone of the user's voice and the content of their speech.

[0635] "Anomaly detection means" refers to a technology or process used to determine whether there are abnormalities in a user's health or emotional state based on analyzed data.

[0636] A "notification method" is a technology or process for sending warnings or messages to pre-designated parties when an anomaly is detected.

[0637] The system for carrying out the present invention mainly consists of a terminal and a server. The terminal is equipped with a microphone and plays the role of capturing the voice of an elderly person as a means of receiving voice. For example, suppose the terminal receives a voice message from the user saying, "I haven't been sleeping well lately."

[0638] The speech recognition system within the device converts received speech information into text information. This conversion can utilize speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text information represents the user's spoken content in digital form.

[0639] Next, the device performs emotion analysis. This involves analyzing the tone of the voice using emotion analysis libraries such as IBM Watson Tone Analyzer and evaluating the emotional state. Based on this evaluation, emotional states such as "anxiety" and "stress" are inferred.

[0640] Text information and emotional state data are transmitted to the server in real time. The server performs natural language processing on this data to analyze speech fluency and emotional state. Furthermore, anomaly detection measures determine whether there are any abnormalities in the user's state based on the results of the data analysis. For example, if the emotional tone is lower than usual compared to past data, it is detected as an anomaly.

[0641] When an anomaly is detected, the server uses a notification system such as Firebase Cloud Messaging to send alert messages to pre-registered stakeholders. These messages may include specific notifications such as, "It is recommended that you check in based on recent conversations."

[0642] As an example, a user's everyday statement, such as "I'm feeling a little down today," is analyzed, and if the emotional tone is determined to be different from normal, a notification requesting confirmation is sent to family members or caregivers.

[0643] An example of a prompt for a generative AI model is, "Explain how to analyze the health and emotional state of elderly people from their everyday conversations and notify their families if any abnormalities are found." This prompt provides an accurate explanation of the system's processing and operation.

[0644] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0645] Step 1:

[0646] The device receives the user's voice as input. Specifically, the device's microphone captures the user's voice, and this voice information is taken into the device as a digital signal. The output is the digitized voice data.

[0647] Step 2:

[0648] The device's speech recognition system takes the audio data obtained in step 1 as input and converts it into text information. This process utilizes the Google Cloud Speech-to-Text API to perform data calculations that convert the audio signal into string-formatted data. The output is text information representing the user's speech.

[0649] Step 3:

[0650] The terminal performs sentiment analysis using the text information and audio data generated in step 2 as input. Specifically, IBM Watson Tone Analyzer is used to process the data and evaluate emotions based on the tone and content of the audio. The output is an evaluation result regarding the user's emotional state.

[0651] Step 4:

[0652] The device sends the text information and sentiment data obtained in step 3 to the server. This involves transferring data over the network. The server receives this data and prepares it for analysis.

[0653] Step 5:

[0654] The server performs natural language processing using the text information and sentiment data received in step 4 as input. The server analyzes the fluency and sentiment changes of the data and performs data calculations to determine whether or not there are any anomalies. The analysis results are obtained as output.

[0655] Step 6:

[0656] The server uses the analysis results from step 5 as input to perform anomaly detection. A comparison operation is performed to determine whether the data is anomaly by comparing it with past data. As a result, data in which anomalies have been detected is output.

[0657] Step 7:

[0658] If an anomaly is detected in step 6, the server sends an alert message to registered stakeholders using the notification system. Here, Firebase Cloud Messaging is used to ensure that the message is delivered quickly, resulting in a notification to the stakeholders.

[0659] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0660] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0661] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0662] [Fourth Embodiment]

[0663] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0664] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0665] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0666] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0667] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0668] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0669] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0670] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0671] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0672] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0675] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] Embodiments for carrying out the present invention will now be described. This system is designed to monitor the health status of elderly people in remote locations, and users can receive health checks through everyday conversations.

[0677] System Configuration

[0678] This system consists of a smartphone used by an elderly person (hereinafter referred to as the terminal) and a cloud-based analysis platform (hereinafter referred to as the server). The terminal is equipped with voice receiving and voice recognition means and inputs the voice spoken by the user. The server is equipped with natural language processing means and anomaly detection means and analyzes the voice data to evaluate the user's health status.

[0679] Details of the program's processing

[0680] 1. Voice capture and text conversion

[0681] The device captures the user's voice in real time. When the user speaks to the device, saying "Good morning," it receives this as an audio signal.

[0682] The speech recognition system instantly converts this audio signal into text information. For example, the phrase "Good morning" is converted directly into text.

[0683] 2. Data transmission and analysis

[0684] The terminal sends the converted text to the server. The server uses the received text to perform natural language processing.

[0685] The server analyzes the received text data and evaluates the fluency and contextual appropriateness of the speech. For example, it checks for hesitations and pronunciation errors.

[0686] 3. Anomaly detection and notification

[0687] The server detects abnormal patterns by comparing them with past data from a healthy state. If an abnormality is detected, it records the details and sends a warning to the user's family or caregiver via a notification system.

[0688] For example, you can send a message to your family saying, "We noticed an unusual pattern in your conversation this morning. We recommend you consult a medical professional as a precaution."

[0689] This system enables rapid anomaly detection by routinely collecting and analyzing data on the user's health status. Furthermore, because it is conducted through natural communication with elderly users, health monitoring can be seamlessly integrated into their daily lives. This approach allows elderly individuals to live their daily lives with peace of mind, and also provides peace of mind for their families.

[0690] The following describes the processing flow.

[0691] Step 1:

[0692] The device captures the voice spoken by the user. The user starts a normal conversation into the smartphone, and the device's microphone picks up the voice signal.

[0693] Step 2:

[0694] The device's voice recognition system converts the captured audio signal into text information. The voice recognition engine performs real-time audio analysis and converts it into text such as "Good morning."

[0695] Step 3:

[0696] The device sends the converted text information to the server. The text data is sent to a cloud-based server via the network.

[0697] Step 4:

[0698] The server analyzes the received text data using natural language processing techniques. The analysis checks for speech fluency, grammatical consistency, and conversational coherence.

[0699] Step 5:

[0700] The server detects anomalies by comparing the analysis results with past data. It analyzes the differences from past normal speech patterns and determines whether an anomaly exists.

[0701] Step 6:

[0702] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. For example, it might send a message such as, "An anomaly was detected in this morning's conversation. We recommend consulting a medical professional."

[0703] Step 7:

[0704] The device notifies the user of the analysis results as feedback. If there are no problems, it displays a message such as "Have a great day!"

[0705] (Example 1)

[0706] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0707] There is a challenge in remotely monitoring the health status of elderly people in real time. In particular, when elderly people stumble over their words or exhibit unusual voice patterns, prompt action is required, but there is a lack of efficient and natural methods to achieve this.

[0708] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0709] In this invention, the server includes voice input means, voice analysis means, language processing means, anomaly detection means, and notification means. This makes it possible to naturally monitor the health status of a user in a remote location through everyday conversation and to quickly notify them if an anomaly is detected.

[0710] "Voice input means" refers to a device or technology for receiving voice data from a user located in a remote location.

[0711] "Speech analysis means" refers to a device or technique that converts received audio material into textual information.

[0712] "Language processing means" refers to a device or method for analyzing textual information and evaluating the fluency and contextual appropriateness of speech.

[0713] An "anomaly detection method" is a device or technique for detecting anomalies based on evaluation results and comparing them with past health data.

[0714] "Notification means" refers to a device or method for sending a warning to the necessary recipients when an anomaly is detected.

[0715] This invention relates to a system for monitoring the health status of a user located in a remote location via voice. This system consists of the user's smartphone or other voice-enabled device (hereinafter referred to as the terminal) and an analysis platform on the cloud (hereinafter referred to as the server).

[0716] The device captures the user's voice using a microphone as a voice input method. A specific example is the microphone built into a smartphone. The device is equipped with a voice analysis system that converts the voice signal into text information. This is done using a speech recognition engine (e.g., Google Cloud Speech-to-Text API).

[0717] The server analyzes the received text information using language processing tools to evaluate the fluency and contextual appropriateness of the utterance. This evaluation utilizes a generative AI model (e.g., GPT-3). This analysis triggers an anomaly detection mechanism that detects deviations from normal speech patterns. If an anomaly is detected, the server sends a warning to the user's family or care staff via a notification mechanism. Communication takes place via email or a dedicated application.

[0718] For example, if a user casually says, "The weather was nice today, so I went for a walk," the content of that statement is analyzed to check for any abnormalities. In particular, the fluency of the speech and the consistency of the content are evaluated.

[0719] Examples of prompt statements include the following:

[0720] "How does this system support the daily lives of the elderly?"

[0721] "What are the advantages of health checks conducted through everyday conversation?"

[0722] This system allows elderly individuals to have their health monitored naturally through everyday conversations, and enables a rapid response if any abnormalities are detected.

[0723] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0724] Step 1:

[0725] Audio Capture

[0726] The device collects voice data from the user in real time. The input is the voice the user speaks into the device's microphone. The device captures this voice using a voice input device and transmits it as an audio signal to an internal computer. This operation usually occurs without the user pressing any specific trigger.

[0727] Step 2:

[0728] Speech recognition and text conversion

[0729] The device converts the audio signal into text data. The input for this conversion is the audio signal acquired in step 1. A speech recognition engine (e.g., Google Cloud Speech-to-Text API) is used as the speech analysis tool to convert the audio signal into text information. This conversion results in the audio "Good morning" being output as the text string "Good morning".

[0730] Step 3:

[0731] Send text

[0732] The terminal sends the character information converted in step 2 to the server. The input is the character data obtained by speech recognition, and the output is the transmission of data to the server. The terminal uploads the character data to the server via an internet connection and applies encrypted communication to ensure the security of the data.

[0733] Step 4:

[0734] Natural Language Processing and Analysis

[0735] The server applies natural language processing to the received text information to evaluate the fluency and contextual appropriateness of the utterance. The input is the text information received from step 3. A generative AI model (e.g., GPT-3) is used to analyze the text and evaluate the naturalness and semantic content of the utterance. The output is the evaluation result, which determines whether there are any abnormalities in the utterance.

[0736] Step 5:

[0737] Anomaly detection and notification

[0738] The server determines whether or not there is an abnormality based on the evaluation results obtained in step 4. The input is the evaluation results after analysis, and the abnormality detection means compares it with past health data to identify abnormal patterns. If an abnormality is detected, the server records the information and generates a warning message as output. Using the notification means, it sends warnings via email or app to the user's family or care staff as needed.

[0739] (Application Example 1)

[0740] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0741] There is a need to effectively monitor the health status of elderly people living in remote locations and to respond quickly when abnormalities occur. However, conventional methods often require the wearing of devices or special operation, which can be burdensome for the elderly, making them resistant to such methods. Therefore, there is a need to develop a system that monitors health status through natural everyday conversation and promptly notifies users when abnormalities occur.

[0742] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0743] In this invention, the server includes voice input means for receiving the user's voice in real time, voice recognition means for converting the voice into text format, natural language processing means for analyzing the text data and evaluating the health status, and notification means for detecting abnormalities based on the evaluation results and sending warnings. This makes it possible to monitor the health status of the elderly while reducing their burden, and to effectively detect and notify abnormalities.

[0744] "Voice input means" refers to a device or process for receiving voice signals from a user in real time.

[0745] "Speech recognition means" refers to a device or software that has the function of converting received speech into digital text.

[0746] A "natural language processing tool" is a program or process that analyzes text data and evaluates the user's health status based on its content.

[0747] A "notification system" is a mechanism that issues an alarm when an abnormality is detected in the user's health condition and communicates the warning to the relevant parties.

[0748] The system implementing this invention utilizes voice analysis technology that allows users to monitor their health status through everyday conversations. This system mainly consists of three sections.

[0749] First, the device receives the user's everyday conversation in real time through a "voice input method." The smartphone's microphone is used as the voice input method. The device converts the recognized voice into digital text using a "speech recognition method." Cloud-based services such as the Google Cloud Speech-to-Text API can be used for speech recognition.

[0750] Next, the server uses "natural language processing tools" to process the converted text data and evaluate the user's health status. The analysis employs a natural language processing (NLP) model built using the Hugging Face Transformers library. This evaluates the user's utterances and speech fluency, and detects abnormal patterns.

[0751] Finally, if the server detects an anomaly, it will use a "notification system" to send an alert to the relevant parties. This notification utilizes the Twilio API to send SMS or emails to family members and caregivers. For example, if a user says, "I'm not feeling very well today," this is considered an abnormal pattern, and a message is sent to the family saying, "Your parent is showing an unusual health pattern. Please check on them."

[0752] Such a system allows for monitoring of users' health status through natural, everyday conversations, providing an environment where elderly people can continue to live with peace of mind. Furthermore, it enables prompt and appropriate notification to relevant parties, allowing for early intervention.

[0753] An example of a prompt that utilizes a generative AI model is: "Please describe the steps to analyze the user's voice and check for any abnormalities in their health condition."

[0754] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0755] Step 1:

[0756] The device receives the user's voice in real time through a "voice input means." This input is voice data captured using the smartphone's microphone. The collected voice data is then sent directly to the next processing step.

[0757] Step 2:

[0758] The device converts the received audio data into digital text using a "speech recognition tool." The Google Cloud Speech-to-Text API is used for the speech recognition process. In this process, the audio data is converted into natural language text data, and the converted text information is transferred to the server.

[0759] Step 3:

[0760] The server analyzes the received text information using "natural language processing tools." It uses the Hugging Face Transformers library to analyze the meaning of the text and assess the health status. This analysis identifies abnormal patterns based on keywords and context within the text, obtaining health status assessment data.

[0761] Step 4:

[0762] The server uses the analysis results to determine if there is an abnormality in the health status. This determination is made by comparing it with past health status data, and if an abnormality is detected, the server determines the specific nature of the abnormality.

[0763] Step 5:

[0764] The server immediately issues an alarm if an anomaly is detected by the "notification mechanism." Specifically, it uses the Twilio API to send warning messages via SMS or email to the user's family or caregivers. In this way, it supports relevant parties in responding quickly.

[0765] This enables a system that monitors users' health status through natural, everyday conversations and allows for rapid response in case of abnormalities.

[0766] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0767] This invention relates to a system for monitoring the health and emotional state of elderly individuals in remote locations. The system aims to support the user's peace of mind and safety by detecting abnormalities in health and emotional state through the user's daily conversations and providing appropriate notifications.

[0768] System Configuration

[0769] This system consists of a terminal used by the user and a server that processes and analyzes the data. The terminal is equipped with voice receiving and voice recognition capabilities, capturing the user's voice and converting it into text information in real time. It also incorporates an emotion engine that analyzes the user's emotions from their voice tone and speech content. The converted text information and emotion data are sent to the server.

[0770] Details of the program processing

[0771] 1. Voice capture and text conversion

[0772] The device receives the user's voice via its microphone. The user says to the device, "I'm a little tired today."

[0773] The speech recognition system converts this speech signal into text information and generates the text, "I'm a little tired today."

[0774] 2. Analysis of emotions

[0775] The device's emotion engine analyzes the tone of voice (for example, if the way of speaking is heavier or slower than usual) and evaluates the emotional state. In this case, it determines that the user is experiencing fatigue or stress.

[0776] 3. Data transmission and analysis

[0777] The device sends text information and sentiment data to the server.

[0778] The server uses natural language processing to analyze text information and evaluate the fluency and abnormalities of speech.

[0779] 4. Anomaly detection and notification

[0780] The server determines anomalies based on the analysis results and emotional assessment. Here, the emotional engine detects an anomaly based on the assessment that the user is "more tired than usual."

[0781] If an anomaly is detected, a warning message will be sent to the user's family or caregiver using the notification system. For example, a message such as, "You seemed tired during our conversation this morning, so we recommend checking in," might be sent.

[0782] This system allows for comprehensive monitoring of users' health and emotions through everyday conversations, enabling prompt responses to any abnormalities. This allows users and their families to live their daily lives with greater peace of mind.

[0783] The following describes the processing flow.

[0784] Step 1:

[0785] The device captures the user's voice. The user speaks into the smartphone, saying everyday greetings or conversations, such as "I'm a little tired today." The device's microphone receives this voice signal.

[0786] Step 2:

[0787] The device's voice recognition system converts the received audio signal into text information. The voice recognition engine analyzes this audio in real time and generates the sentence "I'm a little tired today" as text.

[0788] Step 3:

[0789] The device's emotion engine begins analyzing the voice to evaluate its characteristics. It analyzes the tone, speed, and intonation of the voice, and in this case, evaluates emotions such as "feeling tired."

[0790] Step 4:

[0791] The device sends text information and sentiment data to the server. The collected data is transferred to a cloud-based server using secure data communication methods.

[0792] Step 5:

[0793] The server analyzes the received text information using natural language processing techniques. The analysis evaluates whether the utterance is fluent and consistent in content.

[0794] Step 6:

[0795] The server analyzes the current emotional data and speech content by comparing it with past data. By comparing it to past normal states, it determines that "a particular level of fatigue is being recognized this time."

[0796] Step 7:

[0797] If the server detects an anomaly, it will use a notification system to send a warning to the user's family or caregiver. The notification message may include something like, "You appeared more fatigued than usual during today's conversation. We recommend checking on you."

[0798] Step 8:

[0799] The device provides the user with feedback based on the analysis results. For example, if there is a problem, it will display a message such as "Please rest well today." This allows the user to become more aware of their own health status.

[0800] (Example 2)

[0801] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0802] This system addresses the technical challenges of providing peace of mind to users and their families by closely monitoring the health and emotional state of users in remote locations through everyday conversations and promptly notifying them of any abnormalities. In particular, accurately detecting emotional abnormalities based on voice intonation and speech content is crucial.

[0803] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0804] In this invention, the server includes terminal means for receiving acoustic information transmitted by a user, recognition means for converting the acoustic information into linguistic information, and analysis means for analyzing the linguistic information and the intonation of the voice to evaluate the emotional state. This makes it possible to effectively monitor the user's emotions and health status and to quickly notify them if an abnormality is detected.

[0805] "Terminal means" refers to a device for receiving acoustic information transmitted by a user. This device may include microphones, sensors, and other similar devices.

[0806] "Recognition means" refers to technologies and algorithms for converting acoustic information into text. This may include speech recognition software and hardware.

[0807] "Analysis means" refers to methods for evaluating a user's emotional state by analyzing text information and speech intonation. This includes natural language processing and sentiment analysis algorithms.

[0808] "Determination means" refers to a mechanism for determining an anomaly based on the analyzed evaluation results and for issuing notifications based on those determination results. This may include an anomaly detection algorithm and a notification system.

[0809] A "notification system" refers to a system used to send warnings to the user's associates or supporters when an anomaly is detected. Email or messaging services are commonly used for this purpose.

[0810] This invention relates to a system that monitors the health and emotional state of users located remotely and notifies them of abnormalities as needed. This system mainly consists of terminal means and a server, and a detailed embodiment thereof is shown below.

[0811] Embodiment of terminal means

[0812] A terminal is a device equipped with a microphone and various sensors to receive speech from the user. The user's voice is first converted into text information by the terminal's speech recognition system. This conversion often utilizes general-purpose speech recognition software. Specifically, it employs a technology that combines a speech signal processing module and a conversion algorithm.

[0813] Server Embodiment

[0814] The server receives text information and speech intonation data transmitted from the terminal and uses analysis tools to evaluate the user's emotional state based on this information. This analysis utilizes natural language processing technology and sentiment analysis algorithms. Based on the evaluation results, a judgment tool identifies any abnormalities and, if necessary, sends a warning message to the user's associates or supporters using a notification tool.

[0815] Specific example

[0816] For example, when a user says to their device, "I'm a little tired today," this voice is instantly converted into text. Based on this text data and the tone of the voice, the device's emotion analysis engine detects "fatigue." If the server detects an anomaly, a message is sent saying, "We sensed fatigue in your conversation this morning; we recommend you check it."

[0817] Example of a prompt

[0818] "If the user is elderly, what kind of voice tone or text would indicate that they are experiencing fatigue?"

[0819] In this way, the entire system works together to realize advanced monitoring functions that provide peace of mind and security to users and their stakeholders. This configuration supports the technical scope defined in the claims and provides concrete, implementable details.

[0820] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0821] Step 1:

[0822] The device receives the audio signal emitted by the user via a microphone. The input is the user's raw voice, which is converted into digital audio data. Specifically, sound waves are converted into electrical signals, and then processed into a format that can be processed through digital signal processing.

[0823] Step 2:

[0824] The terminal's speech recognition means converts digital audio data into text information. The input is the digital audio data obtained in step 1, and the output is the converted text information. Specifically, this involves a process of sequentially converting speech into text using an acoustic model and a language model.

[0825] Step 3:

[0826] The device's emotion analysis engine analyzes text information and acoustic data to evaluate the emotional state. The input is the text information and acoustic data obtained in step 2, and the evaluation result as an emotional state is output. Specifically, an emotion score is calculated based on voice tone, speaking speed, and keywords in the text.

[0827] Step 4:

[0828] The terminal sends text information and emotional state data to the server. In this process, the input is the emotional analysis results and text data, and the server receives the data as output. Specifically, encrypted data packets are sent via network communication.

[0829] Step 5:

[0830] The server analyzes the received data. The input consists of text information and sentiment data received in step 4. The server uses natural language processing to evaluate the fluency and abnormalities of the speech and obtains the analysis results as output. This process involves keyword extraction and contextual analysis to detect abnormal patterns.

[0831] Step 6:

[0832] The server's judgment mechanism determines whether an anomaly exists based on the analysis results and decides whether notification is necessary. The input is the analysis results obtained in step 5, and the output is the judgment result. Specifically, it evaluates whether an anomaly exists by comparing it with predefined criteria, and if an anomaly is found, it activates the notification process.

[0833] Step 7:

[0834] If an anomaly is detected, the server sends a warning to the user's stakeholders or supporters using a notification mechanism. The input is the determination result obtained in step 6, and the output is the notification message to the stakeholders. Specifically, the message is sent via email or a messaging service.

[0835] (Application Example 2)

[0836] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0837] There is a need for a system that can effectively monitor the health and emotional state of elderly people remotely. However, conventional technology has struggled to accurately detect changes in emotions and to quickly notify family members or caregivers in the event of an anomaly. Therefore, the challenge is to provide a more advanced monitoring system that ensures the safety and security of users and enables a rapid response.

[0838] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0839] In this invention, the server includes means for receiving voice information transmitted by a user in a remote location, means for converting the voice information into text information, and means for analyzing the text information and the tone of the voice to evaluate the fluency of speech and emotional state. This enables high-precision monitoring of the user's health and emotional state and accurate notification in the event of an abnormality.

[0840] "Remote location" refers to a place that is physically distant from the user, and the concept includes environments where observation and interaction take place via communication technology.

[0841] "User" refers to an individual whose health and emotional state are monitored through this system.

[0842] "Audio information" refers to digital or analog data that includes all words and sounds emitted by the user.

[0843] "Voice receiving means" refers to hardware or software used to acquire voice information emitted by a user and incorporate it into the system.

[0844] "Speech recognition means" refers to a technology or process for analyzing received speech information and converting it into corresponding text information.

[0845] "Text information" refers to data in string format converted by speech recognition technology, representing the content of the user's speech.

[0846] "Natural language processing means" refers to technologies or processes for analyzing text information and determining its meaning and context.

[0847] "Speech fluency" is an index that evaluates how natural and smooth a user's speech is, based on the analysis of text information.

[0848] "Emotional state" refers to the psychological or emotional state that can be inferred from the tone of the user's voice and the content of their speech.

[0849] "Anomaly detection means" refers to a technology or process used to determine whether there are abnormalities in a user's health or emotional state based on analyzed data.

[0850] A "notification method" is a technology or process for sending warnings or messages to pre-designated parties when an anomaly is detected.

[0851] The system for carrying out the present invention mainly consists of a terminal and a server. The terminal is equipped with a microphone and plays the role of capturing the voice of an elderly person as a means of receiving voice. For example, suppose the terminal receives a voice message from the user saying, "I haven't been sleeping well lately."

[0852] The speech recognition system within the device converts received speech information into text information. This conversion can utilize speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text information represents the user's spoken content in digital form.

[0853] Next, the device performs emotion analysis. This involves analyzing the tone of the voice using emotion analysis libraries such as IBM Watson Tone Analyzer and evaluating the emotional state. Based on this evaluation, emotional states such as "anxiety" and "stress" are inferred.

[0854] Text information and emotional state data are transmitted to the server in real time. The server performs natural language processing on this data to analyze speech fluency and emotional state. Furthermore, anomaly detection measures determine whether there are any abnormalities in the user's state based on the results of the data analysis. For example, if the emotional tone is lower than usual compared to past data, it is detected as an anomaly.

[0855] When an anomaly is detected, the server uses a notification system such as Firebase Cloud Messaging to send alert messages to pre-registered stakeholders. These messages may include specific notifications such as, "It is recommended that you check in based on recent conversations."

[0856] As an example, a user's everyday statement, such as "I'm feeling a little down today," is analyzed, and if the emotional tone is determined to be different from normal, a notification requesting confirmation is sent to family members or caregivers.

[0857] An example of a prompt for a generative AI model is, "Explain how to analyze the health and emotional state of elderly people from their everyday conversations and notify their families if any abnormalities are found." This prompt provides an accurate explanation of the system's processing and operation.

[0858] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0859] Step 1:

[0860] The device receives the user's voice as input. Specifically, the device's microphone captures the user's voice, and this voice information is taken into the device as a digital signal. The output is the digitized voice data.

[0861] Step 2:

[0862] The device's speech recognition system takes the audio data obtained in step 1 as input and converts it into text information. This process utilizes the Google Cloud Speech-to-Text API to perform data calculations that convert the audio signal into string-formatted data. The output is text information representing the user's speech.

[0863] Step 3:

[0864] The terminal performs sentiment analysis using the text information and audio data generated in step 2 as input. Specifically, IBM Watson Tone Analyzer is used to process the data and evaluate emotions based on the tone and content of the audio. The output is an evaluation result regarding the user's emotional state.

[0865] Step 4:

[0866] The device sends the text information and sentiment data obtained in step 3 to the server. This involves transferring data over the network. The server receives this data and prepares it for analysis.

[0867] Step 5:

[0868] The server performs natural language processing using the text information and sentiment data received in step 4 as input. The server analyzes the fluency and sentiment changes of the data and performs data calculations to determine whether or not there are any anomalies. The analysis results are obtained as output.

[0869] Step 6:

[0870] The server uses the analysis results from step 5 as input to perform anomaly detection. A comparison operation is performed to determine whether the data is anomaly by comparing it with past data. As a result, data in which anomalies have been detected is output.

[0871] Step 7:

[0872] If an anomaly is detected in step 6, the server sends an alert message to registered stakeholders using the notification system. Here, Firebase Cloud Messaging is used to ensure that the message is delivered quickly, resulting in a notification to the stakeholders.

[0873] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0874] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0875] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0876] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0877] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0878] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0879] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0880] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0881] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0882] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0883] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0884] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0885] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0886] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0887] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0888] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0889] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0890] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0891] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0892] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0893] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0894] The following is further disclosed regarding the embodiments described above.

[0895] (Claim 1)

[0896] A voice receiving means for receiving voice signals transmitted by a user in a remote location,

[0897] A speech recognition means that converts the aforementioned audio signal into text information,

[0898] A natural language processing means that analyzes the aforementioned text information and evaluates the fluency of the utterance,

[0899] An anomaly detection means that detects and notifies an anomaly based on the evaluation results,

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, characterized in that the anomaly detection means compares speech patterns based on past evaluation results.

[0903] (Claim 3)

[0904] The system according to claim 1, characterized in that the notification means sends a warning to the user's family or caregiver when an abnormality is detected.

[0905] "Example 1"

[0906] (Claim 1)

[0907] A voice input method for receiving audio materials transmitted by a user in a remote location,

[0908] A speech analysis means for converting the aforementioned audio material into text information,

[0909] A language processing means that analyzes the aforementioned textual information and evaluates the fluency and contextual appropriateness of the utterance,

[0910] An anomaly detection means that detects anomalies based on the evaluation results and compares them with past health data,

[0911] A notification means that records the aforementioned abnormality and issues a warning as necessary,

[0912] A system that includes this.

[0913] (Claim 2)

[0914] The system according to claim 1, characterized in that the anomaly detection means performs analysis using a generated AI model and acquires new information.

[0915] (Claim 3)

[0916] The system according to claim 1, characterized in that the notification means communicates a warning to the user's relatives or care staff when an abnormal pattern is detected.

[0917] "Application Example 1"

[0918] (Claim 1)

[0919] A voice input method that receives the user's voice in real time,

[0920] A speech recognition means for converting the audio into text format,

[0921] A natural language processing means for analyzing the aforementioned text data and evaluating the health status,

[0922] A notification means that detects an anomaly based on the evaluation result and sends a warning,

[0923] A system that includes this.

[0924] (Claim 2)

[0925] The system according to claim 1, characterized in that the natural language processing means determines the health status based on the content of the utterance and identifies abnormal patterns.

[0926] (Claim 3)

[0927] The system according to claim 1, characterized in that the notification means transmits an alarm via an external device or external network when an abnormality is detected.

[0928] "Example 2 of combining an emotion engine"

[0929] (Claim 1)

[0930] A terminal means for receiving audio information transmitted by a user,

[0931] A recognition means for converting the aforementioned acoustic information into linguistic information,

[0932] An analysis means for analyzing the aforementioned linguistic information and speech intonation to evaluate the emotional state,

[0933] A determination means that determines and notifies of an abnormality based on the evaluation results,

[0934] A system that includes this.

[0935] (Claim 2)

[0936] The system according to claim 1, characterized in that the determination means compares speech and emotion patterns based on past evaluation results.

[0937] (Claim 3)

[0938] The system according to claim 1, characterized in that the notification means sends a warning to a person related to or supporting the user when it detects an abnormality.

[0939] "Application example 2 when combining with an emotional engine"

[0940] (Claim 1)

[0941] A voice receiving means for receiving voice information transmitted by a user in a remote location,

[0942] A speech recognition means that converts the aforementioned speech information into text information,

[0943] A natural language processing means that analyzes the aforementioned text information and speech tone to evaluate speech fluency and emotional state,

[0944] An anomaly detection means that detects an anomaly based on the evaluation results and notifies the registered contact person,

[0945] A system that includes this.

[0946] (Claim 2)

[0947] The system according to claim 1, characterized in that the anomaly detection means compares speech patterns and emotional patterns based on past evaluation results.

[0948] (Claim 3)

[0949] The system according to claim 1, characterized in that the notification means sends a warning to the user's relevant parties when an abnormality is detected. [Explanation of symbols]

[0950] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A voice input method that receives the user's voice in real time, A speech recognition means for converting the audio into text format, A natural language processing means for analyzing the aforementioned text data and evaluating the health status, A notification means that detects an anomaly based on the evaluation result and sends a warning, A system that includes this.

2. The system according to claim 1, characterized in that the natural language processing means determines the health status based on the content of the utterance and identifies abnormal patterns.

3. The system according to claim 1, characterized in that the notification means transmits an alarm via an external device or external network when an abnormality is detected.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A