system

The integrated system addresses the challenges of caring for dementia patients by using sensor data, eye movement, and audio analysis to provide personalized care measures, enhancing care quality and reducing caregiver burden.

JP2026074850APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Caring for dementia patients is burdensome due to their diverse behaviors and psychological symptoms, and existing methods struggle to accurately grasp the progression of dementia, leading to deteriorating care quality and impaired quality of life.

Method used

A system that integrates sensor data analysis, eye movement evaluation, audio data analysis, and generative AI to provide personalized care measures by detecting abnormal conditions, assessing cognitive function, and generating synthesized content tailored to individual needs.

Benefits of technology

The system improves care quality by accurately understanding dementia progression and reducing caregiver burden through real-time monitoring and personalized responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074850000001_ABST
    Figure 2026074850000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for receiving sensor data and detecting abnormal conditions, A means of acquiring eye movement data and evaluating cognitive function, A means for analyzing audio data and identifying specific speech patterns, A means of generating synthetic content using past visual information, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Caring for dementia patients is a great burden for caregivers because their behaviors and psychological symptoms are diverse. Also, with conventional methods, it has been difficult to accurately grasp the progression of dementia and promptly provide appropriate countermeasures based on it. As a result, there has been a problem that the quality of care deteriorates and the quality of life of the care recipients is also impaired.

Means for Solving the Problems

[0005] This invention provides a system comprising means for receiving sensor data and detecting abnormal conditions, means for evaluating cognitive function from eye movements, means for analyzing audio data and identifying speech patterns, and means for generating synthesized content using past visual information. This allows for a more accurate understanding of the condition of dementia patients and provides caregivers with appropriate countermeasures tailored to each individual's condition, thereby reducing the burden of care and improving the quality of care.

[0006] "Sensor data" refers to data acquired by sensors used to measure the state of the environment and living organisms, and includes temperature, humidity, heart rate, blood pressure, etc.

[0007] "Detecting abnormal conditions" refers to the process of identifying data that deviates from normal patterns and notifying users or systems.

[0008] "Eye movement data" refers to data obtained by tracking the eye movements and fixation points of the person receiving care, and cognitive function is evaluated based on this data.

[0009] "Cognitive function assessment" is a method of quantifying and scoring a care recipient's cognitive abilities using eye movements and other data.

[0010] "Audio data" refers to sound signals acquired through a microphone, including conversations and spoken words.

[0011] "Identification of speech patterns" is the process of extracting specific linguistic or phonetic features from speech data to determine the progression and state of dementia.

[0012] "Past visual information" refers to visual data such as photographs and videos related to events and memories experienced by the person receiving care in the past.

[0013] "Synthetic content" refers to new visual content that has been processed or generated using AI or other technologies, and is based on past visual information. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system designed for dementia patients that integrates multiple data collection and analysis technologies. This system aggregates sensor data, eye movement data, and voice data, and analyzes them using AI to suggest appropriate countermeasures to caregivers.

[0036] Sensor data collection and analysis

[0037] The device acquires data from multiple environmental sensors and wearable devices. This includes the care recipient's heart rate and blood pressure, as well as room temperature and humidity. The data is immediately sent to the server, and if any abnormal values ​​are detected, an alert is generated in real time.

[0038] Assessment of cognitive function

[0039] The user uses a dedicated eye-tracking device. This device tracks the eye movements of the person being cared for while they perform specific tasks and provides this data to a server. The server then applies an algorithm to score cognitive function based on the data obtained from these movements.

[0040] Analysis of audio data

[0041] The device records the care recipient's conversational audio via a smartphone app. The audio data is sent to a server, where an AI model identifies speech patterns and changes. For example, it can detect speech fluency and vocabulary changes.

[0042] Generating Synthetic Content

[0043] The server uses user-provided past photos and related visual data to create synthesized content using generative AI. This synthesized content helps care recipients visually re-experience past memories. This process contributes to emotional stability and promotes better communication.

[0044] As a concrete example, if a person receiving care is experiencing stress at a particular time, the system infers the cause from sensor data and voice data. For instance, if an increase in heart rate coincides with changes in voice, it presents music or a synthesized story related to past memories to alleviate the condition. Then, through automated feedback mechanisms, the system also records user emotional reactions in real time and performs further analysis and adjustments as needed. This ensures the best possible care plan, reduces the burden on caregivers, and improves the quality of life for patients.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The device acquires vital data in real time from environmental sensors and wearable devices. This includes data such as temperature, humidity, heart rate, and blood pressure.

[0048] Step 2:

[0049] The terminal collects data, bundles it into data packets, and sends them to the server over the network. This transmission is performed periodically to ensure uninterrupted data flow.

[0050] Step 3:

[0051] The server stores the received data in a database. After storage, an AI model is used to perform initial data analysis and detect anomalies and patterns.

[0052] Step 4:

[0053] The user uses an eye-tracking device to perform a designated visual task. This generates data on the care recipient's eye movements.

[0054] Step 5:

[0055] The device instantly transmits eye movement data to the server. This data transmission is performed very quickly to ensure accuracy.

[0056] Step 6:

[0057] The server analyzes the received eye movement data and calculates a score to assess cognitive function. The assessed results become part of the feedback provided to the caregiver.

[0058] Step 7:

[0059] Users record voice responses using their smartphone's voice recording function. This includes answers to questions and everyday speech.

[0060] Step 8:

[0061] The terminal processes the recorded audio data, such as removing noise, before sending it to the server.

[0062] Step 9:

[0063] The server uses an AI model to analyze voice data in detail, identifying speech patterns and changes. This allows it to determine signs and progression of dementia.

[0064] Step 10:

[0065] The server uses generative AI to create synthesized content based on past photographs and visual data provided. This is a process that generates a visual story that is emotionally meaningful to the person receiving care.

[0066] Step 11:

[0067] The device presents the generated content to the user and records the user's response. The recorded data is used in the next feedback loop.

[0068] (Example 1)

[0069] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0070] In an aging society, improving the quality of life for dementia patients requires a system that accurately assesses the physical and mental condition of those receiving care and promptly provides appropriate care methods. However, current systems struggle to comprehensively analyze multiple data points and respond flexibly to the individual needs of each user.

[0071] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0072] In this invention, the server includes a configuration for receiving sensor information and detecting abnormal conditions, a configuration for acquiring gaze data and evaluating cognitive function, a configuration for analyzing voice information and identifying specific speech patterns, and a configuration for adjusting the generation algorithm using feedback obtained from the user. This enables more accurate monitoring of dementia patients and the provision of personalized care support measures.

[0073] "Sensor information" refers to information obtained from devices and equipment used to collect data about the environment and individual biological systems.

[0074] An "abnormal condition" refers to a physical or environmental situation that deviates from the normal range, and is a condition that requires particular attention or action, especially in the context of caregiving.

[0075] "Eye-gaze data" refers to data about an individual's visual focus and eye movements, and is important information used to evaluate cognitive function.

[0076] A "system for evaluating cognitive function" is a set of procedures and algorithms for analyzing eye-tracking data and other information to evaluate an individual's cognitive abilities and behavioral patterns.

[0077] "Vocal information" refers to data obtained from the speech and conversations of the person receiving care, and is fundamental information for understanding speech patterns.

[0078] "Speech patterns" refer to the characteristics and structure of an individual's way of speaking, and are used to identify specific changes or tendencies.

[0079] "Synthetic content" refers to content such as videos and stories generated based on past visual data, and is used with the aim of promoting emotional stability and memory recall in those receiving care.

[0080] A "generative algorithm" is a computational procedure that utilizes past data and feedback to create new content and suggestions.

[0081] "User feedback" refers to information about the reactions and evaluations that care recipients and caregivers give to the proposed synthetic content and system, and this information can be used to improve the system.

[0082] This invention is an advanced data analysis system designed to support dementia patients. The system aims to propose appropriate care strategies to caregivers by integrating information collected from various sensors and devices and performing in-depth analysis using artificial intelligence (AI).

[0083] Sensor data collection

[0084] The terminal uses environmental sensors and wearable devices to acquire vital information such as the care recipient's heart rate and blood pressure, as well as environmental information such as room temperature and humidity, in real time. Common wearable devices and sensor modules are used for this purpose. The collected data is immediately transmitted to a server, and an alert is generated if an abnormal condition is detected.

[0085] Assessment of cognitive function

[0086] The user tracks the eye movements of the person being cared for while they perform tasks on the screen, using a dedicated eye-tracking device. This device utilizes commonly used eye-tracking technology and sends the eye movement data to a server, which then uses this data to score cognitive function.

[0087] Analysis of audio data

[0088] The device uses a smartphone app to record the care recipient's conversations and sends the audio data to a server. The server analyzes the collected audio data using an AI model to identify specific speech patterns and changes. This makes it possible to understand the fluency of speech and changes in vocabulary in detail.

[0089] Generating Synthetic Content

[0090] The server utilizes past photos and related visual data provided by the user and develops synthesized content using a generative AI model. This content is used to allow the person receiving care to visually re-experience past memories, promoting emotional stability. For example, it can generate stories based on relaxing landscape images and verbalize them through prompts.

[0091] As a concrete example, based on the prompt "Generate relaxing composite content using past images that the care recipient finds comforting," it is possible to create a composite movie using past travel photos of the care recipient. Throughout this process, the system collects user feedback and uses it as data for continuous improvement.

[0092] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0093] Step 1:

[0094] The terminal acquires data from multiple environmental sensors and wearable devices. It receives data from sensors such as heart rate sensors and room temperature sensors as input and sends it to the server. The output is raw data from each sensor, sent to the server via Bluetooth or Wi-Fi. The terminal periodically checks whether the sensors are functioning correctly to prevent data loss.

[0095] Step 2:

[0096] The server receives sensor data transmitted from the terminal and performs analysis to detect anomalies. It receives heart rate and room temperature data from the sensor as input and compares it to normal range data recorded in the database. The output is a warning message if an anomaly is detected, which is notified to the caregiver's terminal in real time. The server improves the accuracy of anomaly detection through time-series analysis by comparing the current data with past data.

[0097] Step 3:

[0098] The user uses an eye-tracking device to collect data on the care recipient's eye movements. The input is the capture of gaze data while the care recipient is looking at a screen, and this data is sent to the server. The output is a cognitive function assessment score based on the gaze movements. The user periodically adjusts the device's position and calibration to ensure accurate tracking.

[0099] Step 4:

[0100] The server analyzes eye-tracking data sent by the user to assess cognitive function. It receives eye-tracking data as input and evaluates its movement using an analysis algorithm. Output includes an evaluation score and feedback for the caregiver. The server utilizes an AI-based model to identify eye-tracking patterns when performing different tasks.

[0101] Step 5:

[0102] The device records audio data via a smartphone app and sends the recording to a server. The input is captured audio of a conversation and saved in a digital format. The output is an audio data file, which is sent to the server for analysis. The device uses noise reduction filtering technology to improve recording quality.

[0103] Step 6:

[0104] The server receives audio data and analyzes it using an AI model. It receives audio data files as input and processes the data to identify speech patterns and changes. The output is the analysis result based on specific speech patterns, which is then used to provide notifications and advice to caregivers. The server identifies trends and changes by comparing the current audio data with past audio data.

[0105] Step 7:

[0106] The server generates synthetic content using a generative AI model. It utilizes past photographs and related visual data as input, and determines content based on prompt text. The output is generated as visualized video content or stories and delivered to the user. The server incorporates user feedback during the generation process to provide more personalized content.

[0107] (Application Example 1)

[0108] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0109] Effective care for dementia patients requires real-time monitoring of abnormal conditions and changes in cognitive function, and appropriate responses. However, it is difficult to provide prompt and accurate responses while reducing the burden on caregivers, and problems stemming from information overload and ambiguity in countermeasures are particularly pronounced in facilities with a large number of dementia patients. Therefore, support for flexible and efficient care methods tailored to individual patients is necessary.

[0110] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0111] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing voice data and identifying specific speech patterns, means for integrating information from multiple data sources in real time and generating guidelines for implementing proposed care methods, and means for communicating warnings and suggestions to caregivers using push notifications. This enables caregivers to understand the situation of dementia patients and provide prompt and accurate care.

[0112] "Sensor data" is a general term for multiple measurement pieces of information that indicate the environment and vital signs, and is dynamic data including body temperature, heart rate, room temperature, and humidity.

[0113] "Detection of abnormal conditions" is the act of finding unique patterns or signs that differ from normal conditions in the received sensor data.

[0114] "Eye movement data" refers to information that shows how an individual tracks and focuses on visual information, and is used to assess cognitive function.

[0115] "Cognitive function assessment" is a process of analyzing and scoring an individual's brain functions, such as language, judgment, and memory, using collected eye movement data.

[0116] "Voice data" refers to information that records individual conversations and utterances, and serves as fundamental data for analyzing language patterns and changes in speech.

[0117] "Identification of specific speech patterns" refers to the function of extracting and recognizing linguistic and rhythmic features within analyzed speech data.

[0118] "Past visual information" refers to photographs and videos related to the user's past, and is visual data used to activate memories.

[0119] "Generating synthetic content" is the process of creating new images and videos using computer generation technology based on past visual information.

[0120] "Push notification" is a communication method that transmits information from the sender to the recipient instantly at a specific time.

[0121] "Real-time information integration" refers to the process of immediately combining and analyzing information obtained from multiple data sources, thereby supporting rapid decision-making.

[0122] This invention is a care support system for dementia patients. The system is implemented with the following configuration.

[0123] The server receives data from multiple sensors at each terminal and uses this data to detect abnormal conditions. The sensors are mainly installed around the person being cared for and collect vital data and environmental data in real time. This provides a function that immediately sends an alert to the caregiver if there is an abnormality in heart rate or room temperature. Widely available Wi-Fi connected sensors are used as the hardware.

[0124] Users assess their cognitive function using a dedicated eye-tracking device. The eye-tracking device meticulously records the eye movements of the person being cared for, and this data is transmitted to a server. Based on this information, the server-side algorithm scores changes in cognitive function and provides feedback to the caregiver.

[0125] Additionally, voice data is collected by smart devices. The devices record everyday conversations and send the data to a server for voice analysis. An AI model in the cloud analyzes changes in speech patterns, which can serve as an indicator of dementia progression.

[0126] The server integrates this data and uses a generative AI model to suggest synthesized content to caregivers. The generated content serves as a guide for caregivers to implement the most appropriate individualized support for the patient. For example, if a patient has comforting memories from the past, images of those memories can be synthesized and displayed for stress reduction purposes.

[0127] For example, if a patient becomes restless in the afternoon, the system will detect an abnormal increase in heart rate during that time and notify the caregiver's device with a relaxing, synthesized video from the past. An example of this prompt message would be: "Please generate a calming synthesized story using past photos. This content should help the specific patient relax."

[0128] This allows caregivers to effectively understand the condition of dementia patients and provide real-time support tailored to their individual needs.

[0129] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0130] Step 1:

[0131] The device collects sensor data such as the care recipient's body temperature, heart rate, and room temperature in real time. It takes data from each sensor as input, organizes it within the device, and then sends it to the server. The output is a set of the organized sensor data sent to the server.

[0132] Step 2:

[0133] The server analyzes the received sensor data to detect abnormal conditions. The input is sensor data sent from the terminal, and abnormalities are visualized by comparing it with normal data stored in the database. If an abnormality is detected through this process, alert information is generated.

[0134] Step 3:

[0135] The user uses an eye-tracking device to capture the eye movements of the person being cared for. The eye movement data is recorded as input on the device and sent to the server. As output, the eye movement data is formatted and becomes data for analysis on the server.

[0136] Step 4:

[0137] The server evaluates cognitive function based on the received eye movement data. Using eye movement data as input, it applies an algorithm to calculate a cognitive function score. The output is the cognitive function score as the evaluation result, which is provided to the caregiver.

[0138] Step 5:

[0139] The device records everyday conversations and transfers the audio data to a server. The input is the voice of the person receiving care, which is recorded clearly and then sent to the server. The audio file is then formatted and used as material for analysis.

[0140] Step 6:

[0141] The server analyzes audio data and identifies specific speech patterns. Using audio data as input, it extracts language patterns using a generative AI model. The output generates hints for caregivers based on the identified speech patterns.

[0142] Step 7:

[0143] The server generates synthetic content using past visual information and generative AI. It analyzes the input visual data using the prompt: "Generate a calming synthetic story using past photographs. This content should help a specific patient relax." The output is the generated synthetic content, provided as text and images appropriate to the patient's condition.

[0144] Step 8:

[0145] The terminal receives synthesized content and alerts sent from the server and sends push notifications to the caregiver. The input is notification data from the server, which is immediately displayed on the caregiver's device to enable early response by the caregiver. The output is the alerts and synthesized content displayed on the actual device.

[0146] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0147] This invention provides caregivers with more effective response strategies by combining an emotion engine with a system designed to support the care of dementia patients. The system analyzes sensor data, eye movement data, and voice data, as well as performing emotion recognition and providing feedback based on the user's psychological state.

[0148] Collection and analysis of emotional data

[0149] The device monitors the user's facial expressions through its built-in camera and acquires data in real time. This data is input into an emotion engine to determine emotional states such as smiles, anger, and sadness.

[0150] Furthermore, the device records the user's voice tone and speaking speed, and uses these voice characteristics to estimate emotions and make more accurate emotional judgments.

[0151] Emotion-based content delivery

[0152] The server generates synthesized content tailored to the user's emotions based on the results obtained from the emotion engine. For example, if the user is feeling anxious, it will provide a visual story or music with a relaxing effect.

[0153] The device presents this content to the user, aiming for the displayed content to calm the user's state. Simultaneously, it collects user feedback again, continuing to form a feedback loop.

[0154] As a concrete example, suppose a user is watching television in the living room with an anxious expression. The system identifies this and, via the server, provides video and audio based on past experiences where the user felt reassured. As a result, the user's expression softens, and their heart rate and other vital data stabilize. Through this feedback mechanism, the system provides caregivers with an effective and consistent care strategy.

[0155] The following describes the processing flow.

[0156] Step 1:

[0157] The device uses its built-in camera to monitor the user's face and periodically captures facial expression data and features. This data includes the movement and position of each part of the face.

[0158] Step 2:

[0159] The device transmits the acquired facial expression data to the emotion engine. The emotion engine analyzes the received data and applies an algorithm to determine emotions such as smiles, surprise, and anger.

[0160] Step 3:

[0161] The device records the user's voice using a microphone and extracts the tone and speaking speed of the voice. This audio data is used to identify variations in the volume and tempo of speech.

[0162] Step 4:

[0163] The device inputs voice data into an emotion engine, which then estimates emotions from the voice characteristics. This allows for a more comprehensive identification of emotions.

[0164] Step 5:

[0165] The server identifies the user's emotional state based on the analysis results and generates appropriate synthesized content. It selects the appropriate visual story and music templates based on the emotional state.

[0166] Step 6:

[0167] The device presents synthesized content sent from the server to the user and observes the impact the content has on the user's emotions.

[0168] Step 7:

[0169] The device records the user's response again during the presentation and sends that data to the server for analysis.

[0170] Step 8:

[0171] Based on the newly obtained response data, the server adjusts the next proposed countermeasures and content, forming a feedback loop.

[0172] This process allows the system to continuously monitor users' emotions and provide care tailored to their individual needs.

[0173] (Example 2)

[0174] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0175] In an aging society, there is a need to provide efficient and effective psychological care for dementia patients. Traditional methods rely heavily on the caregiver's experience and intuition, making it difficult to respond appropriately to the patient's emotions and psychological state. Therefore, new technologies are needed to accurately recognize the patient's emotions and provide appropriate synthetic information to promote their psychological stability.

[0176] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0177] In this invention, the server includes means for collecting facial expression information and performing emotion recognition, means for analyzing voice characteristics and evaluating psychological state, and means for generating synthesized information based on the emotion recognition results. This makes it possible to provide appropriate care based on the patient's emotional state and promote psychological stability.

[0178] "Facial expression information" refers to visual data obtained from the user's face and is used to determine their emotional state.

[0179] "Emotion recognition" is a technology that analyzes facial expressions and voice characteristics to identify a user's emotional state.

[0180] "Vocal features" refer to data such as tone, speech rate, and volume obtained from speech, and are used to evaluate the user's psychological state.

[0181] "Synthetic information" refers to visual or auditory content generated in response to the user's emotional state, and is provided for the purpose of supporting psychological stabilization.

[0182] "Feedback" is the process of recording user responses again using sensors, evaluating the effectiveness of the system, and making adjustments as needed.

[0183] "Vital information" refers to data that indicates the user's physical condition, such as heart rate, blood pressure, and body temperature.

[0184] "Environmental information" refers to data that indicates external conditions that affect the user's psychological state, such as ambient noise, light intensity, and temperature.

[0185] This invention relates to a system for supporting the psychological care of dementia patients, aiming to monitor the user's emotional state in real time and provide appropriate care. This system can be implemented as follows:

[0186] Data collection and analysis

[0187] The device collects user facial expressions and voice characteristics using its built-in camera and microphone. Facial expressions are captured using image analysis software with real-time data processing capabilities to capture facial movements and subtle changes in facial muscles. Voice characteristics are analyzed using voice analysis software to measure tone, speed, and volume based on the collected voice data.

[0188] The server aggregates this data and uses a generative AI model to perform emotion recognition. This model includes various pre-trained emotion analysis algorithms and can accurately distinguish between diverse emotional states.

[0189] Content generation and feedback

[0190] Based on the results of emotion recognition, the server generates synthetic information appropriate to the user's psychological state. For example, if the user is showing signs of anxiety, it will create relaxing music or calming visuals. It is also possible to adjust the content based on past data.

[0191] The device provides the generated composite information to the user, presenting it through the screen and speaker. The device also records the user's new responses and sends them to the server, forming a feedback loop.

[0192] Specific example

[0193] For example, if a user appears bored, the system can detect this and generate and present images of refreshing natural scenery or music. This approach allows users to regain a sense of calm and improve their quality of daily life.

[0194] Examples of prompts for generative AI models

[0195] "Regarding the development of a psychological care support system for dementia patients, please explain the specific methods for sensing emotions and generating appropriate content based on those results."

[0196] This system configuration makes it possible to support dementia patients in leading more stable lives.

[0197] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0198] Step 1:

[0199] The device activates its built-in camera and microphone to capture the user's facial expressions and voice. Image data is acquired as facial information, and voice data is recorded as voice features. These form the initial dataset as sensor inputs. This dataset is transmitted to the server in real time.

[0200] Step 2:

[0201] The server processes the received facial information through image analysis software to detect the user's facial movements and changes in facial muscles, thereby determining their emotional state. The output of this process is a list of possible emotions the user may be exhibiting, each with a confidence score.

[0202] Step 3:

[0203] The server processes the audio data using speech analysis software. It analyzes speech features (tone, speed, volume) and evaluates the user's psychological state. This analysis complements the list of emotions obtained in the previous step, improving the accuracy of emotion recognition.

[0204] Step 4:

[0205] The server uses a generative AI model to generate synthetic content based on the aforementioned emotional states. This content is designed to provide relaxation and a sense of security. The input data is a list of emotions, and the output is synthetic visual and auditory content.

[0206] Step 5:

[0207] The terminal receives synthesized content sent from the server and presents it to the user. The content is output via the terminal's display and speakers. The user's new responses are again recorded by the terminal and sent back to the server for the next processing cycle to form a feedback loop.

[0208] (Application Example 2)

[0209] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0210] As society ages, caring for dementia patients has become a critical issue. In particular, accurately understanding the emotional and psychological changes of dementia patients and providing appropriate care accordingly is extremely difficult in care facilities. Solving this problem is essential.

[0211] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0212] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing audio data and identifying specific speech patterns, means for generating synthesized content using past visual information, and means for analyzing the emotions of dementia patients in real time and presenting care suggestions on a visual display device. This enables care staff to provide appropriate care based on the emotional state of dementia patients quickly and effectively.

[0213] "Sensor data" refers to information about the environment and the state of the object obtained from various sensors.

[0214] "Detecting abnormal conditions" means recognizing a state that is different from the normal state or a problematic state.

[0215] "Eye movement data" refers to information about visual movements such as gaze and blinking.

[0216] "Cognitive function" is a general term for functions involved in human intellectual activities, such as thinking, memory, judgment, and learning.

[0217] "Audio data" refers to physical sound signals that contain information about vocalization.

[0218] "Speech patterns" refer to a set of characteristics in speech, such as word choice, rhythm and speed of language use, etc.

[0219] "Visual information" refers to information that is represented as images or videos.

[0220] "Synthetic content" refers to visually and aurally appealing content created by combining multiple sources of information.

[0221] A "dementia patient" refers to a person who has a medical condition characterized primarily by memory loss and confused thinking.

[0222] "Real-time emotional analysis" means instantly judging emotions and immediately evaluating their changes.

[0223] A "visual display device" is a technological device used to visually represent information.

[0224] "Care proposals" refer to suggestions that indicate appropriate actions and countermeasures in caregiving.

[0225] This invention is a system for streamlining the care of dementia patients in nursing care facilities. The system analyzes the patient's emotions in real time using sensor data, eye movement data, and voice data. A specific embodiment of this system is described here.

[0226] The server receives sensor data, and if an abnormal condition is detected, immediate action is taken. Eye movements are tracked using a camera, and cognitive function is evaluated using visual information displayed on smart glasses or a headset. Speech recognition software analyzes the patient's voice data and identifies specific speech patterns. It also generates relaxing synthetic content based on previously collected visual information. This synthetic content includes music and visual stories.

[0227] The device uses an emotion analysis engine to assess the patient's emotions in real time and presents care suggestions based on the results to the caregiver via a visual display. For example, if the device detects that the patient is anxious, it will display appropriate care methods such as "Stay calmly by the patient's side and listen to them."

[0228] As a concrete example, let's consider its use in a nursing home. In this facility, care staff wear smart glasses, and if it detects that a dementia patient is irritable, it automatically suggests "playing calming music while taking a short walk with the patient." In this way, specific care based on the patient's emotional state is provided quickly.

[0229] An example of a prompt might be the instruction, "Design an application that tracks the emotional state of users in this facility in real time and provides appropriate advice to staff."

[0230] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0231] Step 1:

[0232] The server receives sensor data, eye movement data, and audio data from the terminal. This input data is treated as initial information for understanding the patient's condition and is stored in the server's database.

[0233] Step 2:

[0234] The server analyzes the received data. It detects abnormal conditions based on sensor data and evaluates cognitive function using eye movement data. Audio data is analyzed using speech recognition software to identify specific speech patterns. As a result of this analysis, various status indicators are output and recorded on the server.

[0235] Step 3:

[0236] The server uses a generative AI model to analyze the patient's emotional state. It provides the model with inputs such as abnormal conditions, cognitive function assessment results, and speech patterns, and determines the patient's emotions in real time. The resulting output represents the patient's emotional state, which forms the basis for creating care recommendations.

[0237] Step 4:

[0238] The server creates synthesized content based on the analysis results. Using past visual information as a reference, it generates music and visual stories best suited to the current emotional state. This content is then sent to the device.

[0239] Step 5:

[0240] The terminal displays the received synthesized content on a visual display device, visually showing care suggestions that are emotionally responsive to the caregiver. These suggestions encourage specific actions that are helpful in face-to-face care.

[0241] Step 6:

[0242] The caregiver, as the user, provides appropriate care to the patient based on the information obtained from the device. If necessary, additional data and feedback can be sent to the server via the device, enabling continuous improvement of the care provided.

[0243] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0244] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0245] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0246] [Second Embodiment]

[0247] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0248] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0249] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0250] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0251] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0252] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0253] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0254] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0255] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0256] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0257] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0258] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0259] This invention is a system designed for dementia patients that integrates multiple data collection and analysis technologies. This system aggregates sensor data, eye movement data, and voice data, and analyzes them using AI to suggest appropriate countermeasures to caregivers.

[0260] Sensor data collection and analysis

[0261] The device acquires data from multiple environmental sensors and wearable devices. This includes the care recipient's heart rate and blood pressure, as well as room temperature and humidity. The data is immediately sent to the server, and if any abnormal values ​​are detected, an alert is generated in real time.

[0262] Assessment of cognitive function

[0263] The user uses a dedicated eye-tracking device. This device tracks the eye movements of the person being cared for while they perform specific tasks and provides this data to a server. The server then applies an algorithm to score cognitive function based on the data obtained from these movements.

[0264] Analysis of audio data

[0265] The device records the care recipient's conversational audio via a smartphone app. The audio data is sent to a server, where an AI model identifies speech patterns and changes. For example, it can detect speech fluency and vocabulary changes.

[0266] Generating Synthetic Content

[0267] The server uses user-provided past photos and related visual data to create synthesized content using generative AI. This synthesized content helps care recipients visually re-experience past memories. This process contributes to emotional stability and promotes better communication.

[0268] As a concrete example, if a person receiving care is experiencing stress at a particular time, the system infers the cause from sensor data and voice data. For instance, if an increase in heart rate coincides with changes in voice, it presents music or a synthesized story related to past memories to alleviate the condition. Then, through automated feedback mechanisms, the system also records user emotional reactions in real time and performs further analysis and adjustments as needed. This ensures the best possible care plan, reduces the burden on caregivers, and improves the quality of life for patients.

[0269] The following describes the processing flow.

[0270] Step 1:

[0271] The device acquires vital data in real time from environmental sensors and wearable devices. This includes data such as temperature, humidity, heart rate, and blood pressure.

[0272] Step 2:

[0273] The terminal collects data, bundles it into data packets, and sends them to the server over the network. This transmission is performed periodically to ensure uninterrupted data flow.

[0274] Step 3:

[0275] The server stores the received data in a database. After storage, an AI model is used to perform initial data analysis and detect anomalies and patterns.

[0276] Step 4:

[0277] The user uses an eye-tracking device to perform a designated visual task. This generates data on the care recipient's eye movements.

[0278] Step 5:

[0279] The device instantly transmits eye movement data to the server. This data transmission is performed very quickly to ensure accuracy.

[0280] Step 6:

[0281] The server analyzes the received eye movement data and calculates a score to assess cognitive function. The assessed results become part of the feedback provided to the caregiver.

[0282] Step 7:

[0283] The user uses the voice recording function of the smartphone to record voice responses. This includes answers to questions and daily conversations.

[0284] Step 8:

[0285] After the terminal performs preprocessing such as noise removal on the recorded voice data, it transmits the data to the server.

[0286] Step 9: [[ID=1​​​​​​​​​​​​​​​​​​​​​​​​​​​​​In an aging society, improving the quality of life for dementia patients requires a system that accurately assesses the physical and mental condition of those receiving care and promptly provides appropriate care methods. However, current systems struggle to comprehensively analyze multiple data points and respond flexibly to the individual needs of each user.

[0295] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0296] In this invention, the server includes a configuration for receiving sensor information and detecting abnormal conditions, a configuration for acquiring gaze data and evaluating cognitive function, a configuration for analyzing voice information and identifying specific speech patterns, and a configuration for adjusting the generation algorithm using feedback obtained from the user. This enables more accurate monitoring of dementia patients and the provision of personalized care support measures.

[0297] "Sensor information" refers to information obtained from devices and equipment used to collect data about the environment and individual biological systems.

[0298] An "abnormal condition" refers to a physical or environmental situation that deviates from the normal range, and is a condition that requires particular attention or action, especially in the context of caregiving.

[0299] "Eye-gaze data" refers to data about an individual's visual focus and eye movements, and is important information used to evaluate cognitive function.

[0300] A "system for evaluating cognitive function" is a set of procedures and algorithms for analyzing eye-tracking data and other information to evaluate an individual's cognitive abilities and behavioral patterns.

[0301] "Vocal information" refers to data obtained from the speech and conversations of the person receiving care, and is fundamental information for understanding speech patterns.

[0302] "Speech pattern" refers to the characteristics and structure of an individual's way of speaking, and is used to identify specific changes and trends.

[0303] "Synthetic content" refers to content such as videos and stories generated based on past visual data, etc., and is used for the purpose of promoting the emotional stability and memory recall of care recipients.

[0304] "Generation algorithm" is a computational procedure for creating new content and proposals by utilizing past data and feedback.

[0305] "Feedback obtained from users" refers to information on the reactions and evaluations shown by care recipients and caregivers to synthetic content and system proposals, and is information that can be used to improve the system.

[0306] The present invention is an advanced data analysis system designed to support dementia patients. This system integrates information collected from various sensors and devices, and deeply analyzes it using artificial intelligence (AI), aiming to propose appropriate care measures to caregivers.

[0307] Collection of sensor data

[0308] The terminal uses environmental sensors and wearable devices to obtain vital information such as the heart rate and blood pressure of the care recipient, and environmental information such as the indoor temperature and humidity in real time. General wearable devices and sensor modules are used for this. The collected data is immediately transmitted to the server, and an alert is generated if an abnormal state is detected.

[0309] Evaluation of cognitive function

[0310] The user tracks the eye movements of the person being cared for while they perform tasks on the screen, using a dedicated eye-tracking device. This device utilizes commonly used eye-tracking technology and sends the eye movement data to a server, which then uses this data to score cognitive function.

[0311] Analysis of audio data

[0312] The device uses a smartphone app to record the care recipient's conversations and sends the audio data to a server. The server analyzes the collected audio data using an AI model to identify specific speech patterns and changes. This makes it possible to understand the fluency of speech and changes in vocabulary in detail.

[0313] Generating Synthetic Content

[0314] The server utilizes past photos and related visual data provided by the user and develops synthesized content using a generative AI model. This content is used to allow the person receiving care to visually re-experience past memories, promoting emotional stability. For example, it can generate stories based on relaxing landscape images and verbalize them through prompts.

[0315] As a concrete example, based on the prompt "Generate relaxing composite content using past images that the care recipient finds comforting," it is possible to create a composite movie using past travel photos of the care recipient. Throughout this process, the system collects user feedback and uses it as data for continuous improvement.

[0316] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0317] Step 1:

[0318] The terminal acquires data from multiple environmental sensors and wearable devices. It receives data from sensors such as heart rate sensors and room temperature sensors as input and sends it to the server. The output is raw data from each sensor, sent to the server via Bluetooth or Wi-Fi. The terminal periodically checks whether the sensors are functioning correctly to prevent data loss.

[0319] Step 2:

[0320] The server receives sensor data transmitted from the terminal and performs analysis to detect anomalies. It receives heart rate and room temperature data from the sensor as input and compares it to normal range data recorded in the database. The output is a warning message if an anomaly is detected, which is notified to the caregiver's terminal in real time. The server improves the accuracy of anomaly detection through time-series analysis by comparing the current data with past data.

[0321] Step 3:

[0322] The user uses an eye-tracking device to collect data on the care recipient's eye movements. The input is the capture of gaze data while the care recipient is looking at a screen, and this data is sent to the server. The output is a cognitive function assessment score based on the gaze movements. The user periodically adjusts the device's position and calibration to ensure accurate tracking.

[0323] Step 4:

[0324] The server analyzes eye-tracking data sent by the user to assess cognitive function. It receives eye-tracking data as input and evaluates its movement using an analysis algorithm. Output includes an evaluation score and feedback for the caregiver. The server utilizes an AI-based model to identify eye-tracking patterns when performing different tasks.

[0325] Step 5:

[0326] The device records audio data via a smartphone app and sends the recording to a server. The input is captured audio of a conversation and saved in a digital format. The output is an audio data file, which is sent to the server for analysis. The device uses noise reduction filtering technology to improve recording quality.

[0327] Step 6:

[0328] The server receives audio data and analyzes it using an AI model. It receives audio data files as input and processes the data to identify speech patterns and changes. The output is the analysis result based on specific speech patterns, which is then used to provide notifications and advice to caregivers. The server identifies trends and changes by comparing the current audio data with past audio data.

[0329] Step 7:

[0330] The server generates synthetic content using a generative AI model. It utilizes past photographs and related visual data as input, and determines content based on prompt text. The output is generated as visualized video content or stories and delivered to the user. The server incorporates user feedback during the generation process to provide more personalized content.

[0331] (Application Example 1)

[0332] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0333] Effective care for dementia patients requires real-time monitoring of abnormal conditions and changes in cognitive function, and appropriate responses. However, it is difficult to provide prompt and accurate responses while reducing the burden on caregivers, and problems stemming from information overload and ambiguity in countermeasures are particularly pronounced in facilities with a large number of dementia patients. Therefore, support for flexible and efficient care methods tailored to individual patients is necessary.

[0334] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0335] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing voice data and identifying specific speech patterns, means for integrating information from multiple data sources in real time and generating guidelines for implementing proposed care methods, and means for communicating warnings and suggestions to caregivers using push notifications. This enables caregivers to understand the situation of dementia patients and provide prompt and accurate care.

[0336] "Sensor data" is a general term for multiple measurement pieces of information that indicate the environment and vital signs, and is dynamic data including body temperature, heart rate, room temperature, and humidity.

[0337] "Detection of abnormal conditions" is the act of finding unique patterns or signs that differ from normal conditions in the received sensor data.

[0338] "Eye movement data" refers to information that shows how an individual tracks and focuses on visual information, and is used to assess cognitive function.

[0339] "Cognitive function assessment" is a process of analyzing and scoring an individual's brain functions, such as language, judgment, and memory, using collected eye movement data.

[0340] "Voice data" refers to information that records individual conversations and utterances, and serves as fundamental data for analyzing language patterns and changes in speech.

[0341] "Identification of specific speech patterns" refers to the function of extracting and recognizing linguistic and rhythmic features within analyzed speech data.

[0342] "Past visual information" refers to photographs and videos related to the user's past, and is visual data used to activate memories.

[0343] "Generating synthetic content" is the process of creating new images and videos using computer generation technology based on past visual information.

[0344] "Push notification" is a communication method that transmits information from the sender to the recipient instantly at a specific time.

[0345] "Real-time information integration" refers to the process of immediately combining and analyzing information obtained from multiple data sources, thereby supporting rapid decision-making.

[0346] This invention is a care support system for dementia patients. The system is implemented with the following configuration.

[0347] The server receives data from multiple sensors at each terminal and uses this data to detect abnormal conditions. The sensors are mainly installed around the person being cared for and collect vital data and environmental data in real time. This provides a function that immediately sends an alert to the caregiver if there is an abnormality in heart rate or room temperature. Widely available Wi-Fi connected sensors are used as the hardware.

[0348] Users assess their cognitive function using a dedicated eye-tracking device. The eye-tracking device meticulously records the eye movements of the person being cared for, and this data is transmitted to a server. Based on this information, the server-side algorithm scores changes in cognitive function and provides feedback to the caregiver.

[0349] Additionally, voice data is collected by smart devices. The devices record everyday conversations and send the data to a server for voice analysis. An AI model in the cloud analyzes changes in speech patterns, which can serve as an indicator of dementia progression.

[0350] The server integrates this data and uses a generative AI model to suggest synthesized content to caregivers. The generated content serves as a guide for caregivers to implement the most appropriate individualized support for the patient. For example, if a patient has comforting memories from the past, images of those memories can be synthesized and displayed for stress reduction purposes.

[0351] For example, if a patient becomes restless in the afternoon, the system will detect an abnormal increase in heart rate during that time and notify the caregiver's device with a relaxing, synthesized video from the past. An example of this prompt message would be: "Please generate a calming synthesized story using past photos. This content should help the specific patient relax."

[0352] This allows caregivers to effectively understand the condition of dementia patients and provide real-time support tailored to their individual needs.

[0353] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0354] Step 1:

[0355] The device collects sensor data such as the care recipient's body temperature, heart rate, and room temperature in real time. It takes data from each sensor as input, organizes it within the device, and then sends it to the server. The output is a set of the organized sensor data sent to the server.

[0356] Step 2:

[0357] The server analyzes the received sensor data to detect abnormal conditions. The input is sensor data sent from the terminal, and abnormalities are visualized by comparing it with normal data stored in the database. If an abnormality is detected through this process, alert information is generated.

[0358] Step 3:

[0359] The user uses an eye-tracking device to capture the eye movements of the person being cared for. The eye movement data is recorded as input on the device and sent to the server. As output, the eye movement data is formatted and becomes data for analysis on the server.

[0360] Step 4:

[0361] The server evaluates cognitive function based on the received eye movement data. Using eye movement data as input, it applies an algorithm to calculate a cognitive function score. The output is the cognitive function score as the evaluation result, which is provided to the caregiver.

[0362] Step 5:

[0363] The device records everyday conversations and transfers the audio data to a server. The input is the voice of the person receiving care, which is recorded clearly and then sent to the server. The audio file is then formatted and used as material for analysis.

[0364] Step 6:

[0365] The server analyzes audio data and identifies specific speech patterns. Using audio data as input, it extracts language patterns using a generative AI model. The output generates hints for caregivers based on the identified speech patterns.

[0366] Step 7:

[0367] The server generates synthetic content using past visual information and generative AI. It analyzes the input visual data using the prompt: "Generate a calming synthetic story using past photographs. This content should help a specific patient relax." The output is the generated synthetic content, provided as text and images appropriate to the patient's condition.

[0368] Step 8:

[0369] The terminal receives synthesized content and alerts sent from the server and sends push notifications to the caregiver. The input is notification data from the server, which is immediately displayed on the caregiver's device to enable early response by the caregiver. The output is the alerts and synthesized content displayed on the actual device.

[0370] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0371] This invention provides caregivers with more effective response strategies by combining an emotion engine with a system designed to support the care of dementia patients. The system analyzes sensor data, eye movement data, and voice data, as well as performing emotion recognition and providing feedback based on the user's psychological state.

[0372] Collection and analysis of emotional data

[0373] The device monitors the user's facial expressions through its built-in camera and acquires data in real time. This data is input into an emotion engine to determine emotional states such as smiles, anger, and sadness.

[0374] Furthermore, the device records the user's voice tone and speaking speed, and uses these voice characteristics to estimate emotions and make more accurate emotional judgments.

[0375] Emotion-based content delivery

[0376] The server generates synthesized content tailored to the user's emotions based on the results obtained from the emotion engine. For example, if the user is feeling anxious, it will provide a visual story or music with a relaxing effect.

[0377] The device presents this content to the user, aiming for the displayed content to calm the user's state. Simultaneously, it collects user feedback again, continuing to form a feedback loop.

[0378] As a concrete example, suppose a user is watching television in the living room with an anxious expression. The system identifies this and, via the server, provides video and audio based on past experiences where the user felt reassured. As a result, the user's expression softens, and their heart rate and other vital data stabilize. Through this feedback mechanism, the system provides caregivers with an effective and consistent care strategy.

[0379] The following describes the processing flow.

[0380] Step 1:

[0381] The device uses its built-in camera to monitor the user's face and periodically captures facial expression data and features. This data includes the movement and position of each part of the face.

[0382] Step 2:

[0383] The device transmits the acquired facial expression data to the emotion engine. The emotion engine analyzes the received data and applies an algorithm to determine emotions such as smiles, surprise, and anger.

[0384] Step 3:

[0385] The device records the user's voice using a microphone and extracts the tone and speaking speed of the voice. This audio data is used to identify variations in the volume and tempo of speech.

[0386] Step 4:

[0387] The device inputs voice data into an emotion engine, which then estimates emotions from the voice characteristics. This allows for a more comprehensive identification of emotions.

[0388] Step 5:

[0389] The server identifies the user's emotional state based on the analysis results and generates appropriate synthesized content. It selects the appropriate visual story and music templates based on the emotional state.

[0390] Step 6:

[0391] The device presents synthesized content sent from the server to the user and observes the impact the content has on the user's emotions.

[0392] Step 7:

[0393] The device records the user's response again during the presentation and sends that data to the server for analysis.

[0394] Step 8:

[0395] Based on the newly obtained response data, the server adjusts the next proposed countermeasures and content, forming a feedback loop.

[0396] This process allows the system to continuously monitor users' emotions and provide care tailored to their individual needs.

[0397] (Example 2)

[0398] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0399] In an aging society, there is a need to provide efficient and effective psychological care for dementia patients. Traditional methods rely heavily on the caregiver's experience and intuition, making it difficult to respond appropriately to the patient's emotions and psychological state. Therefore, new technologies are needed to accurately recognize the patient's emotions and provide appropriate synthetic information to promote their psychological stability.

[0400] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0401] In this invention, the server includes means for collecting facial expression information and performing emotion recognition, means for analyzing voice characteristics and evaluating psychological state, and means for generating synthesized information based on the emotion recognition results. This makes it possible to provide appropriate care based on the patient's emotional state and promote psychological stability.

[0402] "Facial expression information" refers to visual data obtained from the user's face and is used to determine their emotional state.

[0403] "Emotion recognition" is a technology that analyzes facial expressions and voice characteristics to identify a user's emotional state.

[0404] "Vocal features" refer to data such as tone, speech rate, and volume obtained from speech, and are used to evaluate the user's psychological state.

[0405] "Synthetic information" refers to visual or auditory content generated in response to the user's emotional state, and is provided for the purpose of supporting psychological stabilization.

[0406] "Feedback" is the process of recording user responses again using sensors, evaluating the effectiveness of the system, and making adjustments as needed.

[0407] "Vital information" refers to data that indicates the user's physical condition, such as heart rate, blood pressure, and body temperature.

[0408] "Environmental information" refers to data that indicates external conditions that affect the user's psychological state, such as ambient noise, light intensity, and temperature.

[0409] This invention relates to a system for supporting the psychological care of dementia patients, aiming to monitor the user's emotional state in real time and provide appropriate care. This system can be implemented as follows:

[0410] Data collection and analysis

[0411] The device collects user facial expressions and voice characteristics using its built-in camera and microphone. Facial expressions are captured using image analysis software with real-time data processing capabilities to capture facial movements and subtle changes in facial muscles. Voice characteristics are analyzed using voice analysis software to measure tone, speed, and volume based on the collected voice data.

[0412] The server aggregates this data and uses a generative AI model to perform emotion recognition. This model includes various pre-trained emotion analysis algorithms and can accurately distinguish between diverse emotional states.

[0413] Content generation and feedback

[0414] Based on the results of emotion recognition, the server generates synthetic information appropriate to the user's psychological state. For example, if the user is showing signs of anxiety, it will create relaxing music or calming visuals. It is also possible to adjust the content based on past data.

[0415] The device provides the generated composite information to the user, presenting it through the screen and speaker. The device also records the user's new responses and sends them to the server, forming a feedback loop.

[0416] Specific example

[0417] For example, if a user appears bored, the system can detect this and generate and present images of refreshing natural scenery or music. This approach allows users to regain a sense of calm and improve their quality of daily life.

[0418] Examples of prompts for generative AI models

[0419] "Regarding the development of a psychological care support system for dementia patients, please explain the specific methods for sensing emotions and generating appropriate content based on those results."

[0420] This system configuration makes it possible to support dementia patients in leading more stable lives.

[0421] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0422] Step 1:

[0423] The device activates its built-in camera and microphone to capture the user's facial expressions and voice. Image data is acquired as facial information, and voice data is recorded as voice features. These form the initial dataset as sensor inputs. This dataset is transmitted to the server in real time.

[0424] Step 2:

[0425] The server processes the received facial information through image analysis software to detect the user's facial movements and changes in facial muscles, thereby determining their emotional state. The output of this process is a list of possible emotions the user may be exhibiting, each with a confidence score.

[0426] Step 3:

[0427] The server processes the audio data using speech analysis software. It analyzes speech features (tone, speed, volume) and evaluates the user's psychological state. This analysis complements the list of emotions obtained in the previous step, improving the accuracy of emotion recognition.

[0428] Step 4:

[0429] The server uses a generative AI model to generate synthetic content based on the aforementioned emotional states. This content is designed to provide relaxation and a sense of security. The input data is a list of emotions, and the output is synthetic visual and auditory content.

[0430] Step 5:

[0431] The terminal receives synthesized content sent from the server and presents it to the user. The content is output via the terminal's display and speakers. The user's new responses are again recorded by the terminal and sent back to the server for the next processing cycle to form a feedback loop.

[0432] (Application Example 2)

[0433] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0434] As society ages, caring for dementia patients has become a critical issue. In particular, accurately understanding the emotional and psychological changes of dementia patients and providing appropriate care accordingly is extremely difficult in care facilities. Solving this problem is essential.

[0435] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0436] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing audio data and identifying specific speech patterns, means for generating synthesized content using past visual information, and means for analyzing the emotions of dementia patients in real time and presenting care suggestions on a visual display device. This enables care staff to provide appropriate care based on the emotional state of dementia patients quickly and effectively.

[0437] "Sensor data" refers to information about the environment and the state of the object obtained from various sensors.

[0438] "Detecting abnormal conditions" means recognizing a state that is different from the normal state or a problematic state.

[0439] "Eye movement data" refers to information about visual movements such as gaze and blinking.

[0440] "Cognitive function" is a general term for functions involved in human intellectual activities, such as thinking, memory, judgment, and learning.

[0441] "Audio data" refers to physical sound signals that contain information about vocalization.

[0442] "Speech patterns" refer to a set of characteristics in speech, such as word choice, rhythm and speed of language use, etc.

[0443] "Visual information" refers to information that is represented as images or videos.

[0444] "Synthetic content" refers to visually and aurally appealing content created by combining multiple sources of information.

[0445] A "dementia patient" refers to a person who has a medical condition characterized primarily by memory loss and confused thinking.

[0446] "Real-time emotional analysis" means instantly judging emotions and immediately evaluating their changes.

[0447] A "visual display device" is a technological device used to visually represent information.

[0448] "Care proposals" refer to suggestions that indicate appropriate actions and countermeasures in caregiving.

[0449] This invention is a system for streamlining the care of dementia patients in nursing care facilities. The system analyzes the patient's emotions in real time using sensor data, eye movement data, and voice data. A specific embodiment of this system is described here.

[0450] The server receives sensor data, and if an abnormal condition is detected, immediate action is taken. Eye movements are tracked using a camera, and cognitive function is evaluated using visual information displayed on smart glasses or a headset. Speech recognition software analyzes the patient's voice data and identifies specific speech patterns. It also generates relaxing synthetic content based on previously collected visual information. This synthetic content includes music and visual stories.

[0451] The device uses an emotion analysis engine to assess the patient's emotions in real time and presents care suggestions based on the results to the caregiver via a visual display. For example, if the device detects that the patient is anxious, it will display appropriate care methods such as "Stay calmly by the patient's side and listen to them."

[0452] As a concrete example, let's consider its use in a nursing home. In this facility, care staff wear smart glasses, and if it detects that a dementia patient is irritable, it automatically suggests "playing calming music while taking a short walk with the patient." In this way, specific care based on the patient's emotional state is provided quickly.

[0453] An example of a prompt might be the instruction, "Design an application that tracks the emotional state of users in this facility in real time and provides appropriate advice to staff."

[0454] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0455] Step 1:

[0456] The server receives sensor data, eye movement data, and audio data from the terminal. This input data is treated as initial information for understanding the patient's condition and is stored in the server's database.

[0457] Step 2:

[0458] The server analyzes the received data. It detects abnormal conditions based on sensor data and evaluates cognitive function using eye movement data. Audio data is analyzed using speech recognition software to identify specific speech patterns. As a result of this analysis, various status indicators are output and recorded on the server.

[0459] Step 3:

[0460] The server uses a generative AI model to analyze the patient's emotional state. It provides the model with inputs such as abnormal conditions, cognitive function assessment results, and speech patterns, and determines the patient's emotions in real time. The resulting output represents the patient's emotional state, which forms the basis for creating care recommendations.

[0461] Step 4:

[0462] The server creates synthesized content based on the analysis results. Using past visual information as a reference, it generates music and visual stories best suited to the current emotional state. This content is then sent to the device.

[0463] Step 5:

[0464] The terminal displays the received synthesized content on a visual display device, visually showing care suggestions that are emotionally responsive to the caregiver. These suggestions encourage specific actions that are helpful in face-to-face care.

[0465] Step 6:

[0466] The caregiver, as the user, provides appropriate care to the patient based on the information obtained from the device. If necessary, additional data and feedback can be sent to the server via the device, enabling continuous improvement of the care provided.

[0467] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0468] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0469] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0470] [Third Embodiment]

[0471] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0472] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0473] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0474] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0475] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0476] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0477] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0478] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0479] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0480] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0481] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0482] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0483] This invention is a system designed for dementia patients that integrates multiple data collection and analysis technologies. This system aggregates sensor data, eye movement data, and voice data, and analyzes them using AI to suggest appropriate countermeasures to caregivers.

[0484] Sensor data collection and analysis

[0485] The device acquires data from multiple environmental sensors and wearable devices. This includes the care recipient's heart rate and blood pressure, as well as room temperature and humidity. The data is immediately sent to the server, and if any abnormal values ​​are detected, an alert is generated in real time.

[0486] Assessment of cognitive function

[0487] The user uses a dedicated eye-tracking device. This device tracks the eye movements of the person being cared for while they perform specific tasks and provides this data to a server. The server then applies an algorithm to score cognitive function based on the data obtained from these movements.

[0488] Analysis of audio data

[0489] The device records the care recipient's conversational audio via a smartphone app. The audio data is sent to a server, where an AI model identifies speech patterns and changes. For example, it can detect speech fluency and vocabulary changes.

[0490] Generating Synthetic Content

[0491] The server uses user-provided past photos and related visual data to create synthesized content using generative AI. This synthesized content helps care recipients visually re-experience past memories. This process contributes to emotional stability and promotes better communication.

[0492] As a concrete example, if a person receiving care is experiencing stress at a particular time, the system infers the cause from sensor data and voice data. For instance, if an increase in heart rate coincides with changes in voice, it presents music or a synthesized story related to past memories to alleviate the condition. Then, through automated feedback mechanisms, the system also records user emotional reactions in real time and performs further analysis and adjustments as needed. This ensures the best possible care plan, reduces the burden on caregivers, and improves the quality of life for patients.

[0493] The following describes the processing flow.

[0494] Step 1:

[0495] The device acquires vital data in real time from environmental sensors and wearable devices. This includes data such as temperature, humidity, heart rate, and blood pressure.

[0496] Step 2:

[0497] The terminal collects data, bundles it into data packets, and sends them to the server over the network. This transmission is performed periodically to ensure uninterrupted data flow.

[0498] Step 3:

[0499] The server stores the received data in a database. After storage, an AI model is used to perform initial data analysis and detect anomalies and patterns.

[0500] Step 4:

[0501] The user uses an eye-tracking device to perform a designated visual task. This generates data on the care recipient's eye movements.

[0502] Step 5:

[0503] The device instantly transmits eye movement data to the server. This data transmission is performed very quickly to ensure accuracy.

[0504] Step 6:

[0505] The server analyzes the received eye movement data and calculates a score to assess cognitive function. The assessed results become part of the feedback provided to the caregiver.

[0506] Step 7:

[0507] Users record voice responses using their smartphone's voice recording function. This includes answers to questions and everyday speech.

[0508] Step 8:

[0509] The terminal processes the recorded audio data, such as removing noise, before sending it to the server.

[0510] Step 9:

[0511] The server uses an AI model to analyze voice data in detail, identifying speech patterns and changes. This allows it to determine signs and progression of dementia.

[0512] Step 10:

[0513] The server uses generative AI to create synthesized content based on past photographs and visual data provided. This is a process that generates a visual story that is emotionally meaningful to the person receiving care.

[0514] Step 11:

[0515] The device presents the generated content to the user and records the user's response. The recorded data is used in the next feedback loop.

[0516] (Example 1)

[0517] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0518] In an aging society, improving the quality of life for dementia patients requires a system that accurately assesses the physical and mental condition of those receiving care and promptly provides appropriate care methods. However, current systems struggle to comprehensively analyze multiple data points and respond flexibly to the individual needs of each user.

[0519] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0520] In this invention, the server includes a configuration for receiving sensor information and detecting abnormal conditions, a configuration for acquiring gaze data and evaluating cognitive function, a configuration for analyzing voice information and identifying specific speech patterns, and a configuration for adjusting the generation algorithm using feedback obtained from the user. This enables more accurate monitoring of dementia patients and the provision of personalized care support measures.

[0521] "Sensor information" refers to information obtained from devices and equipment used to collect data about the environment and individual biological systems.

[0522] An "abnormal condition" refers to a physical or environmental situation that deviates from the normal range, and is a condition that requires particular attention or action, especially in the context of caregiving.

[0523] "Eye-gaze data" refers to data about an individual's visual focus and eye movements, and is important information used to evaluate cognitive function.

[0524] A "system for evaluating cognitive function" is a set of procedures and algorithms for analyzing eye-tracking data and other information to evaluate an individual's cognitive abilities and behavioral patterns.

[0525] "Vocal information" refers to data obtained from the speech and conversations of the person receiving care, and is fundamental information for understanding speech patterns.

[0526] "Speech patterns" refer to the characteristics and structure of an individual's way of speaking, and are used to identify specific changes or tendencies.

[0527] "Synthetic content" refers to content such as videos and stories generated based on past visual data, and is used with the aim of promoting emotional stability and memory recall in those receiving care.

[0528] A "generative algorithm" is a computational procedure that utilizes past data and feedback to create new content and suggestions.

[0529] "User feedback" refers to information about the reactions and evaluations that care recipients and caregivers give to the proposed synthetic content and system, and this information can be used to improve the system.

[0530] This invention is an advanced data analysis system designed to support dementia patients. The system aims to propose appropriate care strategies to caregivers by integrating information collected from various sensors and devices and performing in-depth analysis using artificial intelligence (AI).

[0531] Sensor data collection

[0532] The terminal uses environmental sensors and wearable devices to acquire vital information such as the care recipient's heart rate and blood pressure, as well as environmental information such as room temperature and humidity, in real time. Common wearable devices and sensor modules are used for this purpose. The collected data is immediately transmitted to a server, and an alert is generated if an abnormal condition is detected.

[0533] Assessment of cognitive function

[0534] The user tracks the eye movements of the person being cared for while they perform tasks on the screen, using a dedicated eye-tracking device. This device utilizes commonly used eye-tracking technology and sends the eye movement data to a server, which then uses this data to score cognitive function.

[0535] Analysis of audio data

[0536] The device uses a smartphone app to record the care recipient's conversations and sends the audio data to a server. The server analyzes the collected audio data using an AI model to identify specific speech patterns and changes. This makes it possible to understand the fluency of speech and changes in vocabulary in detail.

[0537] Generating Synthetic Content

[0538] The server utilizes past photos and related visual data provided by the user and develops synthesized content using a generative AI model. This content is used to allow the person receiving care to visually re-experience past memories, promoting emotional stability. For example, it can generate stories based on relaxing landscape images and verbalize them through prompts.

[0539] As a concrete example, based on the prompt "Generate relaxing composite content using past images that the care recipient finds comforting," it is possible to create a composite movie using past travel photos of the care recipient. Throughout this process, the system collects user feedback and uses it as data for continuous improvement.

[0540] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0541] Step 1:

[0542] The terminal acquires data from multiple environmental sensors and wearable devices. It receives data from sensors such as heart rate sensors and room temperature sensors as input and sends it to the server. The output is raw data from each sensor, sent to the server via Bluetooth or Wi-Fi. The terminal periodically checks whether the sensors are functioning correctly to prevent data loss.

[0543] Step 2:

[0544] The server receives sensor data transmitted from the terminal and performs analysis to detect anomalies. It receives heart rate and room temperature data from the sensor as input and compares it to normal range data recorded in the database. The output is a warning message if an anomaly is detected, which is notified to the caregiver's terminal in real time. The server improves the accuracy of anomaly detection through time-series analysis by comparing the current data with past data.

[0545] Step 3:

[0546] The user uses an eye-tracking device to collect data on the care recipient's eye movements. The input is the capture of gaze data while the care recipient is looking at a screen, and this data is sent to the server. The output is a cognitive function assessment score based on the gaze movements. The user periodically adjusts the device's position and calibration to ensure accurate tracking.

[0547] Step 4:

[0548] The server analyzes eye-tracking data sent by the user to assess cognitive function. It receives eye-tracking data as input and evaluates its movement using an analysis algorithm. Output includes an evaluation score and feedback for the caregiver. The server utilizes an AI-based model to identify eye-tracking patterns when performing different tasks.

[0549] Step 5:

[0550] The device records audio data via a smartphone app and sends the recording to a server. The input is captured audio of a conversation and saved in a digital format. The output is an audio data file, which is sent to the server for analysis. The device uses noise reduction filtering technology to improve recording quality.

[0551] Step 6:

[0552] The server receives audio data and analyzes it using an AI model. It receives audio data files as input and processes the data to identify speech patterns and changes. The output is the analysis result based on specific speech patterns, which is then used to provide notifications and advice to caregivers. The server identifies trends and changes by comparing the current audio data with past audio data.

[0553] Step 7:

[0554] The server generates synthetic content using a generative AI model. It utilizes past photographs and related visual data as input, and determines content based on prompt text. The output is generated as visualized video content or stories and delivered to the user. The server incorporates user feedback during the generation process to provide more personalized content.

[0555] (Application Example 1)

[0556] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0557] Effective care for dementia patients requires real-time monitoring of abnormal conditions and changes in cognitive function, and appropriate responses. However, it is difficult to provide prompt and accurate responses while reducing the burden on caregivers, and problems stemming from information overload and ambiguity in countermeasures are particularly pronounced in facilities with a large number of dementia patients. Therefore, support for flexible and efficient care methods tailored to individual patients is necessary.

[0558] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0559] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing voice data and identifying specific speech patterns, means for integrating information from multiple data sources in real time and generating guidelines for implementing proposed care methods, and means for communicating warnings and suggestions to caregivers using push notifications. This enables caregivers to understand the situation of dementia patients and provide prompt and accurate care.

[0560] "Sensor data" is a general term for multiple measurement pieces of information that indicate the environment and vital signs, and is dynamic data including body temperature, heart rate, room temperature, and humidity.

[0561] "Detection of abnormal conditions" is the act of finding unique patterns or signs that differ from normal conditions in the received sensor data.

[0562] "Eye movement data" refers to information that shows how an individual tracks and focuses on visual information, and is used to assess cognitive function.

[0563] "Cognitive function assessment" is a process of analyzing and scoring an individual's brain functions, such as language, judgment, and memory, using collected eye movement data.

[0564] "Voice data" refers to information that records individual conversations and utterances, and serves as fundamental data for analyzing language patterns and changes in speech.

[0565] "Identification of specific speech patterns" refers to the function of extracting and recognizing linguistic and rhythmic features within analyzed speech data.

[0566] "Past visual information" refers to photographs and videos related to the user's past, and is visual data used to activate memories.

[0567] "Generating synthetic content" is the process of creating new images and videos using computer generation technology based on past visual information.

[0568] "Push notification" is a communication method that transmits information from the sender to the recipient instantly at a specific time.

[0569] "Real-time information integration" refers to the process of immediately combining and analyzing information obtained from multiple data sources, thereby supporting rapid decision-making.

[0570] This invention is a care support system for dementia patients. The system is implemented with the following configuration.

[0571] The server receives data from multiple sensors at each terminal and uses this data to detect abnormal conditions. The sensors are mainly installed around the person being cared for and collect vital data and environmental data in real time. This provides a function that immediately sends an alert to the caregiver if there is an abnormality in heart rate or room temperature. Widely available Wi-Fi connected sensors are used as the hardware.

[0572] Users assess their cognitive function using a dedicated eye-tracking device. The eye-tracking device meticulously records the eye movements of the person being cared for, and this data is transmitted to a server. Based on this information, the server-side algorithm scores changes in cognitive function and provides feedback to the caregiver.

[0573] Additionally, voice data is collected by smart devices. The devices record everyday conversations and send the data to a server for voice analysis. An AI model in the cloud analyzes changes in speech patterns, which can serve as an indicator of dementia progression.

[0574] The server integrates this data and uses a generative AI model to suggest synthesized content to caregivers. The generated content serves as a guide for caregivers to implement the most appropriate individualized support for the patient. For example, if a patient has comforting memories from the past, images of those memories can be synthesized and displayed for stress reduction purposes.

[0575] For example, if a patient becomes restless in the afternoon, the system will detect an abnormal increase in heart rate during that time and notify the caregiver's device with a relaxing, synthesized video from the past. An example of this prompt message would be: "Please generate a calming synthesized story using past photos. This content should help the specific patient relax."

[0576] This allows caregivers to effectively understand the condition of dementia patients and provide real-time support tailored to their individual needs.

[0577] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0578] Step 1:

[0579] The device collects sensor data such as the care recipient's body temperature, heart rate, and room temperature in real time. It takes data from each sensor as input, organizes it within the device, and then sends it to the server. The output is a set of the organized sensor data sent to the server.

[0580] Step 2:

[0581] The server analyzes the received sensor data to detect abnormal conditions. The input is sensor data sent from the terminal, and abnormalities are visualized by comparing it with normal data stored in the database. If an abnormality is detected through this process, alert information is generated.

[0582] Step 3:

[0583] The user uses an eye-tracking device to capture the eye movements of the person being cared for. The eye movement data is recorded as input on the device and sent to the server. As output, the eye movement data is formatted and becomes data for analysis on the server.

[0584] Step 4:

[0585] The server evaluates cognitive function based on the received eye movement data. Using eye movement data as input, it applies an algorithm to calculate a cognitive function score. The output is the cognitive function score as the evaluation result, which is provided to the caregiver.

[0586] Step 5:

[0587] The device records everyday conversations and transfers the audio data to a server. The input is the voice of the person receiving care, which is recorded clearly and then sent to the server. The audio file is then formatted and used as material for analysis.

[0588] Step 6:

[0589] The server analyzes audio data and identifies specific speech patterns. Using audio data as input, it extracts language patterns using a generative AI model. The output generates hints for caregivers based on the identified speech patterns.

[0590] Step 7:

[0591] The server generates synthetic content using past visual information and generative AI. It analyzes the input visual data using the prompt: "Generate a calming synthetic story using past photographs. This content should help a specific patient relax." The output is the generated synthetic content, provided as text and images appropriate to the patient's condition.

[0592] Step 8:

[0593] The terminal receives synthesized content and alerts sent from the server and sends push notifications to the caregiver. The input is notification data from the server, which is immediately displayed on the caregiver's device to enable early response by the caregiver. The output is the alerts and synthesized content displayed on the actual device.

[0594] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0595] This invention provides caregivers with more effective response strategies by combining an emotion engine with a system designed to support the care of dementia patients. The system analyzes sensor data, eye movement data, and voice data, as well as performing emotion recognition and providing feedback based on the user's psychological state.

[0596] Collection and analysis of emotional data

[0597] The device monitors the user's facial expressions through its built-in camera and acquires data in real time. This data is input into an emotion engine to determine emotional states such as smiles, anger, and sadness.

[0598] Furthermore, the device records the user's voice tone and speaking speed, and uses these voice characteristics to estimate emotions and make more accurate emotional judgments.

[0599] Emotion-based content delivery

[0600] The server generates synthesized content tailored to the user's emotions based on the results obtained from the emotion engine. For example, if the user is feeling anxious, it will provide a visual story or music with a relaxing effect.

[0601] The device presents this content to the user, aiming for the displayed content to calm the user's state. Simultaneously, it collects user feedback again, continuing to form a feedback loop.

[0602] As a concrete example, suppose a user is watching television in the living room with an anxious expression. The system identifies this and, via the server, provides video and audio based on past experiences where the user felt reassured. As a result, the user's expression softens, and their heart rate and other vital data stabilize. Through this feedback mechanism, the system provides caregivers with an effective and consistent care strategy.

[0603] The following describes the processing flow.

[0604] Step 1:

[0605] The device uses its built-in camera to monitor the user's face and periodically captures facial expression data and features. This data includes the movement and position of each part of the face.

[0606] Step 2:

[0607] The device transmits the acquired facial expression data to the emotion engine. The emotion engine analyzes the received data and applies an algorithm to determine emotions such as smiles, surprise, and anger.

[0608] Step 3:

[0609] The device records the user's voice using a microphone and extracts the tone and speaking speed of the voice. This audio data is used to identify variations in the volume and tempo of speech.

[0610] Step 4:

[0611] The device inputs voice data into an emotion engine, which then estimates emotions from the voice characteristics. This allows for a more comprehensive identification of emotions.

[0612] Step 5:

[0613] The server identifies the user's emotional state based on the analysis results and generates appropriate synthesized content. It selects the appropriate visual story and music templates based on the emotional state.

[0614] Step 6:

[0615] The device presents synthesized content sent from the server to the user and observes the impact the content has on the user's emotions.

[0616] Step 7:

[0617] The device records the user's response again during the presentation and sends that data to the server for analysis.

[0618] Step 8:

[0619] Based on the newly obtained response data, the server adjusts the next proposed countermeasures and content, forming a feedback loop.

[0620] This process allows the system to continuously monitor users' emotions and provide care tailored to their individual needs.

[0621] (Example 2)

[0622] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0623] In an aging society, there is a need to provide efficient and effective psychological care for dementia patients. Traditional methods rely heavily on the caregiver's experience and intuition, making it difficult to respond appropriately to the patient's emotions and psychological state. Therefore, new technologies are needed to accurately recognize the patient's emotions and provide appropriate synthetic information to promote their psychological stability.

[0624] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0625] In this invention, the server includes means for collecting facial expression information and performing emotion recognition, means for analyzing voice characteristics and evaluating psychological state, and means for generating synthesized information based on the emotion recognition results. This makes it possible to provide appropriate care based on the patient's emotional state and promote psychological stability.

[0626] "Facial expression information" refers to visual data obtained from the user's face and is used to determine their emotional state.

[0627] "Emotion recognition" is a technology that analyzes facial expressions and voice characteristics to identify a user's emotional state.

[0628] "Vocal features" refer to data such as tone, speech rate, and volume obtained from speech, and are used to evaluate the user's psychological state.

[0629] "Synthetic information" refers to visual or auditory content generated in response to the user's emotional state, and is provided for the purpose of supporting psychological stabilization.

[0630] "Feedback" is the process of recording user responses again using sensors, evaluating the effectiveness of the system, and making adjustments as needed.

[0631] "Vital information" refers to data that indicates the user's physical condition, such as heart rate, blood pressure, and body temperature.

[0632] "Environmental information" refers to data that indicates external conditions that affect the user's psychological state, such as ambient noise, light intensity, and temperature.

[0633] This invention relates to a system for supporting the psychological care of dementia patients, aiming to monitor the user's emotional state in real time and provide appropriate care. This system can be implemented as follows:

[0634] Data collection and analysis

[0635] The device collects user facial expressions and voice characteristics using its built-in camera and microphone. Facial expressions are captured using image analysis software with real-time data processing capabilities to capture facial movements and subtle changes in facial muscles. Voice characteristics are analyzed using voice analysis software to measure tone, speed, and volume based on the collected voice data.

[0636] The server aggregates this data and uses a generative AI model to perform emotion recognition. This model includes various pre-trained emotion analysis algorithms and can accurately distinguish between diverse emotional states.

[0637] Content generation and feedback

[0638] Based on the results of emotion recognition, the server generates synthetic information appropriate to the user's psychological state. For example, if the user is showing signs of anxiety, it will create relaxing music or calming visuals. It is also possible to adjust the content based on past data.

[0639] The device provides the generated composite information to the user, presenting it through the screen and speaker. The device also records the user's new responses and sends them to the server, forming a feedback loop.

[0640] Specific example

[0641] For example, if a user appears bored, the system can detect this and generate and present images of refreshing natural scenery or music. This approach allows users to regain a sense of calm and improve their quality of daily life.

[0642] Examples of prompts for generative AI models

[0643] "Regarding the development of a psychological care support system for dementia patients, please explain the specific methods for sensing emotions and generating appropriate content based on those results."

[0644] This system configuration makes it possible to support dementia patients in leading more stable lives.

[0645] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0646] Step 1:

[0647] The device activates its built-in camera and microphone to capture the user's facial expressions and voice. Image data is acquired as facial information, and voice data is recorded as voice features. These form the initial dataset as sensor inputs. This dataset is transmitted to the server in real time.

[0648] Step 2:

[0649] The server processes the received facial information through image analysis software to detect the user's facial movements and changes in facial muscles, thereby determining their emotional state. The output of this process is a list of possible emotions the user may be exhibiting, each with a confidence score.

[0650] Step 3:

[0651] The server processes the audio data using speech analysis software. It analyzes speech features (tone, speed, volume) and evaluates the user's psychological state. This analysis complements the list of emotions obtained in the previous step, improving the accuracy of emotion recognition.

[0652] Step 4:

[0653] The server uses a generative AI model to generate synthetic content based on the aforementioned emotional states. This content is designed to provide relaxation and a sense of security. The input data is a list of emotions, and the output is synthetic visual and auditory content.

[0654] Step 5:

[0655] The terminal receives synthesized content sent from the server and presents it to the user. The content is output via the terminal's display and speakers. The user's new responses are again recorded by the terminal and sent back to the server for the next processing cycle to form a feedback loop.

[0656] (Application Example 2)

[0657] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0658] As society ages, caring for dementia patients has become a critical issue. In particular, accurately understanding the emotional and psychological changes of dementia patients and providing appropriate care accordingly is extremely difficult in care facilities. Solving this problem is essential.

[0659] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0660] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing audio data and identifying specific speech patterns, means for generating synthesized content using past visual information, and means for analyzing the emotions of dementia patients in real time and presenting care suggestions on a visual display device. This enables care staff to provide appropriate care based on the emotional state of dementia patients quickly and effectively.

[0661] "Sensor data" refers to information about the environment and the state of the object obtained from various sensors.

[0662] "Detecting abnormal conditions" means recognizing a state that is different from the normal state or a problematic state.

[0663] "Eye movement data" refers to information about visual movements such as gaze and blinking.

[0664] "Cognitive function" is a general term for functions involved in human intellectual activities, such as thinking, memory, judgment, and learning.

[0665] "Audio data" refers to physical sound signals that contain information about vocalization.

[0666] "Speech patterns" refer to a set of characteristics in speech, such as word choice, rhythm and speed of language use, etc.

[0667] "Visual information" refers to information that is represented as images or videos.

[0668] "Synthetic content" refers to visually and aurally appealing content created by combining multiple sources of information.

[0669] A "dementia patient" refers to a person who has a medical condition characterized primarily by memory loss and confused thinking.

[0670] "Real-time emotional analysis" means instantly judging emotions and immediately evaluating their changes.

[0671] A "visual display device" is a technological device used to visually represent information.

[0672] "Care proposals" refer to suggestions that indicate appropriate actions and countermeasures in caregiving.

[0673] This invention is a system for streamlining the care of dementia patients in nursing care facilities. The system analyzes the patient's emotions in real time using sensor data, eye movement data, and voice data. A specific embodiment of this system is described here.

[0674] The server receives sensor data, and if an abnormal condition is detected, immediate action is taken. Eye movements are tracked using a camera, and cognitive function is evaluated using visual information displayed on smart glasses or a headset. Speech recognition software analyzes the patient's voice data and identifies specific speech patterns. It also generates relaxing synthetic content based on previously collected visual information. This synthetic content includes music and visual stories.

[0675] The device uses an emotion analysis engine to assess the patient's emotions in real time and presents care suggestions based on the results to the caregiver via a visual display. For example, if the device detects that the patient is anxious, it will display appropriate care methods such as "Stay calmly by the patient's side and listen to them."

[0676] As a concrete example, let's consider its use in a nursing home. In this facility, care staff wear smart glasses, and if it detects that a dementia patient is irritable, it automatically suggests "playing calming music while taking a short walk with the patient." In this way, specific care based on the patient's emotional state is provided quickly.

[0677] An example of a prompt might be the instruction, "Design an application that tracks the emotional state of users in this facility in real time and provides appropriate advice to staff."

[0678] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0679] Step 1:

[0680] The server receives sensor data, eye movement data, and audio data from the terminal. This input data is treated as initial information for understanding the patient's condition and is stored in the server's database.

[0681] Step 2:

[0682] The server analyzes the received data. It detects abnormal conditions based on sensor data and evaluates cognitive function using eye movement data. Audio data is analyzed using speech recognition software to identify specific speech patterns. As a result of this analysis, various status indicators are output and recorded on the server.

[0683] Step 3:

[0684] The server uses a generative AI model to analyze the patient's emotional state. It provides the model with inputs such as abnormal conditions, cognitive function assessment results, and speech patterns, and determines the patient's emotions in real time. The resulting output represents the patient's emotional state, which forms the basis for creating care recommendations.

[0685] Step 4:

[0686] The server creates synthesized content based on the analysis results. Using past visual information as a reference, it generates music and visual stories best suited to the current emotional state. This content is then sent to the device.

[0687] Step 5:

[0688] The terminal displays the received synthesized content on a visual display device, visually showing care suggestions that are emotionally responsive to the caregiver. These suggestions encourage specific actions that are helpful in face-to-face care.

[0689] Step 6:

[0690] The caregiver, as the user, provides appropriate care to the patient based on the information obtained from the device. If necessary, additional data and feedback can be sent to the server via the device, enabling continuous improvement of the care provided.

[0691] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0692] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0693] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0694] [Fourth Embodiment]

[0695] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0696] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0697] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0698] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0699] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0700] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0701] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0702] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0703] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0704] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0705] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0706] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0707] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0708] This invention is a system designed for dementia patients that integrates multiple data collection and analysis technologies. This system aggregates sensor data, eye movement data, and voice data, and analyzes them using AI to suggest appropriate countermeasures to caregivers.

[0709] Sensor data collection and analysis

[0710] The device acquires data from multiple environmental sensors and wearable devices. This includes the care recipient's heart rate and blood pressure, as well as room temperature and humidity. The data is immediately sent to the server, and if any abnormal values ​​are detected, an alert is generated in real time.

[0711] Assessment of cognitive function

[0712] The user uses a dedicated eye-tracking device. This device tracks the eye movements of the person being cared for while they perform specific tasks and provides this data to a server. The server then applies an algorithm to score cognitive function based on the data obtained from these movements.

[0713] Analysis of audio data

[0714] The device records the care recipient's conversational audio via a smartphone app. The audio data is sent to a server, where an AI model identifies speech patterns and changes. For example, it can detect speech fluency and vocabulary changes.

[0715] Generating Synthetic Content

[0716] The server uses user-provided past photos and related visual data to create synthesized content using generative AI. This synthesized content helps care recipients visually re-experience past memories. This process contributes to emotional stability and promotes better communication.

[0717] As a concrete example, if a person receiving care is experiencing stress at a particular time, the system infers the cause from sensor data and voice data. For instance, if an increase in heart rate coincides with changes in voice, it presents music or a synthesized story related to past memories to alleviate the condition. Then, through automated feedback mechanisms, the system also records user emotional reactions in real time and performs further analysis and adjustments as needed. This ensures the best possible care plan, reduces the burden on caregivers, and improves the quality of life for patients.

[0718] The following describes the processing flow.

[0719] Step 1:

[0720] The device acquires vital data in real time from environmental sensors and wearable devices. This includes data such as temperature, humidity, heart rate, and blood pressure.

[0721] Step 2:

[0722] The terminal collects data, bundles it into data packets, and sends them to the server over the network. This transmission is performed periodically to ensure uninterrupted data flow.

[0723] Step 3:

[0724] The server stores the received data in a database. After storage, an AI model is used to perform initial data analysis and detect anomalies and patterns.

[0725] Step 4:

[0726] The user uses an eye-tracking device to perform a designated visual task. This generates data on the care recipient's eye movements.

[0727] Step 5:

[0728] The device instantly transmits eye movement data to the server. This data transmission is performed very quickly to ensure accuracy.

[0729] Step 6:

[0730] The server analyzes the received eye movement data and calculates a score to assess cognitive function. The assessed results become part of the feedback provided to the caregiver.

[0731] Step 7:

[0732] Users record voice responses using their smartphone's voice recording function. This includes answers to questions and everyday speech.

[0733] Step 8:

[0734] The terminal processes the recorded audio data, such as removing noise, before sending it to the server.

[0735] Step 9:

[0736] The server uses an AI model to analyze voice data in detail, identifying speech patterns and changes. This allows it to determine signs and progression of dementia.

[0737] Step 10:

[0738] The server uses generative AI to create synthesized content based on past photographs and visual data provided. This is a process that generates a visual story that is emotionally meaningful to the person receiving care.

[0739] Step 11:

[0740] The device presents the generated content to the user and records the user's response. The recorded data is used in the next feedback loop.

[0741] (Example 1)

[0742] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0743] In an aging society, improving the quality of life for dementia patients requires a system that accurately assesses the physical and mental condition of those receiving care and promptly provides appropriate care methods. However, current systems struggle to comprehensively analyze multiple data points and respond flexibly to the individual needs of each user.

[0744] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0745] In this invention, the server includes a configuration for receiving sensor information and detecting abnormal conditions, a configuration for acquiring gaze data and evaluating cognitive function, a configuration for analyzing voice information and identifying specific speech patterns, and a configuration for adjusting the generation algorithm using feedback obtained from the user. This enables more accurate monitoring of dementia patients and the provision of personalized care support measures.

[0746] "Sensor information" refers to information obtained from devices and equipment used to collect data about the environment and individual biological systems.

[0747] An "abnormal condition" refers to a physical or environmental situation that deviates from the normal range, and is a condition that requires particular attention or action, especially in the context of caregiving.

[0748] "Eye-gaze data" refers to data about an individual's visual focus and eye movements, and is important information used to evaluate cognitive function.

[0749] A "system for evaluating cognitive function" is a set of procedures and algorithms for analyzing eye-tracking data and other information to evaluate an individual's cognitive abilities and behavioral patterns.

[0750] "Vocal information" refers to data obtained from the speech and conversations of the person receiving care, and is fundamental information for understanding speech patterns.

[0751] "Speech patterns" refer to the characteristics and structure of an individual's way of speaking, and are used to identify specific changes or tendencies.

[0752] "Synthetic content" refers to content such as videos and stories generated based on past visual data, and is used with the aim of promoting emotional stability and memory recall in those receiving care.

[0753] A "generative algorithm" is a computational procedure that utilizes past data and feedback to create new content and suggestions.

[0754] "User feedback" refers to information about the reactions and evaluations that care recipients and caregivers give to the proposed synthetic content and system, and this information can be used to improve the system.

[0755] This invention is an advanced data analysis system designed to support dementia patients. The system aims to propose appropriate care strategies to caregivers by integrating information collected from various sensors and devices and performing in-depth analysis using artificial intelligence (AI).

[0756] Sensor data collection

[0757] The terminal uses environmental sensors and wearable devices to acquire vital information such as the care recipient's heart rate and blood pressure, as well as environmental information such as room temperature and humidity, in real time. Common wearable devices and sensor modules are used for this purpose. The collected data is immediately transmitted to a server, and an alert is generated if an abnormal condition is detected.

[0758] Assessment of cognitive function

[0759] The user tracks the eye movements of the person being cared for while they perform tasks on the screen, using a dedicated eye-tracking device. This device utilizes commonly used eye-tracking technology and sends the eye movement data to a server, which then uses this data to score cognitive function.

[0760] Analysis of audio data

[0761] The device uses a smartphone app to record the care recipient's conversations and sends the audio data to a server. The server analyzes the collected audio data using an AI model to identify specific speech patterns and changes. This makes it possible to understand the fluency of speech and changes in vocabulary in detail.

[0762] Generating Synthetic Content

[0763] The server utilizes past photos and related visual data provided by the user and develops synthesized content using a generative AI model. This content is used to allow the person receiving care to visually re-experience past memories, promoting emotional stability. For example, it can generate stories based on relaxing landscape images and verbalize them through prompts.

[0764] As a concrete example, based on the prompt "Generate relaxing composite content using past images that the care recipient finds comforting," it is possible to create a composite movie using past travel photos of the care recipient. Throughout this process, the system collects user feedback and uses it as data for continuous improvement.

[0765] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0766] Step 1:

[0767] The terminal acquires data from multiple environmental sensors and wearable devices. It receives data from sensors such as heart rate sensors and room temperature sensors as input and sends it to the server. The output is raw data from each sensor, sent to the server via Bluetooth or Wi-Fi. The terminal periodically checks whether the sensors are functioning correctly to prevent data loss.

[0768] Step 2:

[0769] The server receives sensor data transmitted from the terminal and performs analysis to detect anomalies. It receives heart rate and room temperature data from the sensor as input and compares it to normal range data recorded in the database. The output is a warning message if an anomaly is detected, which is notified to the caregiver's terminal in real time. The server improves the accuracy of anomaly detection through time-series analysis by comparing the current data with past data.

[0770] Step 3:

[0771] The user uses an eye-tracking device to collect data on the care recipient's eye movements. The input is the capture of gaze data while the care recipient is looking at a screen, and this data is sent to the server. The output is a cognitive function assessment score based on the gaze movements. The user periodically adjusts the device's position and calibration to ensure accurate tracking.

[0772] Step 4:

[0773] The server analyzes eye-tracking data sent by the user to assess cognitive function. It receives eye-tracking data as input and evaluates its movement using an analysis algorithm. Output includes an evaluation score and feedback for the caregiver. The server utilizes an AI-based model to identify eye-tracking patterns when performing different tasks.

[0774] Step 5:

[0775] The device records audio data via a smartphone app and sends the recording to a server. The input is captured audio of a conversation and saved in a digital format. The output is an audio data file, which is sent to the server for analysis. The device uses noise reduction filtering technology to improve recording quality.

[0776] Step 6:

[0777] The server receives audio data and analyzes it using an AI model. It receives audio data files as input and processes the data to identify speech patterns and changes. The output is the analysis result based on specific speech patterns, which is then used to provide notifications and advice to caregivers. The server identifies trends and changes by comparing the current audio data with past audio data.

[0778] Step 7:

[0779] The server generates synthetic content using a generative AI model. It utilizes past photographs and related visual data as input, and determines content based on prompt text. The output is generated as visualized video content or stories and delivered to the user. The server incorporates user feedback during the generation process to provide more personalized content.

[0780] (Application Example 1)

[0781] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0782] Effective care for dementia patients requires real-time monitoring of abnormal conditions and changes in cognitive function, and appropriate responses. However, it is difficult to provide prompt and accurate responses while reducing the burden on caregivers, and problems stemming from information overload and ambiguity in countermeasures are particularly pronounced in facilities with a large number of dementia patients. Therefore, support for flexible and efficient care methods tailored to individual patients is necessary.

[0783] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0784] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing voice data and identifying specific speech patterns, means for integrating information from multiple data sources in real time and generating guidelines for implementing proposed care methods, and means for communicating warnings and suggestions to caregivers using push notifications. This enables caregivers to understand the situation of dementia patients and provide prompt and accurate care.

[0785] "Sensor data" is a general term for multiple measurement pieces of information that indicate the environment and vital signs, and is dynamic data including body temperature, heart rate, room temperature, and humidity.

[0786] "Detection of abnormal conditions" is the act of finding unique patterns or signs that differ from normal conditions in the received sensor data.

[0787] "Eye movement data" refers to information that shows how an individual tracks and focuses on visual information, and is used to assess cognitive function.

[0788] "Cognitive function assessment" is a process of analyzing and scoring an individual's brain functions, such as language, judgment, and memory, using collected eye movement data.

[0789] "Voice data" refers to information that records individual conversations and utterances, and serves as fundamental data for analyzing language patterns and changes in speech.

[0790] "Identification of specific speech patterns" refers to the function of extracting and recognizing linguistic and rhythmic features within analyzed speech data.

[0791] "Past visual information" refers to photographs and videos related to the user's past, and is visual data used to activate memories.

[0792] "Generating synthetic content" is the process of creating new images and videos using computer generation technology based on past visual information.

[0793] "Push notification" is a communication method that transmits information from the sender to the recipient instantly at a specific time.

[0794] "Real-time information integration" refers to the process of immediately combining and analyzing information obtained from multiple data sources, thereby supporting rapid decision-making.

[0795] This invention is a care support system for dementia patients. The system is implemented with the following configuration.

[0796] The server receives data from multiple sensors at each terminal and uses this data to detect abnormal conditions. The sensors are mainly installed around the person being cared for and collect vital data and environmental data in real time. This provides a function that immediately sends an alert to the caregiver if there is an abnormality in heart rate or room temperature. Widely available Wi-Fi connected sensors are used as the hardware.

[0797] Users assess their cognitive function using a dedicated eye-tracking device. The eye-tracking device meticulously records the eye movements of the person being cared for, and this data is transmitted to a server. Based on this information, the server-side algorithm scores changes in cognitive function and provides feedback to the caregiver.

[0798] Additionally, voice data is collected by smart devices. The devices record everyday conversations and send the data to a server for voice analysis. An AI model in the cloud analyzes changes in speech patterns, which can serve as an indicator of dementia progression.

[0799] The server integrates this data and uses a generative AI model to suggest synthesized content to caregivers. The generated content serves as a guide for caregivers to implement the most appropriate individualized support for the patient. For example, if a patient has comforting memories from the past, images of those memories can be synthesized and displayed for stress reduction purposes.

[0800] For example, if a patient becomes restless in the afternoon, the system will detect an abnormal increase in heart rate during that time and notify the caregiver's device with a relaxing, synthesized video from the past. An example of this prompt message would be: "Please generate a calming synthesized story using past photos. This content should help the specific patient relax."

[0801] This allows caregivers to effectively understand the condition of dementia patients and provide real-time support tailored to their individual needs.

[0802] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0803] Step 1:

[0804] The device collects sensor data such as the care recipient's body temperature, heart rate, and room temperature in real time. It takes data from each sensor as input, organizes it within the device, and then sends it to the server. The output is a set of the organized sensor data sent to the server.

[0805] Step 2:

[0806] The server analyzes the received sensor data to detect abnormal conditions. The input is sensor data sent from the terminal, and abnormalities are visualized by comparing it with normal data stored in the database. If an abnormality is detected through this process, alert information is generated.

[0807] Step 3:

[0808] The user uses an eye-tracking device to capture the eye movements of the person being cared for. The eye movement data is recorded as input on the device and sent to the server. As output, the eye movement data is formatted and becomes data for analysis on the server.

[0809] Step 4:

[0810] The server evaluates cognitive function based on the received eye movement data. Using eye movement data as input, it applies an algorithm to calculate a cognitive function score. The output is the cognitive function score as the evaluation result, which is provided to the caregiver.

[0811] Step 5:

[0812] The device records everyday conversations and transfers the audio data to a server. The input is the voice of the person receiving care, which is recorded clearly and then sent to the server. The audio file is then formatted and used as material for analysis.

[0813] Step 6:

[0814] The server analyzes audio data and identifies specific speech patterns. Using audio data as input, it extracts language patterns using a generative AI model. The output generates hints for caregivers based on the identified speech patterns.

[0815] Step 7:

[0816] The server generates synthetic content using past visual information and generative AI. It analyzes the input visual data using the prompt: "Generate a calming synthetic story using past photographs. This content should help a specific patient relax." The output is the generated synthetic content, provided as text and images appropriate to the patient's condition.

[0817] Step 8:

[0818] The terminal receives synthesized content and alerts sent from the server and sends push notifications to the caregiver. The input is notification data from the server, which is immediately displayed on the caregiver's device to enable early response by the caregiver. The output is the alerts and synthesized content displayed on the actual device.

[0819] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0820] This invention provides caregivers with more effective response strategies by combining an emotion engine with a system designed to support the care of dementia patients. The system analyzes sensor data, eye movement data, and voice data, as well as performing emotion recognition and providing feedback based on the user's psychological state.

[0821] Collection and analysis of emotional data

[0822] The device monitors the user's facial expressions through its built-in camera and acquires data in real time. This data is input into an emotion engine to determine emotional states such as smiles, anger, and sadness.

[0823] Furthermore, the device records the user's voice tone and speaking speed, and uses these voice characteristics to estimate emotions and make more accurate emotional judgments.

[0824] Emotion-based content delivery

[0825] The server generates synthesized content tailored to the user's emotions based on the results obtained from the emotion engine. For example, if the user is feeling anxious, it will provide a visual story or music with a relaxing effect.

[0826] The device presents this content to the user, aiming for the displayed content to calm the user's state. Simultaneously, it collects user feedback again, continuing to form a feedback loop.

[0827] As a concrete example, suppose a user is watching television in the living room with an anxious expression. The system identifies this and, via the server, provides video and audio based on past experiences where the user felt reassured. As a result, the user's expression softens, and their heart rate and other vital data stabilize. Through this feedback mechanism, the system provides caregivers with an effective and consistent care strategy.

[0828] The following describes the processing flow.

[0829] Step 1:

[0830] The device uses its built-in camera to monitor the user's face and periodically captures facial expression data and features. This data includes the movement and position of each part of the face.

[0831] Step 2:

[0832] The device transmits the acquired facial expression data to the emotion engine. The emotion engine analyzes the received data and applies an algorithm to determine emotions such as smiles, surprise, and anger.

[0833] Step 3:

[0834] The device records the user's voice using a microphone and extracts the tone and speaking speed of the voice. This audio data is used to identify variations in the volume and tempo of speech.

[0835] Step 4:

[0836] The device inputs voice data into an emotion engine, which then estimates emotions from the voice characteristics. This allows for a more comprehensive identification of emotions.

[0837] Step 5:

[0838] The server identifies the user's emotional state based on the analysis results and generates appropriate synthesized content. It selects the appropriate visual story and music templates based on the emotional state.

[0839] Step 6:

[0840] The device presents synthesized content sent from the server to the user and observes the impact the content has on the user's emotions.

[0841] Step 7:

[0842] The device records the user's response again during the presentation and sends that data to the server for analysis.

[0843] Step 8:

[0844] Based on the newly obtained response data, the server adjusts the next proposed countermeasures and content, forming a feedback loop.

[0845] This process allows the system to continuously monitor users' emotions and provide care tailored to their individual needs.

[0846] (Example 2)

[0847] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0848] In an aging society, there is a need to provide efficient and effective psychological care for dementia patients. Traditional methods rely heavily on the caregiver's experience and intuition, making it difficult to respond appropriately to the patient's emotions and psychological state. Therefore, new technologies are needed to accurately recognize the patient's emotions and provide appropriate synthetic information to promote their psychological stability.

[0849] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0850] In this invention, the server includes means for collecting facial expression information and performing emotion recognition, means for analyzing voice characteristics and evaluating psychological state, and means for generating synthesized information based on the emotion recognition results. This makes it possible to provide appropriate care based on the patient's emotional state and promote psychological stability.

[0851] "Facial expression information" refers to visual data obtained from the user's face and is used to determine their emotional state.

[0852] "Emotion recognition" is a technology that analyzes facial expressions and voice characteristics to identify a user's emotional state.

[0853] "Vocal features" refer to data such as tone, speech rate, and volume obtained from speech, and are used to evaluate the user's psychological state.

[0854] "Synthetic information" refers to visual or auditory content generated in response to the user's emotional state, and is provided for the purpose of supporting psychological stabilization.

[0855] "Feedback" is the process of recording user responses again using sensors, evaluating the effectiveness of the system, and making adjustments as needed.

[0856] "Vital information" refers to data that indicates the user's physical condition, such as heart rate, blood pressure, and body temperature.

[0857] "Environmental information" refers to data that indicates external conditions that affect the user's psychological state, such as ambient noise, light intensity, and temperature.

[0858] This invention relates to a system for supporting the psychological care of dementia patients, aiming to monitor the user's emotional state in real time and provide appropriate care. This system can be implemented as follows:

[0859] Data collection and analysis

[0860] The device collects user facial expressions and voice characteristics using its built-in camera and microphone. Facial expressions are captured using image analysis software with real-time data processing capabilities to capture facial movements and subtle changes in facial muscles. Voice characteristics are analyzed using voice analysis software to measure tone, speed, and volume based on the collected voice data.

[0861] The server aggregates this data and uses a generative AI model to perform emotion recognition. This model includes various pre-trained emotion analysis algorithms and can accurately distinguish between diverse emotional states.

[0862] Content generation and feedback

[0863] Based on the results of emotion recognition, the server generates synthetic information appropriate to the user's psychological state. For example, if the user is showing signs of anxiety, it will create relaxing music or calming visuals. It is also possible to adjust the content based on past data.

[0864] The device provides the generated composite information to the user, presenting it through the screen and speaker. The device also records the user's new responses and sends them to the server, forming a feedback loop.

[0865] Specific example

[0866] For example, if a user appears bored, the system can detect this and generate and present images of refreshing natural scenery or music. This approach allows users to regain a sense of calm and improve their quality of daily life.

[0867] Examples of prompts for generative AI models

[0868] "Regarding the development of a psychological care support system for dementia patients, please explain the specific methods for sensing emotions and generating appropriate content based on those results."

[0869] This system configuration makes it possible to support dementia patients in leading more stable lives.

[0870] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0871] Step 1:

[0872] The device activates its built-in camera and microphone to capture the user's facial expressions and voice. Image data is acquired as facial information, and voice data is recorded as voice features. These form the initial dataset as sensor inputs. This dataset is transmitted to the server in real time.

[0873] Step 2:

[0874] The server processes the received facial information through image analysis software to detect the user's facial movements and changes in facial muscles, thereby determining their emotional state. The output of this process is a list of possible emotions the user may be exhibiting, each with a confidence score.

[0875] Step 3:

[0876] The server processes the audio data using speech analysis software. It analyzes speech features (tone, speed, volume) and evaluates the user's psychological state. This analysis complements the list of emotions obtained in the previous step, improving the accuracy of emotion recognition.

[0877] Step 4:

[0878] The server uses a generative AI model to generate synthetic content based on the aforementioned emotional states. This content is designed to provide relaxation and a sense of security. The input data is a list of emotions, and the output is synthetic visual and auditory content.

[0879] Step 5:

[0880] The terminal receives synthesized content sent from the server and presents it to the user. The content is output via the terminal's display and speakers. The user's new responses are again recorded by the terminal and sent back to the server for the next processing cycle to form a feedback loop.

[0881] (Application Example 2)

[0882] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0883] As society ages, caring for dementia patients has become a critical issue. In particular, accurately understanding the emotional and psychological changes of dementia patients and providing appropriate care accordingly is extremely difficult in care facilities. Solving this problem is essential.

[0884] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0885] In this invention, the server includes means for receiving sensor data and detecting abnormal conditions, means for acquiring eye movement data and evaluating cognitive function, means for analyzing audio data and identifying specific speech patterns, means for generating synthesized content using past visual information, and means for analyzing the emotions of dementia patients in real time and presenting care suggestions on a visual display device. This enables care staff to provide appropriate care based on the emotional state of dementia patients quickly and effectively.

[0886] "Sensor data" refers to information about the environment and the state of the object obtained from various sensors.

[0887] "Detecting abnormal conditions" means recognizing a state that is different from the normal state or a problematic state.

[0888] "Eye movement data" refers to information about visual movements such as gaze and blinking.

[0889] "Cognitive function" is a general term for functions involved in human intellectual activities, such as thinking, memory, judgment, and learning.

[0890] "Audio data" refers to physical sound signals that contain information about vocalization.

[0891] "Speech patterns" refer to a set of characteristics in speech, such as word choice, rhythm and speed of language use, etc.

[0892] "Visual information" refers to information that is represented as images or videos.

[0893] "Synthetic content" refers to visually and aurally appealing content created by combining multiple sources of information.

[0894] A "dementia patient" refers to a person who has a medical condition characterized primarily by memory loss and confused thinking.

[0895] "Real-time emotional analysis" means instantly judging emotions and immediately evaluating their changes.

[0896] A "visual display device" is a technological device used to visually represent information.

[0897] "Care proposals" refer to suggestions that indicate appropriate actions and countermeasures in caregiving.

[0898] This invention is a system for streamlining the care of dementia patients in nursing care facilities. The system analyzes the patient's emotions in real time using sensor data, eye movement data, and voice data. A specific embodiment of this system is described here.

[0899] The server receives sensor data, and if an abnormal condition is detected, immediate action is taken. Eye movements are tracked using a camera, and cognitive function is evaluated using visual information displayed on smart glasses or a headset. Speech recognition software analyzes the patient's voice data and identifies specific speech patterns. It also generates relaxing synthetic content based on previously collected visual information. This synthetic content includes music and visual stories.

[0900] The device uses an emotion analysis engine to assess the patient's emotions in real time and presents care suggestions based on the results to the caregiver via a visual display. For example, if the device detects that the patient is anxious, it will display appropriate care methods such as "Stay calmly by the patient's side and listen to them."

[0901] As a concrete example, let's consider its use in a nursing home. In this facility, care staff wear smart glasses, and if it detects that a dementia patient is irritable, it automatically suggests "playing calming music while taking a short walk with the patient." In this way, specific care based on the patient's emotional state is provided quickly.

[0902] An example of a prompt might be the instruction, "Design an application that tracks the emotional state of users in this facility in real time and provides appropriate advice to staff."

[0903] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0904] Step 1:

[0905] The server receives sensor data, eye movement data, and audio data from the terminal. This input data is treated as initial information for understanding the patient's condition and is stored in the server's database.

[0906] Step 2:

[0907] The server analyzes the received data. It detects abnormal conditions based on sensor data and evaluates cognitive function using eye movement data. Audio data is analyzed using speech recognition software to identify specific speech patterns. As a result of this analysis, various status indicators are output and recorded on the server.

[0908] Step 3:

[0909] The server uses a generative AI model to analyze the patient's emotional state. It provides the model with inputs such as abnormal conditions, cognitive function assessment results, and speech patterns, and determines the patient's emotions in real time. The resulting output represents the patient's emotional state, which forms the basis for creating care recommendations.

[0910] Step 4:

[0911] The server creates synthesized content based on the analysis results. Using past visual information as a reference, it generates music and visual stories best suited to the current emotional state. This content is then sent to the device.

[0912] Step 5:

[0913] The terminal displays the received synthesized content on a visual display device, visually showing care suggestions that are emotionally responsive to the caregiver. These suggestions encourage specific actions that are helpful in face-to-face care.

[0914] Step 6:

[0915] The caregiver, as the user, provides appropriate care to the patient based on the information obtained from the device. If necessary, additional data and feedback can be sent to the server via the device, enabling continuous improvement of the care provided.

[0916] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0917] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0918] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0919] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0920] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0921] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0922] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0923] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0924] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0925] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0926] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0927] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0928] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0929] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0930] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0931] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0932] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0933] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0934] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0935] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0936] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0937] The following is further disclosed regarding the embodiments described above.

[0938] (Claim 1)

[0939] A means for receiving sensor data and detecting abnormal conditions,

[0940] A means of acquiring eye movement data and evaluating cognitive function,

[0941] A means for analyzing audio data and identifying specific speech patterns,

[0942] A means of generating synthetic content using past visual information,

[0943] A system that includes this.

[0944] (Claim 2)

[0945] The system according to claim 1, further comprising means for transmitting vital data and environmental data to a server.

[0946] (Claim 3)

[0947] The system according to claim 1, further comprising means for recording user responses and evaluating the effectiveness of proposed countermeasures.

[0948] "Example 1"

[0949] (Claim 1)

[0950] A configuration that receives sensor information and detects abnormal conditions,

[0951] A configuration that acquires eye-tracking data and evaluates cognitive function,

[0952] A configuration that analyzes audio information and identifies specific speech patterns,

[0953] A configuration that generates composite content using past visual data,

[0954] A configuration that adjusts the generation algorithm using feedback obtained from users,

[0955] A system that includes this.

[0956] (Claim 2)

[0957] The system according to claim 1, further comprising a configuration for transmitting vital information and environmental information to a central processing unit.

[0958] (Claim 3)

[0959] The system according to claim 1, further comprising a configuration for recording user responses, evaluating the effectiveness of proposed countermeasures, and making appropriate adjustments.

[0960] "Application Example 1"

[0961] (Claim 1)

[0962] A means for receiving sensor data and detecting abnormal conditions,

[0963] A means of acquiring eye movement data and evaluating cognitive function,

[0964] A means for analyzing audio data and identifying specific speech patterns,

[0965] A means of generating synthetic content using past visual information,

[0966] A means of integrating information from multiple data sources in real time and generating guidelines for implementing proposed care methods,

[0967] A means of communicating warnings and suggestions to caregivers using push notifications,

[0968] A system that includes this.

[0969] (Claim 2)

[0970] The system according to claim 1, further comprising means for transmitting vital data and environmental data to a server.

[0971] (Claim 3)

[0972] The system according to claim 1, further comprising means for recording user responses and evaluating the effectiveness of proposed countermeasures.

[0973] "Example 2 of combining an emotion engine"

[0974] (Claim 1)

[0975] A means of collecting facial expression information and performing emotion recognition,

[0976] A method for analyzing voice characteristics to evaluate psychological state,

[0977] A means for generating synthetic information based on the emotion recognition results,

[0978] A means for distributing the aforementioned synthesized information and providing feedback to stabilize the user's state,

[0979] A system that includes this.

[0980] (Claim 2)

[0981] The system according to claim 1, further comprising means for transmitting vital information and environmental information to a core device.

[0982] (Claim 3)

[0983] The system according to claim 1, further comprising means for recording user responses and forming a feedback loop to measure the effectiveness of proposed countermeasures.

[0984] "Application example 2 when combining with an emotional engine"

[0985] (Claim 1)

[0986] A means for receiving sensor data and detecting abnormal conditions,

[0987] A means of acquiring eye movement data and evaluating cognitive function,

[0988] A means for analyzing audio data and identifying specific speech patterns,

[0989] A means of generating synthetic content using past visual information,

[0990] A means of analyzing the emotions of dementia patients in real time and presenting care suggestions on a visual display device,

[0991] A system that includes this.

[0992] (Claim 2)

[0993] The system according to claim 1, further comprising means for transmitting vital data and environmental data to a server.

[0994] (Claim 3)

[0995] The system according to claim 1, further comprising means for recording user responses and evaluating the effectiveness of proposed countermeasures. [Explanation of symbols]

[0996] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving sensor data and detecting abnormal conditions, A means of acquiring eye movement data and evaluating cognitive function, A means for analyzing audio data and identifying specific speech patterns, A means of generating synthetic content using past visual information, A system that includes this.

2. The system according to claim 1, further comprising means for transmitting vital data and environmental data to a server.

3. The system according to claim 1, further comprising means for recording user responses and evaluating the effectiveness of proposed countermeasures.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A