system
A generative AI model-based system addresses the challenge of real-time health and emotional state monitoring for the elderly, enhancing care quality and reducing caregiver burden through speech and image analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Conventional care systems for the elderly struggle to accurately grasp individual health and emotional states in real time, leading to a burden on caregivers and inadequate quality of life for the elderly.
A system utilizing a generative AI model that integrates speech recognition, image analysis, and conversation generation to provide personalized care, including speech recognition means, analysis means for emotional information, conversation generation, response presentation, storage, and learning means to enhance the model's accuracy.
The system provides personalized care by accurately responding to individual needs, reducing caregiver burden and improving the quality of life for the elderly.
Smart Images

Figure 2026068426000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the care of the elderly, the shortage of labor force has become a serious problem as the aging population with low birth rates progresses. In conventional care systems, it is difficult to accurately grasp the individual health and emotional states of the elderly in real time, and there is a need to improve the quality of life of the elderly while reducing the burden on caregivers. The problem of the present invention is to provide a personal care system utilizing a generative AI model, improve communication with the elderly, and reduce the burden on caregivers.
Means for Solving the Problems
[0005] This invention includes a speech recognition means that receives voice data from a user and converts it into text data. Furthermore, it includes an analysis means that analyzes facial expressions from the voice data and image data and acquires emotional information. Based on the text data and emotional information, a conversation generation means is used to generate appropriate responses, providing communication optimized for the elderly. The system also includes a response presentation means that presents the generated responses to the user, a storage means that stores all user interactions in a database, and a learning means that uses this data to strengthen the generation model. This configuration makes it possible to respond to the individual needs of the elderly, reduce the workload of caregivers, and improve the quality of life for the elderly.
[0006] "Voice recognition means" refers to a device or software that has the function of acquiring a user's voice information as digital data and converting it into text data.
[0007] "Analysis means" refers to a device or software that has the function of analyzing acquired audio and image data to identify emotional information such as facial expressions and tone of voice.
[0008] A "conversation generation system" is a system that generates appropriate and natural language responses for the user based on text data and emotional information.
[0009] A "response presentation means" is a device or software that has the function of providing the generated linguistic response to the user as audio or text.
[0010] A "storage device" refers to a device or software that stores all data related to user interactions in a database, making it available for later analysis and learning.
[0011] "Learning methods" refer to algorithms and processes used to enhance generative models with the aim of improving the accuracy and adaptability of the overall system response, using accumulated data. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0018] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention is a system for supporting personalized care for the elderly. It optimizes communication with the elderly using a generative AI model and monitors their health status using speech recognition and image analysis technologies. The implementation of this system aims to provide individualized care services to the elderly through collaboration among multiple technical departments.
[0034] First, the user speaks or makes facial expressions towards the device's camera. The audio and image data are appropriately captured by the device. The device sends the audio data to a server, where it is converted into text through speech recognition. Simultaneously, the image data is sent to the server, where facial expressions and emotions are detected through image analysis.
[0035] The server determines the emotions and state of elderly individuals based on the results of speech recognition and image analysis. For example, if the voice detects a tone of concern about recent sleep problems and the facial expression is analyzed as anxious, the system utilizes response generation tools to prepare an appropriate response. For instance, the server might generate advice such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is sent to the terminal and presented to the user via the terminal in either voice or text format.
[0036] Furthermore, the system continuously accumulates data obtained through daily interactions. This data is sequentially analyzed by the server and used to improve the accuracy of the overall system generative model, as well as to aid in medical evaluations to detect abnormalities in specific health conditions.
[0037] In this way, the system provides personalized support to individual elderly individuals, enabling better health management and communication. For example, if a user says, "I don't feel quite right today," the server can detect this and offer kind words such as, "Don't push yourself, take a rest." This provides reassurance to the elderly and reduces the workload of caregivers.
[0038] The following describes the processing flow.
[0039] Step 1:
[0040] The user speaks to the device or makes facial expressions in front of the camera. The device records these as audio and image data.
[0041] Step 2:
[0042] The device sends voice data to the server, where a speech recognition engine converts the voice into text data.
[0043] Step 3:
[0044] The terminal sends the acquired image data to the server, which uses image analysis technology to extract facial expression data as emotional information.
[0045] Step 4:
[0046] The server evaluates the elderly person's condition based on the converted text data and analyzed emotional information, and generates appropriate conversational responses using a generative AI model.
[0047] Step 5:
[0048] The server sends the generated response to the terminal. The terminal then presents this response to the user in either audio or text format.
[0049] Step 6:
[0050] The server stores interaction data in a database, uses the stored data to enhance the generative model, and aims to improve the quality of care.
[0051] (Example 1)
[0052] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0053] In monitoring the health status and providing personalized care for the elderly, it is difficult to accurately and quickly detect changes in emotions and health conditions, and to provide appropriate, individualized responses. Therefore, there is a growing need for automated support systems that can provide a sense of security, especially for the elderly, while reducing the burden on caregivers. Furthermore, accumulating daily interaction data and utilizing it to continuously improve the accuracy of the system is also a challenge.
[0054] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0055] In this invention, the server includes a speech recognition means that receives voice information from a user and converts it into text information; an analysis means that analyzes facial expressions from the voice information and image information and obtains emotional information; and a judgment means that uses a generative AI model to determine the emotional state and detect abnormalities in specific health conditions. This makes it possible to provide appropriate responses to elderly people and to detect abnormalities in their health conditions at an early stage.
[0056] "Speech recognition means" refers to technologies and devices that receive speech information as input and convert that speech into text information.
[0057] "Analysis means" refers to technologies and devices that analyze audio and image information to obtain emotional information from facial expressions, tone of voice, etc.
[0058] "Conversation generation means" refers to technologies and devices for generating appropriate responses based on acquired text information and emotional information.
[0059] A "response presentation means" refers to a technology or device for presenting a generated response to a user in audio or text format.
[0060] "Storage means" refers to technologies and devices for recording and storing user interactions and analysis results in a database.
[0061] "Learning methods" refer to technologies and devices that use accumulated data to strengthen generative models and improve their accuracy.
[0062] "Decision-making tools" refer to technologies and devices that utilize generative AI models to judge emotional states and detect abnormalities in specific health conditions.
[0063] This invention relates to a system for monitoring the health status of elderly individuals and supporting personalized care. Specifically, it optimizes the daily communication of elderly individuals using a generative AI model, and utilizes speech recognition and image analysis technologies to detect changes in their health status early and provide appropriate responses.
[0064] First, the user speaks to the device or makes facial expressions towards the camera. The hardware used includes smartphones and tablets that capture audio and video. A dedicated application runs on these devices to process the user's interaction.
[0065] The device sends the captured audio to the server, where this audio data is converted to text by speech recognition software. Typically, services such as Google Cloud Speech-to-Text or Amazon Transcribe are used for speech recognition. Simultaneously, the device also sends image data to the server, which is then analyzed for facial expressions using tools such as OpenCV or Azure Face API.
[0066] The server uses a generative AI model based on these analysis results to determine the emotional state of the elderly person. Specifically, it generates gentle words such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is presented to the user via the device as either audio or text. Amazon Polly or Google Text-to-Speech are often used for speech synthesis.
[0067] Furthermore, the system uses data accumulated from daily interactions to improve the accuracy and detection capabilities of its generative models. This enables anomaly detection based on analysis, allowing for the prevention of major health problems before they occur.
[0068] For example, if a user says, "I'm not feeling well today," the server will detect this and create a gentle response such as, "Don't push yourself, take a rest." An example of a prompt might be, "The user said, 'I'm not feeling well today.' Please think of a gentle response to this."
[0069] Thus, this system has the effect of providing support and a sense of security to the elderly while reducing the burden on caregivers.
[0070] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0071] Step 1:
[0072] The user speaks to the device and makes facial expressions towards the camera. During this process, audio and images are captured using a smartphone or tablet. The input consists of the user's voice and image data, which are then fed into the system.
[0073] Step 2:
[0074] The terminal sends the captured audio data to the server. Based on the audio data input, speech recognition software is applied. This converts the audio data into text data. The output is the corresponding text data for the audio data.
[0075] Step 3:
[0076] Simultaneously, the device sends image data to the server, and image analysis technology is used to analyze the user's facial expressions. The input is image data, and an image analysis algorithm is applied. The analysis outputs information about the user's emotions.
[0077] Step 4:
[0078] The server receives text data from speech recognition and emotion information from image analysis, and uses a generative AI model to determine the user's emotional state. The input consists of text data and emotion information, which are then applied to the generative AI model. A prompt is used to generate a response appropriate to the user. The output is the generated response.
[0079] Step 5:
[0080] The server sends the response generated by the generative AI model to the terminal. The response is typically sent in text format.
[0081] Step 6:
[0082] The terminal presents the received response to the user in either voice or text. In this process, when using speech synthesis technology, the input is the generated response text, and the output is the synthesized voice. Various speech synthesis engines are used for speech synthesis.
[0083] Step 7:
[0084] The server stores daily interaction data in a database, which is then used to improve subsequent generative models. Here, the input is all dialogue data, and the output is stored as training data. This data will be used to improve the accuracy of future anomaly detection and generative AI models.
[0085] (Application Example 1)
[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0087] In today's factory environment, efficiently managing worker health and safety is essential, but challenges remain in providing real-time status updates and appropriate feedback to individual workers. Furthermore, there is a lack of systems that accurately monitor workers' physiological and emotional states and provide appropriate breaks and advice.
[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0089] In this invention, the server includes: speech recognition means for receiving voice information from a user and converting it into text information; analysis means for analyzing facial expressions from the voice information and visual information and acquiring emotional information; conversation generation means for generating an appropriate response based on the text information and emotional information; response presentation means for presenting the generated response to the user; storage means for storing all interactions from the user in a storage device; learning means for strengthening the generation model using the stored information; work monitoring means having the function of monitoring the physiological state of the worker and encouraging appropriate breaks based on the emotional information; and response generation means for determining the state of fatigue and stress from the worker's voice and facial information and generating appropriate advice. This makes it possible to analyze the emotions and physiological state of individual workers in real time and provide appropriate feedback quickly.
[0090] "Voice recognition means" refers to a device or software that receives voice information from a user and converts it into text information.
[0091] "Analysis means" refers to a device or system for analyzing facial expressions from audio and visual information and obtaining emotional information.
[0092] "Conversation generation means" refers to a device or algorithm that generates an appropriate response based on textual information and emotional information.
[0093] A "response presentation means" is a device or interface for presenting a generated response to a user.
[0094] A "storage means" is a device or database system for storing all user interactions in a storage device.
[0095] A "learning tool" is an algorithm or software that uses accumulated information to enhance a generative model.
[0096] "Work monitoring means" refers to a device or system that has the function of monitoring the physiological state of a worker based on emotional information and prompting appropriate breaks.
[0097] A "response generation means" is a device or software that determines the level of fatigue and stress from the worker's voice and facial expression information and generates appropriate advice.
[0098] In this invention, the entire system utilizes a highly integrated algorithm and various sensors to monitor and respond to the health and emotional state of workers in the work environment in real time.
[0099] The server consists of multiple modules. First, for speech recognition, the server uses speech recognition software to convert speech data into text data. Specific software options include open-source speech recognition tools. For image analysis, software is used to analyze image data acquired from a camera, extracting emotional information from facial expressions. Machine learning algorithms are used for this; for example, existing facial recognition models can be customized and used.
[0100] Next, the terminal uses a generation AI model to generate responses, creating appropriate responses based on the worker's state. These responses are presented to the worker in either voice or text. The server generates these responses and executes a response presentation mechanism to present them to the worker through the terminal.
[0101] The system further stores worker interaction data in memory and uses it to learn and improve its generative model. This allows the server to improve response accuracy in subsequent interactions.
[0102] For example, if a factory worker mutters, "I'm a little tired today," the system can analyze the voice, and if it determines from the worker's facial expression that they are fatigued, it can automatically provide advice such as, "Why don't you take a short break?"
[0103] An example of a prompt message could be a language input such as, "Please suggest a method to monitor worker fatigue in a factory in real time using voice and facial expressions, and encourage appropriate breaks."
[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0105] Step 1:
[0106] The server receives voice data from the user via a speech recognition system. This voice data is used as input, and data processing is performed by converting it into text data using speech recognition software. The converted text data is then output.
[0107] Step 2:
[0108] The server uses an analysis tool to receive the user's visual data acquired from the terminal's camera. This visual data is used as input, and an image analysis algorithm is used to analyze facial expressions and extract the user's emotional information. This emotional information is then obtained as output.
[0109] Step 3:
[0110] The server inputs text data and extracted sentiment information into the conversation generation system and uses a generation AI model. Based on this data, it performs data calculations to generate an appropriate response and outputs the generated response.
[0111] Step 4:
[0112] The terminal presents the generated response received from the server to the user using a response presentation mechanism. The output response is presented to the user as audio or text using a speech synthesis system or text display software.
[0113] Step 5:
[0114] The server records all user interactions and stores them in memory using storage means. This stored data is used as input and analyzed using learning means to improve the accuracy of future generative models, which helps in building new generative AI models.
[0115] Step 6:
[0116] The server operates the work monitoring means based on the aforementioned emotional information. It checks the user's physiological state, executes a response generation means that generates advice to revise the work pace as needed, and prompts the user to take appropriate breaks.
[0117] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0118] This invention provides a system that integrates a generative AI model and an emotion engine to improve personalized care for the elderly. This enables more refined and individualized care based on the user's emotional and health condition.
[0119] First, the user speaks to the device and makes facial expressions in front of the camera. This allows the device to acquire audio and image data. When the device sends this data to the server, it filters and optimizes the data before sending it. The audio data is converted to text by speech recognition, and the image data is analyzed to extract emotional information.
[0120] The server uses analytical tools, particularly an emotion engine, to identify the user's emotional state from voice and images. For example, if the user is smiling or showing signs of sadness, the emotion engine quickly detects this information. This emotional information is also reflected in automatically generated conversations, and the AI model generates responses appropriate to the user's state. For example, if the user says, "I'm feeling a little lonely today," the system will generate encouraging words such as, "That's a shame. Why don't you make plans to meet someone?"
[0121] The generated responses are presented to the user by the device. Voice responses are output clearly through the speaker. The server also records all of these user interactions and stores them in a database. This stored data is used to improve the generative model and the emotion engine. It also serves as valuable information for experts when evaluating the user's health status.
[0122] This system will not only significantly improve the quality of care but also reduce the burden on caregivers, allowing them to dedicate more time and resources to sustained care. As a result, older adults will be able to live their days in a more comfortable and supportive environment.
[0123] The following describes the processing flow.
[0124] Step 1:
[0125] The user speaks into the device, saying, "I'm feeling a little sad today," or makes a gloomy expression in front of the camera. The device captures this audio and facial expression as digital data.
[0126] Step 2:
[0127] The terminal sends the captured audio data to the server. At this time, pre-processing is performed to remove noise and transcribe the audio clearly into text.
[0128] Step 3:
[0129] The device simultaneously sends image data to a server for use in facial expression analysis algorithms. This image data is optimized to facilitate the extraction of facial features.
[0130] Step 4:
[0131] The server uses speech recognition to convert speech data into text data. It then identifies the user's emotions from the speech content.
[0132] Step 5:
[0133] The server uses an emotion engine to determine if the user is actually sad based on voice tone and facial expression analysis. This allows for a more detailed understanding of the user's emotional state.
[0134] Step 6:
[0135] Based on the acquired emotional information, the server uses a generative AI model to generate appropriate conversational responses. For example, it might create a response like, "You seem a little lonely today, how about watching a movie that interests you?"
[0136] Step 7:
[0137] The server sends the generated response to the terminal. The terminal presents the response in an easily audible voice or easy-to-read text format.
[0138] Step 8:
[0139] The server stores all interaction data in a database. This data is used to improve future interactions and for expert health assessments.
[0140] (Example 2)
[0141] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0142] In personal care for the elderly, there is a need to provide individualized support based on emotional and health conditions. However, conventional systems have been insufficient in properly analyzing acquired data and generating responses, making it difficult to respond in a way that is in line with the user's feelings. As a result, it has been difficult to improve the quality of care and reduce the burden on caregivers.
[0143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0144] In this invention, the server includes acquisition means for acquiring audio and video information from the user, analysis means for converting the audio information into text information and analyzing the video information to identify the emotional state, and generation means for generating appropriate dialogue using a generation AI model based on the text information and the emotional state. This makes it possible to accurately grasp the emotional needs of the user and generate situation-appropriate responses with high efficiency.
[0145] "Acquisition means" refers to the interface and technology used to acquire audio and video information from users.
[0146] "Analysis means" refers to the process and function of converting acquired audio information into text information and identifying emotional states from video information.
[0147] "Generative means" refers to functions and systems that create appropriate dialogues using generative AI models based on text information and emotional states.
[0148] "Presentation means" refers to hardware or software used to present the generated dialogue to the user in audio or video format.
[0149] "Recording means" refers to a system or device for recording all interactions with users and storing them on a recording medium.
[0150] "Learning methods" refer to methods and mechanisms for improving the performance of generative AI models and emotion engines by utilizing accumulated records.
[0151] This invention is implemented as a system that improves personal care for the elderly using audio and video information. Specifically, it uses a server and terminals to acquire the user's audio and video, and integrates the functions of analysis, response, recording, and learning.
[0152] The user interacts with the system using a terminal equipped with a microphone and a camera. This terminal collects audio and video information by capturing the user's voice with the microphone and capturing video with the camera. This information is then transmitted from the terminal to the server.
[0153] The server uses speech recognition software to convert the user's voice information into text, and then uses an emotion engine to analyze video information. The emotional state obtained from the analyzed data is input into a generative AI model, which is used to generate appropriate dialogue.
[0154] The generative AI model uses these analysis results to generate dialogue that is tailored to the user's emotional state. For example, if a user says, "I'm feeling a little lonely today," the generated response might be something like, "That's a shame. Why don't you try making plans to meet someone?"
[0155] The generated dialogue is presented to the user via the device as audio or video. At this time, the speaker and display are used to provide the user with clear information.
[0156] The server also records all user interactions, continuously refining its generative AI models and emotion engine through learning algorithms. This recorded data is also used as a valuable source of information for experts to assess the user's health status.
[0157] As an example of a prompt, if the instruction is given, "Provide the best words of encouragement to give when a user is feeling lonely," the generative AI model will produce an appropriate response.
[0158] The implementation of this system is expected to enable care tailored to the emotional and physical condition of the elderly, thereby improving the quality of care while simultaneously reducing the burden on caregivers.
[0159] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0160] Step 1:
[0161] The user speaks to the device and makes facial expressions in front of the camera. The user's voice information is captured by the microphone, and video information is captured by the camera. The input is the user's voice and facial expressions, and the output is their digital data.
[0162] Step 2:
[0163] The device optimizes the acquired audio and video information. Specifically, it performs filtering such as noise reduction and image quality improvement. This process results in clean audio data and high-quality image data being output.
[0164] Step 3:
[0165] The terminal sends optimized audio and video data to the server. The audio data is converted into text data by speech recognition software on the server; the input is audio data, and the output is text data.
[0166] Step 4:
[0167] The server inputs video data into the emotion engine, which performs analysis to identify the user's emotional state. This results in outputting information about the emotional state, which is then provided to the generative AI model.
[0168] Step 5:
[0169] The server generates appropriate dialogue using a generative AI model based on text data and emotional states. The input is text data and emotional information, and the output is a response message to the user.
[0170] Step 6:
[0171] The terminal presents the response message received from the server to the user via speech synthesis or display. Specifically, it provides the message via voice using the speaker or displays it as text using the display.
[0172] Step 7:
[0173] The server records all user interactions and stores them in a recording medium. This data is used to continuously improve the generative AI model and emotion engine. The input is the history of user interactions, and the output is the training data store.
[0174] (Application Example 2)
[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0176] Conventional personal care systems for the elderly have struggled to grasp individual emotional states in real time and provide appropriate services based on that understanding. Furthermore, the services and responses provided were uniform, failing to adequately address the needs of each user. Moreover, they lacked the ability to effectively analyze users' emotions and states and propose appropriate services based on that analysis.
[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0178] In this invention, the server includes speech recognition means for converting speech data into text data, analysis means for analyzing emotional information from speech and image data, conversation generation means for generating appropriate responses based on the acquired emotional information, service suggestion means for presenting service suggestions according to the user's emotional state, portable display device linkage means for displaying the generated responses and service suggestions via a portable display device, storage means for accumulating user interactions in a database, and learning means for strengthening the generation model using the stored data. This enables the provision of quick and accurate services tailored to the individual emotional state of the user.
[0179] "Speech recognition means" refers to technology that converts voice data from users into text data in real time.
[0180] "Analysis means" refers to a technology that analyzes the user's facial expressions and voice tone from audio and image data to obtain emotional information.
[0181] A "conversation generation method" is a technology that generates appropriate responses tailored to the user based on acquired text data and emotional information.
[0182] A "response presentation method" is a technology that presents a generated response to the user and conveys it in an easily understandable format.
[0183] A "service proposal method" is a technology that proposes the most suitable service based on the user's emotional state and presents the user with choices.
[0184] "Portable display device integration means" refers to a technology that displays responses and service proposals generated through a portable display device worn by the user, and provides real-time information.
[0185] A "storage method" is a technology that records all interactions in a database and stores data in preparation for future analysis and learning.
[0186] "Learning methods" refer to technologies that use accumulated data to enhance the response generation capabilities and accuracy of generative AI models, thereby improving the overall system performance.
[0187] This system is designed to provide personalized services tailored to the user's emotional state. The server receives voice and image data from the user, analyzes it, and generates responses. The following hardware and software are used to implement each function.
[0188] First, the device used by the user is a portable display device such as a smartphone, tablet, or smart glasses. This device is equipped with a microphone to collect voice data and a camera to capture facial expressions.
[0189] For speech recognition and text conversion, the server utilizes recognition technologies such as "Google Cloud Speech-to-Text API" and "IBM Watson® Speech to Text." For analysis, it uses "Microsoft® Azure Face API" and "Google Cloud Vision API" to identify facial expressions and emotional information from image data.
[0190] The conversation generation uses generative AI models such as "OpenAI® GPT-3®". This generates appropriate and personalized responses based on the user's emotional information. The generated conversation is displayed as text on the device's screen and also includes service suggestions. Voice responses are played back through the device's speaker.
[0191] Furthermore, through a portable display device integration mechanism, the response information and suggestions displayed on smart glasses and other devices are presented in a way that matches the user's gaze. This feedback function allows service providers and family members to understand the user's condition in real time and take appropriate action.
[0192] As a concrete example, if a user says, "I'm feeling a little stressed today," the system will generate and display suggestions for relaxation methods and places to seek advice. An example of a prompt would be, "Generate appropriate suggestions if the user says, 'I'm lonely.'"
[0193] This invention will significantly advance care systems for the elderly, creating an environment where users can live their daily lives more comfortably and with greater peace of mind.
[0194] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0195] Step 1:
[0196] When a user speaks into the device, the device captures the audio data in real time. This audio data becomes the input. Using the microphone inside the device, high-quality audio input is obtained and prepared for transmission to the server.
[0197] Step 2:
[0198] The device uses its camera to capture image data of the user's face and sends it to the server as input data. From this image data, basic data is collected to understand the user's emotional state.
[0199] Step 3:
[0200] The server converts the received audio data into text data using APIs such as "Google Cloud Speech-to-Text API". This process converts the audio signal into appropriate text information. The converted text data is then output.
[0201] Step 4:
[0202] Simultaneously, the server uses the "Microsoft Azure Face API" to analyze the user's facial features from the image data and extract emotional information. A feature extraction algorithm is applied to the input image data, and the emotional state is output.
[0203] Step 5:
[0204] The server sends the acquired text data and sentiment information as input to a generative AI model, which then generates appropriate responses and service suggestions. This process uses generative AI models such as "OpenAI GPT-3," which output customized responses based on the input information.
[0205] Step 6:
[0206] The terminal receives output from the server and displays the generated response and service proposal on its screen. The displayed information is provided in a visually easy-to-understand format for the user, and voice responses are also output through the speaker.
[0207] Step 7:
[0208] Once the user interaction is complete, the device sends all data to the server, where it is stored in the database. This stored data will be used for future data analysis and training of generative models.
[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0212] [Second Embodiment]
[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0225] This invention is a system for supporting personalized care for the elderly. It optimizes communication with the elderly using a generative AI model and monitors their health status using speech recognition and image analysis technologies. The implementation of this system aims to provide individualized care services to the elderly through collaboration among multiple technical departments.
[0226] First, the user speaks or makes facial expressions towards the device's camera. The audio and image data are appropriately captured by the device. The device sends the audio data to a server, where it is converted into text through speech recognition. Simultaneously, the image data is sent to the server, where facial expressions and emotions are detected through image analysis.
[0227] The server determines the emotions and state of elderly individuals based on the results of speech recognition and image analysis. For example, if the voice detects a tone of concern about recent sleep problems and the facial expression is analyzed as anxious, the system utilizes response generation tools to prepare an appropriate response. For instance, the server might generate advice such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is sent to the terminal and presented to the user via the terminal in either voice or text format.
[0228] Furthermore, the system continuously accumulates data obtained through daily interactions. This data is sequentially analyzed by the server and used to improve the accuracy of the overall system generative model, as well as to aid in medical evaluations to detect abnormalities in specific health conditions.
[0229] In this way, the system provides personalized support to individual elderly individuals, enabling better health management and communication. For example, if a user says, "I don't feel quite right today," the server can detect this and offer kind words such as, "Don't push yourself, take a rest." This provides reassurance to the elderly and reduces the workload of caregivers.
[0230] The following describes the processing flow.
[0231] Step 1:
[0232] The user speaks to the device or makes facial expressions in front of the camera. The device records these as audio and image data.
[0233] Step 2:
[0234] The device sends voice data to the server, where a speech recognition engine converts the voice into text data.
[0235] Step 3:
[0236] The terminal sends the acquired image data to the server, which uses image analysis technology to extract facial expression data as emotional information.
[0237] Step 4:
[0238] The server evaluates the elderly person's condition based on the converted text data and analyzed emotional information, and generates appropriate conversational responses using a generative AI model.
[0239] Step 5:
[0240] The server sends the generated response to the terminal. The terminal then presents this response to the user in either audio or text format.
[0241] Step 6:
[0242] The server stores interaction data in a database, uses the stored data to enhance the generative model, and aims to improve the quality of care.
[0243] (Example 1)
[0244] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0245] In monitoring the health status and providing personalized care for the elderly, it is difficult to accurately and quickly detect changes in emotions and health conditions, and to provide appropriate, individualized responses. Therefore, there is a growing need for automated support systems that can provide a sense of security, especially for the elderly, while reducing the burden on caregivers. Furthermore, accumulating daily interaction data and utilizing it to continuously improve the accuracy of the system is also a challenge.
[0246] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0247] In this invention, the server includes a speech recognition means that receives voice information from a user and converts it into text information; an analysis means that analyzes facial expressions from the voice information and image information and obtains emotional information; and a judgment means that uses a generative AI model to determine the emotional state and detect abnormalities in specific health conditions. This makes it possible to provide appropriate responses to elderly people and to detect abnormalities in their health conditions at an early stage.
[0248] "Speech recognition means" refers to technologies and devices that receive speech information as input and convert that speech into text information.
[0249] "Analysis means" refers to technologies and devices that analyze audio and image information to obtain emotional information from facial expressions, tone of voice, etc.
[0250] "Conversation generation means" refers to technologies and devices for generating appropriate responses based on acquired text information and emotional information.
[0251] A "response presentation means" refers to a technology or device for presenting a generated response to a user in audio or text format.
[0252] "Storage means" refers to technologies and devices for recording and storing user interactions and analysis results in a database.
[0253] "Learning methods" refer to technologies and devices that use accumulated data to strengthen generative models and improve their accuracy.
[0254] "Decision-making tools" refer to technologies and devices that utilize generative AI models to judge emotional states and detect abnormalities in specific health conditions.
[0255] This invention relates to a system for monitoring the health status of elderly individuals and supporting personalized care. Specifically, it optimizes the daily communication of elderly individuals using a generative AI model, and utilizes speech recognition and image analysis technologies to detect changes in their health status early and provide appropriate responses.
[0256] First, the user speaks to the device or makes facial expressions towards the camera. The hardware used includes smartphones and tablets that capture audio and video. A dedicated application runs on these devices to process the user's interaction.
[0257] The device sends the captured audio to a server, where this audio data is converted to text by speech recognition software. Typically, services such as Google Cloud Speech-to-Text or Amazon Transcribe are used for speech recognition. Simultaneously, the device also sends image data to the server, which is then used for facial expression analysis using tools like OpenCV or Azure Face API.
[0258] The server uses a generative AI model based on these analysis results to determine the emotional state of the elderly person. Specifically, it generates gentle words such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is presented to the user via the device as either audio or text. Amazon Polly or Google Text-to-Speech are often used for speech synthesis.
[0259] Furthermore, the system uses data accumulated from daily interactions to improve the accuracy and detection capabilities of its generative models. This enables anomaly detection based on analysis, allowing for the prevention of major health problems before they occur.
[0260] For example, if a user says, "I'm not feeling well today," the server will detect this and create a gentle response such as, "Don't push yourself, take a rest." An example of a prompt might be, "The user said, 'I'm not feeling well today.' Please think of a gentle response to this."
[0261] Thus, this system has the effect of providing support and a sense of security to the elderly while reducing the burden on caregivers.
[0262] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0263] Step 1:
[0264] The user speaks to the device and makes facial expressions towards the camera. During this process, audio and images are captured using a smartphone or tablet. The input consists of the user's voice and image data, which are then fed into the system.
[0265] Step 2:
[0266] The terminal sends the captured audio data to the server. Based on the audio data input, speech recognition software is applied. This converts the audio data into text data. The output is the corresponding text data for the audio data.
[0267] Step 3:
[0268] Simultaneously, the device sends image data to the server, and image analysis technology is used to analyze the user's facial expressions. The input is image data, and an image analysis algorithm is applied. The analysis outputs information about the user's emotions.
[0269] Step 4:
[0270] The server receives text data from speech recognition and emotion information from image analysis, and uses a generative AI model to determine the user's emotional state. The input consists of text data and emotion information, which are then applied to the generative AI model. A prompt is used to generate a response appropriate to the user. The output is the generated response.
[0271] Step 5:
[0272] The server sends the response generated by the generative AI model to the terminal. The response is typically sent in text format.
[0273] Step 6:
[0274] The terminal presents the received response to the user in either voice or text. In this process, when using speech synthesis technology, the input is the generated response text, and the output is the synthesized voice. Various speech synthesis engines are used for speech synthesis.
[0275] Step 7:
[0276] The server stores daily interaction data in a database, which is then used to improve subsequent generative models. Here, the input is all dialogue data, and the output is stored as training data. This data will be used to improve the accuracy of future anomaly detection and generative AI models.
[0277] (Application Example 1)
[0278] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0279] In today's factory environment, efficiently managing worker health and safety is essential, but challenges remain in providing real-time status updates and appropriate feedback to individual workers. Furthermore, there is a lack of systems that accurately monitor workers' physiological and emotional states and provide appropriate breaks and advice.
[0280] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0281] In this invention, the server includes: speech recognition means for receiving voice information from a user and converting it into text information; analysis means for analyzing facial expressions from the voice information and visual information and acquiring emotional information; conversation generation means for generating an appropriate response based on the text information and emotional information; response presentation means for presenting the generated response to the user; storage means for storing all interactions from the user in a storage device; learning means for strengthening the generation model using the stored information; work monitoring means having the function of monitoring the physiological state of the worker and encouraging appropriate breaks based on the emotional information; and response generation means for determining the state of fatigue and stress from the worker's voice and facial information and generating appropriate advice. This makes it possible to analyze the emotions and physiological state of individual workers in real time and provide appropriate feedback quickly.
[0282] "Voice recognition means" refers to a device or software that receives voice information from a user and converts it into text information.
[0283] "Analysis means" refers to a device or system for analyzing facial expressions from audio and visual information and obtaining emotional information.
[0284] The "conversation generation means" is a device or algorithm that generates appropriate responses based on character information and emotional information.
[0285] The "response presentation means" is a device or interface for presenting the generated response to the user.
[0286] The "storage means" is a device or database system for storing all interactions from the user in a storage device.
[0287] The "learning means" is an algorithm or software for strengthening the generation model using the stored information.
[0288] The "work monitoring means" is a device or system that monitors the physiological state of the worker based on emotional information and has a function to prompt appropriate breaks.
[0289] The "response generation means" is a device or software that determines the state of fatigue and stress from the voice and facial expression information of the worker and generates appropriate advice.
[0290] In this invention, by using a highly integrated algorithm and various sensors for the entire system, the health state and emotional state of the worker in the working environment are monitored in real time and corresponding actions are taken.
[0291] The server is composed of multiple modules. First, for speech recognition, the server uses speech recognition software to convert speech data into character data. As specific software to be used, open-source speech recognition tools can be considered. Also, for image analysis, software that analyzes the image data acquired from a camera is used to obtain emotional information from the facial expression. A machine learning algorithm is used for this, and for example, an existing facial expression recognition model can be customized and used.
[0292] Next, the terminal uses a generation AI model to generate responses, creating appropriate responses based on the worker's state. These responses are presented to the worker in either voice or text. The server generates these responses and executes a response presentation mechanism to present them to the worker through the terminal.
[0293] The system further stores worker interaction data in memory and uses it to learn and improve its generative model. This allows the server to improve response accuracy in subsequent interactions.
[0294] For example, if a factory worker mutters, "I'm a little tired today," the system can analyze the voice, and if it determines from the worker's facial expression that they are fatigued, it can automatically provide advice such as, "Why don't you take a short break?"
[0295] An example of a prompt message could be a language input such as, "Please suggest a method to monitor worker fatigue in a factory in real time using voice and facial expressions, and encourage appropriate breaks."
[0296] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0297] Step 1:
[0298] The server receives voice data from the user via a speech recognition system. This voice data is used as input, and data processing is performed by converting it into text data using speech recognition software. The converted text data is then output.
[0299] Step 2:
[0300] The server uses an analysis tool to receive the user's visual data acquired from the terminal's camera. This visual data is used as input, and an image analysis algorithm is used to analyze facial expressions and extract the user's emotional information. This emotional information is then obtained as output.
[0301] Step 3:
[0302] The server inputs the character data and the extracted emotion information into the conversation generation means and uses the generation AI model. Based on this data, it performs data operations to generate an appropriate response and outputs the generated response.
[0303] Step 4:
[0304] The terminal presents the generated response received from the server to the user using the response presentation means. The output response is presented to the user in voice or text using a speech synthesis system or text display software.
[0305] Step 5:
[0306] The server records all interactions from the user and stores them in the storage device using the storage means. Using this stored data as input, it analyzes using the learning means for improving the accuracy of the future generation model and uses it to build a new generation AI model.
[0307] Step 6:
[0308] The server operates the work monitoring means based on the above-mentioned emotion information. It checks the user's physiological state, executes the response generation means for generating advice to review the work pace as needed, and prompts the user to take appropriate breaks.
[0309] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion specific model 59 and perform specific processing using the user's emotion.
[0310] This invention provides a system that integrates a generation AI model and an emotion engine to improve personal care for the elderly. Thereby, it realizes more refined and individualized care based on the user's emotional and health states.
[0311] First, the user speaks to the device and makes facial expressions in front of the camera. This allows the device to acquire audio and image data. When the device sends this data to the server, it filters and optimizes the data before sending it. The audio data is converted to text by speech recognition, and the image data is analyzed to extract emotional information.
[0312] The server uses analytical tools, particularly an emotion engine, to identify the user's emotional state from voice and images. For example, if the user is smiling or showing signs of sadness, the emotion engine quickly detects this information. This emotional information is also reflected in automatically generated conversations, and the AI model generates responses appropriate to the user's state. For example, if the user says, "I'm feeling a little lonely today," the system will generate encouraging words such as, "That's a shame. Why don't you make plans to meet someone?"
[0313] The generated responses are presented to the user by the device. Voice responses are output clearly through the speaker. The server also records all of these user interactions and stores them in a database. This stored data is used to improve the generative model and the emotion engine. It also serves as valuable information for experts when evaluating the user's health status.
[0314] This system will not only significantly improve the quality of care but also reduce the burden on caregivers, allowing them to dedicate more time and resources to sustained care. As a result, older adults will be able to live their days in a more comfortable and supportive environment.
[0315] The following describes the processing flow.
[0316] Step 1:
[0317] The user speaks into the device, saying, "I'm feeling a little sad today," or makes a gloomy expression in front of the camera. The device captures this audio and facial expression as digital data.
[0318] Step 2:
[0319] The terminal sends the captured audio data to the server. At this time, pre-processing is performed to remove noise and transcribe the audio clearly into text.
[0320] Step 3:
[0321] The device simultaneously sends image data to a server for use in facial expression analysis algorithms. This image data is optimized to facilitate the extraction of facial features.
[0322] Step 4:
[0323] The server uses speech recognition to convert speech data into text data. It then identifies the user's emotions from the speech content.
[0324] Step 5:
[0325] The server uses an emotion engine to determine if the user is actually sad based on voice tone and facial expression analysis. This allows for a more detailed understanding of the user's emotional state.
[0326] Step 6:
[0327] Based on the acquired emotional information, the server uses a generative AI model to generate appropriate conversational responses. For example, it might create a response like, "You seem a little lonely today, how about watching a movie that interests you?"
[0328] Step 7:
[0329] The server sends the generated response to the terminal. The terminal presents the response in an easily audible voice or easy-to-read text format.
[0330] Step 8:
[0331] The server stores all interaction data in a database. This data is used to improve future interactions and for expert health assessments.
[0332] (Example 2)
[0333] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0334] In personal care for the elderly, there is a need to provide individualized support based on emotional and health conditions. However, conventional systems have been insufficient in properly analyzing acquired data and generating responses, making it difficult to respond in a way that is in line with the user's feelings. As a result, it has been difficult to improve the quality of care and reduce the burden on caregivers.
[0335] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0336] In this invention, the server includes acquisition means for acquiring audio and video information from the user, analysis means for converting the audio information into text information and analyzing the video information to identify the emotional state, and generation means for generating appropriate dialogue using a generation AI model based on the text information and the emotional state. This makes it possible to accurately grasp the emotional needs of the user and generate situation-appropriate responses with high efficiency.
[0337] "Acquisition means" refers to the interface and technology used to acquire audio and video information from users.
[0338] "Analysis means" refers to the process and function of converting acquired audio information into text information and identifying emotional states from video information.
[0339] "Generative means" refers to functions and systems that create appropriate dialogues using generative AI models based on text information and emotional states.
[0340] "Presentation means" refers to hardware or software used to present the generated dialogue to the user in audio or video format.
[0341] "Recording means" refers to a system or device for recording all interactions with users and storing them on a recording medium.
[0342] "Learning methods" refer to methods and mechanisms for improving the performance of generative AI models and emotion engines by utilizing accumulated records.
[0343] This invention is implemented as a system that improves personal care for the elderly using audio and video information. Specifically, it uses a server and terminals to acquire the user's audio and video, and integrates the functions of analysis, response, recording, and learning.
[0344] The user interacts with the system using a terminal equipped with a microphone and a camera. This terminal collects audio and video information by capturing the user's voice with the microphone and capturing video with the camera. This information is then transmitted from the terminal to the server.
[0345] The server uses speech recognition software to convert the user's voice information into text, and then uses an emotion engine to analyze video information. The emotional state obtained from the analyzed data is input into a generative AI model, which is used to generate appropriate dialogue.
[0346] The generative AI model uses these analysis results to generate dialogue that is tailored to the user's emotional state. For example, if a user says, "I'm feeling a little lonely today," the generated response might be something like, "That's a shame. Why don't you try making plans to meet someone?"
[0347] The generated dialogue is presented to the user via the device as audio or video. At this time, the speaker and display are used to provide the user with clear information.
[0348] The server also records all user interactions, continuously refining its generative AI models and emotion engine through learning algorithms. This recorded data is also used as a valuable source of information for experts to assess the user's health status.
[0349] As an example of a prompt, if the instruction is given, "Provide the best words of encouragement to give when a user is feeling lonely," the generative AI model will produce an appropriate response.
[0350] The implementation of this system is expected to enable care tailored to the emotional and physical condition of the elderly, thereby improving the quality of care while simultaneously reducing the burden on caregivers.
[0351] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0352] Step 1:
[0353] The user speaks to the device and makes facial expressions in front of the camera. The user's voice information is captured by the microphone, and video information is captured by the camera. The input is the user's voice and facial expressions, and the output is their digital data.
[0354] Step 2:
[0355] The device optimizes the acquired audio and video information. Specifically, it performs filtering such as noise reduction and image quality improvement. This process results in clean audio data and high-quality image data being output.
[0356] Step 3:
[0357] The terminal sends optimized audio and video data to the server. The audio data is converted into text data by speech recognition software on the server; the input is audio data, and the output is text data.
[0358] Step 4:
[0359] The server inputs video data into the emotion engine, which performs analysis to identify the user's emotional state. This results in outputting information about the emotional state, which is then provided to the generative AI model.
[0360] Step 5:
[0361] The server generates appropriate dialogue using a generative AI model based on text data and emotional states. The input is text data and emotional information, and the output is a response message to the user.
[0362] Step 6:
[0363] The terminal presents the response message received from the server to the user via speech synthesis or display. Specifically, it provides the message via voice using the speaker or displays it as text using the display.
[0364] Step 7:
[0365] The server records all user interactions and stores them in a recording medium. This data is used to continuously improve the generative AI model and emotion engine. The input is the history of user interactions, and the output is the training data store.
[0366] (Application Example 2)
[0367] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0368] Conventional personal care systems for the elderly have struggled to grasp individual emotional states in real time and provide appropriate services based on that understanding. Furthermore, the services and responses provided were uniform, failing to adequately address the needs of each user. Moreover, they lacked the ability to effectively analyze users' emotions and states and propose appropriate services based on that analysis.
[0369] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0370] In this invention, the server includes speech recognition means for converting speech data into text data, analysis means for analyzing emotional information from speech and image data, conversation generation means for generating appropriate responses based on the acquired emotional information, service suggestion means for presenting service suggestions according to the user's emotional state, portable display device linkage means for displaying the generated responses and service suggestions via a portable display device, storage means for accumulating user interactions in a database, and learning means for strengthening the generation model using the stored data. This enables the provision of quick and accurate services tailored to the individual emotional state of the user.
[0371] "Speech recognition means" refers to technology that converts voice data from users into text data in real time.
[0372] "Analysis means" refers to a technology that analyzes the user's facial expressions and voice tone from audio and image data to obtain emotional information.
[0373] A "conversation generation method" is a technology that generates appropriate responses tailored to the user based on acquired text data and emotional information.
[0374] A "response presentation method" is a technology that presents a generated response to the user and conveys it in an easily understandable format.
[0375] A "service proposal method" is a technology that proposes the most suitable service based on the user's emotional state and presents the user with choices.
[0376] "Portable display device integration means" refers to a technology that displays responses and service proposals generated through a portable display device worn by the user, and provides real-time information.
[0377] A "storage method" is a technology that records all interactions in a database and stores data in preparation for future analysis and learning.
[0378] "Learning methods" refer to technologies that use accumulated data to enhance the response generation capabilities and accuracy of generative AI models, thereby improving the overall system performance.
[0379] This system is designed to provide personalized services tailored to the user's emotional state. The server receives voice and image data from the user, analyzes it, and generates responses. The following hardware and software are used to implement each function.
[0380] First, the device used by the user is a portable display device such as a smartphone, tablet, or smart glasses. This device is equipped with a microphone to collect voice data and a camera to capture facial expressions.
[0381] For speech recognition and text conversion, the server utilizes recognition technologies such as "Google Cloud Speech-to-Text API" and "IBM Watson Speech to Text." For analysis, it uses "Microsoft Azure Face API" and "Google Cloud Vision API" to identify facial expressions and emotional information from image data.
[0382] The system uses generative AI models such as "OpenAI GPT-3" to generate conversations. This allows for the generation of appropriate and personalized responses based on the user's emotional information. The generated conversations are displayed as text on the device's screen and also include service suggestions. Voice responses are played back through the device's speaker.
[0383] Furthermore, through a portable display device integration mechanism, the response information and suggestions displayed on smart glasses and other devices are presented in a way that matches the user's gaze. This feedback function allows service providers and family members to understand the user's condition in real time and take appropriate action.
[0384] As a concrete example, if a user says, "I'm feeling a little stressed today," the system will generate and display suggestions for relaxation methods and places to seek advice. An example of a prompt would be, "Generate appropriate suggestions if the user says, 'I'm lonely.'"
[0385] This invention will significantly advance care systems for the elderly, creating an environment where users can live their daily lives more comfortably and with greater peace of mind.
[0386] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0387] Step 1:
[0388] When a user speaks into the device, the device captures the audio data in real time. This audio data becomes the input. Using the microphone inside the device, high-quality audio input is obtained and prepared for transmission to the server.
[0389] Step 2:
[0390] The device uses its camera to capture image data of the user's face and sends it to the server as input data. From this image data, basic data is collected to understand the user's emotional state.
[0391] Step 3:
[0392] The server converts the received audio data into text data using APIs such as "Google Cloud Speech-to-Text API". This process converts the audio signal into appropriate text information. The converted text data is then output.
[0393] Step 4:
[0394] Simultaneously, the server uses the "Microsoft Azure Face API" to analyze the user's facial features from the image data and extract emotional information. A feature extraction algorithm is applied to the input image data, and the emotional state is output.
[0395] Step 5:
[0396] The server sends the acquired text data and sentiment information as input to a generative AI model, which then generates appropriate responses and service suggestions. This process uses generative AI models such as "OpenAI GPT-3," which output customized responses based on the input information.
[0397] Step 6:
[0398] The terminal receives output from the server and displays the generated response and service proposal on its screen. The displayed information is provided in a visually easy-to-understand format for the user, and voice responses are also output through the speaker.
[0399] Step 7:
[0400] Once the user interaction is complete, the device sends all data to the server, where it is stored in the database. This stored data will be used for future data analysis and training of generative models.
[0401] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0402] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0403] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0404] [Third Embodiment]
[0405] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0406] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0407] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0408] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0409] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0410] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0411] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0412] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0413] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0414] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0415] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0416] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0417] This invention is a system for supporting personalized care for the elderly. It optimizes communication with the elderly using a generative AI model and monitors their health status using speech recognition and image analysis technologies. The implementation of this system aims to provide individualized care services to the elderly through collaboration among multiple technical departments.
[0418] First, the user speaks or makes facial expressions towards the device's camera. The audio and image data are appropriately captured by the device. The device sends the audio data to a server, where it is converted into text through speech recognition. Simultaneously, the image data is sent to the server, where facial expressions and emotions are detected through image analysis.
[0419] The server determines the emotions and state of elderly individuals based on the results of speech recognition and image analysis. For example, if the voice detects a tone of concern about recent sleep problems and the facial expression is analyzed as anxious, the system utilizes response generation tools to prepare an appropriate response. For instance, the server might generate advice such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is sent to the terminal and presented to the user via the terminal in either voice or text format.
[0420] Furthermore, the system continuously accumulates data obtained through daily interactions. This data is sequentially analyzed by the server and used to improve the accuracy of the overall system generative model, as well as to aid in medical evaluations to detect abnormalities in specific health conditions.
[0421] In this way, the system provides personalized support to individual elderly individuals, enabling better health management and communication. For example, if a user says, "I don't feel quite right today," the server can detect this and offer kind words such as, "Don't push yourself, take a rest." This provides reassurance to the elderly and reduces the workload of caregivers.
[0422] The following describes the processing flow.
[0423] Step 1:
[0424] The user speaks to the device or makes facial expressions in front of the camera. The device records these as audio and image data.
[0425] Step 2:
[0426] The device sends voice data to the server, where a speech recognition engine converts the voice into text data.
[0427] Step 3:
[0428] The terminal sends the acquired image data to the server, which uses image analysis technology to extract facial expression data as emotional information.
[0429] Step 4:
[0430] The server evaluates the elderly person's condition based on the converted text data and analyzed emotional information, and generates appropriate conversational responses using a generative AI model.
[0431] Step 5:
[0432] The server sends the generated response to the terminal. The terminal then presents this response to the user in either audio or text format.
[0433] Step 6:
[0434] The server stores interaction data in a database, uses the stored data to enhance the generative model, and aims to improve the quality of care.
[0435] (Example 1)
[0436] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0437] In monitoring the health status and providing personalized care for the elderly, it is difficult to accurately and quickly detect changes in emotions and health conditions, and to provide appropriate, individualized responses. Therefore, there is a growing need for automated support systems that can provide a sense of security, especially for the elderly, while reducing the burden on caregivers. Furthermore, accumulating daily interaction data and utilizing it to continuously improve the accuracy of the system is also a challenge.
[0438] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0439] In this invention, the server includes a speech recognition means that receives voice information from a user and converts it into text information; an analysis means that analyzes facial expressions from the voice information and image information and obtains emotional information; and a judgment means that uses a generative AI model to determine the emotional state and detect abnormalities in specific health conditions. This makes it possible to provide appropriate responses to elderly people and to detect abnormalities in their health conditions at an early stage.
[0440] "Speech recognition means" refers to technologies and devices that receive speech information as input and convert that speech into text information.
[0441] "Analysis means" refers to technologies and devices that analyze audio and image information to obtain emotional information from facial expressions, tone of voice, etc.
[0442] "Conversation generation means" refers to technologies and devices for generating appropriate responses based on acquired text information and emotional information.
[0443] A "response presentation means" refers to a technology or device for presenting a generated response to a user in audio or text format.
[0444] "Storage means" refers to technologies and devices for recording and storing user interactions and analysis results in a database.
[0445] "Learning methods" refer to technologies and devices that use accumulated data to strengthen generative models and improve their accuracy.
[0446] "Decision-making tools" refer to technologies and devices that utilize generative AI models to judge emotional states and detect abnormalities in specific health conditions.
[0447] This invention relates to a system for monitoring the health status of elderly individuals and supporting personalized care. Specifically, it optimizes the daily communication of elderly individuals using a generative AI model, and utilizes speech recognition and image analysis technologies to detect changes in their health status early and provide appropriate responses.
[0448] First, the user speaks to the device or makes facial expressions towards the camera. The hardware used includes smartphones and tablets that capture audio and video. A dedicated application runs on these devices to process the user's interaction.
[0449] The device sends the captured audio to a server, where this audio data is converted to text by speech recognition software. Typically, services such as Google Cloud Speech-to-Text or Amazon Transcribe are used for speech recognition. Simultaneously, the device also sends image data to the server, which is then used for facial expression analysis using tools like OpenCV or Azure Face API.
[0450] The server uses a generative AI model based on these analysis results to determine the emotional state of the elderly person. Specifically, it generates gentle words such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is presented to the user via the device as either audio or text. Amazon Polly or Google Text-to-Speech are often used for speech synthesis.
[0451] Furthermore, the system uses data accumulated from daily interactions to improve the accuracy and detection capabilities of its generative models. This enables anomaly detection based on analysis, allowing for the prevention of major health problems before they occur.
[0452] For example, if a user says, "I'm not feeling well today," the server will detect this and create a gentle response such as, "Don't push yourself, take a rest." An example of a prompt might be, "The user said, 'I'm not feeling well today.' Please think of a gentle response to this."
[0453] Thus, this system has the effect of providing support and a sense of security to the elderly while reducing the burden on caregivers.
[0454] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0455] Step 1:
[0456] The user speaks to the device and makes facial expressions towards the camera. During this process, audio and images are captured using a smartphone or tablet. The input consists of the user's voice and image data, which are then fed into the system.
[0457] Step 2:
[0458] The terminal sends the captured audio data to the server. Based on the audio data input, speech recognition software is applied. This converts the audio data into text data. The output is the corresponding text data for the audio data.
[0459] Step 3:
[0460] Simultaneously, the device sends image data to the server, and image analysis technology is used to analyze the user's facial expressions. The input is image data, and an image analysis algorithm is applied. The analysis outputs information about the user's emotions.
[0461] Step 4:
[0462] The server receives text data from speech recognition and emotion information from image analysis, and uses a generative AI model to determine the user's emotional state. The input consists of text data and emotion information, which are then applied to the generative AI model. A prompt is used to generate a response appropriate to the user. The output is the generated response.
[0463] Step 5:
[0464] The server sends the response generated by the generative AI model to the terminal. The response is typically sent in text format.
[0465] Step 6:
[0466] The terminal presents the received response to the user in either voice or text. In this process, when using speech synthesis technology, the input is the generated response text, and the output is the synthesized voice. Various speech synthesis engines are used for speech synthesis.
[0467] Step 7:
[0468] The server stores daily interaction data in a database, which is then used to improve subsequent generative models. Here, the input is all dialogue data, and the output is stored as training data. This data will be used to improve the accuracy of future anomaly detection and generative AI models.
[0469] (Application Example 1)
[0470] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0471] In today's factory environment, efficiently managing worker health and safety is essential, but challenges remain in providing real-time status updates and appropriate feedback to individual workers. Furthermore, there is a lack of systems that accurately monitor workers' physiological and emotional states and provide appropriate breaks and advice.
[0472] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0473] In this invention, the server includes: speech recognition means for receiving voice information from a user and converting it into text information; analysis means for analyzing facial expressions from the voice information and visual information and acquiring emotional information; conversation generation means for generating an appropriate response based on the text information and emotional information; response presentation means for presenting the generated response to the user; storage means for storing all interactions from the user in a storage device; learning means for strengthening the generation model using the stored information; work monitoring means having the function of monitoring the physiological state of the worker and encouraging appropriate breaks based on the emotional information; and response generation means for determining the state of fatigue and stress from the worker's voice and facial information and generating appropriate advice. This makes it possible to analyze the emotions and physiological state of individual workers in real time and provide appropriate feedback quickly.
[0474] "Voice recognition means" refers to a device or software that receives voice information from a user and converts it into text information.
[0475] "Analysis means" refers to a device or system for analyzing facial expressions from audio and visual information and obtaining emotional information.
[0476] "Conversation generation means" refers to a device or algorithm that generates an appropriate response based on textual information and emotional information.
[0477] A "response presentation means" is a device or interface for presenting a generated response to a user.
[0478] A "storage means" is a device or database system for storing all user interactions in a storage device.
[0479] A "learning tool" is an algorithm or software that uses accumulated information to enhance a generative model.
[0480] "Work monitoring means" refers to a device or system that has the function of monitoring the physiological state of a worker based on emotional information and prompting appropriate breaks.
[0481] A "response generation means" is a device or software that determines the level of fatigue and stress from the worker's voice and facial expression information and generates appropriate advice.
[0482] In this invention, the entire system utilizes a highly integrated algorithm and various sensors to monitor and respond to the health and emotional state of workers in the work environment in real time.
[0483] The server consists of multiple modules. First, for speech recognition, the server uses speech recognition software to convert speech data into text data. Specific software options include open-source speech recognition tools. For image analysis, software is used to analyze image data acquired from a camera, extracting emotional information from facial expressions. Machine learning algorithms are used for this; for example, existing facial recognition models can be customized and used.
[0484] Next, the terminal uses a generation AI model to generate responses, creating appropriate responses based on the worker's state. These responses are presented to the worker in either voice or text. The server generates these responses and executes a response presentation mechanism to present them to the worker through the terminal.
[0485] The system further stores worker interaction data in memory and uses it to learn and improve its generative model. This allows the server to improve response accuracy in subsequent interactions.
[0486] For example, if a factory worker mutters, "I'm a little tired today," the system can analyze the voice, and if it determines from the worker's facial expression that they are fatigued, it can automatically provide advice such as, "Why don't you take a short break?"
[0487] An example of a prompt message could be a language input such as, "Please suggest a method to monitor worker fatigue in a factory in real time using voice and facial expressions, and encourage appropriate breaks."
[0488] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0489] Step 1:
[0490] The server receives voice data from the user via a speech recognition system. This voice data is used as input, and data processing is performed by converting it into text data using speech recognition software. The converted text data is then output.
[0491] Step 2:
[0492] The server uses an analysis tool to receive the user's visual data acquired from the terminal's camera. This visual data is used as input, and an image analysis algorithm is used to analyze facial expressions and extract the user's emotional information. This emotional information is then obtained as output.
[0493] Step 3:
[0494] The server inputs text data and extracted sentiment information into the conversation generation system and uses a generation AI model. Based on this data, it performs data calculations to generate an appropriate response and outputs the generated response.
[0495] Step 4:
[0496] The terminal presents the generated response received from the server to the user using a response presentation mechanism. The output response is presented to the user as audio or text using a speech synthesis system or text display software.
[0497] Step 5:
[0498] The server records all user interactions and stores them in memory using storage means. This stored data is used as input and analyzed using learning means to improve the accuracy of future generative models, which helps in building new generative AI models.
[0499] Step 6:
[0500] The server operates the work monitoring means based on the aforementioned emotional information. It checks the user's physiological state, executes a response generation means that generates advice to revise the work pace as needed, and prompts the user to take appropriate breaks.
[0501] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0502] This invention provides a system that integrates a generative AI model and an emotion engine to improve personalized care for the elderly. This enables more refined and individualized care based on the user's emotional and health condition.
[0503] First, the user speaks to the device and makes facial expressions in front of the camera. This allows the device to acquire audio and image data. When the device sends this data to the server, it filters and optimizes the data before sending it. The audio data is converted to text by speech recognition, and the image data is analyzed to extract emotional information.
[0504] The server uses analytical tools, particularly an emotion engine, to identify the user's emotional state from voice and images. For example, if the user is smiling or showing signs of sadness, the emotion engine quickly detects this information. This emotional information is also reflected in automatically generated conversations, and the AI model generates responses appropriate to the user's state. For example, if the user says, "I'm feeling a little lonely today," the system will generate encouraging words such as, "That's a shame. Why don't you make plans to meet someone?"
[0505] The generated responses are presented to the user by the device. Voice responses are output clearly through the speaker. The server also records all of these user interactions and stores them in a database. This stored data is used to improve the generative model and the emotion engine. It also serves as valuable information for experts when evaluating the user's health status.
[0506] This system will not only significantly improve the quality of care but also reduce the burden on caregivers, allowing them to dedicate more time and resources to sustained care. As a result, older adults will be able to live their days in a more comfortable and supportive environment.
[0507] The following describes the processing flow.
[0508] Step 1:
[0509] The user speaks into the device, saying, "I'm feeling a little sad today," or makes a gloomy expression in front of the camera. The device captures this audio and facial expression as digital data.
[0510] Step 2:
[0511] The terminal sends the captured audio data to the server. At this time, pre-processing is performed to remove noise and transcribe the audio clearly into text.
[0512] Step 3:
[0513] The device simultaneously sends image data to a server for use in facial expression analysis algorithms. This image data is optimized to facilitate the extraction of facial features.
[0514] Step 4:
[0515] The server uses speech recognition to convert speech data into text data. It then identifies the user's emotions from the speech content.
[0516] Step 5:
[0517] The server uses an emotion engine to determine if the user is actually sad based on voice tone and facial expression analysis. This allows for a more detailed understanding of the user's emotional state.
[0518] Step 6:
[0519] Based on the acquired emotional information, the server uses a generative AI model to generate appropriate conversational responses. For example, it might create a response like, "You seem a little lonely today, how about watching a movie that interests you?"
[0520] Step 7:
[0521] The server sends the generated response to the terminal. The terminal presents the response in an easily audible voice or easy-to-read text format.
[0522] Step 8:
[0523] The server stores all interaction data in a database. This data is used to improve future interactions and for expert health assessments.
[0524] (Example 2)
[0525] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0526] In personal care for the elderly, there is a need to provide individualized support based on emotional and health conditions. However, conventional systems have been insufficient in properly analyzing acquired data and generating responses, making it difficult to respond in a way that is in line with the user's feelings. As a result, it has been difficult to improve the quality of care and reduce the burden on caregivers.
[0527] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0528] In this invention, the server includes acquisition means for acquiring audio and video information from the user, analysis means for converting the audio information into text information and analyzing the video information to identify the emotional state, and generation means for generating appropriate dialogue using a generation AI model based on the text information and the emotional state. This makes it possible to accurately grasp the emotional needs of the user and generate situation-appropriate responses with high efficiency.
[0529] "Acquisition means" refers to the interface and technology used to acquire audio and video information from users.
[0530] "Analysis means" refers to the process and function of converting acquired audio information into text information and identifying emotional states from video information.
[0531] "Generative means" refers to functions and systems that create appropriate dialogues using generative AI models based on text information and emotional states.
[0532] "Presentation means" refers to hardware or software used to present the generated dialogue to the user in audio or video format.
[0533] "Recording means" refers to a system or device for recording all interactions with users and storing them on a recording medium.
[0534] "Learning methods" refer to methods and mechanisms for improving the performance of generative AI models and emotion engines by utilizing accumulated records.
[0535] This invention is implemented as a system that improves personal care for the elderly using audio and video information. Specifically, it uses a server and terminals to acquire the user's audio and video, and integrates the functions of analysis, response, recording, and learning.
[0536] The user interacts with the system using a terminal equipped with a microphone and a camera. This terminal collects audio and video information by capturing the user's voice with the microphone and capturing video with the camera. This information is then transmitted from the terminal to the server.
[0537] The server uses speech recognition software to convert the user's voice information into text, and then uses an emotion engine to analyze video information. The emotional state obtained from the analyzed data is input into a generative AI model, which is used to generate appropriate dialogue.
[0538] The generative AI model uses these analysis results to generate dialogue that is tailored to the user's emotional state. For example, if a user says, "I'm feeling a little lonely today," the generated response might be something like, "That's a shame. Why don't you try making plans to meet someone?"
[0539] The generated dialogue is presented to the user via the device as audio or video. At this time, the speaker and display are used to provide the user with clear information.
[0540] The server also records all user interactions, continuously refining its generative AI models and emotion engine through learning algorithms. This recorded data is also used as a valuable source of information for experts to assess the user's health status.
[0541] As an example of a prompt, if the instruction is given, "Provide the best words of encouragement to give when a user is feeling lonely," the generative AI model will produce an appropriate response.
[0542] The implementation of this system is expected to enable care tailored to the emotional and physical condition of the elderly, thereby improving the quality of care while simultaneously reducing the burden on caregivers.
[0543] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0544] Step 1:
[0545] The user speaks to the device and makes facial expressions in front of the camera. The user's voice information is captured by the microphone, and video information is captured by the camera. The input is the user's voice and facial expressions, and the output is their digital data.
[0546] Step 2:
[0547] The device optimizes the acquired audio and video information. Specifically, it performs filtering such as noise reduction and image quality improvement. This process results in clean audio data and high-quality image data being output.
[0548] Step 3:
[0549] The terminal sends optimized audio and video data to the server. The audio data is converted into text data by speech recognition software on the server; the input is audio data, and the output is text data.
[0550] Step 4:
[0551] The server inputs video data into the emotion engine, which performs analysis to identify the user's emotional state. This results in outputting information about the emotional state, which is then provided to the generative AI model.
[0552] Step 5:
[0553] The server generates appropriate dialogue using a generative AI model based on text data and emotional states. The input is text data and emotional information, and the output is a response message to the user.
[0554] Step 6:
[0555] The terminal presents the response message received from the server to the user via speech synthesis or display. Specifically, it provides the message via voice using the speaker or displays it as text using the display.
[0556] Step 7:
[0557] The server records all user interactions and stores them in a recording medium. This data is used to continuously improve the generative AI model and emotion engine. The input is the history of user interactions, and the output is the training data store.
[0558] (Application Example 2)
[0559] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0560] Conventional personal care systems for the elderly have struggled to grasp individual emotional states in real time and provide appropriate services based on that understanding. Furthermore, the services and responses provided were uniform, failing to adequately address the needs of each user. Moreover, they lacked the ability to effectively analyze users' emotions and states and propose appropriate services based on that analysis.
[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0562] In this invention, the server includes speech recognition means for converting speech data into text data, analysis means for analyzing emotional information from speech and image data, conversation generation means for generating appropriate responses based on the acquired emotional information, service suggestion means for presenting service suggestions according to the user's emotional state, portable display device linkage means for displaying the generated responses and service suggestions via a portable display device, storage means for accumulating user interactions in a database, and learning means for strengthening the generation model using the stored data. This enables the provision of quick and accurate services tailored to the individual emotional state of the user.
[0563] "Speech recognition means" refers to technology that converts voice data from users into text data in real time.
[0564] "Analysis means" refers to a technology that analyzes the user's facial expressions and voice tone from audio and image data to obtain emotional information.
[0565] A "conversation generation method" is a technology that generates appropriate responses tailored to the user based on acquired text data and emotional information.
[0566] A "response presentation method" is a technology that presents a generated response to the user and conveys it in an easily understandable format.
[0567] A "service proposal method" is a technology that proposes the most suitable service based on the user's emotional state and presents the user with choices.
[0568] "Portable display device integration means" refers to a technology that displays responses and service proposals generated through a portable display device worn by the user, and provides real-time information.
[0569] A "storage method" is a technology that records all interactions in a database and stores data in preparation for future analysis and learning.
[0570] "Learning methods" refer to technologies that use accumulated data to enhance the response generation capabilities and accuracy of generative AI models, thereby improving the overall system performance.
[0571] This system is designed to provide personalized services tailored to the user's emotional state. The server receives voice and image data from the user, analyzes it, and generates responses. The following hardware and software are used to implement each function.
[0572] First, the device used by the user is a portable display device such as a smartphone, tablet, or smart glasses. This device is equipped with a microphone to collect voice data and a camera to capture facial expressions.
[0573] For speech recognition and text conversion, the server utilizes recognition technologies such as "Google Cloud Speech-to-Text API" and "IBM Watson Speech to Text." For analysis, it uses "Microsoft Azure Face API" and "Google Cloud Vision API" to identify facial expressions and emotional information from image data.
[0574] The system uses generative AI models such as "OpenAI GPT-3" to generate conversations. This allows for the generation of appropriate and personalized responses based on the user's emotional information. The generated conversations are displayed as text on the device's screen and also include service suggestions. Voice responses are played back through the device's speaker.
[0575] Furthermore, through a portable display device integration mechanism, the response information and suggestions displayed on smart glasses and other devices are presented in a way that matches the user's gaze. This feedback function allows service providers and family members to understand the user's condition in real time and take appropriate action.
[0576] As a concrete example, if a user says, "I'm feeling a little stressed today," the system will generate and display suggestions for relaxation methods and places to seek advice. An example of a prompt would be, "Generate appropriate suggestions if the user says, 'I'm lonely.'"
[0577] This invention will significantly advance care systems for the elderly, creating an environment where users can live their daily lives more comfortably and with greater peace of mind.
[0578] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0579] Step 1:
[0580] When a user speaks into the device, the device captures the audio data in real time. This audio data becomes the input. Using the microphone inside the device, high-quality audio input is obtained and prepared for transmission to the server.
[0581] Step 2:
[0582] The device uses its camera to capture image data of the user's face and sends it to the server as input data. From this image data, basic data is collected to understand the user's emotional state.
[0583] Step 3:
[0584] The server converts the received audio data into text data using APIs such as "Google Cloud Speech-to-Text API". This process converts the audio signal into appropriate text information. The converted text data is then output.
[0585] Step 4:
[0586] Simultaneously, the server uses the "Microsoft Azure Face API" to analyze the user's facial features from the image data and extract emotional information. A feature extraction algorithm is applied to the input image data, and the emotional state is output.
[0587] Step 5:
[0588] The server sends the acquired text data and sentiment information as input to a generative AI model, which then generates appropriate responses and service suggestions. This process uses generative AI models such as "OpenAI GPT-3," which output customized responses based on the input information.
[0589] Step 6:
[0590] The terminal receives output from the server and displays the generated response and service proposal on its screen. The displayed information is provided in a visually easy-to-understand format for the user, and voice responses are also output through the speaker.
[0591] Step 7:
[0592] Once the user interaction is complete, the device sends all data to the server, where it is stored in the database. This stored data will be used for future data analysis and training of generative models.
[0593] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0594] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0595] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0596] [Fourth Embodiment]
[0597] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0598] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0599] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0600] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0601] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0602] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0603] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0604] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0605] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0606] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0607] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0608] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0609] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0610] This invention is a system for supporting personalized care for the elderly. It optimizes communication with the elderly using a generative AI model and monitors their health status using speech recognition and image analysis technologies. The implementation of this system aims to provide individualized care services to the elderly through collaboration among multiple technical departments.
[0611] First, the user speaks or makes facial expressions towards the device's camera. The audio and image data are appropriately captured by the device. The device sends the audio data to a server, where it is converted into text through speech recognition. Simultaneously, the image data is sent to the server, where facial expressions and emotions are detected through image analysis.
[0612] The server determines the emotions and state of elderly individuals based on the results of speech recognition and image analysis. For example, if the voice detects a tone of concern about recent sleep problems and the facial expression is analyzed as anxious, the system utilizes response generation tools to prepare an appropriate response. For instance, the server might generate advice such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is sent to the terminal and presented to the user via the terminal in either voice or text format.
[0613] Furthermore, the system continuously accumulates data obtained through daily interactions. This data is sequentially analyzed by the server and used to improve the accuracy of the overall system generative model, as well as to aid in medical evaluations to detect abnormalities in specific health conditions.
[0614] In this way, the system provides personalized support to individual elderly individuals, enabling better health management and communication. For example, if a user says, "I don't feel quite right today," the server can detect this and offer kind words such as, "Don't push yourself, take a rest." This provides reassurance to the elderly and reduces the workload of caregivers.
[0615] The following describes the processing flow.
[0616] Step 1:
[0617] The user speaks to the device or makes facial expressions in front of the camera. The device records these as audio and image data.
[0618] Step 2:
[0619] The device sends voice data to the server, where a speech recognition engine converts the voice into text data.
[0620] Step 3:
[0621] The terminal sends the acquired image data to the server, which uses image analysis technology to extract facial expression data as emotional information.
[0622] Step 4:
[0623] The server evaluates the elderly person's condition based on the converted text data and analyzed emotional information, and generates appropriate conversational responses using a generative AI model.
[0624] Step 5:
[0625] The server sends the generated response to the terminal. The terminal then presents this response to the user in either audio or text format.
[0626] Step 6:
[0627] The server stores interaction data in a database, uses the stored data to enhance the generative model, and aims to improve the quality of care.
[0628] (Example 1)
[0629] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0630] In monitoring the health status and providing personalized care for the elderly, it is difficult to accurately and quickly detect changes in emotions and health conditions, and to provide appropriate, individualized responses. Therefore, there is a growing need for automated support systems that can provide a sense of security, especially for the elderly, while reducing the burden on caregivers. Furthermore, accumulating daily interaction data and utilizing it to continuously improve the accuracy of the system is also a challenge.
[0631] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0632] In this invention, the server includes a speech recognition means that receives voice information from a user and converts it into text information; an analysis means that analyzes facial expressions from the voice information and image information and obtains emotional information; and a judgment means that uses a generative AI model to determine the emotional state and detect abnormalities in specific health conditions. This makes it possible to provide appropriate responses to elderly people and to detect abnormalities in their health conditions at an early stage.
[0633] "Speech recognition means" refers to technologies and devices that receive speech information as input and convert that speech into text information.
[0634] "Analysis means" refers to technologies and devices that analyze audio and image information to obtain emotional information from facial expressions, tone of voice, etc.
[0635] "Conversation generation means" refers to technologies and devices for generating appropriate responses based on acquired text information and emotional information.
[0636] A "response presentation means" refers to a technology or device for presenting a generated response to a user in audio or text format.
[0637] "Storage means" refers to technologies and devices for recording and storing user interactions and analysis results in a database.
[0638] "Learning methods" refer to technologies and devices that use accumulated data to strengthen generative models and improve their accuracy.
[0639] "Decision-making tools" refer to technologies and devices that utilize generative AI models to judge emotional states and detect abnormalities in specific health conditions.
[0640] This invention relates to a system for monitoring the health status of elderly individuals and supporting personalized care. Specifically, it optimizes the daily communication of elderly individuals using a generative AI model, and utilizes speech recognition and image analysis technologies to detect changes in their health status early and provide appropriate responses.
[0641] First, the user speaks to the device or makes facial expressions towards the camera. The hardware used includes smartphones and tablets that capture audio and video. A dedicated application runs on these devices to process the user's interaction.
[0642] The device sends the captured audio to a server, where this audio data is converted to text by speech recognition software. Typically, services such as Google Cloud Speech-to-Text or Amazon Transcribe are used for speech recognition. Simultaneously, the device also sends image data to the server, which is then used for facial expression analysis using tools like OpenCV or Azure Face API.
[0643] The server uses a generative AI model based on these analysis results to determine the emotional state of the elderly person. Specifically, it generates gentle words such as, "It seems you've been having trouble sleeping. How about trying some tea to relax a little?" This generated response is presented to the user via the device as either audio or text. Amazon Polly or Google Text-to-Speech are often used for speech synthesis.
[0644] Furthermore, the system uses data accumulated from daily interactions to improve the accuracy and detection capabilities of its generative models. This enables anomaly detection based on analysis, allowing for the prevention of major health problems before they occur.
[0645] For example, if a user says, "I'm not feeling well today," the server will detect this and create a gentle response such as, "Don't push yourself, take a rest." An example of a prompt might be, "The user said, 'I'm not feeling well today.' Please think of a gentle response to this."
[0646] Thus, this system has the effect of providing support and a sense of security to the elderly while reducing the burden on caregivers.
[0647] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0648] Step 1:
[0649] The user speaks to the device and makes facial expressions towards the camera. During this process, audio and images are captured using a smartphone or tablet. The input consists of the user's voice and image data, which are then fed into the system.
[0650] Step 2:
[0651] The terminal sends the captured audio data to the server. Based on the audio data input, speech recognition software is applied. This converts the audio data into text data. The output is the corresponding text data for the audio data.
[0652] Step 3:
[0653] Simultaneously, the device sends image data to the server, and image analysis technology is used to analyze the user's facial expressions. The input is image data, and an image analysis algorithm is applied. The analysis outputs information about the user's emotions.
[0654] Step 4:
[0655] The server receives text data from speech recognition and emotion information from image analysis, and uses a generative AI model to determine the user's emotional state. The input consists of text data and emotion information, which are then applied to the generative AI model. A prompt is used to generate a response appropriate to the user. The output is the generated response.
[0656] Step 5:
[0657] The server sends the response generated by the generative AI model to the terminal. The response is typically sent in text format.
[0658] Step 6:
[0659] The terminal presents the received response to the user in either voice or text. In this process, when using speech synthesis technology, the input is the generated response text, and the output is the synthesized voice. Various speech synthesis engines are used for speech synthesis.
[0660] Step 7:
[0661] The server stores daily interaction data in a database, which is then used to improve subsequent generative models. Here, the input is all dialogue data, and the output is stored as training data. This data will be used to improve the accuracy of future anomaly detection and generative AI models.
[0662] (Application Example 1)
[0663] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0664] In today's factory environment, efficiently managing worker health and safety is essential, but challenges remain in providing real-time status updates and appropriate feedback to individual workers. Furthermore, there is a lack of systems that accurately monitor workers' physiological and emotional states and provide appropriate breaks and advice.
[0665] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0666] In this invention, the server includes: speech recognition means for receiving voice information from a user and converting it into text information; analysis means for analyzing facial expressions from the voice information and visual information and acquiring emotional information; conversation generation means for generating an appropriate response based on the text information and emotional information; response presentation means for presenting the generated response to the user; storage means for storing all interactions from the user in a storage device; learning means for strengthening the generation model using the stored information; work monitoring means having the function of monitoring the physiological state of the worker and encouraging appropriate breaks based on the emotional information; and response generation means for determining the state of fatigue and stress from the worker's voice and facial information and generating appropriate advice. This makes it possible to analyze the emotions and physiological state of individual workers in real time and provide appropriate feedback quickly.
[0667] "Voice recognition means" refers to a device or software that receives voice information from a user and converts it into text information.
[0668] "Analysis means" refers to a device or system for analyzing facial expressions from audio and visual information and obtaining emotional information.
[0669] "Conversation generation means" refers to a device or algorithm that generates an appropriate response based on textual information and emotional information.
[0670] A "response presentation means" is a device or interface for presenting a generated response to a user.
[0671] A "storage means" is a device or database system for storing all user interactions in a storage device.
[0672] A "learning tool" is an algorithm or software that uses accumulated information to enhance a generative model.
[0673] "Work monitoring means" refers to a device or system that has the function of monitoring the physiological state of a worker based on emotional information and prompting appropriate breaks.
[0674] A "response generation means" is a device or software that determines the level of fatigue and stress from the worker's voice and facial expression information and generates appropriate advice.
[0675] In this invention, the entire system utilizes a highly integrated algorithm and various sensors to monitor and respond to the health and emotional state of workers in the work environment in real time.
[0676] The server consists of multiple modules. First, for speech recognition, the server uses speech recognition software to convert speech data into text data. Specific software options include open-source speech recognition tools. For image analysis, software is used to analyze image data acquired from a camera, extracting emotional information from facial expressions. Machine learning algorithms are used for this; for example, existing facial recognition models can be customized and used.
[0677] Next, the terminal uses a generation AI model to generate responses, creating appropriate responses based on the worker's state. These responses are presented to the worker in either voice or text. The server generates these responses and executes a response presentation mechanism to present them to the worker through the terminal.
[0678] The system further stores worker interaction data in memory and uses it to learn and improve its generative model. This allows the server to improve response accuracy in subsequent interactions.
[0679] For example, if a factory worker mutters, "I'm a little tired today," the system can analyze the voice, and if it determines from the worker's facial expression that they are fatigued, it can automatically provide advice such as, "Why don't you take a short break?"
[0680] An example of a prompt message could be a language input such as, "Please suggest a method to monitor worker fatigue in a factory in real time using voice and facial expressions, and encourage appropriate breaks."
[0681] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0682] Step 1:
[0683] The server receives voice data from the user via a speech recognition system. This voice data is used as input, and data processing is performed by converting it into text data using speech recognition software. The converted text data is then output.
[0684] Step 2:
[0685] The server uses an analysis tool to receive the user's visual data acquired from the terminal's camera. This visual data is used as input, and an image analysis algorithm is used to analyze facial expressions and extract the user's emotional information. This emotional information is then obtained as output.
[0686] Step 3:
[0687] The server inputs text data and extracted sentiment information into the conversation generation system and uses a generation AI model. Based on this data, it performs data calculations to generate an appropriate response and outputs the generated response.
[0688] Step 4:
[0689] The terminal presents the generated response received from the server to the user using a response presentation mechanism. The output response is presented to the user as audio or text using a speech synthesis system or text display software.
[0690] Step 5:
[0691] The server records all user interactions and stores them in memory using storage means. This stored data is used as input and analyzed using learning means to improve the accuracy of future generative models, which helps in building new generative AI models.
[0692] Step 6:
[0693] The server operates the work monitoring means based on the aforementioned emotional information. It checks the user's physiological state, executes a response generation means that generates advice to revise the work pace as needed, and prompts the user to take appropriate breaks.
[0694] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0695] This invention provides a system that integrates a generative AI model and an emotion engine to improve personalized care for the elderly. This enables more refined and individualized care based on the user's emotional and health condition.
[0696] First, the user speaks to the device and makes facial expressions in front of the camera. This allows the device to acquire audio and image data. When the device sends this data to the server, it filters and optimizes the data before sending it. The audio data is converted to text by speech recognition, and the image data is analyzed to extract emotional information.
[0697] The server uses analytical tools, particularly an emotion engine, to identify the user's emotional state from voice and images. For example, if the user is smiling or showing signs of sadness, the emotion engine quickly detects this information. This emotional information is also reflected in automatically generated conversations, and the AI model generates responses appropriate to the user's state. For example, if the user says, "I'm feeling a little lonely today," the system will generate encouraging words such as, "That's a shame. Why don't you make plans to meet someone?"
[0698] The generated responses are presented to the user by the device. Voice responses are output clearly through the speaker. The server also records all of these user interactions and stores them in a database. This stored data is used to improve the generative model and the emotion engine. It also serves as valuable information for experts when evaluating the user's health status.
[0699] This system will not only significantly improve the quality of care but also reduce the burden on caregivers, allowing them to dedicate more time and resources to sustained care. As a result, older adults will be able to live their days in a more comfortable and supportive environment.
[0700] The following describes the processing flow.
[0701] Step 1:
[0702] The user speaks into the device, saying, "I'm feeling a little sad today," or makes a gloomy expression in front of the camera. The device captures this audio and facial expression as digital data.
[0703] Step 2:
[0704] The terminal sends the captured audio data to the server. At this time, pre-processing is performed to remove noise and transcribe the audio clearly into text.
[0705] Step 3:
[0706] The device simultaneously sends image data to a server for use in facial expression analysis algorithms. This image data is optimized to facilitate the extraction of facial features.
[0707] Step 4:
[0708] The server uses speech recognition to convert speech data into text data. It then identifies the user's emotions from the speech content.
[0709] Step 5:
[0710] The server uses an emotion engine to determine if the user is actually sad based on voice tone and facial expression analysis. This allows for a more detailed understanding of the user's emotional state.
[0711] Step 6:
[0712] Based on the acquired emotional information, the server uses a generative AI model to generate appropriate conversational responses. For example, it might create a response like, "You seem a little lonely today, how about watching a movie that interests you?"
[0713] Step 7:
[0714] The server sends the generated response to the terminal. The terminal presents the response in an easily audible voice or easy-to-read text format.
[0715] Step 8:
[0716] The server stores all interaction data in a database. This data is used to improve future interactions and for expert health assessments.
[0717] (Example 2)
[0718] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0719] In personal care for the elderly, there is a need to provide individualized support based on emotional and health conditions. However, conventional systems have been insufficient in properly analyzing acquired data and generating responses, making it difficult to respond in a way that is in line with the user's feelings. As a result, it has been difficult to improve the quality of care and reduce the burden on caregivers.
[0720] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0721] In this invention, the server includes acquisition means for acquiring audio and video information from the user, analysis means for converting the audio information into text information and analyzing the video information to identify the emotional state, and generation means for generating appropriate dialogue using a generation AI model based on the text information and the emotional state. This makes it possible to accurately grasp the emotional needs of the user and generate situation-appropriate responses with high efficiency.
[0722] "Acquisition means" refers to the interface and technology used to acquire audio and video information from users.
[0723] "Analysis means" refers to the process and function of converting acquired audio information into text information and identifying emotional states from video information.
[0724] "Generative means" refers to functions and systems that create appropriate dialogues using generative AI models based on text information and emotional states.
[0725] "Presentation means" refers to hardware or software used to present the generated dialogue to the user in audio or video format.
[0726] "Recording means" refers to a system or device for recording all interactions with users and storing them on a recording medium.
[0727] "Learning methods" refer to methods and mechanisms for improving the performance of generative AI models and emotion engines by utilizing accumulated records.
[0728] This invention is implemented as a system that improves personal care for the elderly using audio and video information. Specifically, it uses a server and terminals to acquire the user's audio and video, and integrates the functions of analysis, response, recording, and learning.
[0729] The user interacts with the system using a terminal equipped with a microphone and a camera. This terminal collects audio and video information by capturing the user's voice with the microphone and capturing video with the camera. This information is then transmitted from the terminal to the server.
[0730] The server uses speech recognition software to convert the user's voice information into text, and then uses an emotion engine to analyze video information. The emotional state obtained from the analyzed data is input into a generative AI model, which is used to generate appropriate dialogue.
[0731] The generative AI model uses these analysis results to generate dialogue that is tailored to the user's emotional state. For example, if a user says, "I'm feeling a little lonely today," the generated response might be something like, "That's a shame. Why don't you try making plans to meet someone?"
[0732] The generated dialogue is presented to the user via the device as audio or video. At this time, the speaker and display are used to provide the user with clear information.
[0733] The server also records all user interactions, continuously refining its generative AI models and emotion engine through learning algorithms. This recorded data is also used as a valuable source of information for experts to assess the user's health status.
[0734] As an example of a prompt, if the instruction is given, "Provide the best words of encouragement to give when a user is feeling lonely," the generative AI model will produce an appropriate response.
[0735] The implementation of this system is expected to enable care tailored to the emotional and physical condition of the elderly, thereby improving the quality of care while simultaneously reducing the burden on caregivers.
[0736] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0737] Step 1:
[0738] The user speaks to the device and makes facial expressions in front of the camera. The user's voice information is captured by the microphone, and video information is captured by the camera. The input is the user's voice and facial expressions, and the output is their digital data.
[0739] Step 2:
[0740] The device optimizes the acquired audio and video information. Specifically, it performs filtering such as noise reduction and image quality improvement. This process results in clean audio data and high-quality image data being output.
[0741] Step 3:
[0742] The terminal sends optimized audio and video data to the server. The audio data is converted into text data by speech recognition software on the server; the input is audio data, and the output is text data.
[0743] Step 4:
[0744] The server inputs video data into the emotion engine, which performs analysis to identify the user's emotional state. This results in outputting information about the emotional state, which is then provided to the generative AI model.
[0745] Step 5:
[0746] The server generates appropriate dialogue using a generative AI model based on text data and emotional states. The input is text data and emotional information, and the output is a response message to the user.
[0747] Step 6:
[0748] The terminal presents the response message received from the server to the user via speech synthesis or display. Specifically, it provides the message via voice using the speaker or displays it as text using the display.
[0749] Step 7:
[0750] The server records all user interactions and stores them in a recording medium. This data is used to continuously improve the generative AI model and emotion engine. The input is the history of user interactions, and the output is the training data store.
[0751] (Application Example 2)
[0752] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0753] Conventional personal care systems for the elderly have struggled to grasp individual emotional states in real time and provide appropriate services based on that understanding. Furthermore, the services and responses provided were uniform, failing to adequately address the needs of each user. Moreover, they lacked the ability to effectively analyze users' emotions and states and propose appropriate services based on that analysis.
[0754] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0755] In this invention, the server includes speech recognition means for converting speech data into text data, analysis means for analyzing emotional information from speech and image data, conversation generation means for generating appropriate responses based on the acquired emotional information, service suggestion means for presenting service suggestions according to the user's emotional state, portable display device linkage means for displaying the generated responses and service suggestions via a portable display device, storage means for accumulating user interactions in a database, and learning means for strengthening the generation model using the stored data. This enables the provision of quick and accurate services tailored to the individual emotional state of the user.
[0756] "Speech recognition means" refers to technology that converts voice data from users into text data in real time.
[0757] "Analysis means" refers to a technology that analyzes the user's facial expressions and voice tone from audio and image data to obtain emotional information.
[0758] A "conversation generation method" is a technology that generates appropriate responses tailored to the user based on acquired text data and emotional information.
[0759] A "response presentation method" is a technology that presents a generated response to the user and conveys it in an easily understandable format.
[0760] A "service proposal method" is a technology that proposes the most suitable service based on the user's emotional state and presents the user with choices.
[0761] "Portable display device integration means" refers to a technology that displays responses and service proposals generated through a portable display device worn by the user, and provides real-time information.
[0762] A "storage method" is a technology that records all interactions in a database and stores data in preparation for future analysis and learning.
[0763] "Learning methods" refer to technologies that use accumulated data to enhance the response generation capabilities and accuracy of generative AI models, thereby improving the overall system performance.
[0764] This system is designed to provide personalized services tailored to the user's emotional state. The server receives voice and image data from the user, analyzes it, and generates responses. The following hardware and software are used to implement each function.
[0765] First, the device used by the user is a portable display device such as a smartphone, tablet, or smart glasses. This device is equipped with a microphone to collect voice data and a camera to capture facial expressions.
[0766] For speech recognition and text conversion, the server utilizes recognition technologies such as "Google Cloud Speech-to-Text API" and "IBM Watson Speech to Text." For analysis, it uses "Microsoft Azure Face API" and "Google Cloud Vision API" to identify facial expressions and emotional information from image data.
[0767] The system uses generative AI models such as "OpenAI GPT-3" to generate conversations. This allows for the generation of appropriate and personalized responses based on the user's emotional information. The generated conversations are displayed as text on the device's screen and also include service suggestions. Voice responses are played back through the device's speaker.
[0768] Furthermore, through a portable display device integration mechanism, the response information and suggestions displayed on smart glasses and other devices are presented in a way that matches the user's gaze. This feedback function allows service providers and family members to understand the user's condition in real time and take appropriate action.
[0769] As a concrete example, if a user says, "I'm feeling a little stressed today," the system will generate and display suggestions for relaxation methods and places to seek advice. An example of a prompt would be, "Generate appropriate suggestions if the user says, 'I'm lonely.'"
[0770] This invention will significantly advance care systems for the elderly, creating an environment where users can live their daily lives more comfortably and with greater peace of mind.
[0771] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0772] Step 1:
[0773] When a user speaks into the device, the device captures the audio data in real time. This audio data becomes the input. Using the microphone inside the device, high-quality audio input is obtained and prepared for transmission to the server.
[0774] Step 2:
[0775] The device uses its camera to capture image data of the user's face and sends it to the server as input data. From this image data, basic data is collected to understand the user's emotional state.
[0776] Step 3:
[0777] The server converts the received audio data into text data using APIs such as "Google Cloud Speech-to-Text API". This process converts the audio signal into appropriate text information. The converted text data is then output.
[0778] Step 4:
[0779] Simultaneously, the server uses the "Microsoft Azure Face API" to analyze the user's facial features from the image data and extract emotional information. A feature extraction algorithm is applied to the input image data, and the emotional state is output.
[0780] Step 5:
[0781] The server sends the acquired text data and sentiment information as input to a generative AI model, which then generates appropriate responses and service suggestions. This process uses generative AI models such as "OpenAI GPT-3," which output customized responses based on the input information.
[0782] Step 6:
[0783] The terminal receives output from the server and displays the generated response and service proposal on its screen. The displayed information is provided in a visually easy-to-understand format for the user, and voice responses are also output through the speaker.
[0784] Step 7:
[0785] Once the user interaction is complete, the device sends all data to the server, where it is stored in the database. This stored data will be used for future data analysis and training of generative models.
[0786] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0787] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0788] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0789] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0790] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0791] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0792] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0793] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0794] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0795] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0796] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0797] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0798] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0799] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0800] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0801] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0802] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0803] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0804] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0805] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0806] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0807] The following is further disclosed regarding the embodiments described above.
[0808] (Claim 1)
[0809] A speech recognition means that receives voice data from a user and converts it into text data,
[0810] An analysis means for analyzing facial expressions from the aforementioned audio data and image data and obtaining emotional information,
[0811] A conversation generation means that generates an appropriate response based on the aforementioned text data and emotional information,
[0812] A response presentation means for presenting the generated response to the user,
[0813] A storage means for storing all interactions from the aforementioned user in a database,
[0814] A learning means to enhance the generative model using the aforementioned accumulated data,
[0815] A system that includes this.
[0816] (Claim 2)
[0817] The system according to claim 1, wherein the analysis means analyzes not only facial expressions but also the tone of voice to obtain the emotional information.
[0818] (Claim 3)
[0819] The system according to claim 1, wherein the data accumulated by the storage means may be used by experts to evaluate the health status.
[0820] "Example 1"
[0821] (Claim 1)
[0822] A speech recognition means that receives voice information from a user and converts it into text information,
[0823] An analysis means for analyzing facial expressions from the aforementioned audio and image information and obtaining emotional information,
[0824] A conversation generation means that generates an appropriate response based on the aforementioned text information and emotional information,
[0825] A response presentation means for presenting the generated response to the user,
[0826] Storage means for storing all interactions from the user in a storage device,
[0827] A learning means to enhance the generative model using the aforementioned accumulated data,
[0828] A decision-making method that uses a generative AI model to determine emotional states and detect abnormalities in specific health conditions,
[0829] A system that includes this.
[0830] (Claim 2)
[0831] The system according to claim 1, wherein the analysis means analyzes not only facial expressions but also voice tone to obtain emotional information.
[0832] (Claim 3)
[0833] The system according to claim 1, wherein the data accumulated by the storage means may be used by experts to evaluate the health status.
[0834] "Application Example 1"
[0835] (Claim 1)
[0836] A speech recognition means that receives voice information from a user and converts it into text information,
[0837] An analysis means for analyzing facial expressions from the aforementioned audio and visual information and obtaining emotional information,
[0838] A conversation generation means that generates an appropriate response based on the aforementioned textual information and emotional information,
[0839] A response presentation means for presenting the generated response to the user,
[0840] A storage means for storing all interactions from the user in a storage device,
[0841] A learning means for strengthening the generative model using the aforementioned accumulated information,
[0842] A work monitoring means having the function of monitoring the physiological state of the worker and encouraging appropriate breaks based on the aforementioned emotional information,
[0843] A response generation means that determines the level of fatigue and stress from the worker's voice and facial expression information and generates appropriate advice,
[0844] A system that includes this.
[0845] (Claim 2)
[0846] The system according to claim 1, wherein the analysis means analyzes not only facial expressions but also the tone of voice to obtain emotional information and physiological state information.
[0847] (Claim 3)
[0848] The system according to claim 1, wherein the information stored by the storage device may be used by experts for health assessment and optimization of the work environment.
[0849] "Example 2 of combining an emotion engine"
[0850] (Claim 1)
[0851] A means for acquiring audio and video information from users,
[0852] An analysis means that converts the aforementioned audio information into text information and analyzes the aforementioned video information to identify the emotional state,
[0853] A generation means that generates an appropriate dialogue using a generative AI model based on the aforementioned text information and emotional state,
[0854] A presentation means for presenting the generated dialogue to the user in audio or video,
[0855] A recording means for recording all interactions with the aforementioned user and storing them in a recording medium,
[0856] Using the aforementioned accumulated records, the generative AI model is enhanced, and the emotion engine optimization learning means is provided.
[0857] A system that includes this.
[0858] (Claim 2)
[0859] The system according to claim 1, wherein the analysis means analyzes the tone of the voice and the facial expression of the image to identify the emotional state.
[0860] (Claim 3)
[0861] The system according to claim 1, wherein the records accumulated by the recording means may be used by experts as material for evaluating a person's health status.
[0862] "Application example 2 when combining with an emotional engine"
[0863] (Claim 1)
[0864] A speech recognition means that receives voice data from a user and converts it into text data,
[0865] An analysis means for analyzing facial expressions from the aforementioned audio data and image data and obtaining emotional information,
[0866] A conversation generation means that generates an appropriate response based on the aforementioned text data and emotional information,
[0867] A response presentation means for presenting the generated response to the user,
[0868] A service proposal means that presents service proposals according to the emotional state of the user,
[0869] A portable display device linkage means that presents the generated response and service proposal via a portable display device worn by the user,
[0870] A storage means for storing all interactions from the aforementioned user in a database,
[0871] A learning means to enhance the generative model using the aforementioned accumulated data,
[0872] A system that includes this.
[0873] (Claim 2)
[0874] The system according to claim 1, wherein the analysis means analyzes not only facial expressions but also the tone of voice to obtain the emotional information.
[0875] (Claim 3)
[0876] The system according to claim 1, wherein the data accumulated by the storage means may be used by experts to evaluate the health status. [Explanation of Symbols]
[0877] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A speech recognition means that receives voice data from a user and converts it into text data, An analysis means for analyzing facial expressions from the aforementioned audio data and image data and obtaining emotional information, A conversation generation means that generates an appropriate response based on the aforementioned text data and emotional information, A response presentation means for presenting the generated response to the user, A storage means for storing all interactions from the aforementioned user in a database, A learning means to enhance the generative model using the aforementioned accumulated data, A system that includes this.
2. The system according to claim 1, wherein the analysis means analyzes not only facial expressions but also the tone of voice to obtain the emotional information.
3. The system according to claim 1, wherein the data accumulated by the storage means may be used by experts to evaluate the health status.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A