system
The system addresses the challenge of understanding emotions in elderly and hearing-impaired communication by using emotion recognition and data analysis to provide accurate and personalized responses, improving care quality and reducing caregiver stress.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing systems struggle to accurately understand the emotions of the elderly and hearing-impaired individuals in communication, leading to difficulties in grasping important information and potentially causing stress for caregivers, which can result in decreased care quality and social isolation.
A system equipped with emotion recognition capabilities that acquires video and audio data, identifies facial expressions and tone of voice, estimates emotions, and provides automatic responses, summarizes conversations, and predicts questions to facilitate effective communication.
The system enhances communication quality by accurately estimating emotions, reducing caregiver stress, and facilitating smooth interactions through personalized responses and efficient information management.
Smart Images

Figure 2026073375000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In communication with the elderly and the hearing-impaired, there is a problem that it is difficult to accurately understand their emotions. Also, it is difficult to grasp important information in a long conversation, and in some cases, the caregiver may feel stressed. Due to such problems, the quality of care may decline and social isolation may occur, so it is necessary to solve these problems.
Means for Solving the Problems
[0005] This invention solves the problem by providing a system with emotion recognition capabilities. This system includes means for acquiring video data and identifying facial expressions, and means for acquiring audio data and identifying tone of voice. Furthermore, it is possible to estimate emotions based on this data and automatically determine an appropriate response. In addition, by providing functions that can convert audio into text data, analyze it, and summarize important points, and functions that can understand the context of a conversation and predict the next expected question, the system aims to facilitate communication in caregiving settings.
[0006] "Video data" refers to data that captures visual information in digital format and is acquired using devices such as cameras.
[0007] "Audio data" refers to data that records the waveform of sound in digital format, and is acquired using devices such as microphones.
[0008] "Acquisition means" refers to devices or processes for capturing and collecting specific data, particularly video and audio.
[0009] "Identification means" refers to technologies and devices used to analyze data and recognize specific patterns or features.
[0010] An "estimation method" is a process or technique used to calculate or determine a certain state or characteristic using acquired data.
[0011] A "decision-making tool" is a device or algorithm used to integrate multiple data sets and make decisions based on specific conditions.
[0012] A "conversion means" is a process or device for converting data in one format to another.
[0013] "Extraction means" refers to techniques or devices used to select important or relevant information from data.
[0014] The "generation means" is a process or technology that creates new information or text based on data and analysis results.
[0015] The "understanding means" is a technology or device for analyzing text and other data to grasp its content and background.
[0016] The "prediction means" is a process or technology for predicting future events and questions based on current data.
Brief Description of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map where multiple emotions are mapped. [Figure 10] It shows an emotion map where multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the language used in the following description will be explained.
[0020] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] This invention relates to a system for improving the quality of communication in caregiving settings. This system functions by acquiring user video and audio data in real time and transmitting it to a server. The server analyzes this data to accurately estimate the user's emotions and automatically determine situations requiring intervention.
[0039] Specifically, the camera on the device captures the user's facial expressions, and the microphone collects audio. This data is converted into digital signals and transmitted to a server via the internet. The server processes the received video data and applies a facial recognition algorithm to identify facial features. For the audio data, spectral analysis is performed to extract tone and pitch information.
[0040] Based on these analyses, the server uses estimation methods to determine the user's emotions. For example, if a smile is recognized through facial analysis and the voice tone is high, the emotion of "joy" is estimated. Conversely, if the facial expression is stern and the voice tone is low, "anger" may be estimated.
[0041] Furthermore, the speech-to-text and summarization functions allow for the automatic generation of key points from long conversations. In this process, the server converts speech to text using speech recognition technology and then summarizes it using natural language processing. For example, it can transcribe meeting audio and extract meeting conclusions and action items.
[0042] To understand the context of a conversation, the server analyzes the input text data to understand the context and predict the next expected question or topic. This allows for natural responses to spontaneous questions from users, such as "Why?".
[0043] This system can also generate personalized responses by learning from the past history based on each care recipient's profile, thereby providing individualized responses to each user. This improves the user experience.
[0044] Regarding information recording and sharing, the system automatically records important data during caregiving and securely shares it with family members and medical professionals. This data management is primarily done on servers and centralized in a database to maintain consistency in care.
[0045] In this way, this system integrates and operates multiple functions to meet the diverse needs of users, thereby reducing the burden on caregivers and facilitating smooth communication.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0049] Step 2:
[0050] The device transmits the acquired video and audio data to the server via the internet.
[0051] Step 3:
[0052] The server analyzes the received video data using an image processing algorithm to detect the user's face.
[0053] Step 4:
[0054] The server applies an expression recognition algorithm after face detection to identify facial features. This extracts features such as smiles and frowns.
[0055] Step 5:
[0056] The server analyzes the received audio data using an audio processing algorithm to extract the tone and pitch of the voice.
[0057] Step 6:
[0058] The server integrates identified facial and vocal features to estimate the user's emotions. For example, it might estimate "joy" from a smile and a high-pitched voice.
[0059] Step 7:
[0060] The server converts the audio data into text using automatic speech recognition technology.
[0061] Step 8:
[0062] The server analyzes the converted text using natural language processing techniques to extract keywords and topics.
[0063] Step 9:
[0064] The server generates a summary based on the extracted keywords and topics, and extracts the key points.
[0065] Step 10:
[0066] The server analyzes text data to understand the context of the conversation and predicts the next expected question or topic.
[0067] Step 11:
[0068] The server generates an appropriate response to the predicted question and sends it to the terminal.
[0069] Step 12:
[0070] The terminal displays the response received from the server to the user.
[0071] Step 13:
[0072] The server stores important data and information recorded during caregiving in a database and shares it with family members and medical professionals as needed.
[0073] Through these steps, the system accurately grasps the user's emotions, enabling smooth communication and efficient information management in caregiving settings.
[0074] (Example 1)
[0075] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0076] In caregiving settings, smooth communication between care recipients and caregivers is essential. However, conventional systems have difficulty accurately grasping users' emotions, making immediate responses challenging. Furthermore, utilizing accumulated information for caregiving purposes and securely sharing it have also been identified as issues.
[0077] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0078] In this invention, the server includes acquisition means for acquiring video information, acquisition means for acquiring audio information, and estimation means for analyzing the acquired information to estimate emotions. This makes it possible to grasp the user's emotions in real time and automate appropriate responses. Furthermore, the quality and efficiency of care can be improved through personalized recording and secure sharing of the collected information.
[0079] "Acquisition means" refers to devices and methods for collecting user video and audio information.
[0080] "Identification means" refers to a device or method that analyzes acquired information to identify the user's facial expressions and voice characteristics.
[0081] An "estimation means" is a device or method for determining a user's emotions or situation based on identified data.
[0082] A "means of judgment" refers to a device or method for determining the necessary response based on presumed emotions.
[0083] "Conversion means" refers to devices or methods for converting audio information into text information.
[0084] An "extraction means" is a device or method for finding useful information from converted character information.
[0085] "Generating means" refers to devices or methods for creating summaries or responses based on extracted information.
[0086] "Means of understanding" refer to devices and methods for grasping the context of text data and understanding the next course of action to take.
[0087] "Inference tools" are devices or methods that predict future actions or questions based on an understood context.
[0088] "Management means" refers to devices and methods for properly recording and protecting information and sharing it as necessary.
[0089] In this invention, the following hardware and software are used to implement a system that improves the quality of communication in caregiving settings.
[0090] The device is equipped with a camera and microphone, and captures the user's video and audio in real time. The device converts the captured video data into a digital signal and transmits it to a server via the internet. Similarly, the audio data acquired by the microphone is also converted into a digital signal and transmitted to the server.
[0091] The server uses a facial recognition algorithm to process the received video and analyze the facial expression data. This identifies feature points on the user's face and identifies the characteristics of their facial expressions. For audio data, spectral analysis is performed to extract the tone and pitch of the voice. Based on this data, a generative AI model is used to estimate emotions.
[0092] As a concrete example, if a user is smiling and their voice tone is high, the generative AI model will estimate that this indicates an emotion of "joy." The server will then communicate this information to the caregiver in real time to help them take the necessary actions.
[0093] Furthermore, the system converts the acquired audio using speech recognition technology into text data. The converted text data is then processed using natural language processing to extract key points and generate a summary, enabling efficient comprehension of long conversations.
[0094] An example of a prompt message might be: "If the user is smiling and their voice is high-pitched, assume that this emotion is joy. Also, generate words of encouragement appropriate to the situation." This system will make daily care smoother and of higher quality.
[0095] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0096] Step 1:
[0097] The device uses a camera to capture the user's video and a microphone to collect audio. It converts the video and audio information into digital format and then transmits it to a server via the internet. This requires maintaining the quality of the video information while converting it to a transmission-compatible format.
[0098] Step 2:
[0099] The server analyzes the received video information using a facial recognition algorithm. It extracts feature points of the user's face from the input video data and identifies their patterns. As output of this identification, it obtains data on the user's facial features. Specifically, it calculates and digitizes things like the angle of the mouth and the position of the eyebrows.
[0100] Step 3:
[0101] The server performs spectral analysis on the received audio information. It extracts frequency components from the input audio data and analyzes the tone and pitch of the voice. This analysis outputs parameters such as the pitch and volume of the user's voice. In particular, the tone of voice is important information for estimating emotions.
[0102] Step 4:
[0103] The server uses a generative AI model to estimate the user's emotions based on the output of facial recognition and voice analysis. Combining the input facial expression data and voice tone data, the emotion estimation algorithm classifies emotions such as "joy" and "anger." The resulting emotion estimation data is then output.
[0104] Step 5:
[0105] The server generates the necessary response based on the emotion estimation result. In this process, it creates response measures and action plans based on the specific emotion. For example, if "joy" is estimated, it outputs a prompt message that generates words of encouragement as the response.
[0106] Step 6:
[0107] The server uses speech recognition technology to convert speech information into text data. It transcribes the input speech data into text, generating text data as a result. This text data serves as the basis for further analysis.
[0108] Step 7:
[0109] The server performs natural language processing based on text data to extract important information. It identifies key points of a conversation from the input text data and outputs that information as a summary. This allows for the efficient extraction of important parts from long conversations.
[0110] Step 8:
[0111] The server uses historical data and individual profiles to generate personalized responses using a generative AI model. It takes prompts as input, prepares and outputs responses tailored to each user. This response generation assists caregivers in naturally interacting with users.
[0112] (Application Example 1)
[0113] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0114] There is a need to monitor the health status and mental stress levels of workers in factories and other workplaces in real time, and to provide appropriate work support and break suggestions based on that information. However, currently, these decisions rely on individual workers' self-reporting, and there is a problem in that support based on objective data is lacking.
[0115] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0116] In this invention, the server includes a device for acquiring video information, a device for acquiring audio information, a device for analyzing the acquired video information to identify facial expressions, a device for analyzing the acquired audio information to identify audio characteristics, a device for estimating emotions based on the identified facial expressions and audio characteristics, a device for judging the situation based on the estimated emotions, a device for monitoring the worker's condition and detecting stress and fatigue, and a device for suggesting work improvements or breaks based on the detection results. This makes it possible to objectively evaluate the worker's health and emotions and provide appropriate support and suggestions.
[0117] "Visual information" refers to visual data acquired using cameras and visual sensors.
[0118] "Auditory information" refers to auditory data acquired using microphones or sound sensors.
[0119] A "facial expression recognition device" is a device that analyzes acquired video information to identify the characteristics of individual facial expressions.
[0120] A "device for identifying speech characteristics" is a device that analyzes acquired speech information to identify the tone and other acoustic characteristics of the speech.
[0121] A "device for estimating emotions" is a device that integrates identified facial and vocal characteristics to estimate an individual's emotional state.
[0122] A "situation assessment device" is a device that determines whether appropriate action or response is necessary based on estimated emotions.
[0123] "Monitoring" refers to the entire activity of continuously observing the condition of workers, collecting information, and analyzing it.
[0124] A "stress and fatigue detection device" is a device that uses collected data to identify the level of stress and fatigue experienced by workers.
[0125] A "device for suggesting work improvements and breaks" is a device that suggests improvement measures and necessary breaks, taking into account the health and efficiency of workers.
[0126] This invention aims to realize a system that monitors the health status and emotions of workers in real time in workplaces such as factories and provides appropriate feedback. It basically consists of a server, terminals, and users.
[0127] The server plays a central role in data analysis and decision-making, specifically handling the analysis of video and audio information. The server utilizes visual analysis software such as OpenCV and Dlib to acquire video information and identify facial expressions. It also converts audio information into text using the Google Cloud Speech-to-Text API and performs spectral analysis to identify speech features. Natural language processing libraries such as NLTK and spaCy are used to understand the context of conversations and extract important elements. Furthermore, the server assesses the worker's health status based on estimated emotions and suggests breaks or work improvements as needed. These suggestions are customized to individual worker profiles based on predetermined criteria.
[0128] The terminal is equipped with hardware devices such as cameras and microphones, and is responsible for collecting information in the field. The data obtained from these devices is transmitted to a server via the internet. The terminal also receives feedback from the server and displays it appropriately to the user.
[0129] Workers, as users, can utilize the information and suggestions provided by the system to improve work efficiency and manage their health. Specifically, they can receive advice on appropriate timing for work breaks and daily stress management. Furthermore, the system has the ability to personalize subsequent suggestions based on the worker's responses.
[0130] As a concrete example, consider a case at a metalworking factory. When a worker is performing complex machine operations, the system detects his fatigue level and suggests, "This task requires concentration, so we recommend you take a 10-minute break." This kind of feedback helps workers prevent mistakes caused by excessive fatigue.
[0131] An example of a prompt for a generative AI model would be: "Based on the following dataset, analyze the emotional state of the workers and suggest appropriate breaks if stress is detected."
[0132] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0133] Step 1:
[0134] The terminal uses cameras and microphones installed at the work site to acquire video and audio information of workers in real time. This information is output as raw video data (e.g., video stream) and audio data (e.g., audio stream). Next, this data is converted into a digital format and sent to the server.
[0135] Step 2:
[0136] The server processes the received video information and uses OpenCV and Dlib libraries to identify facial expressions. Specifically, it extracts facial features from the video data and classifies expressions such as smiles and surprise. The input is digital video data, and the output is category information of the identified facial expressions.
[0137] Step 3:
[0138] The server converts the audio information into text using the Google Cloud Speech-to-Text API and then performs spectral analysis of the audio. It identifies features related to the pitch and intensity of the speech and measures the pitch of the tone. The input for this step is digital audio data, and the output is the converted text information and speech tone features.
[0139] Step 4:
[0140] The server integrates the identified facial expression category information and voice tone characteristics, and uses an estimation algorithm to estimate the worker's emotional state. For example, a smiling expression and a high but gentle tone are likely to indicate "satisfied." The input for this step is the facial expression and voice tone characteristics, and the output is the estimated emotional state.
[0141] Step 5:
[0142] The server assesses the worker's stress level and fatigue based on their estimated emotional state, and provides suggestions for breaks and work improvements as needed. This involves using a generative AI model to create appropriate advice. The input for this step is the result of the emotional state assessment, and the output is specific suggestions for work improvements and breaks.
[0143] Step 6:
[0144] The terminal displays suggestions received from the server to the worker in real time. The user reviews the suggestions and adjusts their actions as needed. The input for this step is suggestion information from the server, and the output is feedback messages provided to the worker.
[0145] In this way, the system can continuously monitor the health status of workers and provide appropriate feedback in real time.
[0146] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0147] This invention provides a highly accurate system for recognizing user emotions. The system incorporates an emotion engine into a server to perform emotion recognition, and estimates the user's emotional state in real time based on video and audio data.
[0148] The device captures the user's current video and audio through its camera and microphone. This data is then immediately transmitted to a server via the internet. The server analyzes the received video data to extract the user's facial features and simultaneously processes the audio data using spectral analysis to identify the tone of voice.
[0149] The emotion engine embedded in the server estimates emotions based on extracted facial features and voice tone. This emotion engine utilizes machine learning algorithms and can improve estimation accuracy by referencing emotion patterns based on the user's past data. For example, if a user has previously displayed certain facial expressions and voice patterns in specific situations, the engine can identify emotional states such as "stress" or "relaxation" based on that data.
[0150] Furthermore, the emotion engine can monitor the user's emotional changes in real time and determine the appropriate timing for intervention based on the analysis results. For example, if a trend of the user gradually becoming depressed is detected, the server will send an early notification to the caregiver to encourage proactive action.
[0151] This system can also effectively provide users with the information they need by converting audio data into text and summarizing key points. The server achieves this function by converting audio data into text using automatic speech recognition and summarizing it using natural language processing.
[0152] Thus, the emotion recognition system of the present invention establishes a highly accurate understanding of the user's emotions, enabling high-quality communication in caregiving settings and providing responses that meet the user's needs.
[0153] The following describes the processing flow.
[0154] Step 1:
[0155] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0156] Step 2:
[0157] The terminal transmits the acquired video and audio data to the server as digital signals.
[0158] Step 3:
[0159] The server analyzes the received video data using an image processing algorithm to extract the user's facial features.
[0160] Step 4:
[0161] The server analyzes the received audio data using an audio processing algorithm to identify the tone and pitch of the voice.
[0162] Step 5:
[0163] The emotion engine embedded in the server estimates the user's emotions based on extracted facial features and voice tone. This emotion engine uses machine learning models to improve its estimation accuracy.
[0164] Step 6:
[0165] The server analyzes the estimated emotions and monitors the trend of emotional changes in real time.
[0166] Step 7:
[0167] Based on trends in emotional changes, the server determines the appropriate timing for intervention for the user and sends notifications to caregivers as needed.
[0168] Step 8:
[0169] The server uses speech recognition technology to convert the audio data into text data.
[0170] Step 9:
[0171] The server analyzes the converted text data using natural language processing techniques, extracts the key points of the conversation, and generates a summary.
[0172] Step 10:
[0173] The server sends summary information to the terminal and presents it to the user.
[0174] Through this process, the system can recognize the user's emotions with high accuracy, enabling a quick and appropriate response.
[0175] (Example 2)
[0176] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0177] Conventional systems struggled to accurately recognize users' emotions and respond appropriately based on them. Furthermore, they were inadequate at grasping emotional changes in real time through the analysis of voice and video data, and at determining the timing of necessary interventions. As a result, they were unable to respond in a way that aligned with users' emotions and needs, making it difficult to provide satisfactory support, particularly in areas such as caregiving and communication.
[0178] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0179] In this invention, the server includes means for acquiring video information, means for acquiring audio information, and means for analyzing the acquired video information to identify facial features. This enables high-precision estimation of the user's emotional state in real time and allows for quick and accurate responses as needed.
[0180] "Visual information" refers to visual data acquired by cameras and other imaging devices.
[0181] "Acoustic information" refers to audio and sound data acquired by sound-collecting devices such as microphones.
[0182] "Device means" refers to a mechanical or electrical device designed to perform a specific function.
[0183] "Analysis" refers to the process of examining data in detail and identifying specific characteristics or trends.
[0184] "Facial features" refer to a person's expression and the appearance of their face, and are used as elements that indicate emotions and reactions.
[0185] "Voice tone" refers to characteristics such as the tone, pitch, and intensity of a voice that reflect the speaker's emotions and intentions.
[0186] "To estimate" refers to predicting or judging unconfirmed matters or situations based on obtained data.
[0187] "Real-time" refers to data processing and system responses occurring immediately, meaning timely execution without delay.
[0188] The "timing of intervention" refers to the point at which specific actions or responses are required, indicating the point that leads to effective support and resolution.
[0189] This invention provides a system that recognizes a user's emotions in real time with high accuracy. The entire system mainly consists of a terminal and a server, and evaluates emotions by utilizing the user's video and audio information.
[0190] The device uses a camera and microphone to record the user's current status. The camera captures video information from the user's face, and the microphone captures audio information from the user's voice. This captured information is transmitted to a server via a communication network.
[0191] The server analyzes the transmitted video information and uses image processing algorithms to identify the user's facial features. It also employs acoustic analysis software to identify the tone of voice by performing spectral analysis on acoustic information. The server incorporates a generative AI model that estimates the user's emotions based on these analysis results. This generative AI model utilizes machine learning algorithms and improves estimation accuracy by referencing emotion patterns based on past data.
[0192] This system also has the function of monitoring emotional changes in real time and determining the appropriate timing for intervention. For example, if a user shows signs of rapidly becoming stressed, the system will notify the caregiver of this information and instruct them on appropriate measures. It also converts acoustic information into text and provides feedback to the user by summarizing the key points.
[0193] For example, if the system determines that a user may be feeling down, it will compare this with past data and recommend relaxing music for that user. Furthermore, using a generative AI model, an example of a prompt message could be: "Estimate the user's emotions in real time. The video data shows a smiling expression, and the audio data indicates a calm tone of voice."
[0194] Thus, the system of the present invention is designed with the aim of deeply understanding the user's emotions and providing effective responses accordingly.
[0195] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0196] Step 1:
[0197] The device acquires the user's video information via its camera. The input is the user's real-time visual situation, which is output as high-resolution digital video data. This video data includes the user's facial expressions, facial features, and movements. The device captures accurate video while maintaining a stable frame rate.
[0198] Step 2:
[0199] The device acquires the user's acoustic information via a microphone. The input is the user's voice, which is output as acoustic sample data. This data captures the tone, pitch, speed, and emotional nuances of the voice in detail. The device uses filters to reduce background noise and ensure clear sound quality.
[0200] Step 3:
[0201] The terminal transmits the acquired video and audio information to the server via the internet. The input is the output data from steps 1 and 2, and this data is transmitted encrypted. The terminal uses a communication protocol to enable secure and rapid transmission of data.
[0202] Step 4:
[0203] The server analyzes the received video information and uses machine learning algorithms to identify facial features. The input is encrypted video data, and the output is elements extracted as identified facial features and expressions. The server identifies key features of the expressions and matches them against known expression patterns registered in the database.
[0204] Step 5:
[0205] The server performs spectral analysis on the received acoustic information to identify the tone of the voice. The input is acoustic sample data, and the output is the analyzed voice tone and emotional characteristics. The server analyzes the frequency components of the voice waveform to identify tones related to the user's emotional state.
[0206] Step 6:
[0207] The server uses a generative AI model to estimate the user's emotions from the identified facial features and tone of voice. The input is the output data from steps 4 and 5, and the output is the estimated emotional state. The server compares the estimation results with past data to improve accuracy and recognize the user's current emotions with high precision.
[0208] Step 7:
[0209] The server monitors estimated emotional changes in real time and determines the appropriate timing for intervention as needed. The input is continuously estimated emotional data, and the output is alerts and notifications for intervention. The server performs real-time analysis based on certain criteria and prompts necessary actions.
[0210] (Application Example 2)
[0211] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0212] In modern society, it is crucial to understand an individual's emotional state in real time and proactively detect potential dangers and stressors. However, conventional technologies lack the means to instantly grasp emotional changes and provide appropriate responses, making it difficult to create a safe and comfortable environment.
[0213] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0214] In this invention, the server includes means for acquiring video information, means for acquiring audio information, means for analyzing the acquired video information to identify facial expressions, means for analyzing the acquired audio information to identify voice tone, means for estimating emotions based on the identified facial expressions and voice tone, and means for determining situations requiring safety measures based on the estimated emotions and sending notifications to the individual to encourage relaxation. This enables real-time emotion monitoring tailored to individual situations and appropriate interventions to ensure safety.
[0215] "Video information" refers to visual data, including the user's facial expressions and movements, and is digital data acquired via a camera device.
[0216] "Acoustic information" refers to auditory data, including the user's voice and ambient sounds, and is digital data acquired using a microphone device.
[0217] "Means for identifying facial expressions" refers to a system component that analyzes facial features from acquired video information and, based on that, identifies facial expressions that indicate the user's emotions.
[0218] "Means for identifying the tone of voice" refers to a component of a system that analyzes the tone and intonation of a voice from acquired acoustic information and identifies the characteristics of a voice that indicate emotion based on that analysis.
[0219] "Means for estimating emotions" refers to a system component that utilizes machine learning algorithms, based on identified facial expression and voice tone data, to infer the user's current emotional state.
[0220] "Means for determining situations requiring safety assurance" refers to a system component that has the function of determining whether or not a user's psychological state requires vigilance, based on estimated emotions.
[0221] "Means of sending notifications to encourage relaxation" refers to system components that have the functionality to send messages to users urging them to relieve stress or rest.
[0222] The specific system for realizing this invention analyzes video and audio data, estimates the user's emotional state in real time, and intervenes to improve safety. Processing is primarily carried out using smart devices and servers.
[0223] The cameras and microphones built into smart devices (such as smart glasses and smartphones) capture the user's video and audio information. This data is transmitted to a server via the network.
[0224] On the server, video information is analyzed using the OpenCV library to identify the user's facial expressions. Furthermore, acoustic information is analyzed using the LibROSA library to identify the tone of the voice. The resulting facial and voice feature data is input into an emotion recognition model using TENSORFLOW® to estimate the user's emotions.
[0225] Based on the emotion assessment, the server determines whether a situation requiring safety measures is present. Specifically, if the system determines that the user is experiencing stress, it sends a notification to the user encouraging relaxation based on that information. This notification is displayed on the user's smart device to prompt timely intervention.
[0226] For example, if a user on public transport feels anxious, the server can instantly recognize this and send a message to the user's device such as, "Take a deep breath to relax." This helps the user instantly recognize their own state and manage themselves.
[0227] An example of a prompt would be: "Generate a scenario in which an emotion monitoring security app monitors a user's emotional state in real time and provides notifications to enhance safety. Consider the specific situation of anxiety experienced while using public transportation."
[0228] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0229] Step 1:
[0230] The device uses a camera and microphone to acquire video and audio information of the user. The input data consists of video data captured by the camera and audio data collected by the microphone. This data is then prepared for transmission to a server via the network.
[0231] Step 2:
[0232] The server analyzes the received video information using the OpenCV library. The input is video data transmitted from the terminal, and the output is feature data representing the user's facial expressions. Specifically, it calculates the movement and changes of each part of the face and quantifies the facial expression features.
[0233] Step 3:
[0234] The server analyzes the received acoustic information using the LibROSA library. The input is audio data transmitted from the terminal, and the output is feature data indicating the tone of the voice. Specifically, it performs spectral analysis of the audio waveform and extracts features such as pitch, intonation, and tempo.
[0235] Step 4:
[0236] The server inputs the obtained facial expression feature data and voice tone feature data into a generative AI model using TensorFlow to estimate emotions. The output is data indicating the user's emotional state. Specifically, the model uses pattern recognition based on past data to determine whether the emotion corresponds to "stress" or "relaxation."
[0237] Step 5:
[0238] The server generates prompt messages to determine if the estimated emotional state requires safety measures. The input is emotional state data, and the output is a notification message to the user. It generates messages such as "Take a deep breath to relax" to assess the need for intervention.
[0239] Step 6:
[0240] The server sends a notification to the terminal to encourage the user to relax if it determines that security measures are necessary. The input is the generated notification message, and the output is the relaxation notification displayed on the terminal screen. This allows the user to take necessary actions in a timely manner.
[0241] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0242] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0243] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0244] [Second Embodiment]
[0245] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0246] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0247] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0248] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0249] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0250] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0251] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0252] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0253] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0254] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0255] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0256] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0257] This invention relates to a system for improving the quality of communication in caregiving settings. This system functions by acquiring user video and audio data in real time and transmitting it to a server. The server analyzes this data to accurately estimate the user's emotions and automatically determine situations requiring intervention.
[0258] Specifically, the camera on the device captures the user's facial expressions, and the microphone collects audio. This data is converted into digital signals and transmitted to a server via the internet. The server processes the received video data and applies a facial recognition algorithm to identify facial features. For the audio data, spectral analysis is performed to extract tone and pitch information.
[0259] Based on these analyses, the server uses estimation methods to determine the user's emotions. For example, if a smile is recognized through facial analysis and the voice tone is high, the emotion of "joy" is estimated. Conversely, if the facial expression is stern and the voice tone is low, "anger" may be estimated.
[0260] Furthermore, the speech-to-text and summarization functions allow for the automatic generation of key points from long conversations. In this process, the server converts speech to text using speech recognition technology and then summarizes it using natural language processing. For example, it can transcribe meeting audio and extract meeting conclusions and action items.
[0261] To understand the context of a conversation, the server analyzes the input text data to understand the context and predict the next expected question or topic. This allows for natural responses to spontaneous questions from users, such as "Why?".
[0262] This system can also generate personalized responses by learning from the past history based on each care recipient's profile, thereby providing individualized responses to each user. This improves the user experience.
[0263] Regarding information recording and sharing, the system automatically records important data during caregiving and securely shares it with family members and medical professionals. This data management is primarily done on servers and centralized in a database to maintain consistency in care.
[0264] In this way, this system integrates and operates multiple functions to meet the diverse needs of users, thereby reducing the burden on caregivers and facilitating smooth communication.
[0265] The following describes the processing flow.
[0266] Step 1:
[0267] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0268] Step 2:
[0269] The device transmits the acquired video and audio data to the server via the internet.
[0270] Step 3:
[0271] The server analyzes the received video data using an image processing algorithm to detect the user's face.
[0272] Step 4:
[0273] The server applies an expression recognition algorithm after face detection to identify facial features. This extracts features such as smiles and frowns.
[0274] Step 5:
[0275] The server analyzes the received audio data using an audio processing algorithm to extract the tone and pitch of the voice.
[0276] Step 6:
[0277] The server integrates identified facial and vocal features to estimate the user's emotions. For example, it might estimate "joy" from a smile and a high-pitched voice.
[0278] Step 7:
[0279] The server converts the voice data into text using automatic speech recognition technology.
[0280] Step 8:
[0281] The server analyzes the converted text using natural language processing technology and extracts keywords and topics.
[0282] Step 9:
[0283] The server generates a summary based on the extracted keywords and topics and extracts important points.
[0284] Step 10:
[0285] The server analyzes the text data to understand the context of the conversation and then predicts the next expected questions and topics.
[0286] Step 11:
[0287] The server generates an appropriate response to the predicted questions and sends it to the terminal.
[0288] Step 12:
[0289] The terminal presents the response received from the server to the user.
[0290] Step 13:
[0291] The server stores the important data and information recorded during the care in a database and shares it with family members and medical experts as needed.
[0292] Through these steps, the system accurately grasps the user's emotions and realizes smooth communication and efficient information management at the care site.
[0293] (Example 1)
[0294] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0295] In caregiving settings, smooth communication between care recipients and caregivers is essential. However, conventional systems have difficulty accurately grasping users' emotions, making immediate responses challenging. Furthermore, utilizing accumulated information for caregiving purposes and securely sharing it have also been identified as issues.
[0296] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0297] In this invention, the server includes acquisition means for acquiring video information, acquisition means for acquiring audio information, and estimation means for analyzing the acquired information to estimate emotions. This makes it possible to grasp the user's emotions in real time and automate appropriate responses. Furthermore, the quality and efficiency of care can be improved through personalized recording and secure sharing of the collected information.
[0298] "Acquisition means" refers to devices and methods for collecting user video and audio information.
[0299] "Identification means" refers to a device or method that analyzes acquired information to identify the user's facial expressions and voice characteristics.
[0300] An "estimation means" is a device or method for determining a user's emotions or situation based on identified data.
[0301] A "means of judgment" refers to a device or method for determining the necessary response based on presumed emotions.
[0302] "Conversion means" refers to devices or methods for converting audio information into text information.
[0303] An "extraction means" is a device or method for finding useful information from converted character information.
[0304] The "generation means" is a device or method for creating a summary or response based on the extracted information.
[0305] The "understanding means" is a device or method for grasping the context of text data and understanding the actions to be taken next.
[0306] The "inference means" is a device or method for predicting future actions and questions based on the understood context.
[0307] The "management means" is a device or method for appropriately recording and protecting information and sharing it as necessary.
[0308] In the present invention, in order to implement a system for improving the quality of communication at the care site, the following hardware and software are used.
[0309] The terminal is equipped with a camera and a microphone, and acquires the user's video and audio in real time. This terminal converts the captured video data into a digital signal and transmits it to the server via the Internet. Also, the audio data acquired by the microphone is similarly converted into a digital signal and transmitted to the server.
[0310] The server uses a face recognition algorithm to process the received video and analyzes the expression data. Thereby, the feature points on the user's face are specified and the features of the expression are identified. Also, spectral analysis is performed on the audio data to extract the tone and pitch of the voice. Based on these data, emotion estimation is performed using a generated AI model.
[0311] As a specific example, when the user is smiling and the tone of voice is also high, the generated AI model estimates that the emotion is "joy". The server transmits this information to the caregiver in real time and supports the ability to take necessary actions.
[0312] Furthermore, the system converts the acquired audio using speech recognition technology into text data. The converted text data is then processed using natural language processing to extract key points and generate a summary, enabling efficient comprehension of long conversations.
[0313] An example of a prompt message might be: "If the user is smiling and their voice is high-pitched, assume that this emotion is joy. Also, generate words of encouragement appropriate to the situation." This system will make daily care smoother and of higher quality.
[0314] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0315] Step 1:
[0316] The device uses a camera to capture the user's video and a microphone to collect audio. It converts the video and audio information into digital format and then transmits it to a server via the internet. This requires maintaining the quality of the video information while converting it to a transmission-compatible format.
[0317] Step 2:
[0318] The server analyzes the received video information using a facial recognition algorithm. It extracts feature points of the user's face from the input video data and identifies their patterns. As output of this identification, it obtains data on the user's facial features. Specifically, it calculates and digitizes things like the angle of the mouth and the position of the eyebrows.
[0319] Step 3:
[0320] The server performs spectral analysis on the received audio information. It extracts frequency components from the input audio data and analyzes the tone and pitch of the voice. This analysis outputs parameters such as the pitch and volume of the user's voice. In particular, the tone of voice is important information for estimating emotions.
[0321] Step 4:
[0322] The server uses a generative AI model to estimate the user's emotions based on the output of facial recognition and voice analysis. Combining the input facial expression data and voice tone data, the emotion estimation algorithm classifies emotions such as "joy" and "anger." The resulting emotion estimation data is then output.
[0323] Step 5:
[0324] The server generates the necessary response based on the emotion estimation result. In this process, it creates response measures and action plans based on the specific emotion. For example, if "joy" is estimated, it outputs a prompt message that generates words of encouragement as the response.
[0325] Step 6:
[0326] The server uses speech recognition technology to convert speech information into text data. It transcribes the input speech data into text, generating text data as a result. This text data serves as the basis for further analysis.
[0327] Step 7:
[0328] The server performs natural language processing based on text data to extract important information. It identifies key points of a conversation from the input text data and outputs that information as a summary. This allows for the efficient extraction of important parts from long conversations.
[0329] Step 8:
[0330] The server uses historical data and individual profiles to generate personalized responses using a generative AI model. It takes prompts as input, prepares and outputs responses tailored to each user, and this response generation assists caregivers in naturally interacting with users.
[0331] (Application Example 1)
[0332] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0333] There is a need to monitor the health status and mental stress levels of workers in factories and other workplaces in real time, and to provide appropriate work support and break suggestions based on that information. However, currently, these decisions rely on individual workers' self-reporting, and there is a problem in that support based on objective data is lacking.
[0334] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0335] In this invention, the server includes a device for acquiring video information, a device for acquiring audio information, a device for analyzing the acquired video information to identify facial expressions, a device for analyzing the acquired audio information to identify audio characteristics, a device for estimating emotions based on the identified facial expressions and audio characteristics, a device for judging the situation based on the estimated emotions, a device for monitoring the worker's condition and detecting stress and fatigue, and a device for suggesting work improvements or breaks based on the detection results. This makes it possible to objectively evaluate the worker's health and emotions and provide appropriate support and suggestions.
[0336] "Visual information" refers to visual data acquired using cameras and visual sensors.
[0337] "Auditory information" refers to auditory data acquired using microphones or sound sensors.
[0338] A "facial expression recognition device" is a device that analyzes acquired video information to identify the characteristics of individual facial expressions.
[0339] A "device for identifying speech characteristics" is a device that analyzes acquired speech information to identify the tone and other acoustic characteristics of the speech.
[0340] A "device for estimating emotions" is a device that integrates identified facial and vocal characteristics to estimate an individual's emotional state.
[0341] A "situation assessment device" is a device that determines whether appropriate action or response is necessary based on estimated emotions.
[0342] "Monitoring" refers to the entire activity of continuously observing the condition of workers, collecting information, and analyzing it.
[0343] A "stress and fatigue detection device" is a device that uses collected data to identify the level of stress and fatigue experienced by workers.
[0344] A "device for suggesting work improvements and breaks" is a device that suggests improvement measures and necessary breaks, taking into account the health and efficiency of workers.
[0345] This invention aims to realize a system that monitors the health status and emotions of workers in real time in workplaces such as factories and provides appropriate feedback. It basically consists of a server, terminals, and users.
[0346] The server plays a central role in data analysis and decision-making, specifically handling the analysis of video and audio information. The server utilizes visual analysis software such as OpenCV and Dlib to acquire video information and identify facial expressions. It also converts audio information into text using the Google Cloud Speech-to-Text API and performs spectral analysis to identify speech features. Natural language processing libraries such as NLTK and spaCy are used to understand the context of conversations and extract important elements. Furthermore, the server assesses the worker's health status based on estimated emotions and suggests breaks or work improvements as needed. These suggestions are customized to individual worker profiles based on predefined criteria.
[0347] The terminal is equipped with hardware devices such as cameras and microphones, and is responsible for collecting information in the field. The data obtained from these devices is transmitted to a server via the internet. The terminal also receives feedback from the server and displays it appropriately to the user.
[0348] Workers, as users, can utilize the information and suggestions provided by the system to improve work efficiency and manage their health. Specifically, they can receive advice on appropriate timing for work breaks and daily stress management. Furthermore, the system has the ability to personalize subsequent suggestions based on the worker's responses.
[0349] As a concrete example, consider a case at a metalworking factory. When a worker is performing complex machine operations, the system detects his fatigue level and suggests, "This task requires concentration, so we recommend you take a 10-minute break." This kind of feedback helps workers prevent mistakes caused by excessive fatigue.
[0350] An example of a prompt for a generative AI model would be: "Based on the following dataset, analyze the emotional state of the workers and suggest appropriate breaks if stress is detected."
[0351] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0352] Step 1:
[0353] The terminal uses cameras and microphones installed at the work site to acquire video and audio information of workers in real time. This information is output as raw video data (e.g., video stream) and audio data (e.g., audio stream). Next, this data is converted into a digital format and sent to the server.
[0354] Step 2:
[0355] The server processes the received video information and uses OpenCV and Dlib libraries to identify facial expressions. Specifically, it extracts facial features from the video data and classifies expressions such as smiles and surprise. The input is digital video data, and the output is category information of the identified facial expressions.
[0356] Step 3:
[0357] The server converts the audio information into text using the Google Cloud Speech-to-Text API and then performs spectral analysis of the audio. It identifies features related to the pitch and intensity of the speech and measures the pitch of the tone. The input for this step is digital audio data, and the output is the converted text information and speech tone features.
[0358] Step 4:
[0359] The server integrates the identified facial expression category information and voice tone characteristics, and uses an estimation algorithm to estimate the worker's emotional state. For example, a smiling expression and a high but gentle tone are likely to indicate "satisfied." The input for this step is the facial expression and voice tone characteristics, and the output is the estimated emotional state.
[0360] Step 5:
[0361] The server assesses the worker's stress level and fatigue based on their estimated emotional state, and provides suggestions for breaks and work improvements as needed. This involves using a generative AI model to create appropriate advice. The input for this step is the result of the emotional state assessment, and the output is specific suggestions for work improvements and breaks.
[0362] Step 6:
[0363] The terminal displays suggestions received from the server to the worker in real time. The user reviews the suggestions and adjusts their actions as needed. The input for this step is suggestion information from the server, and the output is feedback messages provided to the worker.
[0364] In this way, the system can continuously monitor the health status of workers and provide appropriate feedback in real time.
[0365] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0366] This invention provides a highly accurate system for recognizing user emotions. The system incorporates an emotion engine into a server to perform emotion recognition, and estimates the user's emotional state in real time based on video and audio data.
[0367] The device captures the user's current video and audio through its camera and microphone. This data is then immediately transmitted to a server via the internet. The server analyzes the received video data to extract the user's facial features and simultaneously processes the audio data using spectral analysis to identify the tone of voice.
[0368] The emotion engine embedded in the server estimates emotions based on extracted facial features and voice tone. This emotion engine utilizes machine learning algorithms and can improve estimation accuracy by referencing emotion patterns based on the user's past data. For example, if a user has previously displayed certain facial expressions and voice patterns in specific situations, the engine can identify emotional states such as "stress" or "relaxation" based on that data.
[0369] Furthermore, the emotion engine can monitor the user's emotional changes in real time and determine the appropriate timing for intervention based on the analysis results. For example, if a trend of the user gradually becoming depressed is detected, the server will send an early notification to the caregiver to encourage proactive action.
[0370] This system can also effectively provide users with the information they need by converting audio data into text and summarizing key points. The server achieves this function by converting audio data into text using automatic speech recognition and summarizing it using natural language processing.
[0371] Thus, the emotion recognition system of the present invention establishes a highly accurate understanding of the user's emotions, enabling high-quality communication in caregiving settings and providing responses that meet the user's needs.
[0372] The following describes the processing flow.
[0373] Step 1:
[0374] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0375] Step 2:
[0376] The terminal transmits the acquired video and audio data to the server as digital signals.
[0377] Step 3:
[0378] The server analyzes the received video data using an image processing algorithm to extract the user's facial features.
[0379] Step 4:
[0380] The server analyzes the received audio data using an audio processing algorithm to identify the tone and pitch of the voice.
[0381] Step 5:
[0382] The emotion engine embedded in the server estimates the user's emotions based on extracted facial features and voice tone. This emotion engine uses machine learning models to improve its estimation accuracy.
[0383] Step 6:
[0384] The server analyzes the estimated emotions and monitors the trend of emotional changes in real time.
[0385] Step 7:
[0386] Based on trends in emotional changes, the server determines the appropriate timing for intervention for the user and sends notifications to caregivers as needed.
[0387] Step 8:
[0388] The server uses speech recognition technology to convert the audio data into text data.
[0389] Step 9:
[0390] The server analyzes the converted text data using natural language processing techniques, extracts the key points of the conversation, and generates a summary.
[0391] Step 10:
[0392] The server sends summary information to the terminal and presents it to the user.
[0393] Through this process, the system can recognize the user's emotions with high accuracy, enabling a quick and appropriate response.
[0394] (Example 2)
[0395] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0396] Conventional systems struggled to accurately recognize users' emotions and respond appropriately based on them. Furthermore, they were inadequate at grasping emotional changes in real time through the analysis of voice and video data, and at determining the timing of necessary interventions. As a result, they were unable to respond in a way that aligned with users' emotions and needs, making it difficult to provide satisfactory support, particularly in areas such as caregiving and communication.
[0397] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0398] In this invention, the server includes means for acquiring video information, means for acquiring audio information, and means for analyzing the acquired video information to identify facial features. This enables high-precision estimation of the user's emotional state in real time and allows for quick and accurate responses as needed.
[0399] "Visual information" refers to visual data acquired by cameras and other imaging devices.
[0400] "Acoustic information" refers to audio and sound data acquired by sound-collecting devices such as microphones.
[0401] "Device means" refers to a mechanical or electrical device designed to perform a specific function.
[0402] "Analysis" refers to the process of examining data in detail and identifying specific characteristics or trends.
[0403] "Facial features" refer to a person's expression and the appearance of their face, and are used as elements that indicate emotions and reactions.
[0404] "Voice tone" refers to characteristics such as the tone, pitch, and intensity of a voice that reflect the speaker's emotions and intentions.
[0405] "To estimate" refers to predicting or judging unconfirmed matters or situations based on obtained data.
[0406] "Real-time" refers to data processing and system responses occurring immediately, meaning timely execution without delay.
[0407] The "timing of intervention" refers to the point at which specific actions or responses are required, indicating the point that leads to effective support and resolution.
[0408] This invention provides a system that recognizes a user's emotions in real time with high accuracy. The entire system mainly consists of a terminal and a server, and evaluates emotions by utilizing the user's video and audio information.
[0409] The device uses a camera and microphone to record the user's current status. The camera captures video information from the user's face, and the microphone captures audio information from the user's voice. This captured information is transmitted to a server via a communication network.
[0410] The server analyzes the transmitted video information and uses image processing algorithms to identify the user's facial features. It also employs acoustic analysis software to identify the tone of voice by performing spectral analysis on acoustic information. The server incorporates a generative AI model that estimates the user's emotions based on these analysis results. This generative AI model utilizes machine learning algorithms and improves estimation accuracy by referencing emotion patterns based on past data.
[0411] This system also has the function of monitoring emotional changes in real time and determining the appropriate timing for intervention. For example, if a user shows signs of rapidly becoming stressed, the system will notify the caregiver of this information and instruct them on appropriate measures. It also converts acoustic information into text and provides feedback to the user by summarizing the key points.
[0412] For example, if the system determines that a user may be feeling down, it will compare this with past data and recommend relaxing music for that user. Furthermore, using a generative AI model, an example of a prompt message could be: "Estimate the user's emotions in real time. The video data shows a smiling expression, and the audio data indicates a calm tone of voice."
[0413] Thus, the system of the present invention is designed with the aim of deeply understanding the user's emotions and providing effective responses accordingly.
[0414] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0415] Step 1:
[0416] The device acquires the user's video information via its camera. The input is the user's real-time visual situation, which is output as high-resolution digital video data. This video data includes the user's facial expressions, facial features, and movements. The device captures accurate video while maintaining a stable frame rate.
[0417] Step 2:
[0418] The device acquires the user's acoustic information via a microphone. The input is the user's voice, which is output as acoustic sample data. This data captures the tone, pitch, speed, and emotional nuances of the voice in detail. The device uses filters to reduce background noise and ensure clear sound quality.
[0419] Step 3:
[0420] The terminal transmits the acquired video and audio information to the server via the internet. The input is the output data from steps 1 and 2, and this data is transmitted encrypted. The terminal uses a communication protocol to enable secure and rapid transmission of data.
[0421] Step 4:
[0422] The server analyzes the received video information and uses machine learning algorithms to identify facial features. The input is encrypted video data, and the output is elements extracted as identified facial features and expressions. The server identifies key features of the expressions and matches them against known expression patterns registered in the database.
[0423] Step 5:
[0424] The server performs spectral analysis on the received acoustic information to identify the tone of the voice. The input is acoustic sample data, and the output is the analyzed voice tone and emotional characteristics. The server analyzes the frequency components of the voice waveform to identify tones related to the user's emotional state.
[0425] Step 6:
[0426] The server uses a generative AI model to estimate the user's emotions from the identified facial features and tone of voice. The input is the output data from steps 4 and 5, and the output is the estimated emotional state. The server compares the estimation results with past data to improve accuracy and recognize the user's current emotions with high precision.
[0427] Step 7:
[0428] The server monitors estimated emotional changes in real time and determines the appropriate timing for intervention as needed. The input is continuously estimated emotional data, and the output is alerts and notifications for intervention. The server performs real-time analysis based on certain criteria and prompts necessary actions.
[0429] (Application Example 2)
[0430] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0431] In modern society, it is crucial to understand an individual's emotional state in real time and proactively detect potential dangers and stressors. However, conventional technologies lack the means to instantly grasp emotional changes and provide appropriate responses, making it difficult to create a safe and comfortable environment.
[0432] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0433] In this invention, the server includes means for acquiring video information, means for acquiring audio information, means for analyzing the acquired video information to identify facial expressions, means for analyzing the acquired audio information to identify voice tone, means for estimating emotions based on the identified facial expressions and voice tone, and means for determining situations requiring safety measures based on the estimated emotions and sending notifications to the individual to encourage relaxation. This enables real-time emotion monitoring tailored to individual situations and appropriate interventions to ensure safety.
[0434] "Video information" refers to visual data, including the user's facial expressions and movements, and is digital data acquired via a camera device.
[0435] "Acoustic information" refers to auditory data, including the user's voice and ambient sounds, and is digital data acquired using a microphone device.
[0436] "Means for identifying facial expressions" refers to a system component that analyzes facial features from acquired video information and, based on that, identifies facial expressions that indicate the user's emotions.
[0437] "Means for identifying the tone of voice" refers to a component of a system that analyzes the tone and intonation of a voice from acquired acoustic information and identifies the characteristics of a voice that indicate emotion based on that analysis.
[0438] "Means for estimating emotions" refers to a system component that utilizes machine learning algorithms, based on identified facial expression and voice tone data, to infer the user's current emotional state.
[0439] "Means for determining situations requiring safety assurance" refers to a system component that has the function of determining whether or not a user's psychological state requires vigilance, based on estimated emotions.
[0440] "Means of sending notifications to encourage relaxation" refers to a system component that has the function of sending messages to users urging them to relieve stress or rest.
[0441] The specific system for realizing this invention analyzes video and audio data, estimates the user's emotional state in real time, and intervenes to improve safety. Processing is primarily carried out using smart devices and servers.
[0442] The cameras and microphones built into smart devices (such as smart glasses and smartphones) capture the user's video and audio information. This data is transmitted to a server via the network.
[0443] On the server, video information is analyzed using the OpenCV library to identify the user's facial expressions. Furthermore, acoustic information is analyzed using the LibROSA library to identify the tone of the voice. The resulting facial and voice feature data is then input into an emotion recognition model using TensorFlow to estimate the user's emotions.
[0444] Based on the emotion assessment, the server determines whether a situation requiring safety measures is present. Specifically, if the system determines that the user is experiencing stress, it sends a notification to the user encouraging relaxation based on that information. This notification is displayed on the user's smart device to prompt timely intervention.
[0445] For example, if a user on public transport feels anxious, the server can instantly recognize this and send a message to the user's device such as, "Take a deep breath to relax." This helps the user instantly recognize their own state and manage themselves.
[0446] An example of a prompt would be: "Generate a scenario in which an emotion monitoring security app monitors a user's emotional state in real time and provides notifications to enhance safety. Consider the specific situation of anxiety experienced while using public transportation."
[0447] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0448] Step 1:
[0449] The device uses a camera and microphone to acquire video and audio information of the user. The input data consists of video data captured by the camera and audio data collected by the microphone. This data is then prepared for transmission to a server via the network.
[0450] Step 2:
[0451] The server analyzes the received video information using the OpenCV library. The input is video data transmitted from the terminal, and the output is feature data representing the user's facial expressions. Specifically, it calculates the movement and changes of each part of the face and quantifies the facial expression features.
[0452] Step 3:
[0453] The server analyzes the received acoustic information using the LibROSA library. The input is audio data transmitted from the terminal, and the output is feature data indicating the tone of the voice. Specifically, it performs spectral analysis of the audio waveform and extracts features such as pitch, intonation, and tempo.
[0454] Step 4:
[0455] The server inputs the obtained facial expression feature data and voice tone feature data into a generative AI model using TensorFlow to estimate emotions. The output is data indicating the user's emotional state. Specifically, the model uses pattern recognition based on past data to determine whether the emotion corresponds to "stress" or "relaxation."
[0456] Step 5:
[0457] The server generates prompt messages to determine if the estimated emotional state requires safety measures. The input is emotional state data, and the output is a notification message to the user. It generates messages such as "Take a deep breath to relax" to assess the need for intervention.
[0458] Step 6:
[0459] The server sends a notification to the terminal to encourage the user to relax if it determines that security measures are necessary. The input is the generated notification message, and the output is the relaxation notification displayed on the terminal screen. This allows the user to take necessary actions in a timely manner.
[0460] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0461] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0462] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0463] [Third Embodiment]
[0464] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0465] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0466] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0467] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0468] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0470] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0471] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0472] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0473] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0474] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0475] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0476] This invention relates to a system for improving the quality of communication in caregiving settings. This system functions by acquiring user video and audio data in real time and transmitting it to a server. The server analyzes this data to accurately estimate the user's emotions and automatically determine situations requiring intervention.
[0477] Specifically, the camera on the device captures the user's facial expressions, and the microphone collects audio. This data is converted into digital signals and transmitted to a server via the internet. The server processes the received video data and applies a facial recognition algorithm to identify facial features. For the audio data, spectral analysis is performed to extract tone and pitch information.
[0478] Based on these analyses, the server uses estimation methods to determine the user's emotions. For example, if a smile is recognized through facial analysis and the voice tone is high, the emotion of "joy" is estimated. Conversely, if the facial expression is stern and the voice tone is low, "anger" may be estimated.
[0479] Furthermore, the speech-to-text and summarization functions allow for the automatic generation of key points from long conversations. In this process, the server converts speech to text using speech recognition technology and then summarizes it using natural language processing. For example, it can transcribe meeting audio and extract meeting conclusions and action items.
[0480] To understand the context of a conversation, the server analyzes the input text data to understand the context and predict the next expected question or topic. This allows for natural responses to spontaneous questions from users, such as "Why?".
[0481] This system can also generate personalized responses by learning from the past history based on each care recipient's profile, thereby providing individualized responses to each user. This improves the user experience.
[0482] Regarding information recording and sharing, the system automatically records important data during caregiving and securely shares it with family members and medical professionals. This data management is primarily done on servers and centralized in a database to maintain consistency in care.
[0483] In this way, this system integrates and operates multiple functions to meet the diverse needs of users, thereby reducing the burden on caregivers and facilitating smooth communication.
[0484] The following describes the processing flow.
[0485] Step 1:
[0486] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0487] Step 2:
[0488] The device transmits the acquired video and audio data to the server via the internet.
[0489] Step 3:
[0490] The server analyzes the received video data using an image processing algorithm to detect the user's face.
[0491] Step 4:
[0492] The server applies an expression recognition algorithm after face detection to identify facial features. This extracts features such as smiles and frowns.
[0493] Step 5:
[0494] The server analyzes the received audio data using an audio processing algorithm to extract the tone and pitch of the voice.
[0495] Step 6:
[0496] The server integrates identified facial and vocal features to estimate the user's emotions. For example, it might estimate "joy" from a smile and a high-pitched voice.
[0497] Step 7:
[0498] The server converts the audio data into text using automatic speech recognition technology.
[0499] Step 8:
[0500] The server analyzes the converted text using natural language processing techniques to extract keywords and topics.
[0501] Step 9:
[0502] The server generates a summary based on the extracted keywords and topics, and extracts the key points.
[0503] Step 10:
[0504] The server analyzes text data to understand the context of the conversation and predicts the next expected question or topic.
[0505] Step 11:
[0506] The server generates an appropriate response to the predicted question and sends it to the terminal.
[0507] Step 12:
[0508] The terminal displays the response received from the server to the user.
[0509] Step 13:
[0510] The server stores important data and information recorded during caregiving in a database and shares it with family members and medical professionals as needed.
[0511] Through these steps, the system accurately grasps the user's emotions, enabling smooth communication and efficient information management in caregiving settings.
[0512] (Example 1)
[0513] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0514] In caregiving settings, smooth communication between care recipients and caregivers is essential. However, conventional systems have difficulty accurately grasping users' emotions, making immediate responses challenging. Furthermore, utilizing accumulated information for caregiving purposes and securely sharing it have also been identified as issues.
[0515] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0516] In this invention, the server includes acquisition means for acquiring video information, acquisition means for acquiring audio information, and estimation means for analyzing the acquired information to estimate emotions. This makes it possible to grasp the user's emotions in real time and automate appropriate responses. Furthermore, the quality and efficiency of care can be improved through personalized recording and secure sharing of the collected information.
[0517] "Acquisition means" refers to devices and methods for collecting user video and audio information.
[0518] "Identification means" refers to a device or method that analyzes acquired information to identify the user's facial expressions and voice characteristics.
[0519] An "estimation means" is a device or method for determining a user's emotions or situation based on identified data.
[0520] A "means of judgment" refers to a device or method for determining the necessary response based on presumed emotions.
[0521] "Conversion means" refers to devices or methods for converting audio information into text information.
[0522] An "extraction means" is a device or method for finding useful information from converted character information.
[0523] "Generating means" refers to devices or methods for creating summaries or responses based on extracted information.
[0524] "Means of understanding" refer to devices and methods for grasping the context of text data and understanding the next course of action to take.
[0525] "Inference tools" are devices or methods that predict future actions or questions based on an understood context.
[0526] "Management means" refers to devices and methods for properly recording and protecting information and sharing it as necessary.
[0527] In this invention, the following hardware and software are used to implement a system that improves the quality of communication in caregiving settings.
[0528] The device is equipped with a camera and microphone, and captures the user's video and audio in real time. The device converts the captured video data into a digital signal and transmits it to a server via the internet. Similarly, the audio data acquired by the microphone is also converted into a digital signal and transmitted to the server.
[0529] The server uses a facial recognition algorithm to process the received video and analyze the facial expression data. This identifies feature points on the user's face and identifies the characteristics of their facial expressions. For audio data, spectral analysis is performed to extract the tone and pitch of the voice. Based on this data, a generative AI model is used to estimate emotions.
[0530] As a concrete example, if a user is smiling and their voice tone is high, the generative AI model will estimate that this indicates an emotion of "joy." The server will then communicate this information to the caregiver in real time to help them take the necessary actions.
[0531] Furthermore, the system converts the acquired audio using speech recognition technology into text data. The converted text data is then processed using natural language processing to extract key points and generate a summary, enabling efficient comprehension of long conversations.
[0532] An example of a prompt message might be: "If the user is smiling and their voice is high-pitched, assume that this emotion is joy. Also, generate words of encouragement appropriate to the situation." This system will make daily care smoother and of higher quality.
[0533] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0534] Step 1:
[0535] The device uses a camera to capture the user's video and a microphone to collect audio. It converts the video and audio information into digital format and then transmits it to a server via the internet. This requires maintaining the quality of the video information while converting it to a transmission-compatible format.
[0536] Step 2:
[0537] The server analyzes the received video information using a facial recognition algorithm. It extracts feature points of the user's face from the input video data and identifies their patterns. As output of this identification, it obtains data on the user's facial features. Specifically, it calculates and digitizes things like the angle of the mouth and the position of the eyebrows.
[0538] Step 3:
[0539] The server performs spectral analysis on the received audio information. It extracts frequency components from the input audio data and analyzes the tone and pitch of the voice. This analysis outputs parameters such as the pitch and volume of the user's voice. In particular, the tone of voice is important information for estimating emotions.
[0540] Step 4:
[0541] The server uses a generative AI model to estimate the user's emotions based on the output of facial recognition and voice analysis. Combining the input facial expression data and voice tone data, the emotion estimation algorithm classifies emotions such as "joy" and "anger." The resulting emotion estimation data is then output.
[0542] Step 5:
[0543] The server generates the necessary response based on the emotion estimation result. In this process, it creates response measures and action plans based on the specific emotion. For example, if "joy" is estimated, it outputs a prompt message that generates words of encouragement as the response.
[0544] Step 6:
[0545] The server uses speech recognition technology to convert speech information into text data. It transcribes the input speech data into text, generating text data as a result. This text data serves as the basis for further analysis.
[0546] Step 7:
[0547] The server performs natural language processing based on text data to extract important information. It identifies key points of a conversation from the input text data and outputs that information as a summary. This allows for the efficient extraction of important parts from long conversations.
[0548] Step 8:
[0549] The server uses historical data and individual profiles to generate personalized responses using a generative AI model. It takes prompts as input, prepares and outputs responses tailored to each user. This response generation assists caregivers in naturally interacting with users.
[0550] (Application Example 1)
[0551] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0552] There is a need to monitor the health status and mental stress levels of workers in factories and other workplaces in real time, and to provide appropriate work support and break suggestions based on that information. However, currently, these decisions rely on individual workers' self-reporting, and there is a problem in that support based on objective data is lacking.
[0553] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0554] In this invention, the server includes a device for acquiring video information, a device for acquiring audio information, a device for analyzing the acquired video information to identify facial expressions, a device for analyzing the acquired audio information to identify audio characteristics, a device for estimating emotions based on the identified facial expressions and audio characteristics, a device for judging the situation based on the estimated emotions, a device for monitoring the worker's condition and detecting stress and fatigue, and a device for suggesting work improvements or breaks based on the detection results. This makes it possible to objectively evaluate the worker's health and emotions and provide appropriate support and suggestions.
[0555] "Visual information" refers to visual data acquired using cameras and visual sensors.
[0556] "Auditory information" refers to auditory data acquired using microphones or sound sensors.
[0557] A "facial expression recognition device" is a device that analyzes acquired video information to identify the characteristics of individual facial expressions.
[0558] A "device for identifying speech characteristics" is a device that analyzes acquired speech information to identify the tone and other acoustic characteristics of the speech.
[0559] A "device for estimating emotions" is a device that integrates identified facial and vocal characteristics to estimate an individual's emotional state.
[0560] A "situation assessment device" is a device that determines whether appropriate action or response is necessary based on estimated emotions.
[0561] "Monitoring" refers to the entire activity of continuously observing the condition of workers, collecting information, and analyzing it.
[0562] A "stress and fatigue detection device" is a device that uses collected data to identify the level of stress and fatigue experienced by workers.
[0563] A "device for suggesting work improvements and breaks" is a device that suggests improvement measures and necessary breaks, taking into account the health and efficiency of workers.
[0564] This invention aims to realize a system that monitors the health status and emotions of workers in real time in workplaces such as factories and provides appropriate feedback. It basically consists of a server, terminals, and users.
[0565] The server plays a central role in data analysis and decision-making, specifically handling the analysis of video and audio information. The server utilizes visual analysis software such as OpenCV and Dlib to acquire video information and identify facial expressions. It also converts audio information into text using the Google Cloud Speech-to-Text API and performs spectral analysis to identify speech features. Natural language processing libraries such as NLTK and spaCy are used to understand the context of conversations and extract important elements. Furthermore, the server assesses the worker's health status based on estimated emotions and suggests breaks or work improvements as needed. These suggestions are customized to individual worker profiles based on predefined criteria.
[0566] The terminal is equipped with hardware devices such as cameras and microphones, and is responsible for collecting information in the field. The data obtained from these devices is transmitted to a server via the internet. The terminal also receives feedback from the server and displays it appropriately to the user.
[0567] Workers, as users, can utilize the information and suggestions provided by the system to improve work efficiency and manage their health. Specifically, they can receive advice on appropriate timing for work breaks and daily stress management. Furthermore, the system has the ability to personalize subsequent suggestions based on the worker's responses.
[0568] As a concrete example, consider a case at a metalworking factory. When a worker is performing complex machine operations, the system detects his fatigue level and suggests, "This task requires concentration, so we recommend you take a 10-minute break." This kind of feedback helps workers prevent mistakes caused by excessive fatigue.
[0569] An example of a prompt for a generative AI model would be: "Based on the following dataset, analyze the emotional state of the workers and suggest appropriate breaks if stress is detected."
[0570] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0571] Step 1:
[0572] The terminal uses cameras and microphones installed at the work site to acquire video and audio information of workers in real time. This information is output as raw video data (e.g., video stream) and audio data (e.g., audio stream). Next, this data is converted into a digital format and sent to the server.
[0573] Step 2:
[0574] The server processes the received video information and uses OpenCV and Dlib libraries to identify facial expressions. Specifically, it extracts facial features from the video data and classifies expressions such as smiles and surprise. The input is digital video data, and the output is category information of the identified facial expressions.
[0575] Step 3:
[0576] The server converts the audio information into text using the Google Cloud Speech-to-Text API and then performs spectral analysis of the audio. It identifies features related to the pitch and intensity of the speech and measures the pitch of the tone. The input for this step is digital audio data, and the output is the converted text information and speech tone features.
[0577] Step 4:
[0578] The server integrates the identified facial expression category information and voice tone characteristics, and uses an estimation algorithm to estimate the worker's emotional state. For example, a smiling expression and a high but gentle tone are likely to indicate "satisfied." The input for this step is the facial expression and voice tone characteristics, and the output is the estimated emotional state.
[0579] Step 5:
[0580] The server assesses the worker's stress level and fatigue based on their estimated emotional state, and provides suggestions for breaks and work improvements as needed. This involves using a generative AI model to create appropriate advice. The input for this step is the result of the emotional state assessment, and the output is specific suggestions for work improvements and breaks.
[0581] Step 6:
[0582] The terminal displays suggestions received from the server to the worker in real time. The user reviews the suggestions and adjusts their actions as needed. The input for this step is suggestion information from the server, and the output is feedback messages provided to the worker.
[0583] In this way, the system can continuously monitor the health status of workers and provide appropriate feedback in real time.
[0584] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0585] This invention provides a highly accurate system for recognizing user emotions. The system incorporates an emotion engine into a server to perform emotion recognition, and estimates the user's emotional state in real time based on video and audio data.
[0586] The device captures the user's current video and audio through its camera and microphone. This data is then immediately transmitted to a server via the internet. The server analyzes the received video data to extract the user's facial features and simultaneously processes the audio data using spectral analysis to identify the tone of voice.
[0587] The emotion engine embedded in the server estimates emotions based on extracted facial features and voice tone. This emotion engine utilizes machine learning algorithms and can improve estimation accuracy by referencing emotion patterns based on the user's past data. For example, if a user has previously displayed certain facial expressions and voice patterns in specific situations, the engine can identify emotional states such as "stress" or "relaxation" based on that data.
[0588] Furthermore, the emotion engine can monitor the user's emotional changes in real time and determine the appropriate timing for intervention based on the analysis results. For example, if a trend of the user gradually becoming depressed is detected, the server will send an early notification to the caregiver to encourage proactive action.
[0589] This system can also effectively provide users with the information they need by converting audio data into text and summarizing key points. The server achieves this function by converting audio data into text using automatic speech recognition and summarizing it using natural language processing.
[0590] Thus, the emotion recognition system of the present invention establishes a highly accurate understanding of the user's emotions, enabling high-quality communication in caregiving settings and providing responses that meet the user's needs.
[0591] The following describes the processing flow.
[0592] Step 1:
[0593] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0594] Step 2:
[0595] The terminal transmits the acquired video and audio data to the server as digital signals.
[0596] Step 3:
[0597] The server analyzes the received video data using an image processing algorithm to extract the user's facial features.
[0598] Step 4:
[0599] The server analyzes the received audio data using an audio processing algorithm to identify the tone and pitch of the voice.
[0600] Step 5:
[0601] The emotion engine embedded in the server estimates the user's emotions based on extracted facial features and voice tone. This emotion engine uses machine learning models to improve its estimation accuracy.
[0602] Step 6:
[0603] The server analyzes the estimated emotions and monitors the trend of emotional changes in real time.
[0604] Step 7:
[0605] Based on trends in emotional changes, the server determines the appropriate timing for intervention for the user and sends notifications to caregivers as needed.
[0606] Step 8:
[0607] The server uses speech recognition technology to convert the audio data into text data.
[0608] Step 9:
[0609] The server analyzes the converted text data using natural language processing techniques, extracts the key points of the conversation, and generates a summary.
[0610] Step 10:
[0611] The server sends summary information to the terminal and presents it to the user.
[0612] Through this process, the system can recognize the user's emotions with high accuracy, enabling a quick and appropriate response.
[0613] (Example 2)
[0614] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0615] Conventional systems struggled to accurately recognize users' emotions and respond appropriately based on them. Furthermore, they were inadequate at grasping emotional changes in real time through the analysis of voice and video data, and at determining the timing of necessary interventions. As a result, they were unable to respond in a way that aligned with users' emotions and needs, making it difficult to provide satisfactory support, particularly in areas such as caregiving and communication.
[0616] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0617] In this invention, the server includes means for acquiring video information, means for acquiring audio information, and means for analyzing the acquired video information to identify facial features. This enables high-precision estimation of the user's emotional state in real time and allows for quick and accurate responses as needed.
[0618] "Visual information" refers to visual data acquired by cameras and other imaging devices.
[0619] "Acoustic information" refers to audio and sound data acquired by sound-collecting devices such as microphones.
[0620] "Device means" refers to a mechanical or electrical device designed to perform a specific function.
[0621] "Analysis" refers to the process of examining data in detail and identifying specific characteristics or trends.
[0622] "Facial features" refer to a person's expression and the appearance of their face, and are used as elements that indicate emotions and reactions.
[0623] "Voice tone" refers to characteristics such as the tone, pitch, and intensity of a voice that reflect the speaker's emotions and intentions.
[0624] "To estimate" refers to predicting or judging unconfirmed matters or situations based on obtained data.
[0625] "Real-time" refers to data processing and system responses occurring immediately, meaning timely execution without delay.
[0626] The "timing of intervention" refers to the point at which specific actions or responses are required, indicating the point that leads to effective support and resolution.
[0627] This invention provides a system that recognizes a user's emotions in real time with high accuracy. The entire system mainly consists of a terminal and a server, and evaluates emotions by utilizing the user's video and audio information.
[0628] The device uses a camera and microphone to record the user's current status. The camera captures video information from the user's face, and the microphone captures audio information from the user's voice. This captured information is transmitted to a server via a communication network.
[0629] The server analyzes the transmitted video information and uses image processing algorithms to identify the user's facial features. It also employs acoustic analysis software to identify the tone of voice by performing spectral analysis on acoustic information. The server incorporates a generative AI model that estimates the user's emotions based on these analysis results. This generative AI model utilizes machine learning algorithms and improves estimation accuracy by referencing emotion patterns based on past data.
[0630] This system also has the function of monitoring emotional changes in real time and determining the appropriate timing for intervention. For example, if a user shows signs of rapidly becoming stressed, the system will notify the caregiver of this information and instruct them on appropriate measures. It also converts acoustic information into text and provides feedback to the user by summarizing the key points.
[0631] For example, if the system determines that a user may be feeling down, it will compare this with past data and recommend relaxing music for that user. Furthermore, using a generative AI model, an example of a prompt message could be: "Estimate the user's emotions in real time. The video data shows a smiling expression, and the audio data indicates a calm tone of voice."
[0632] Thus, the system of the present invention is designed with the aim of deeply understanding the user's emotions and providing effective responses accordingly.
[0633] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0634] Step 1:
[0635] The device acquires the user's video information via its camera. The input is the user's real-time visual situation, which is output as high-resolution digital video data. This video data includes the user's facial expressions, facial features, and movements. The device captures accurate video while maintaining a stable frame rate.
[0636] Step 2:
[0637] The device acquires the user's acoustic information via a microphone. The input is the user's voice, which is output as acoustic sample data. This data captures the tone, pitch, speed, and emotional nuances of the voice in detail. The device uses filters to reduce background noise and ensure clear sound quality.
[0638] Step 3:
[0639] The terminal transmits the acquired video and audio information to the server via the internet. The input is the output data from steps 1 and 2, and this data is transmitted encrypted. The terminal uses a communication protocol to enable secure and rapid transmission of data.
[0640] Step 4:
[0641] The server analyzes the received video information and uses machine learning algorithms to identify facial features. The input is encrypted video data, and the output is elements extracted as identified facial features and expressions. The server identifies key features of the expressions and matches them against known expression patterns registered in the database.
[0642] Step 5:
[0643] The server performs spectral analysis on the received acoustic information to identify the tone of the voice. The input is acoustic sample data, and the output is the analyzed voice tone and emotional characteristics. The server analyzes the frequency components of the voice waveform to identify tones related to the user's emotional state.
[0644] Step 6:
[0645] The server uses a generative AI model to estimate the user's emotions from the identified facial features and tone of voice. The input is the output data from steps 4 and 5, and the output is the estimated emotional state. The server compares the estimation results with past data to improve accuracy and recognize the user's current emotions with high precision.
[0646] Step 7:
[0647] The server monitors estimated emotional changes in real time and determines the appropriate timing for intervention as needed. The input is continuously estimated emotional data, and the output is alerts and notifications for intervention. The server performs real-time analysis based on certain criteria and prompts necessary actions.
[0648] (Application Example 2)
[0649] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0650] In modern society, it is crucial to understand an individual's emotional state in real time and proactively detect potential dangers and stressors. However, conventional technologies lack the means to instantly grasp emotional changes and provide appropriate responses, making it difficult to create a safe and comfortable environment.
[0651] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0652] In this invention, the server includes means for acquiring video information, means for acquiring audio information, means for analyzing the acquired video information to identify facial expressions, means for analyzing the acquired audio information to identify voice tone, means for estimating emotions based on the identified facial expressions and voice tone, and means for determining situations requiring safety measures based on the estimated emotions and sending notifications to the individual to encourage relaxation. This enables real-time emotion monitoring tailored to individual situations and appropriate interventions to ensure safety.
[0653] "Video information" refers to visual data, including the user's facial expressions and movements, and is digital data acquired via a camera device.
[0654] "Acoustic information" refers to auditory data, including the user's voice and ambient sounds, and is digital data acquired using a microphone device.
[0655] "Means for identifying facial expressions" refers to a system component that analyzes facial features from acquired video information and, based on that, identifies facial expressions that indicate the user's emotions.
[0656] "Means for identifying the tone of voice" refers to a component of a system that analyzes the tone and intonation of a voice from acquired acoustic information and identifies the characteristics of a voice that indicate emotion based on that analysis.
[0657] "Means for estimating emotions" refers to a system component that utilizes machine learning algorithms, based on identified facial expression and voice tone data, to infer the user's current emotional state.
[0658] "Means for determining situations requiring safety assurance" refers to a system component that has the function of determining whether or not a user's psychological state requires vigilance, based on estimated emotions.
[0659] "Means of sending notifications to encourage relaxation" refers to system components that have the functionality to send messages to users urging them to relieve stress or rest.
[0660] The specific system for realizing this invention analyzes video and audio data, estimates the user's emotional state in real time, and intervenes to improve safety. Processing is primarily carried out using smart devices and servers.
[0661] The cameras and microphones built into smart devices (such as smart glasses and smartphones) capture the user's video and audio information. This data is transmitted to a server via the network.
[0662] On the server, video information is analyzed using the OpenCV library to identify the user's facial expressions. Furthermore, acoustic information is analyzed using the LibROSA library to identify the tone of the voice. The resulting facial and voice feature data is then input into an emotion recognition model using TensorFlow to estimate the user's emotions.
[0663] Based on the emotion assessment, the server determines whether a situation requiring safety measures is present. Specifically, if the system determines that the user is experiencing stress, it sends a notification to the user encouraging relaxation based on that information. This notification is displayed on the user's smart device to prompt timely intervention.
[0664] For example, if a user on public transport feels anxious, the server can instantly recognize this and send a message to the user's device such as, "Take a deep breath to relax." This helps the user instantly recognize their own state and manage themselves.
[0665] An example of a prompt would be: "Generate a scenario in which an emotion monitoring security app monitors a user's emotional state in real time and provides notifications to enhance safety. Consider the specific situation of anxiety experienced while using public transportation."
[0666] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0667] Step 1:
[0668] The device uses a camera and microphone to acquire video and audio information of the user. The input data consists of video data captured by the camera and audio data collected by the microphone. This data is then prepared for transmission to a server via the network.
[0669] Step 2:
[0670] The server analyzes the received video information using the OpenCV library. The input is video data transmitted from the terminal, and the output is feature data representing the user's facial expressions. Specifically, it calculates the movement and changes of each part of the face and quantifies the facial expression features.
[0671] Step 3:
[0672] The server analyzes the received acoustic information using the LibROSA library. The input is audio data transmitted from the terminal, and the output is feature data indicating the tone of the voice. Specifically, it performs spectral analysis of the audio waveform and extracts features such as pitch, intonation, and tempo.
[0673] Step 4:
[0674] The server inputs the obtained facial expression feature data and voice tone feature data into a generative AI model using TensorFlow to estimate emotions. The output is data indicating the user's emotional state. Specifically, the model uses pattern recognition based on past data to determine whether the emotion corresponds to "stress" or "relaxation."
[0675] Step 5:
[0676] The server generates prompt messages to determine if the estimated emotional state requires safety measures. The input is emotional state data, and the output is a notification message to the user. It generates messages such as "Take a deep breath to relax" to assess the need for intervention.
[0677] Step 6:
[0678] The server sends a notification to the terminal to encourage the user to relax if it determines that security measures are necessary. The input is the generated notification message, and the output is the relaxation notification displayed on the terminal screen. This allows the user to take necessary actions in a timely manner.
[0679] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0680] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0681] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0682] [Fourth Embodiment]
[0683] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0684] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0685] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0686] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0687] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0688] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0689] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0690] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0691] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0692] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0693] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0694] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0695] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0696] This invention relates to a system for improving the quality of communication in caregiving settings. This system functions by acquiring user video and audio data in real time and transmitting it to a server. The server analyzes this data to accurately estimate the user's emotions and automatically determine situations requiring intervention.
[0697] Specifically, the camera on the device captures the user's facial expressions, and the microphone collects audio. This data is converted into digital signals and transmitted to a server via the internet. The server processes the received video data and applies a facial recognition algorithm to identify facial features. For the audio data, spectral analysis is performed to extract tone and pitch information.
[0698] Based on these analyses, the server uses estimation methods to determine the user's emotions. For example, if a smile is recognized through facial analysis and the voice tone is high, the emotion of "joy" is estimated. Conversely, if the facial expression is stern and the voice tone is low, "anger" may be estimated.
[0699] Furthermore, the speech-to-text and summarization functions allow for the automatic generation of key points from long conversations. In this process, the server converts speech to text using speech recognition technology and then summarizes it using natural language processing. For example, it can transcribe meeting audio and extract meeting conclusions and action items.
[0700] To understand the context of a conversation, the server analyzes the input text data to understand the context and predict the next expected question or topic. This allows for natural responses to spontaneous questions from users, such as "Why?".
[0701] This system can also generate personalized responses by learning from the past history based on each care recipient's profile, thereby providing individualized responses to each user. This improves the user experience.
[0702] Regarding information recording and sharing, the system automatically records important data during caregiving and securely shares it with family members and medical professionals. This data management is primarily done on servers and centralized in a database to maintain consistency in care.
[0703] In this way, this system integrates and operates multiple functions to meet the diverse needs of users, thereby reducing the burden on caregivers and facilitating smooth communication.
[0704] The following describes the processing flow.
[0705] Step 1:
[0706] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0707] Step 2:
[0708] The device transmits the acquired video and audio data to the server via the internet.
[0709] Step 3:
[0710] The server analyzes the received video data using an image processing algorithm to detect the user's face.
[0711] Step 4:
[0712] The server applies an expression recognition algorithm after face detection to identify facial features. This extracts features such as smiles and frowns.
[0713] Step 5:
[0714] The server analyzes the received audio data using an audio processing algorithm to extract the tone and pitch of the voice.
[0715] Step 6:
[0716] The server integrates identified facial and vocal features to estimate the user's emotions. For example, it might estimate "joy" from a smile and a high-pitched voice.
[0717] Step 7:
[0718] The server converts the audio data into text using automatic speech recognition technology.
[0719] Step 8:
[0720] The server analyzes the converted text using natural language processing techniques to extract keywords and topics.
[0721] Step 9:
[0722] The server generates a summary based on the extracted keywords and topics, and extracts the key points.
[0723] Step 10:
[0724] The server analyzes text data to understand the context of the conversation and predicts the next expected question or topic.
[0725] Step 11:
[0726] The server generates an appropriate response to the predicted question and sends it to the terminal.
[0727] Step 12:
[0728] The terminal displays the response received from the server to the user.
[0729] Step 13:
[0730] The server stores important data and information recorded during caregiving in a database and shares it with family members and medical professionals as needed.
[0731] Through these steps, the system accurately grasps the user's emotions, enabling smooth communication and efficient information management in caregiving settings.
[0732] (Example 1)
[0733] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0734] In caregiving settings, smooth communication between care recipients and caregivers is essential. However, conventional systems have difficulty accurately grasping users' emotions, making immediate responses challenging. Furthermore, utilizing accumulated information for caregiving purposes and securely sharing it have also been identified as issues.
[0735] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0736] In this invention, the server includes acquisition means for acquiring video information, acquisition means for acquiring audio information, and estimation means for analyzing the acquired information to estimate emotions. This makes it possible to grasp the user's emotions in real time and automate appropriate responses. Furthermore, the quality and efficiency of care can be improved through personalized recording and secure sharing of the collected information.
[0737] "Acquisition means" refers to devices and methods for collecting user video and audio information.
[0738] "Identification means" refers to a device or method that analyzes acquired information to identify the user's facial expressions and voice characteristics.
[0739] An "estimation means" is a device or method for determining a user's emotions or situation based on identified data.
[0740] A "means of judgment" refers to a device or method for determining the necessary response based on presumed emotions.
[0741] "Conversion means" refers to devices or methods for converting audio information into text information.
[0742] An "extraction means" is a device or method for finding useful information from converted character information.
[0743] "Generating means" refers to devices or methods for creating summaries or responses based on extracted information.
[0744] "Means of understanding" refer to devices and methods for grasping the context of text data and understanding the next course of action to take.
[0745] "Inference tools" are devices or methods that predict future actions or questions based on an understood context.
[0746] "Management means" refers to devices and methods for properly recording and protecting information and sharing it as necessary.
[0747] In this invention, the following hardware and software are used to implement a system that improves the quality of communication in caregiving settings.
[0748] The device is equipped with a camera and microphone, and captures the user's video and audio in real time. The device converts the captured video data into a digital signal and transmits it to a server via the internet. Similarly, the audio data acquired by the microphone is also converted into a digital signal and transmitted to the server.
[0749] The server uses a facial recognition algorithm to process the received video and analyze the facial expression data. This identifies feature points on the user's face and identifies the characteristics of their facial expressions. For audio data, spectral analysis is performed to extract the tone and pitch of the voice. Based on this data, a generative AI model is used to estimate emotions.
[0750] As a concrete example, if a user is smiling and their voice tone is high, the generative AI model will estimate that this indicates an emotion of "joy." The server will then communicate this information to the caregiver in real time to help them take the necessary actions.
[0751] Furthermore, the system converts the acquired audio using speech recognition technology into text data. The converted text data is then processed using natural language processing to extract key points and generate a summary, enabling efficient comprehension of long conversations.
[0752] An example of a prompt message might be: "If the user is smiling and their voice is high-pitched, assume that this emotion is joy. Also, generate words of encouragement appropriate to the situation." This system will make daily care smoother and of higher quality.
[0753] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0754] Step 1:
[0755] The device uses a camera to capture the user's video and a microphone to collect audio. It converts the video and audio information into digital format and then transmits it to a server via the internet. This requires maintaining the quality of the video information while converting it to a transmission-compatible format.
[0756] Step 2:
[0757] The server analyzes the received video information using a facial recognition algorithm. It extracts feature points of the user's face from the input video data and identifies their patterns. As output of this identification, it obtains data on the user's facial features. Specifically, it calculates and digitizes things like the angle of the mouth and the position of the eyebrows.
[0758] Step 3:
[0759] The server performs spectral analysis on the received audio information. It extracts frequency components from the input audio data and analyzes the tone and pitch of the voice. This analysis outputs parameters such as the pitch and volume of the user's voice. In particular, the tone of voice is important information for estimating emotions.
[0760] Step 4:
[0761] The server uses a generative AI model to estimate the user's emotions based on the output of facial recognition and voice analysis. Combining the input facial expression data and voice tone data, the emotion estimation algorithm classifies emotions such as "joy" and "anger." The resulting emotion estimation data is then output.
[0762] Step 5:
[0763] The server generates the necessary response based on the emotion estimation result. In this process, it creates response measures and action plans based on the specific emotion. For example, if "joy" is estimated, it outputs a prompt message that generates words of encouragement as the response.
[0764] Step 6:
[0765] The server uses speech recognition technology to convert speech information into text data. It transcribes the input speech data into text, generating text data as a result. This text data serves as the basis for further analysis.
[0766] Step 7:
[0767] The server performs natural language processing based on text data to extract important information. It identifies key points of a conversation from the input text data and outputs that information as a summary. This allows for the efficient extraction of important parts from long conversations.
[0768] Step 8:
[0769] The server uses historical data and individual profiles to generate personalized responses using a generative AI model. It takes prompts as input, prepares and outputs responses tailored to each user. This response generation assists caregivers in naturally interacting with users.
[0770] (Application Example 1)
[0771] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0772] There is a need to monitor the health status and mental stress levels of workers in factories and other workplaces in real time, and to provide appropriate work support and break suggestions based on that information. However, currently, these decisions rely on individual workers' self-reporting, and there is a problem in that support based on objective data is lacking.
[0773] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0774] In this invention, the server includes a device for acquiring video information, a device for acquiring audio information, a device for analyzing the acquired video information to identify facial expressions, a device for analyzing the acquired audio information to identify audio characteristics, a device for estimating emotions based on the identified facial expressions and audio characteristics, a device for judging the situation based on the estimated emotions, a device for monitoring the worker's condition and detecting stress and fatigue, and a device for suggesting work improvements or breaks based on the detection results. This makes it possible to objectively evaluate the worker's health and emotions and provide appropriate support and suggestions.
[0775] "Visual information" refers to visual data acquired using cameras and visual sensors.
[0776] "Auditory information" refers to auditory data acquired using microphones or sound sensors.
[0777] A "facial expression recognition device" is a device that analyzes acquired video information to identify the characteristics of individual facial expressions.
[0778] A "device for identifying speech characteristics" is a device that analyzes acquired speech information to identify the tone and other acoustic characteristics of the speech.
[0779] A "device for estimating emotions" is a device that integrates identified facial and vocal characteristics to estimate an individual's emotional state.
[0780] A "situation assessment device" is a device that determines whether appropriate action or response is necessary based on estimated emotions.
[0781] "Monitoring" refers to the entire activity of continuously observing the condition of workers, collecting information, and analyzing it.
[0782] A "stress and fatigue detection device" is a device that uses collected data to identify the level of stress and fatigue experienced by workers.
[0783] A "device for suggesting work improvements and breaks" is a device that suggests improvement measures and necessary breaks, taking into account the health and efficiency of workers.
[0784] This invention aims to realize a system that monitors the health status and emotions of workers in real time in workplaces such as factories and provides appropriate feedback. It basically consists of a server, terminals, and users.
[0785] The server plays a central role in data analysis and decision-making, specifically handling the analysis of video and audio information. The server utilizes visual analysis software such as OpenCV and Dlib to acquire video information and identify facial expressions. It also converts audio information into text using the Google Cloud Speech-to-Text API and performs spectral analysis to identify speech features. Natural language processing libraries such as NLTK and spaCy are used to understand the context of conversations and extract important elements. Furthermore, the server assesses the worker's health status based on estimated emotions and suggests breaks or work improvements as needed. These suggestions are customized to individual worker profiles based on predefined criteria.
[0786] The terminal is equipped with hardware devices such as cameras and microphones, and is responsible for collecting information in the field. The data obtained from these devices is transmitted to a server via the internet. The terminal also receives feedback from the server and displays it appropriately to the user.
[0787] Workers, as users, can utilize the information and suggestions provided by the system to improve work efficiency and manage their health. Specifically, they can receive advice on appropriate timing for work breaks and daily stress management. Furthermore, the system has the ability to personalize subsequent suggestions based on the worker's responses.
[0788] As a concrete example, consider a case at a metalworking factory. When a worker is performing complex machine operations, the system detects his fatigue level and suggests, "This task requires concentration, so we recommend you take a 10-minute break." This kind of feedback helps workers prevent mistakes caused by excessive fatigue.
[0789] An example of a prompt for a generative AI model would be: "Based on the following dataset, analyze the emotional state of the workers and suggest appropriate breaks if stress is detected."
[0790] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0791] Step 1:
[0792] The terminal uses cameras and microphones installed at the work site to acquire video and audio information of workers in real time. This information is output as raw video data (e.g., video stream) and audio data (e.g., audio stream). Next, this data is converted into a digital format and sent to the server.
[0793] Step 2:
[0794] The server processes the received video information and uses OpenCV and Dlib libraries to identify facial expressions. Specifically, it extracts facial features from the video data and classifies expressions such as smiles and surprise. The input is digital video data, and the output is category information of the identified facial expressions.
[0795] Step 3:
[0796] The server converts the audio information into text using the Google Cloud Speech-to-Text API and then performs spectral analysis of the audio. It identifies features related to the pitch and intensity of the speech and measures the pitch of the tone. The input for this step is digital audio data, and the output is the converted text information and speech tone features.
[0797] Step 4:
[0798] The server integrates the identified facial expression category information and voice tone characteristics, and uses an estimation algorithm to estimate the worker's emotional state. For example, a smiling expression and a high but gentle tone are likely to indicate "satisfied." The input for this step is the facial expression and voice tone characteristics, and the output is the estimated emotional state.
[0799] Step 5:
[0800] The server assesses the worker's stress level and fatigue based on their estimated emotional state, and provides suggestions for breaks and work improvements as needed. This involves using a generative AI model to create appropriate advice. The input for this step is the result of the emotional state assessment, and the output is specific suggestions for work improvements and breaks.
[0801] Step 6:
[0802] The terminal displays suggestions received from the server to the worker in real time. The user reviews the suggestions and adjusts their actions as needed. The input for this step is suggestion information from the server, and the output is feedback messages provided to the worker.
[0803] In this way, the system can continuously monitor the health status of workers and provide appropriate feedback in real time.
[0804] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0805] This invention provides a highly accurate system for recognizing user emotions. The system incorporates an emotion engine into a server to perform emotion recognition, and estimates the user's emotional state in real time based on video and audio data.
[0806] The device captures the user's current video and audio through its camera and microphone. This data is then immediately transmitted to a server via the internet. The server analyzes the received video data to extract the user's facial features and simultaneously processes the audio data using spectral analysis to identify the tone of voice.
[0807] The emotion engine embedded in the server estimates emotions based on extracted facial features and voice tone. This emotion engine utilizes machine learning algorithms and can improve estimation accuracy by referencing emotion patterns based on the user's past data. For example, if a user has previously displayed certain facial expressions and voice patterns in specific situations, the engine can identify emotional states such as "stress" or "relaxation" based on that data.
[0808] Furthermore, the emotion engine can monitor the user's emotional changes in real time and determine the appropriate timing for intervention based on the analysis results. For example, if a trend of the user gradually becoming depressed is detected, the server will send an early notification to the caregiver to encourage proactive action.
[0809] This system can also effectively provide users with the information they need by converting audio data into text and summarizing key points. The server achieves this function by converting audio data into text using automatic speech recognition and summarizing it using natural language processing.
[0810] Thus, the emotion recognition system of the present invention establishes a highly accurate understanding of the user's emotions, enabling high-quality communication in caregiving settings and providing responses that meet the user's needs.
[0811] The following describes the processing flow.
[0812] Step 1:
[0813] The device uses a camera and microphone to acquire the user's video and audio data in real time.
[0814] Step 2:
[0815] The terminal transmits the acquired video and audio data to the server as digital signals.
[0816] Step 3:
[0817] The server analyzes the received video data using an image processing algorithm to extract the user's facial features.
[0818] Step 4:
[0819] The server analyzes the received audio data using an audio processing algorithm to identify the tone and pitch of the voice.
[0820] Step 5:
[0821] The emotion engine embedded in the server estimates the user's emotions based on extracted facial features and voice tone. This emotion engine uses machine learning models to improve its estimation accuracy.
[0822] Step 6:
[0823] The server analyzes the estimated emotions and monitors the trend of emotional changes in real time.
[0824] Step 7:
[0825] Based on trends in emotional changes, the server determines the appropriate timing for intervention for the user and sends notifications to caregivers as needed.
[0826] Step 8:
[0827] The server uses speech recognition technology to convert the audio data into text data.
[0828] Step 9:
[0829] The server analyzes the converted text data using natural language processing techniques, extracts the key points of the conversation, and generates a summary.
[0830] Step 10:
[0831] The server sends summary information to the terminal and presents it to the user.
[0832] Through this process, the system can recognize the user's emotions with high accuracy, enabling a quick and appropriate response.
[0833] (Example 2)
[0834] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0835] Conventional systems struggled to accurately recognize users' emotions and respond appropriately based on them. Furthermore, they were inadequate at grasping emotional changes in real time through the analysis of voice and video data, and at determining the timing of necessary interventions. As a result, they were unable to respond in a way that aligned with users' emotions and needs, making it difficult to provide satisfactory support, particularly in areas such as caregiving and communication.
[0836] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0837] In this invention, the server includes means for acquiring video information, means for acquiring audio information, and means for analyzing the acquired video information to identify facial features. This enables high-precision estimation of the user's emotional state in real time and allows for quick and accurate responses as needed.
[0838] "Visual information" refers to visual data acquired by cameras and other imaging devices.
[0839] "Acoustic information" refers to audio and sound data acquired by sound-collecting devices such as microphones.
[0840] "Device means" refers to a mechanical or electrical device designed to perform a specific function.
[0841] "Analysis" refers to the process of examining data in detail and identifying specific characteristics or trends.
[0842] "Facial features" refer to a person's expression and the appearance of their face, and are used as elements that indicate emotions and reactions.
[0843] "Voice tone" refers to characteristics such as the tone, pitch, and intensity of a voice that reflect the speaker's emotions and intentions.
[0844] "To estimate" refers to predicting or judging unconfirmed matters or situations based on obtained data.
[0845] "Real-time" refers to data processing and system responses occurring immediately, meaning timely execution without delay.
[0846] The "timing of intervention" refers to the point at which specific actions or responses are required, indicating the point that leads to effective support and resolution.
[0847] This invention provides a system that recognizes a user's emotions in real time with high accuracy. The entire system mainly consists of a terminal and a server, and evaluates emotions by utilizing the user's video and audio information.
[0848] The device uses a camera and microphone to record the user's current status. The camera captures video information from the user's face, and the microphone captures audio information from the user's voice. This captured information is transmitted to a server via a communication network.
[0849] The server analyzes the transmitted video information and uses image processing algorithms to identify the user's facial features. It also employs acoustic analysis software to identify the tone of voice by performing spectral analysis on acoustic information. The server incorporates a generative AI model that estimates the user's emotions based on these analysis results. This generative AI model utilizes machine learning algorithms and improves estimation accuracy by referencing emotion patterns based on past data.
[0850] This system also has the function of monitoring emotional changes in real time and determining the appropriate timing for intervention. For example, if a user shows signs of rapidly becoming stressed, the system will notify the caregiver of this information and instruct them on appropriate measures. It also converts acoustic information into text and provides feedback to the user by summarizing the key points.
[0851] For example, if the system determines that a user may be feeling down, it will compare this with past data and recommend relaxing music for that user. Furthermore, using a generative AI model, an example of a prompt message could be: "Estimate the user's emotions in real time. The video data shows a smiling expression, and the audio data indicates a calm tone of voice."
[0852] Thus, the system of the present invention is designed with the aim of deeply understanding the user's emotions and providing effective responses accordingly.
[0853] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0854] Step 1:
[0855] The device acquires the user's video information via its camera. The input is the user's real-time visual situation, which is output as high-resolution digital video data. This video data includes the user's facial expressions, facial features, and movements. The device captures accurate video while maintaining a stable frame rate.
[0856] Step 2:
[0857] The device acquires the user's acoustic information via a microphone. The input is the user's voice, which is output as acoustic sample data. This data captures the tone, pitch, speed, and emotional nuances of the voice in detail. The device uses filters to reduce background noise and ensure clear sound quality.
[0858] Step 3:
[0859] The terminal transmits the acquired video and audio information to the server via the internet. The input is the output data from steps 1 and 2, and this data is transmitted encrypted. The terminal uses a communication protocol to enable secure and rapid transmission of data.
[0860] Step 4:
[0861] The server analyzes the received video information and uses machine learning algorithms to identify facial features. The input is encrypted video data, and the output is elements extracted as identified facial features and expressions. The server identifies key features of the expressions and matches them against known expression patterns registered in the database.
[0862] Step 5:
[0863] The server performs spectral analysis on the received acoustic information to identify the tone of the voice. The input is acoustic sample data, and the output is the analyzed voice tone and emotional characteristics. The server analyzes the frequency components of the voice waveform to identify tones related to the user's emotional state.
[0864] Step 6:
[0865] The server uses a generative AI model to estimate the user's emotions from the identified facial features and tone of voice. The input is the output data from steps 4 and 5, and the output is the estimated emotional state. The server compares the estimation results with past data to improve accuracy and recognize the user's current emotions with high precision.
[0866] Step 7:
[0867] The server monitors estimated emotional changes in real time and determines the appropriate timing for intervention as needed. The input is continuously estimated emotional data, and the output is alerts and notifications for intervention. The server performs real-time analysis based on certain criteria and prompts necessary actions.
[0868] (Application Example 2)
[0869] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0870] In modern society, it is crucial to understand an individual's emotional state in real time and proactively detect potential dangers and stressors. However, conventional technologies lack the means to instantly grasp emotional changes and provide appropriate responses, making it difficult to create a safe and comfortable environment.
[0871] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0872] In this invention, the server includes means for acquiring video information, means for acquiring audio information, means for analyzing the acquired video information to identify facial expressions, means for analyzing the acquired audio information to identify voice tone, means for estimating emotions based on the identified facial expressions and voice tone, and means for determining situations requiring safety measures based on the estimated emotions and sending notifications to the individual to encourage relaxation. This enables real-time emotion monitoring tailored to individual situations and appropriate interventions to ensure safety.
[0873] "Video information" refers to visual data, including the user's facial expressions and movements, and is digital data acquired via a camera device.
[0874] "Acoustic information" refers to auditory data, including the user's voice and ambient sounds, and is digital data acquired using a microphone device.
[0875] "Means for identifying facial expressions" refers to a system component that analyzes facial features from acquired video information and, based on that, identifies facial expressions that indicate the user's emotions.
[0876] "Means for identifying the tone of voice" refers to a component of a system that analyzes the tone and intonation of a voice from acquired acoustic information and identifies the characteristics of a voice that indicate emotion based on that analysis.
[0877] "Means for estimating emotions" refers to a system component that utilizes machine learning algorithms, based on identified facial expression and voice tone data, to infer the user's current emotional state.
[0878] "Means for determining situations requiring safety assurance" refers to a system component that has the function of determining whether or not a user's psychological state requires vigilance, based on estimated emotions.
[0879] "Means of sending notifications to encourage relaxation" refers to system components that have the functionality to send messages to users urging them to relieve stress or rest.
[0880] The specific system for realizing this invention analyzes video and audio data, estimates the user's emotional state in real time, and intervenes to improve safety. Processing is primarily carried out using smart devices and servers.
[0881] The cameras and microphones built into smart devices (such as smart glasses and smartphones) capture the user's video and audio information. This data is transmitted to a server via the network.
[0882] On the server, video information is analyzed using the OpenCV library to identify the user's facial expressions. Furthermore, acoustic information is analyzed using the LibROSA library to identify the tone of the voice. The resulting facial and voice feature data is then input into an emotion recognition model using TensorFlow to estimate the user's emotions.
[0883] Based on the emotion assessment, the server determines whether a situation requiring safety measures is present. Specifically, if the system determines that the user is experiencing stress, it sends a notification to the user encouraging relaxation based on that information. This notification is displayed on the user's smart device to prompt timely intervention.
[0884] For example, if a user on public transport feels anxious, the server can instantly recognize this and send a message to the user's device such as, "Take a deep breath to relax." This helps the user instantly recognize their own state and manage themselves.
[0885] An example of a prompt would be: "Generate a scenario in which an emotion monitoring security app monitors a user's emotional state in real time and provides notifications to enhance safety. Consider the specific situation of anxiety experienced while using public transportation."
[0886] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0887] Step 1:
[0888] The device uses a camera and microphone to acquire video and audio information of the user. The input data consists of video data captured by the camera and audio data collected by the microphone. This data is then prepared for transmission to a server via the network.
[0889] Step 2:
[0890] The server analyzes the received video information using the OpenCV library. The input is video data transmitted from the terminal, and the output is feature data representing the user's facial expressions. Specifically, it calculates the movement and changes of each part of the face and quantifies the facial expression features.
[0891] Step 3:
[0892] The server analyzes the received acoustic information using the LibROSA library. The input is audio data transmitted from the terminal, and the output is feature data indicating the tone of the voice. Specifically, it performs spectral analysis of the audio waveform and extracts features such as pitch, intonation, and tempo.
[0893] Step 4:
[0894] The server inputs the obtained facial expression feature data and voice tone feature data into a generative AI model using TensorFlow to estimate emotions. The output is data indicating the user's emotional state. Specifically, the model uses pattern recognition based on past data to determine whether the emotion corresponds to "stress" or "relaxation."
[0895] Step 5:
[0896] The server generates prompt messages to determine if the estimated emotional state requires safety measures. The input is emotional state data, and the output is a notification message to the user. It generates messages such as "Take a deep breath to relax" to assess the need for intervention.
[0897] Step 6:
[0898] The server sends a notification to the terminal to encourage the user to relax if it determines that security measures are necessary. The input is the generated notification message, and the output is the relaxation notification displayed on the terminal screen. This allows the user to take necessary actions in a timely manner.
[0899] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0900] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0901] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0902] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0903] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0904] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0905] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0906] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0907] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0908] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0909] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0910] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0911] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0912] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0913] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0914] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0915] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0916] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0917] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0918] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0919] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0920] The following is further disclosed regarding the embodiments described above.
[0921] (Claim 1)
[0922] A means of acquiring video data,
[0923] A means of acquiring audio data,
[0924] An identification means that analyzes acquired video data to identify facial expressions,
[0925] An identification means that analyzes acquired audio data to identify the tone of voice,
[0926] An estimation means for estimating emotions based on identified facial expressions and tone of voice,
[0927] A system that includes a decision-making mechanism to determine situations requiring action based on estimated emotions.
[0928] (Claim 2)
[0929] A conversion method for converting audio data into text data,
[0930] An extraction method for analyzing the converted text data and extracting important points,
[0931] The system according to claim 1, comprising a generation means for generating a summary based on the extracted key points.
[0932] (Claim 3)
[0933] A means of understanding the context by analyzing the text data of conversations,
[0934] A predictive means that predicts the next expected question or topic based on the understood context,
[0935] The system according to claim 1, comprising a generating means for generating appropriate responses to predicted questions.
[0936] "Example 1"
[0937] (Claim 1)
[0938] A means of acquiring video information,
[0939] A means for acquiring audio information,
[0940] An identification means that analyzes acquired video information to identify facial expressions,
[0941] An identification means that analyzes acquired audio information to identify the tone of voice,
[0942] An estimation means for estimating emotions based on identified facial expressions and tone of voice,
[0943] A means of determining situations that require action based on estimated emotions,
[0944] A conversion means for converting audio information into text information,
[0945] An extraction means for analyzing the converted character information and extracting important elements,
[0946] A generation means for generating a summary based on the extracted key elements,
[0947] A means of understanding the context by analyzing the textual information of a conversation,
[0948] A means of inferring the next expected question or topic based on the understood context,
[0949] A generation means for generating an appropriate response to a hypothesized question,
[0950] A generation means that generates personalized responses by utilizing past history based on the individual information of each subject,
[0951] A system that includes management mechanisms for recording and securely sharing information collected during caregiving.
[0952] (Claim 2)
[0953] The system according to claim 1, comprising means for analyzing audio information and extracting its tone and frequency components.
[0954] (Claim 3)
[0955] The system according to claim 1, comprising means for estimating emotions from facial expressions and tone of voice using a generative AI model.
[0956] "Application Example 1"
[0957] (Claim 1)
[0958] A device for acquiring video information,
[0959] A device for acquiring voice information,
[0960] A device that analyzes acquired video information to identify facial expressions,
[0961] A device that analyzes acquired audio information to identify the characteristics of the audio,
[0962] A device that estimates emotions based on identified facial and vocal characteristics,
[0963] A device that judges a situation based on estimated emotions,
[0964] A device that monitors the worker's condition and detects stress and fatigue,
[0965] A system that includes a device that suggests work improvements or breaks based on detection results.
[0966] (Claim 2)
[0967] A device that converts audio information into text information,
[0968] A device that analyzes converted character information and extracts important elements,
[0969] The system according to claim 1, comprising a device for generating a summary based on extracted key elements.
[0970] (Claim 3)
[0971] A device that analyzes the textual information of conversations to understand the context,
[0972] A device that predicts the next expected topic based on the understood context,
[0973] The system according to claim 1, comprising a device for generating appropriate responses to predicted topics.
[0974] "Example 2 of combining an emotion engine"
[0975] (Claim 1)
[0976] A device for acquiring video information,
[0977] A device for acquiring acoustic information,
[0978] A means for analyzing acquired video information to identify facial features,
[0979] A means for analyzing acquired acoustic information to identify the tone of voice,
[0980] A device for estimating emotions based on identified facial features and tone of voice,
[0981] A means to monitor estimated emotional changes in real time and determine the appropriate timing for intervention,
[0982] A system that includes means for determining situations requiring action based on estimated emotions.
[0983] (Claim 2)
[0984] A device and means for converting acoustic information into textual information,
[0985] A device and means for analyzing converted character information and extracting important information,
[0986] The system according to claim 1, comprising a device means for generating a summary based on extracted important information.
[0987] (Claim 3)
[0988] A device and means for analyzing textual information in communications to understand the situation,
[0989] A device that predicts the next expected question or topic based on the understood situation,
[0990] The system according to claim 1, comprising means for a device that generates an appropriate response to a predicted question.
[0991] "Application example 2 when combining with an emotional engine"
[0992] (Claim 1)
[0993] Means for acquiring video information,
[0994] Means for acquiring acoustic information,
[0995] A means for analyzing acquired video information to identify facial expressions,
[0996] A means for analyzing acquired acoustic information to identify the tone of the voice,
[0997] A means for estimating emotions based on identified facial expressions and vocal tone,
[0998] A system that determines situations requiring safety measures based on estimated emotions and includes means of sending notifications to individuals to encourage relaxation.
[0999] (Claim 2)
[1000] A means of converting acoustic information into textual information,
[1001] A means of analyzing the converted text information and extracting important points,
[1002] The system according to claim 1, comprising means for generating a summary based on the extracted key points.
[1003] (Claim 3)
[1004] A means of understanding the situation by analyzing the textual information of the conversation,
[1005] A means of predicting the next expected question or topic based on the context understood,
[1006] The system according to claim 1, comprising means for generating appropriate responses to predicted questions. [Explanation of Symbols]
[1007] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring video data, A means of acquiring audio data, An identification means that analyzes acquired video data to identify facial expressions, An identification means that analyzes acquired audio data to identify the tone of voice, An estimation means for estimating emotions based on identified facial expressions and tone of voice, A system that includes a decision-making mechanism to determine situations requiring action based on estimated emotions.
2. A conversion method for converting audio data into text data, An extraction method for analyzing the converted text data and extracting important points, The system according to claim 1, comprising a generation means for generating a summary based on the extracted key points.
3. A means of understanding the context by analyzing the text data of conversations, A predictive means that predicts the next expected question or topic based on the understood context, The system according to claim 1, comprising a generation means for generating appropriate responses to predicted questions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A