system
The system addresses loneliness and emergency response issues for the elderly by generating virtual characters for interaction and detecting anomalies, ensuring timely support through AI-driven conversation analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
The increasing elderly population faces loneliness and difficulty in receiving immediate assistance during emergencies, with existing technologies failing to provide both natural conversation and early detection of abnormalities effectively.
A system that includes a generation means for creating virtual characters, an analysis means for detecting anomalies in conversations, and a notification means for alerting emergency contacts, utilizing AI models to generate video and audio of virtual persons and analyze user interactions in real-time.
The system alleviates feelings of loneliness and enables rapid responses to emergencies by providing natural interactions and early anomaly detection, ensuring the elderly receive timely assistance.
Smart Images

Figure 2026070254000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] With the increase in the elderly population, problems such as the loneliness they face and the lack of prompt response in case of emergencies are becoming more serious. In particular, elderly people living alone not only feel lonely in their daily lives but also have difficulty receiving immediate assistance in case of emergencies such as sudden changes in physical condition. Solving this problem is also important socially.
Means for Solving the Problems
[0005] The present invention provides a system that includes a generation means for generating video and audio of a virtual person using image data, an analysis means for analyzing the content of conversations with the user and detecting anomalies, and a notification means for notifying pre-set emergency contacts when an anomaly is detected. This allows the user to communicate virtually with family members, and in the event of an emergency, anomalies are automatically detected and notifications are sent to the appropriate contacts, enabling a rapid response.
[0006] "Image data" refers to a collection of visual information used to generate video and audio of a virtual character.
[0007] A "virtual character" is a digital character that resembles a real person, created in real time from image data using a generation method.
[0008] A "generation means" is a component that has the function of generating video and audio of a virtual person using image data.
[0009] A "user" is an individual who operates the system and communicates with virtual characters.
[0010] "Dialogue content" refers to the content and format of the conversation that takes place between the user and the virtual character.
[0011] An "analysis means" is a component that has the function of analyzing the content of the interaction with the user and detecting anomalies.
[0012] An "abnormality" refers to an event that deviates from normal conversation patterns or user behavior, and specifically those that require immediate attention.
[0013] A "notification device" is a component that has the function of notifying pre-configured emergency contacts of the occurrence of an anomaly when an anomaly is detected.
[0014] "Emergency contact information" refers to the contact details of a third party that the user has pre-configured to be notified in the event of an anomaly being detected. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention provides an AI appliance system that alleviates feelings of loneliness among the elderly and enables rapid response in the event of an emergency. This system primarily consists of the interaction between a server, a terminal, and a user. The roles of each component and specific examples are explained below.
[0037] Server Role
[0038] The server is the core of the entire system and is primarily responsible for updating AI models and analyzing data. The server maintains models that generate virtual individuals resembling the user's family members using the latest DeepFake technology, and regularly updates these models with training data. It also analyzes user interactions in real time and uses language processing techniques to detect unnatural conversation patterns and anomalies.
[0039] Terminal role
[0040] The device is a digital photo frame that displays a virtual person using model data transmitted from a server. It recognizes voice input and actions from the user and responds appropriately based on that. By providing voice and visual interaction with the virtual person, the user can feel as if they are always connected to their family. If an abnormal situation is detected, the device quickly sends data to the server, and an emergency contact is made promptly.
[0041] User roles
[0042] Users are the primary users of the system and interact with the device mainly through voice. Users can engage in natural conversations with virtual characters and receive lifestyle advice and reminders from the system. If there are any abnormalities regarding the user's health or activity, the system will make a judgment based on the user's statements and actions and notify them as necessary.
[0043] Examples
[0044] As a concrete example, a user can send a request to their device saying, "I want to talk to my grandchild today." The device analyzes this request, uses a generative model received from the server to generate a virtual person resembling the grandchild, and begins a conversation with the user. While the conversation is ongoing, the server monitors the content and sends an alert to an emergency contact if it detects any unusual patterns.
[0045] This system will improve the quality of life for the elderly and enable a swift and appropriate response in emergencies.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The server generates the latest AI model and updates the data to create virtual people resembling the user's family using Deep Fake technology. The updated model is sent to the device in real time.
[0049] Step 2:
[0050] The terminal installs the generative model received from the server and prepares to display the virtual character. It waits for the user to input what they want to control the terminal by voice.
[0051] Step 3:
[0052] The user speaks to the device and gives a voice command saying, "I want to talk to my grandchild." This voice command is received and recognized by the device.
[0053] Step 4:
[0054] The terminal uses a generative model downloaded from the server based on user commands to begin generating a virtual person resembling the grandchild. During this process, both audio and video are synthesized in real time to provide the user with a conversational environment.
[0055] Step 5:
[0056] The server monitors the interaction between the user and the virtual character in real time. It uses natural language processing to analyze the conversation content and determine whether any anomalies have been detected.
[0057] Step 6:
[0058] If the server detects any unnatural patterns or anomalies in the interaction, it will send an alert to a pre-configured emergency contact. The notification will include a statement that an anomaly may have occurred.
[0059] Step 7:
[0060] After the conversation ends, the device records the conversation history and usage, and sends the data to the server as needed. This data will be used for future model updates and feature improvements.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] There is a need for technologies that can alleviate feelings of loneliness among the elderly and enable rapid responses in emergencies. However, conventional technologies struggle to achieve both natural conversation and early detection of abnormalities, and in particular, the ability to balance flexible conversation through voice interfaces with abnormality detection has not been sufficiently achieved. Solving this problem is necessary to provide a safer and more fulfilling living environment.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes an input means for inputting instructions using voice, a generation means for generating video and audio of a virtual person using generated information, an analysis means for analyzing dialogue information and detecting anomalies, and a notification means for notifying pre-set contacts when an anomaly is detected. This allows elderly people to gain a sense of security through natural dialogue in their daily lives, and enables a quick and appropriate response in emergencies.
[0066] "Sound" is a form of sound, a means by which people communicate information through language.
[0067] "Instructions" refer to commands or requests given to obtain a specific action or result.
[0068] "Input means" refers to devices or methods for importing data or information into a system.
[0069] "Generated information" refers to a collection of data and models used to create the video and audio of a virtual character.
[0070] A "virtual character" refers to a person who does not actually exist but is simulated by a computer.
[0071] "Generation means" refers to devices and methods for creating new images and sounds using data and models.
[0072] "Dialogue information" refers to the content of conversations and communications exchanged between the user and the system.
[0073] "Analytical means" refers to devices and methods for analyzing data and information and extracting useful patterns and insights from them.
[0074] An "abnormality" refers to a state or pattern that is different from the norm and requires attention or action.
[0075] "Notification means" refers to devices or methods for transmitting specific information or messages to other devices or people.
[0076] "Contact information" refers to the contact details of an individual or organization used to communicate in emergencies or for conveying specific information.
[0077] This invention is a virtual dialogue system designed to alleviate feelings of loneliness among the elderly and enable rapid response in emergency situations. It primarily functions by utilizing the interaction between a server, a terminal, and a user.
[0078] Server Role
[0079] The server plays a central role in the entire system, managing AI models and analyzing data. Specifically, it maintains and updates models that generate virtual individuals resembling the user's relatives using DeepFake technology. To this end, it continuously incorporates training data using generative AI models and analyzes user interactions in real time. The server skillfully detects unnatural patterns and anomalies from conversations using natural language processing techniques.
[0080] Terminal role
[0081] The device is a digital photo frame that projects a virtual person onto its screen using model data provided by a server. Users can interact with the virtual person using voice through the device. The device is equipped with voice recognition capabilities, allowing it to recognize and process voice input from the user appropriately. Through conversations with the virtual person, users can experience the feeling of interacting with their family. In addition, if an anomaly is detected, data is quickly sent to the server, and an emergency notification is promptly issued.
[0082] User roles
[0083] Users are the primary users of this system and interact with their devices mainly using voice. Users can enjoy conversations with virtual characters and receive lifestyle advice and reminders from the system. Based on their speech and actions, the system will notify users as needed if it detects any abnormalities in their health or activity.
[0084] Specific example
[0085] For example, a user could request, "I want to talk to my grandchild today." In this case, the device analyzes the request, generates a virtual person from the server, and displays someone who resembles the grandchild. Then, the conversation begins. The server continuously monitors the conversation and immediately sends an alert to an emergency contact if any abnormal patterns are detected.
[0086] Example of a prompt
[0087] A concrete example of a prompt message is input such as, "Please tell me the appropriate procedure for generating a virtual character when a user requests to speak with their grandchild."
[0088] This system reduces feelings of loneliness for the elderly in their daily lives and ensures a swift and appropriate response in emergencies.
[0089] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0090] Step 1:
[0091] The user inputs a request by voice into the device. The user tells the device a specific request in natural language, such as "I want to talk to my grandchild today." This voice input triggers the start of processing.
[0092] Step 2:
[0093] The device receives voice input from the user. Speech recognition software is used to convert the voice data into text. During this process, the device captures the user's voice through the microphone, analyzes it using a speech recognition engine, and outputs it as text data.
[0094] Step 3:
[0095] The terminal parses the transcribed user request and sends a request to the server to generate a virtual person. The terminal extracts key phrases from the text, constructs a prompt, and then sends it to the server as a data packet. This data contains details about the user's request.
[0096] Step 4:
[0097] The server generates a virtual person using a generative AI model based on the request data received from the terminal. The server references stored family data and uses Deep Fake technology to synthesize the virtual person's video and audio. In this process, the model dynamically calculates the data and generates the virtual person's video and audio in real time.
[0098] Step 5:
[0099] The server sends the generated virtual character data to the terminal. The server outputs data containing the virtual character's video and audio to the terminal. Upon receiving this data, the terminal displays it to the user using a display device.
[0100] Step 6:
[0101] The device displays a virtual person and initiates a conversation with the user. The device displays the generated image on the screen and outputs audio through its speaker. This interaction gives the user the feeling of interacting with a family member.
[0102] Step 7:
[0103] The server monitors the conversation between the user and the virtual character and detects anomalies. The server utilizes natural language processing technology to analyze the text data during the conversation, checking for abnormal patterns and changes in emotion.
[0104] Step 8:
[0105] If an anomaly is detected, the server sends an alert to pre-configured contacts. For example, it can send an email or SMS to family members registered as emergency contacts, enabling a quick response. This allows users to receive immediate assistance in emergencies.
[0106] (Application Example 1)
[0107] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0108] Modern seniors face problems such as social isolation and a lack of support in daily life. Furthermore, there is a demand for improved customer experiences and personalized services in physical stores. To address these challenges, an interactive system tailored to individual needs is necessary, but the technology to realize this is still insufficient. This invention aims to provide a system that reduces feelings of loneliness among seniors, improves their safety and quality of life, and simultaneously offers personalized customer experiences in physical stores.
[0109] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0110] In this invention, the server includes a generation means for generating video and audio of a virtual person using image data, an analysis means for analyzing the content of conversations with the user and detecting anomalies, a notification means for notifying pre-set emergency contacts when an anomaly is detected, a recommendation means for making personalized product recommendations based on the user's attribute information, and a display means for presenting virtual objects through visual information. This makes it possible to reduce feelings of loneliness and ensure safety for the elderly, while simultaneously providing a personalized customer experience in physical stores.
[0111] "Image data" refers to a collection of electronically stored visual information, which is used to generate images of virtual characters.
[0112] A "virtual character" is a character that looks like a human being, created using computer generation technology, and is intended to engage in voice and visual interaction with the user.
[0113] "Generation means" refers to a technical device for creating video and audio of a virtual person using image data and other information.
[0114] "Analysis means" refers to a technical device used to analyze the content of conversations with users and detect anomalies or specific patterns.
[0115] "Anomaly" refers to a specific event or situation that is judged to deviate from normal conversation or behavioral patterns.
[0116] A "notification device" is a technical device for transmitting information to a pre-configured emergency contact or other receiving device when an anomaly is detected.
[0117] A "recommendation tool" is a technical device that individually presents appropriate products and information based on the user's attribute information and past behavioral history.
[0118] A "display means" is a device that visually presents virtual objects and related information, and facilitates interaction with the user.
[0119] This invention is a system aimed at improving the quality of life for the elderly and providing personalized customer experiences in physical stores. The configuration of this system is described in detail below.
[0120] The server is responsible for the core functions of the entire system. Using a generative AI model, it analyzes pre-stored image data to generate video and audio of virtual characters. For example, if a user prompts "Tell me today's weather," the server uses natural language processing to analyze the request and construct a virtual character to respond. The server performs processing using cloud services such as Google Cloud Platform and Microsoft Azure.
[0121] The terminal serves as the interface closest to the user. This terminal consists of devices such as smart glasses and head-mounted displays, and displays a virtual person based on a generative model sent from the server. It also analyzes the user's voice commands and gestures using Google Cloud Speech-to-Text and Microsoft Cognitive Services, and provides appropriate feedback.
[0122] For example, if a customer using smart glasses in a physical store says, "Tell me what products would suit these," the device analyzes the voice in real time, and a virtual person recommends products based on information from the server. At this time, the recommendations are personalized based on the customer's past purchase information and pre-registered preferences.
[0123] Users are central to the interaction with this system in their daily lives. Elderly individuals can alleviate feelings of loneliness through daily conversations with virtual characters. For example, if a user asks, "Tell me my exercise record for today," the system will provide the relevant information and offer health advice based on it.
[0124] This system is operated using a high-performance cloud infrastructure and advanced AI algorithms to maintain response speed and reliability. Examples of prompts include: "What are this month's recommended products?" and "I want to check my weekend plans."
[0125] The embodiments of this invention aim to improve the user experience by leveraging personalized and interactive features.
[0126] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0127] Step 1:
[0128] The user enters a voice command into the device. The device captures this audio and uses the Google Cloud Speech-to-Text API to convert the audio data into text. The input is the user's voice command, and the output is the command in text format.
[0129] Step 2:
[0130] The server receives text commands sent from the terminal and performs natural language processing. Specifically, it uses GPT-3(registered trademark) .5 or similar AI models to analyze the user's intent and generate a response from a virtual character. The input is a user command in text format, and the output is the generated virtual character's response text.
[0131] Step 3:
[0132] The server generates video and audio of a virtual character based on the generated response. It uses Unity or a similar platform with a GPU to render the video and synthesize the speech. The input is the virtual character's response text, and the output is the virtual character's video and synthesized speech.
[0133] Step 4:
[0134] The terminal receives video and audio of a virtual person transmitted from the server and presents it to the user using smart glasses or a display device. The input is video and audio data of the virtual person, and the output is a visual and auditory presentation to the user.
[0135] Step 5:
[0136] The user interacts with a virtual person through the device and receives information and product recommendations in response to prompts. In this step, personalized recommendations are presented based on the user's past purchase history and attribute information. Input is the user's prompts and feedback, and output is customized information and product recommendations.
[0137] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0138] This invention provides an AI appliance system that incorporates an emotion engine to recognize the user's emotions, enabling more responsive conversations, thereby reducing loneliness in the daily lives of the elderly and providing notifications in case of abnormalities.
[0139] Server Role
[0140] The server regularly updates its AI model, including the emotion engine algorithm. It utilizes the latest speech recognition and facial expression analysis technologies to extract and analyze emotions from the user's voice and facial expressions in real time. It also has the functionality to detect anomalies based on the analyzed data and notify emergency contacts as needed.
[0141] Terminal role
[0142] The device is equipped with a camera and microphone, allowing it to detect the user's facial expressions and voice. The device sends this data to a server, where an emotion engine analyzes it and adjusts the conversation based on the results. When the user interacts with the system, the device displays a generated virtual persona and changes the tone and content of its responses according to the user's emotional state, providing a more natural and less stressful conversational experience.
[0143] User roles
[0144] The user faces a digital photo frame and begins interacting with it as usual. In addition to daily reminders and questions, the user talks about how they feel that day, allowing the system to recognize their emotional state. For example, if the user sounds tired, the system will respond by suggesting they rest. If an anomaly is detected, the system will automatically contact the necessary people to alleviate the user's burden.
[0145] Examples
[0146] As a concrete example, suppose a user says to their device, "I'm feeling a little down." The device sends their voice and facial expression to the server, where an emotion engine recognizes the emotion of "sadness." The server analyzes this emotional state and generates a response that includes a more comforting and encouraging tone than a normal conversation. Furthermore, if the emotion of "sadness" is recognized frequently, the server can perceive this persistent emotional change as an anomaly and send an alert to emergency contacts.
[0147] This invention enables faster problem resolution while providing a more interactive and responsive user experience.
[0148] The following describes the processing flow.
[0149] Step 1:
[0150] The user speaks into the device, saying, "I'm feeling a little down today." This input is received by the device via the microphone.
[0151] Step 2:
[0152] The device analyzes the acquired audio data and captures the user's facial expressions with its camera. The audio and facial expression data are transmitted to the server in real time.
[0153] Step 3:
[0154] The server uses an emotion engine to analyze the user's emotional state from the transmitted data. It recognizes emotions such as "sadness" from voice tone and facial expressions.
[0155] Step 4:
[0156] The server generates an appropriate response based on the recognized emotional state. In this example, it prepares dialogue that encourages the user or suggests they take a rest.
[0157] Step 5:
[0158] The terminal uses the response sent from the server to generate a virtual character and proceed with the interaction with the user. This includes speech synthesis and video display.
[0159] Step 6:
[0160] The server continuously monitors the user's emotional state even during interactions. If the emotion of "sadness" is detected frequently over a certain period, it can be identified as an anomaly.
[0161] Step 7:
[0162] If an anomaly is detected, the server will notify pre-configured emergency contacts. This notification will include information about changes in the user's emotional state.
[0163] Step 8:
[0164] After the conversation ends, the terminal saves all conversation data and sentiment analysis results as logs and backs them up to the server for use in future analyses and improvements to the AI model.
[0165] (Example 2)
[0166] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0167] This system aims to alleviate feelings of loneliness among the elderly, provide emotional support in daily life, quickly detect emotional abnormalities, and ensure they receive necessary assistance. In particular, it strives to improve the quality of life for the elderly by recognizing emotional changes in real time and responding appropriately through dialogue.
[0168] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0169] In this invention, the server includes generation means for generating visual and auditory information of a virtual person using voice data and image data; analysis means for analyzing the user's voice and facial expressions and identifying their emotional state; dialogue generation means for generating responses based on the emotion analysis; and notification means for notifying pre-configured emergency contacts when emotional changes or persistent emotional abnormalities are detected. This enables real-time understanding of the user's emotions, the provision of dialogues that alleviate feelings of loneliness, and rapid support in the event of an abnormality.
[0170] "Voice data" refers to information recorded in digital format from the user's speech, and is used for analysis.
[0171] "Image data" refers to visual information that digitally records the user's facial expressions and posture, and is data used for analysis.
[0172] A "virtual character" is an artificial human figure generated by a computer, which interacts with the user through visual and auditory information.
[0173] "Generation means" refers to a device or method for generating visual and auditory information of a virtual person using audio data and image data.
[0174] "Analysis means" refers to a device or method that identifies an emotional state by analyzing the user's voice and facial expressions.
[0175] "Dialogue generation means" refers to a device or method that generates an appropriate response based on the results of emotion analysis and conveys it to the user through a virtual character.
[0176] "Notification means" refers to a device or method that notifies pre-set emergency contacts when it detects changes in emotions or abnormal emotions.
[0177] "Emotional state" refers to the state of the user's internal emotional expression, including the type and intensity of emotions identified through analysis.
[0178] An "emergency contact" refers to an external contact designated to receive notifications in the event of an emergency involving the user.
[0179] This invention utilizes an AI system equipped with an emotion analysis engine to alleviate feelings of loneliness in the user's daily life, detect emotional abnormalities, and provide necessary support.
[0180] The server is a computer device that runs the emotion analysis engine and plays a key role. This server receives audio and image data and uses speech recognition software and facial expression analysis software to analyze them, respectively. Specifically, a general speech processing API is used for speech recognition, and an image processing library is used for facial expression analysis. This identifies the user's emotional state. By utilizing an AI model, a natural language response that matches the emotional situation is generated.
[0181] The terminal is a device equipped with a camera and microphone, and its primary role is to capture the user's voice and facial expressions and transmit them to the server. Furthermore, based on the analysis results received from the server, the terminal displays a virtual person's image and interacts with the user visually and audibly. The terminal provides a natural conversational experience by changing the tone and content of the virtual person's responses according to the user's emotional state.
[0182] Users engage in everyday conversations with their devices. They can talk about their feelings for the day when using reminders or seeking advice. For example, by uttering a prompt such as "I'm feeling a little down," the system instantly analyzes their emotions and offers words of comfort and encouragement. If the change in emotions is deemed abnormal, the server automatically notifies emergency contacts.
[0183] This invention aims to reduce feelings of loneliness in the user's living environment and enable prompt support responses through emotional monitoring.
[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0185] Step 1:
[0186] The device acquires the user's voice and facial expression data.
[0187] Specifically, the system uses a camera to capture an image of the user's face and a microphone to record their voice.
[0188] The audio and image data used as input are sent to the server in real time for the next processing step.
[0189] Step 2:
[0190] The server converts the received audio data into text data.
[0191] This process uses speech recognition software to analyze the audio signal and extract it as text information.
[0192] The input is audio data, and the output is text data.
[0193] Step 3:
[0194] The server analyzes image data to recognize the user's emotions.
[0195] This process uses a facial expression analysis algorithm to analyze facial features in an image and identify the emotional state.
[0196] The input is image data, and the output is data indicating emotional state.
[0197] Step 4:
[0198] The server generates an appropriate response based on the analyzed text data and emotional state.
[0199] We utilize generative AI models to create natural language responses based on textual and emotional information.
[0200] The input consists of text data and sentiment data, and the output is a natural language response.
[0201] Step 5:
[0202] The terminal receives a response from the server and communicates the response to the user using a virtual persona.
[0203] Specifically, the system displays a virtual character on the screen and plays back a response generated using speech synthesis software.
[0204] The input is a natural language response sentence, and the output is the response as visual and auditory information.
[0205] Step 6:
[0206] The server continuously monitors emotional data and, if it detects an emotional anomaly, it notifies the designated emergency contact.
[0207] This process analyzes patterns of emotional changes and sends an alert using the notification system if an abnormality is detected.
[0208] The input is continuously collected sentiment data, and the output is the sending of emergency notifications.
[0209] (Application Example 2)
[0210] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0211] In the daily lives of the elderly, there is a need for effective systems that reduce feelings of isolation and can respond quickly to emotional changes and abnormal situations. In particular, a challenge is the inability to detect situations where the elderly require emotional support or where abnormalities are not detected early, potentially jeopardizing their safety. To solve these problems, a system is needed that accurately grasps the user's emotions and enables adaptive responses and rapid action in the event of an abnormality.
[0212] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0213] In this invention, the server includes generation means for generating visual and auditory information of a virtual entity using image data and audio data; analysis means for analyzing the content of conversations with the user and detecting emotions and anomalies; notification means for notifying pre-set emergency contacts and specific locations when an anomaly is detected or when emotions change; and response generation means for providing adaptive responses according to pre-set emotional states. This enables accurate understanding of the user's emotional state, reduces feelings of isolation, and allows for rapid response to anomalies.
[0214] "Image data" refers to data that represents visual information in a digital format.
[0215] "Audio data" refers to data that represents auditory information in a digital format.
[0216] A "virtual entity" is an artificial character or person created using digital technology and represented as visual and auditory information on a computer screen or similar device.
[0217] "Visual information" refers to information obtained through human vision, and is expressed in the form of images, videos, and other visual media.
[0218] "Auditory information" refers to information obtained through human hearing, expressed in the form of sounds, music, and other similar media.
[0219] "Generative means" refers to a technical configuration or device for producing a specific function or result.
[0220] "Analysis means" refers to a technical configuration or device for analyzing input data and extracting specific patterns or information.
[0221] "Detecting emotions" involves analyzing information such as the user's facial expressions and voice to identify the emotions that are expressed.
[0222] "Detecting an anomaly" means detecting a situation or condition that is different from the normal state.
[0223] "Emotional change" refers to the change in the user's emotional state over time.
[0224] "Notification means" refers to a technical configuration or function for informing a third party of specific information.
[0225] "Response generation means" refers to a technical configuration or function for automatically creating an appropriate response based on the user's input and circumstances.
[0226] "Emergency contact information" refers to information indicating how to contact individuals or organizations in the event of an emergency.
[0227] An "adaptive response" is a response that is adjusted according to the user's condition and circumstances.
[0228] To realize this invention, the system consists of a client terminal and a server. The terminal is equipped with a camera and microphone, which can sense the user's facial expressions and voice in real time. The sensed data is sent to the server, which analyzes the data using a generative AI model to identify the user's emotional state.
[0229] Specifically, the server uses OpenCV and dlib as face recognition libraries, and leverages machine learning frameworks such as TENSORFLOW® for speech analysis. The server combines these tools to perform data calculations that extract emotions from facial expressions and speech. Based on the results of the emotion analysis, it generates adaptive responses from pre-configured response patterns. These responses are returned to the user in either voice or text format.
[0230] Furthermore, if a significant change or abnormality is detected in the user's emotional state, the server automatically sends a notification to pre-registered emergency contacts. This helps to ensure the user's safety while also reducing feelings of isolation.
[0231] For example, if a user says, "I'm not feeling very cheerful today," the system will detect "sadness" from the content of their statement and the tone of their voice. The emotion engine will analyze this data and generate a response intended to comfort or support the user. For example, it might say something like, "Why don't you take a short break?"
[0232] An example of a prompt using a generative AI model is, "Please give us some ideas on what to say to comfort a user who appears sad." This prompt can be used to improve the accuracy of response generation.
[0233] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0234] Step 1:
[0235] The device collects user facial expressions and voice data through its camera and microphone. It senses what the user says to the device and their facial expressions in real time and prepares to transmit that data. The input consists of facial image data and voice data, and the output is data to be transmitted to the server.
[0236] Step 2:
[0237] The server receives facial image data and audio data transmitted from the terminal. The server prepares this data for analysis and configures it as a dataset for facial expression analysis and audio analysis. The input is the raw data received from the terminal, and the output is in a data format ready for analysis.
[0238] Step 3:
[0239] The server analyzes facial expression data using OpenCV and dlib, and simultaneously analyzes audio data using TensorFlow to detect emotions. A generative AI model for emotion detection identifies emotions from facial characteristics and voice tone, and obtains analysis results based on this. The input is pre-prepared data for analysis, and the output is the detected emotion information.
[0240] Step 4:
[0241] The server generates responses based on detected emotion information. It uses an AI model to generate appropriate responses according to the emotional state, determining the content of the response to the user. It generates response candidates using prompt text and selects the optimal one. The output is the generated voice or text response.
[0242] Step 5:
[0243] The server sends the generated response to the terminal. The terminal then relays this response to the user, either by providing an audio response through the speaker or by displaying it as text on the screen. The input is the generated response data, and the output is the specific response presented to the user.
[0244] Step 6:
[0245] The server immediately sends a notification to emergency contacts if the detected emotions indicate a significant change or anomaly. Here, it extracts the necessary information based on the anomaly detection and sends an alert to the registered contacts. The output is an emergency contact message.
[0246] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0247] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0248] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0249] [Second Embodiment]
[0250] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0251] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0252] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0253] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0254] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0255] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0256] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0257] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0258] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0259] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0260] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0261] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0262] This invention provides an AI appliance system that alleviates feelings of loneliness among the elderly and enables rapid response in the event of an emergency. This system primarily consists of the interaction between a server, a terminal, and a user. The roles of each component and specific examples are explained below.
[0263] Server Role
[0264] The server is the core of the entire system and is primarily responsible for updating AI models and analyzing data. The server maintains models that generate virtual individuals resembling the user's family members using the latest DeepFake technology, and regularly updates these models with training data. It also analyzes user interactions in real time and uses language processing techniques to detect unnatural conversation patterns and anomalies.
[0265] Terminal role
[0266] The device is a digital photo frame that displays a virtual person using model data transmitted from a server. It recognizes voice input and actions from the user and responds appropriately based on that. By providing voice and visual interaction with the virtual person, the user can feel as if they are always connected to their family. If an abnormal situation is detected, the device quickly sends data to the server, and an emergency contact is made promptly.
[0267] User roles
[0268] Users are the primary users of the system and interact with the device mainly through voice. Users can engage in natural conversations with virtual characters and receive lifestyle advice and reminders from the system. If there are any abnormalities regarding the user's health or activity, the system will make a judgment based on the user's statements and actions and notify them as necessary.
[0269] Examples
[0270] As a concrete example, a user can send a request to their device saying, "I want to talk to my grandchild today." The device analyzes this request, uses a generative model received from the server to generate a virtual person resembling the grandchild, and begins a conversation with the user. While the conversation is ongoing, the server monitors the content and sends an alert to an emergency contact if it detects any unusual patterns.
[0271] This system will improve the quality of life for the elderly and enable a swift and appropriate response in emergencies.
[0272] The following describes the processing flow.
[0273] Step 1:
[0274] The server generates the latest AI model and updates the data to create virtual people resembling the user's family using Deep Fake technology. The updated model is sent to the device in real time.
[0275] Step 2:
[0276] The terminal installs the generative model received from the server and prepares to display the virtual character. It waits for the user to input what they want to control the terminal by voice.
[0277] Step 3:
[0278] The user speaks to the device and gives a voice command saying, "I want to talk to my grandchild." This voice command is received and recognized by the device.
[0279] Step 4:
[0280] The terminal uses a generative model downloaded from the server based on user commands to begin generating a virtual person resembling the grandchild. During this process, both audio and video are synthesized in real time to provide the user with a conversational environment.
[0281] Step 5:
[0282] The server monitors the interaction between the user and the virtual character in real time. It uses natural language processing to analyze the conversation content and determine whether any anomalies have been detected.
[0283] Step 6:
[0284] If the server detects unnatural patterns or abnormalities in the conversation, it sends an alert to a pre-set emergency contact. The notification includes the possibility that an abnormality may have occurred.
[0285] Step 7:
[0286] After the conversation ends, the terminal records the conversation history and usage status, and sends the data to the server as needed. This data is used for future model updates and function improvements.
[0287] (Example 1)
[0288] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] There is a need for technology that can reduce the loneliness of the elderly and enable rapid response in an emergency. However, with conventional technologies, it is difficult to detect abnormalities early while realizing natural conversation, and in particular, the coexistence of flexible conversation and abnormality detection through a voice interface has not been fully achieved. Solving this problem is necessary to provide a more secure and rich living environment.
[0290] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0291] In this invention, the server includes an input means for inputting instructions using voice, a generation means for generating the video and voice of a virtual character using the generated information, an analysis means for analyzing the conversation information and detecting abnormalities, and a notification means for notifying a pre-set contact when an abnormality is detected. As a result, the elderly can obtain a sense of security through natural conversation in their daily lives, and can respond quickly and appropriately in an emergency.
[0292] "Voice" is a form of sound and a means for people to transmit information through language.
[0293] "Instructions" refer to commands or requests given to obtain a specific action or result.
[0294] "Input means" refers to devices or methods for importing data or information into a system.
[0295] "Generated information" refers to a collection of data and models used to create the video and audio of a virtual character.
[0296] A "virtual character" refers to a person who does not actually exist but is simulated by a computer.
[0297] "Generation means" refers to devices and methods for creating new images and sounds using data and models.
[0298] "Dialogue information" refers to the content of conversations and communications exchanged between the user and the system.
[0299] "Analytical means" refers to devices and methods for analyzing data and information and extracting useful patterns and insights from them.
[0300] An "abnormality" refers to a state or pattern that is different from the norm and requires attention or action.
[0301] "Notification means" refers to devices or methods for transmitting specific information or messages to other devices or people.
[0302] "Contact information" refers to the contact details of an individual or organization used to communicate in emergencies or for conveying specific information.
[0303] This invention is a virtual dialogue system designed to alleviate feelings of loneliness among the elderly and enable rapid response in emergency situations. It primarily functions by utilizing the interaction between a server, a terminal, and a user.
[0304] Server Role
[0305] The server plays a central role in the entire system, managing the AI model and analyzing data. Specifically, it maintains and updates a model for generating virtual characters similar to the user's relatives using Deep Fake technology. For this purpose, it continuously incorporates learning data using a generative AI model and analyzes the interaction with the user in real time. The server cleverly detects unnatural patterns and anomalies in conversations by using natural language processing technology.
[0306] Role of the terminal
[0307] The terminal is a photo-frame type device that projects virtual characters onto the screen using the model data provided by the server. The user can interact with the virtual characters using voice through the terminal. The terminal is equipped with a voice recognition function and can recognize voice input from the user and perform appropriate processing. Through interaction with the virtual characters, the user can experience the feeling of being in contact with family. Also, when an anomaly is detected, data is quickly sent to the server and a prompt emergency notification is issued.
[0308] Role of the user
[0309] The user is the main user of this system and mainly communicates with the terminal using voice. While enjoying the interaction with the virtual characters, the user can receive advice and reminders regarding life from the system. If the system senses an anomaly in the user's health condition or activities based on their speech or actions, a notification will be issued as necessary.
[0310] Specific example
[0311] For example, the user can request "I want to talk to my grandson today". In this case, the terminal analyzes the request, generates a virtual character from the server, and displays a person similar to the grandson. Then, the conversation starts. The server continuously monitors the content during the conversation and sends an alert to the emergency contact immediately if there are abnormal patterns.
[0312] Example of prompt text
[0313] A concrete example of a prompt message is input such as, "Please tell me the appropriate procedure for generating a virtual character when a user requests to speak with their grandchild."
[0314] This system reduces feelings of loneliness for the elderly in their daily lives and ensures a swift and appropriate response in emergencies.
[0315] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0316] Step 1:
[0317] The user inputs a request by voice into the device. The user tells the device a specific request in natural language, such as "I want to talk to my grandchild today." This voice input triggers the start of processing.
[0318] Step 2:
[0319] The device receives voice input from the user. Speech recognition software is used to convert the voice data into text. During this process, the device captures the user's voice through the microphone, analyzes it using a speech recognition engine, and outputs it as text data.
[0320] Step 3:
[0321] The terminal parses the transcribed user request and sends a request to the server to generate a virtual person. The terminal extracts key phrases from the text, constructs a prompt, and then sends it to the server as a data packet. This data contains details about the user's request.
[0322] Step 4:
[0323] The server generates a virtual person using a generative AI model based on the request data received from the terminal. The server references stored family data and uses Deep Fake technology to synthesize the virtual person's video and audio. In this process, the model dynamically calculates the data and generates the virtual person's video and audio in real time.
[0324] Step 5:
[0325] The server sends the generated virtual character data to the terminal. The server outputs data containing the virtual character's video and audio to the terminal. Upon receiving this data, the terminal displays it to the user using a display device.
[0326] Step 6:
[0327] The device displays a virtual person and initiates a conversation with the user. The device displays the generated image on the screen and outputs audio through its speaker. This interaction gives the user the feeling of interacting with a family member.
[0328] Step 7:
[0329] The server monitors the conversation between the user and the virtual character and detects anomalies. The server utilizes natural language processing technology to analyze the text data during the conversation, checking for abnormal patterns and changes in emotion.
[0330] Step 8:
[0331] If an anomaly is detected, the server sends an alert to pre-configured contacts. For example, it can send an email or SMS to family members registered as emergency contacts, enabling a quick response. This allows users to receive immediate assistance in emergencies.
[0332] (Application Example 1)
[0333] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0334] Modern seniors face problems such as social isolation and a lack of support in daily life. Furthermore, there is a demand for improved customer experiences and personalized services in physical stores. To address these challenges, an interactive system tailored to individual needs is necessary, but the technology to realize this is still insufficient. This invention aims to provide a system that reduces feelings of loneliness among seniors, improves their safety and quality of life, and simultaneously offers personalized customer experiences in physical stores.
[0335] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0336] In this invention, the server includes a generation means for generating video and audio of a virtual person using image data, an analysis means for analyzing the content of conversations with the user and detecting anomalies, a notification means for notifying pre-set emergency contacts when an anomaly is detected, a recommendation means for making personalized product recommendations based on the user's attribute information, and a display means for presenting virtual objects through visual information. This makes it possible to reduce feelings of loneliness and ensure safety for the elderly, while simultaneously providing a personalized customer experience in physical stores.
[0337] "Image data" refers to a collection of electronically stored visual information, which is used to generate images of virtual characters.
[0338] A "virtual character" is a character that looks like a human being, created using computer generation technology, and is intended to engage in voice and visual interaction with the user.
[0339] "Generation means" refers to a technical device for creating video and audio of a virtual person using image data and other information.
[0340] "Analysis means" refers to a technical device used to analyze the content of conversations with users and detect anomalies or specific patterns.
[0341] "Anomaly" refers to a specific event or situation that is judged to deviate from normal conversation or behavioral patterns.
[0342] A "notification device" is a technical device for transmitting information to a pre-configured emergency contact or other receiving device when an anomaly is detected.
[0343] A "recommendation tool" is a technical device that individually presents appropriate products and information based on the user's attribute information and past behavioral history.
[0344] A "display means" is a device that visually presents virtual objects and related information, and facilitates interaction with the user.
[0345] This invention is a system aimed at improving the quality of life for the elderly and providing personalized customer experiences in physical stores. The configuration of this system is described in detail below.
[0346] The server is responsible for the core functions of the entire system. Using generative AI models, it analyzes pre-stored image data to generate video and audio of virtual characters. For example, if a user prompts "Tell me today's weather," the server uses natural language processing to analyze the request and construct a virtual character to respond. The server performs processing using cloud services such as Google Cloud Platform and Microsoft Azure.
[0347] The terminal serves as the interface closest to the user. This terminal consists of devices such as smart glasses and head-mounted displays, and displays a virtual person based on a generative model sent from the server. It also analyzes the user's voice commands and gestures using Google Cloud Speech-to-Text and Microsoft Cognitive Services, and provides appropriate feedback.
[0348] For example, if a customer using smart glasses in a physical store says, "Tell me what products would suit these," the device analyzes the voice in real time, and a virtual person recommends products based on information from the server. At this time, the recommendations are personalized based on the customer's past purchase information and pre-registered preferences.
[0349] Users are central to the interaction with this system in their daily lives. Elderly individuals can alleviate feelings of loneliness through daily conversations with virtual characters. For example, if a user asks, "Tell me my exercise record for today," the system will provide the relevant information and offer health advice based on it.
[0350] This system is operated using a high-performance cloud infrastructure and advanced AI algorithms to maintain response speed and reliability. Examples of prompts include: "What are this month's recommended products?" and "I want to check my weekend plans."
[0351] The embodiments of this invention aim to improve the user experience by leveraging personalized and interactive features.
[0352] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0353] Step 1:
[0354] The user enters a voice command into the device. The device captures this audio and uses the Google Cloud Speech-to-Text API to convert the audio data into text. The input is the user's voice command, and the output is the command in text format.
[0355] Step 2:
[0356] The server receives text commands sent from the terminal and performs natural language processing. Specifically, it uses GPT-3.5 or similar AI models to analyze the user's intent and generate a response from a virtual character. The input is a user command in text format, and the output is the generated virtual character's response text.
[0357] Step 3:
[0358] The server generates video and audio of a virtual character based on the generated response. It uses Unity or a similar platform with a GPU to render the video and synthesize the speech. The input is the virtual character's response text, and the output is the virtual character's video and synthesized speech.
[0359] Step 4:
[0360] The terminal receives video and audio of a virtual person transmitted from the server and presents it to the user using smart glasses or a display device. The input is video and audio data of the virtual person, and the output is a visual and auditory presentation to the user.
[0361] Step 5:
[0362] The user interacts with a virtual person through the device and receives information and product recommendations in response to prompts. In this step, personalized recommendations are presented based on the user's past purchase history and attribute information. Input is the user's prompts and feedback, and output is customized information and product recommendations.
[0363] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0364] This invention provides an AI appliance system that incorporates an emotion engine to recognize the user's emotions, enabling more responsive conversations, thereby reducing loneliness in the daily lives of the elderly and providing notifications in case of abnormalities.
[0365] Server Role
[0366] The server regularly updates its AI model, including the emotion engine algorithm. It utilizes the latest speech recognition and facial expression analysis technologies to extract and analyze emotions from the user's voice and facial expressions in real time. It also has the functionality to detect anomalies based on the analyzed data and notify emergency contacts as needed.
[0367] Terminal role
[0368] The device is equipped with a camera and microphone, allowing it to detect the user's facial expressions and voice. The device sends this data to a server, where an emotion engine analyzes it and adjusts the conversation based on the results. When the user interacts with the system, the device displays a generated virtual persona and changes the tone and content of its responses according to the user's emotional state, providing a more natural and less stressful conversational experience.
[0369] User roles
[0370] The user faces a digital photo frame and begins interacting with it as usual. In addition to daily reminders and questions, the user talks about how they feel that day, allowing the system to recognize their emotional state. For example, if the user sounds tired, the system will respond by suggesting they rest. If an anomaly is detected, the system will automatically contact the necessary people to alleviate the user's burden.
[0371] Examples
[0372] As a concrete example, suppose a user says to their device, "I'm feeling a little down." The device sends their voice and facial expression to the server, where an emotion engine recognizes the emotion of "sadness." The server analyzes this emotional state and generates a response that includes a more comforting and encouraging tone than a normal conversation. Furthermore, if the emotion of "sadness" is recognized frequently, the server can perceive this persistent emotional change as an anomaly and send an alert to emergency contacts.
[0373] This invention enables faster problem resolution while providing a more interactive and responsive user experience.
[0374] The following describes the processing flow.
[0375] Step 1:
[0376] The user speaks into the device, saying, "I'm feeling a little down today." This input is received by the device via the microphone.
[0377] Step 2:
[0378] The device analyzes the acquired audio data and captures the user's facial expressions with its camera. The audio and facial expression data are transmitted to the server in real time.
[0379] Step 3:
[0380] The server uses an emotion engine to analyze the user's emotional state from the transmitted data. It recognizes emotions such as "sadness" from voice tone and facial expressions.
[0381] Step 4:
[0382] The server generates an appropriate response based on the recognized emotional state. In this example, it prepares dialogue that encourages the user or suggests they take a rest.
[0383] Step 5:
[0384] The terminal uses the response sent from the server to generate a virtual character and proceed with the interaction with the user. This includes speech synthesis and video display.
[0385] Step 6:
[0386] The server continuously monitors the user's emotional state even during interactions. If the emotion of "sadness" is detected frequently over a certain period, it can be identified as an anomaly.
[0387] Step 7:
[0388] If an anomaly is detected, the server will notify pre-configured emergency contacts. This notification will include information about changes in the user's emotional state.
[0389] Step 8:
[0390] After the conversation ends, the terminal saves all conversation data and sentiment analysis results as logs and backs them up to the server for use in future analyses and improvements to the AI model.
[0391] (Example 2)
[0392] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0393] This system aims to alleviate feelings of loneliness among the elderly, provide emotional support in daily life, quickly detect emotional abnormalities, and ensure they receive necessary assistance. In particular, it strives to improve the quality of life for the elderly by recognizing emotional changes in real time and responding appropriately through dialogue.
[0394] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0395] In this invention, the server includes generation means for generating visual and auditory information of a virtual person using voice data and image data; analysis means for analyzing the user's voice and facial expressions and identifying their emotional state; dialogue generation means for generating responses based on the emotion analysis; and notification means for notifying pre-configured emergency contacts when emotional changes or persistent emotional abnormalities are detected. This enables real-time understanding of the user's emotions, the provision of dialogues that alleviate feelings of loneliness, and rapid support in the event of an abnormality.
[0396] "Voice data" refers to information recorded in digital format from the user's speech, and is used for analysis.
[0397] "Image data" refers to visual information that digitally records the user's facial expressions and posture, and is data used for analysis.
[0398] A "virtual character" is an artificial human figure generated by a computer, which interacts with the user through visual and auditory information.
[0399] "Generation means" refers to a device or method for generating visual and auditory information of a virtual person using audio data and image data.
[0400] "Analysis means" refers to a device or method that identifies an emotional state by analyzing the user's voice and facial expressions.
[0401] "Dialogue generation means" refers to a device or method that generates an appropriate response based on the results of emotion analysis and conveys it to the user through a virtual character.
[0402] "Notification means" refers to a device or method that notifies pre-set emergency contacts when it detects changes in emotions or abnormal emotions.
[0403] "Emotional state" refers to the state of the user's internal emotional expression, including the type and intensity of emotions identified through analysis.
[0404] An "emergency contact" refers to an external contact designated to receive notifications in the event of an emergency involving the user.
[0405] This invention utilizes an AI system equipped with an emotion analysis engine to alleviate feelings of loneliness in the user's daily life, detect emotional abnormalities, and provide necessary support.
[0406] The server is a computer device that runs the emotion analysis engine and plays a key role. This server receives audio and image data and uses speech recognition software and facial expression analysis software to analyze them, respectively. Specifically, a general speech processing API is used for speech recognition, and an image processing library is used for facial expression analysis. This identifies the user's emotional state. By utilizing an AI model, a natural language response that matches the emotional situation is generated.
[0407] The terminal is a device equipped with a camera and microphone, and its primary role is to capture the user's voice and facial expressions and transmit them to the server. Furthermore, based on the analysis results received from the server, the terminal displays a virtual person's image and interacts with the user visually and audibly. The terminal provides a natural conversational experience by changing the tone and content of the virtual person's responses according to the user's emotional state.
[0408] Users engage in everyday conversations with their devices. They can talk about their feelings for the day when using reminders or seeking advice. For example, by uttering a prompt such as "I'm feeling a little down," the system instantly analyzes their emotions and offers words of comfort and encouragement. If the change in emotions is deemed abnormal, the server automatically notifies emergency contacts.
[0409] This invention aims to reduce feelings of loneliness in the user's living environment and enable prompt support responses through emotional monitoring.
[0410] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0411] Step 1:
[0412] The device acquires the user's voice and facial expression data.
[0413] Specifically, the system uses a camera to capture an image of the user's face and a microphone to record their voice.
[0414] The audio and image data used as input are sent to the server in real time for the next processing step.
[0415] Step 2:
[0416] The server converts the received audio data into text data.
[0417] This process uses speech recognition software to analyze the audio signal and extract it as text information.
[0418] The input is audio data, and the output is text data.
[0419] Step 3:
[0420] The server analyzes image data to recognize the user's emotions.
[0421] This process uses a facial expression analysis algorithm to analyze facial features in an image and identify the emotional state.
[0422] The input is image data, and the output is data indicating emotional state.
[0423] Step 4:
[0424] The server generates an appropriate response based on the analyzed text data and emotional state.
[0425] We utilize generative AI models to create natural language responses based on textual and emotional information.
[0426] The input consists of text data and sentiment data, and the output is a natural language response.
[0427] Step 5:
[0428] The terminal receives a response from the server and communicates the response to the user using a virtual persona.
[0429] Specifically, the system displays a virtual character on the screen and plays back a response generated using speech synthesis software.
[0430] The input is a natural language response sentence, and the output is the response as visual and auditory information.
[0431] Step 6:
[0432] The server continuously monitors emotional data and, if it detects an emotional anomaly, it notifies the designated emergency contact.
[0433] This process analyzes patterns of emotional changes and sends an alert using the notification system if an abnormality is detected.
[0434] The input is continuously collected sentiment data, and the output is the sending of emergency notifications.
[0435] (Application Example 2)
[0436] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0437] In the daily lives of the elderly, there is a need for effective systems that reduce feelings of isolation and can respond quickly to emotional changes and abnormal situations. In particular, a challenge is the inability to detect situations where the elderly require emotional support or where abnormalities are not detected early, potentially jeopardizing their safety. To solve these problems, a system is needed that accurately grasps the user's emotions and enables adaptive responses and rapid action in the event of an abnormality.
[0438] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0439] In this invention, the server includes generation means for generating visual and auditory information of a virtual entity using image data and audio data; analysis means for analyzing the content of conversations with the user and detecting emotions and anomalies; notification means for notifying pre-set emergency contacts and specific locations when an anomaly is detected or when emotions change; and response generation means for providing adaptive responses according to pre-set emotional states. This enables accurate understanding of the user's emotional state, reduces feelings of isolation, and allows for rapid response to anomalies.
[0440] "Image data" refers to data that represents visual information in a digital format.
[0441] "Audio data" refers to data that represents auditory information in a digital format.
[0442] A "virtual entity" is an artificial character or person created using digital technology and represented as visual and auditory information on a computer screen or similar device.
[0443] "Visual information" refers to information obtained through human vision, and is expressed in the form of images, videos, and other visual media.
[0444] "Auditory information" refers to information obtained through human hearing, expressed in the form of sounds, music, and other similar media.
[0445] "Generative means" refers to a technical configuration or device for producing a specific function or result.
[0446] "Analysis means" refers to a technical configuration or device for analyzing input data and extracting specific patterns or information.
[0447] "Detecting emotions" involves analyzing information such as the user's facial expressions and voice to identify the emotions that are expressed.
[0448] "Detecting an anomaly" means detecting a situation or condition that is different from the normal state.
[0449] "Emotional change" refers to the change in the user's emotional state over time.
[0450] "Notification means" refers to a technical configuration or function for informing a third party of specific information.
[0451] "Response generation means" refers to a technical configuration or function for automatically creating an appropriate response based on the user's input and circumstances.
[0452] "Emergency contact information" refers to information indicating how to contact individuals or organizations in the event of an emergency.
[0453] An "adaptive response" is a response that is adjusted according to the user's condition and circumstances.
[0454] To realize this invention, the system consists of a client terminal and a server. The terminal is equipped with a camera and microphone, which can sense the user's facial expressions and voice in real time. The sensed data is sent to the server, which analyzes the data using a generative AI model to identify the user's emotional state.
[0455] Specifically, the server uses OpenCV and dlib as face recognition libraries, and leverages machine learning frameworks such as TensorFlow for speech analysis. The server combines these tools to perform data calculations that extract emotions from facial expressions and speech. Based on the results of the emotion analysis, it generates adaptive responses from pre-configured response patterns. These responses are returned to the user in either voice or text format.
[0456] Furthermore, if a significant change or abnormality is detected in the user's emotional state, the server automatically sends a notification to pre-registered emergency contacts. This helps to ensure the user's safety while also reducing feelings of isolation.
[0457] For example, if a user says, "I'm not feeling very cheerful today," the system will detect "sadness" from the content of their statement and the tone of their voice. The emotion engine will analyze this data and generate a response intended to comfort or support the user. For example, it might say something like, "Why don't you take a short break?"
[0458] An example of a prompt using a generative AI model is, "Please give us some ideas on what to say to comfort a user who appears sad." This prompt can be used to improve the accuracy of response generation.
[0459] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0460] Step 1:
[0461] The device collects user facial expressions and voice data through its camera and microphone. It senses what the user says to the device and their facial expressions in real time and prepares to transmit that data. The input consists of facial image data and voice data, and the output is data to be transmitted to the server.
[0462] Step 2:
[0463] The server receives facial image data and audio data transmitted from the terminal. The server prepares this data for analysis and configures it as a dataset for facial expression analysis and audio analysis. The input is the raw data received from the terminal, and the output is in a data format ready for analysis.
[0464] Step 3:
[0465] The server analyzes facial expression data using OpenCV and dlib, and simultaneously analyzes audio data using TensorFlow to detect emotions. A generative AI model for emotion detection identifies emotions from facial characteristics and voice tone, and obtains analysis results based on this. The input is pre-prepared data for analysis, and the output is the detected emotion information.
[0466] Step 4:
[0467] The server generates responses based on detected emotion information. It uses an AI model to generate appropriate responses according to the emotional state, determining the content of the response to the user. It generates response candidates using prompt text and selects the optimal one. The output is the generated voice or text response.
[0468] Step 5:
[0469] The server sends the generated response to the terminal. The terminal then relays this response to the user, either by providing an audio response through the speaker or by displaying it as text on the screen. The input is the generated response data, and the output is the specific response presented to the user.
[0470] Step 6:
[0471] The server immediately sends a notification to emergency contacts if the detected emotions indicate a significant change or anomaly. Here, it extracts the necessary information based on the anomaly detection and sends an alert to the registered contacts. The output is an emergency contact message.
[0472] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0473] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0474] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0475] [Third Embodiment]
[0476] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0477] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0478] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0479] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0480] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0481] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0482] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0483] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0484] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0485] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0486] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0487] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0488] This invention provides an AI appliance system that alleviates feelings of loneliness among the elderly and enables rapid response in the event of an emergency. This system primarily consists of the interaction between a server, a terminal, and a user. The roles of each component and specific examples are explained below.
[0489] Server Role
[0490] The server is the core of the entire system and is primarily responsible for updating AI models and analyzing data. The server maintains models that generate virtual individuals resembling the user's family members using the latest DeepFake technology, and regularly updates these models with training data. It also analyzes user interactions in real time and uses language processing techniques to detect unnatural conversation patterns and anomalies.
[0491] Terminal role
[0492] The device is a digital photo frame that displays a virtual person using model data transmitted from a server. It recognizes voice input and actions from the user and responds appropriately based on that. By providing voice and visual interaction with the virtual person, the user can feel as if they are always connected to their family. If an abnormal situation is detected, the device quickly sends data to the server, and an emergency contact is made promptly.
[0493] User roles
[0494] Users are the primary users of the system and interact with the device mainly through voice. Users can engage in natural conversations with virtual characters and receive lifestyle advice and reminders from the system. If there are any abnormalities regarding the user's health or activity, the system will make a judgment based on the user's statements and actions and notify them as necessary.
[0495] Examples
[0496] As a concrete example, a user can send a request to their device saying, "I want to talk to my grandchild today." The device analyzes this request, uses a generative model received from the server to generate a virtual person resembling the grandchild, and begins a conversation with the user. While the conversation is ongoing, the server monitors the content and sends an alert to an emergency contact if it detects any unusual patterns.
[0497] This system will improve the quality of life for the elderly and enable a swift and appropriate response in emergencies.
[0498] The following describes the processing flow.
[0499] Step 1:
[0500] The server generates the latest AI model and updates the data to create virtual people resembling the user's family using Deep Fake technology. The updated model is sent to the device in real time.
[0501] Step 2:
[0502] The terminal installs the generative model received from the server and prepares to display the virtual character. It waits for the user to input what they want to control the terminal by voice.
[0503] Step 3:
[0504] The user speaks to the device and gives a voice command saying, "I want to talk to my grandchild." This voice command is received and recognized by the device.
[0505] Step 4:
[0506] The terminal uses a generative model downloaded from the server based on user commands to begin generating a virtual person resembling the grandchild. During this process, both audio and video are synthesized in real time to provide the user with a conversational environment.
[0507] Step 5:
[0508] The server monitors the interaction between the user and the virtual character in real time. It uses natural language processing to analyze the conversation content and determine whether any anomalies have been detected.
[0509] Step 6:
[0510] If the server detects any unnatural patterns or anomalies in the interaction, it will send an alert to a pre-configured emergency contact. The notification will include a statement that an anomaly may have occurred.
[0511] Step 7:
[0512] After the conversation ends, the device records the conversation history and usage, and sends the data to the server as needed. This data will be used for future model updates and feature improvements.
[0513] (Example 1)
[0514] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0515] There is a need for technologies that can alleviate feelings of loneliness among the elderly and enable rapid responses in emergencies. However, conventional technologies struggle to achieve both natural conversation and early detection of abnormalities, and in particular, the ability to balance flexible conversation through voice interfaces with abnormality detection has not been sufficiently achieved. Solving this problem is necessary to provide a safer and more fulfilling living environment.
[0516] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0517] In this invention, the server includes an input means for inputting instructions using voice, a generation means for generating video and audio of a virtual person using generated information, an analysis means for analyzing dialogue information and detecting anomalies, and a notification means for notifying pre-set contacts when an anomaly is detected. This allows elderly people to gain a sense of security through natural dialogue in their daily lives, and enables a quick and appropriate response in emergencies.
[0518] "Sound" is a form of sound, a means by which people communicate information through language.
[0519] "Instructions" refer to commands or requests given to obtain a specific action or result.
[0520] "Input means" refers to devices or methods for importing data or information into a system.
[0521] "Generated information" refers to a collection of data and models used to create the video and audio of a virtual character.
[0522] A "virtual character" refers to a person who does not actually exist but is simulated by a computer.
[0523] "Generation means" refers to devices and methods for creating new images and sounds using data and models.
[0524] "Dialogue information" refers to the content of conversations and communications exchanged between the user and the system.
[0525] "Analytical means" refers to devices and methods for analyzing data and information and extracting useful patterns and insights from them.
[0526] An "abnormality" refers to a state or pattern that is different from the norm and requires attention or action.
[0527] "Notification means" refers to devices or methods for transmitting specific information or messages to other devices or people.
[0528] "Contact information" refers to the contact details of an individual or organization used to communicate in emergencies or for conveying specific information.
[0529] This invention is a virtual dialogue system designed to alleviate feelings of loneliness among the elderly and enable rapid response in emergency situations. It primarily functions by utilizing the interaction between a server, a terminal, and a user.
[0530] Server Role
[0531] The server plays a central role in the entire system, managing AI models and analyzing data. Specifically, it maintains and updates models that generate virtual individuals resembling the user's relatives using DeepFake technology. To this end, it continuously incorporates training data using generative AI models and analyzes user interactions in real time. The server skillfully detects unnatural patterns and anomalies from conversations using natural language processing techniques.
[0532] Terminal role
[0533] The device is a digital photo frame that projects a virtual person onto its screen using model data provided by a server. Users can interact with the virtual person using voice through the device. The device is equipped with voice recognition capabilities, allowing it to recognize and process voice input from the user appropriately. Through conversations with the virtual person, users can experience the feeling of interacting with their family. In addition, if an anomaly is detected, data is quickly sent to the server, and an emergency notification is promptly issued.
[0534] User roles
[0535] Users are the primary users of this system and interact with their devices mainly using voice. Users can enjoy conversations with virtual characters and receive lifestyle advice and reminders from the system. Based on their speech and actions, the system will notify users as needed if it detects any abnormalities in their health or activity.
[0536] Specific example
[0537] For example, a user could request, "I want to talk to my grandchild today." In this case, the device analyzes the request, generates a virtual person from the server, and displays someone who resembles the grandchild. Then, the conversation begins. The server continuously monitors the conversation and immediately sends an alert to an emergency contact if any abnormal patterns are detected.
[0538] Example of a prompt
[0539] A concrete example of a prompt message is input such as, "Please tell me the appropriate procedure for generating a virtual character when a user requests to speak with their grandchild."
[0540] This system reduces feelings of loneliness for the elderly in their daily lives and ensures a swift and appropriate response in emergencies.
[0541] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0542] Step 1:
[0543] The user inputs a request by voice into the device. The user tells the device a specific request in natural language, such as "I want to talk to my grandchild today." This voice input triggers the start of processing.
[0544] Step 2:
[0545] The device receives voice input from the user. Speech recognition software is used to convert the voice data into text. During this process, the device captures the user's voice through the microphone, analyzes it using a speech recognition engine, and outputs it as text data.
[0546] Step 3:
[0547] The terminal parses the transcribed user request and sends a request to the server to generate a virtual person. The terminal extracts key phrases from the text, constructs a prompt, and then sends it to the server as a data packet. This data contains details about the user's request.
[0548] Step 4:
[0549] The server generates a virtual person using a generative AI model based on the request data received from the terminal. The server references stored family data and uses Deep Fake technology to synthesize the virtual person's video and audio. In this process, the model dynamically calculates the data and generates the virtual person's video and audio in real time.
[0550] Step 5:
[0551] The server sends the generated virtual character data to the terminal. The server outputs data containing the virtual character's video and audio to the terminal. Upon receiving this data, the terminal displays it to the user using a display device.
[0552] Step 6:
[0553] The device displays a virtual person and initiates a conversation with the user. The device displays the generated image on the screen and outputs audio through its speaker. This interaction gives the user the feeling of interacting with a family member.
[0554] Step 7:
[0555] The server monitors the conversation between the user and the virtual character and detects anomalies. The server utilizes natural language processing technology to analyze the text data during the conversation, checking for abnormal patterns and changes in emotion.
[0556] Step 8:
[0557] If an anomaly is detected, the server sends an alert to pre-configured contacts. For example, it can send an email or SMS to family members registered as emergency contacts, enabling a quick response. This allows users to receive immediate assistance in emergencies.
[0558] (Application Example 1)
[0559] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0560] Modern seniors face problems such as social isolation and a lack of support in daily life. Furthermore, there is a demand for improved customer experiences and personalized services in physical stores. To address these challenges, an interactive system tailored to individual needs is necessary, but the technology to realize this is still insufficient. This invention aims to provide a system that reduces feelings of loneliness among seniors, improves their safety and quality of life, and simultaneously offers personalized customer experiences in physical stores.
[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0562] In this invention, the server includes a generation means for generating video and audio of a virtual person using image data, an analysis means for analyzing the content of conversations with the user and detecting anomalies, a notification means for notifying pre-set emergency contacts when an anomaly is detected, a recommendation means for making personalized product recommendations based on the user's attribute information, and a display means for presenting virtual objects through visual information. This makes it possible to reduce feelings of loneliness and ensure safety for the elderly, while simultaneously providing a personalized customer experience in physical stores.
[0563] "Image data" refers to a collection of electronically stored visual information, which is used to generate images of virtual characters.
[0564] A "virtual character" is a character that looks like a human being, created using computer generation technology, and is intended to engage in voice and visual interaction with the user.
[0565] "Generation means" refers to a technical device for creating video and audio of a virtual person using image data and other information.
[0566] "Analysis means" refers to a technical device used to analyze the content of conversations with users and detect anomalies or specific patterns.
[0567] "Anomaly" refers to a specific event or situation that is judged to deviate from normal conversation or behavioral patterns.
[0568] A "notification device" is a technical device for transmitting information to a pre-configured emergency contact or other receiving device when an anomaly is detected.
[0569] A "recommendation tool" is a technical device that individually presents appropriate products and information based on the user's attribute information and past behavioral history.
[0570] A "display means" is a device that visually presents virtual objects and related information, and facilitates interaction with the user.
[0571] This invention is a system aimed at improving the quality of life for the elderly and providing personalized customer experiences in physical stores. The configuration of this system is described in detail below.
[0572] The server is responsible for the core functions of the entire system. Using generative AI models, it analyzes pre-stored image data to generate video and audio of virtual characters. For example, if a user prompts "Tell me today's weather," the server uses natural language processing to analyze the request and construct a virtual character to respond. The server performs processing using cloud services such as Google Cloud Platform and Microsoft Azure.
[0573] The terminal serves as the interface closest to the user. This terminal consists of devices such as smart glasses and head-mounted displays, and displays a virtual person based on a generative model sent from the server. It also analyzes the user's voice commands and gestures using Google Cloud Speech-to-Text and Microsoft Cognitive Services, and provides appropriate feedback.
[0574] For example, if a customer using smart glasses in a physical store says, "Tell me what products would suit these," the device analyzes the voice in real time, and a virtual person recommends products based on information from the server. At this time, the recommendations are personalized based on the customer's past purchase information and pre-registered preferences.
[0575] Users are central to the interaction with this system in their daily lives. Elderly individuals can alleviate feelings of loneliness through daily conversations with virtual characters. For example, if a user asks, "Tell me my exercise record for today," the system will provide the relevant information and offer health advice based on it.
[0576] This system is operated using a high-performance cloud infrastructure and advanced AI algorithms to maintain response speed and reliability. Examples of prompts include: "What are this month's recommended products?" and "I want to check my weekend plans."
[0577] The embodiments of this invention aim to improve the user experience by leveraging personalized and interactive features.
[0578] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0579] Step 1:
[0580] The user enters a voice command into the device. The device captures this audio and uses the Google Cloud Speech-to-Text API to convert the audio data into text. The input is the user's voice command, and the output is the command in text format.
[0581] Step 2:
[0582] The server receives text commands sent from the terminal and performs natural language processing. Specifically, it uses GPT-3.5 or similar AI models to analyze the user's intent and generate a response from a virtual character. The input is a user command in text format, and the output is the generated virtual character's response text.
[0583] Step 3:
[0584] The server generates video and audio of a virtual character based on the generated response. It uses Unity or a similar platform with a GPU to render the video and synthesize the speech. The input is the virtual character's response text, and the output is the virtual character's video and synthesized speech.
[0585] Step 4:
[0586] The terminal receives video and audio of a virtual person transmitted from the server and presents it to the user using smart glasses or a display device. The input is video and audio data of the virtual person, and the output is a visual and auditory presentation to the user.
[0587] Step 5:
[0588] The user interacts with a virtual person through the device and receives information and product recommendations in response to prompts. In this step, personalized recommendations are presented based on the user's past purchase history and attribute information. Input is the user's prompts and feedback, and output is customized information and product recommendations.
[0589] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0590] This invention provides an AI appliance system that incorporates an emotion engine to recognize the user's emotions, enabling more responsive conversations, thereby reducing loneliness in the daily lives of the elderly and providing notifications in case of abnormalities.
[0591] Server Role
[0592] The server regularly updates its AI model, including the emotion engine algorithm. It utilizes the latest speech recognition and facial expression analysis technologies to extract and analyze emotions from the user's voice and facial expressions in real time. It also has the functionality to detect anomalies based on the analyzed data and notify emergency contacts as needed.
[0593] Terminal role
[0594] The device is equipped with a camera and microphone, allowing it to detect the user's facial expressions and voice. The device sends this data to a server, where an emotion engine analyzes it and adjusts the conversation based on the results. When the user interacts with the system, the device displays a generated virtual persona and changes the tone and content of its responses according to the user's emotional state, providing a more natural and less stressful conversational experience.
[0595] User roles
[0596] The user faces a digital photo frame and begins interacting with it as usual. In addition to daily reminders and questions, the user talks about how they feel that day, allowing the system to recognize their emotional state. For example, if the user sounds tired, the system will respond by suggesting they rest. If an anomaly is detected, the system will automatically contact the necessary people to alleviate the user's burden.
[0597] Examples
[0598] As a concrete example, suppose a user says to their device, "I'm feeling a little down." The device sends their voice and facial expression to the server, where an emotion engine recognizes the emotion of "sadness." The server analyzes this emotional state and generates a response that includes a more comforting and encouraging tone than a normal conversation. Furthermore, if the emotion of "sadness" is recognized frequently, the server can perceive this persistent emotional change as an anomaly and send an alert to emergency contacts.
[0599] This invention enables faster problem resolution while providing a more interactive and responsive user experience.
[0600] The following describes the processing flow.
[0601] Step 1:
[0602] The user speaks into the device, saying, "I'm feeling a little down today." This input is received by the device via the microphone.
[0603] Step 2:
[0604] The device analyzes the acquired audio data and captures the user's facial expressions with its camera. The audio and facial expression data are transmitted to the server in real time.
[0605] Step 3:
[0606] The server uses an emotion engine to analyze the user's emotional state from the transmitted data. It recognizes emotions such as "sadness" from voice tone and facial expressions.
[0607] Step 4:
[0608] The server generates an appropriate response based on the recognized emotional state. In this example, it prepares dialogue that encourages the user or suggests they take a rest.
[0609] Step 5:
[0610] The terminal uses the response sent from the server to generate a virtual character and proceed with the interaction with the user. This includes speech synthesis and video display.
[0611] Step 6:
[0612] The server continuously monitors the user's emotional state even during interactions. If the emotion of "sadness" is detected frequently over a certain period, it can be identified as an anomaly.
[0613] Step 7:
[0614] If an anomaly is detected, the server will notify pre-configured emergency contacts. This notification will include information about changes in the user's emotional state.
[0615] Step 8:
[0616] After the conversation ends, the terminal saves all conversation data and sentiment analysis results as logs and backs them up to the server for use in future analyses and improvements to the AI model.
[0617] (Example 2)
[0618] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0619] This system aims to alleviate feelings of loneliness among the elderly, provide emotional support in daily life, quickly detect emotional abnormalities, and ensure they receive necessary assistance. In particular, it strives to improve the quality of life for the elderly by recognizing emotional changes in real time and responding appropriately through dialogue.
[0620] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0621] In this invention, the server includes generation means for generating visual and auditory information of a virtual person using voice data and image data; analysis means for analyzing the user's voice and facial expressions and identifying their emotional state; dialogue generation means for generating responses based on the emotion analysis; and notification means for notifying pre-configured emergency contacts when emotional changes or persistent emotional abnormalities are detected. This enables real-time understanding of the user's emotions, the provision of dialogues that alleviate feelings of loneliness, and rapid support in the event of an abnormality.
[0622] "Voice data" refers to information recorded in digital format from the user's speech, and is used for analysis.
[0623] "Image data" refers to visual information that digitally records the user's facial expressions and posture, and is data used for analysis.
[0624] A "virtual character" is an artificial human figure generated by a computer, which interacts with the user through visual and auditory information.
[0625] "Generation means" refers to a device or method for generating visual and auditory information of a virtual person using audio data and image data.
[0626] "Analysis means" refers to a device or method that identifies an emotional state by analyzing the user's voice and facial expressions.
[0627] "Dialogue generation means" refers to a device or method that generates an appropriate response based on the results of emotion analysis and conveys it to the user through a virtual character.
[0628] "Notification means" refers to a device or method that notifies pre-set emergency contacts when it detects changes in emotions or abnormal emotions.
[0629] "Emotional state" refers to the state of the user's internal emotional expression, including the type and intensity of emotions identified through analysis.
[0630] An "emergency contact" refers to an external contact designated to receive notifications in the event of an emergency involving the user.
[0631] This invention utilizes an AI system equipped with an emotion analysis engine to alleviate feelings of loneliness in the user's daily life, detect emotional abnormalities, and provide necessary support.
[0632] The server is a computer device that runs the emotion analysis engine and plays a key role. This server receives audio and image data and uses speech recognition software and facial expression analysis software to analyze them, respectively. Specifically, a general speech processing API is used for speech recognition, and an image processing library is used for facial expression analysis. This identifies the user's emotional state. By utilizing an AI model, a natural language response that matches the emotional situation is generated.
[0633] The terminal is a device equipped with a camera and microphone, and its primary role is to capture the user's voice and facial expressions and transmit them to the server. Furthermore, based on the analysis results received from the server, the terminal displays a virtual person's image and interacts with the user visually and audibly. The terminal provides a natural conversational experience by changing the tone and content of the virtual person's responses according to the user's emotional state.
[0634] Users engage in everyday conversations with their devices. They can talk about their feelings for the day when using reminders or seeking advice. For example, by uttering a prompt such as "I'm feeling a little down," the system instantly analyzes their emotions and offers words of comfort and encouragement. If the change in emotions is deemed abnormal, the server automatically notifies emergency contacts.
[0635] This invention aims to reduce feelings of loneliness in the user's living environment and enable prompt support responses through emotional monitoring.
[0636] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0637] Step 1:
[0638] The device acquires the user's voice and facial expression data.
[0639] Specifically, the system uses a camera to capture an image of the user's face and a microphone to record their voice.
[0640] The audio and image data used as input are sent to the server in real time for the next processing step.
[0641] Step 2:
[0642] The server converts the received audio data into text data.
[0643] This process uses speech recognition software to analyze the audio signal and extract it as text information.
[0644] The input is audio data, and the output is text data.
[0645] Step 3:
[0646] The server analyzes image data to recognize the user's emotions.
[0647] This process uses a facial expression analysis algorithm to analyze facial features in an image and identify the emotional state.
[0648] The input is image data, and the output is data indicating emotional state.
[0649] Step 4:
[0650] The server generates an appropriate response based on the analyzed text data and emotional state.
[0651] We utilize generative AI models to create natural language responses based on textual and emotional information.
[0652] The input consists of text data and sentiment data, and the output is a natural language response.
[0653] Step 5:
[0654] The terminal receives a response from the server and communicates the response to the user using a virtual persona.
[0655] Specifically, the system displays a virtual character on the screen and plays back a response generated using speech synthesis software.
[0656] The input is a natural language response sentence, and the output is the response as visual and auditory information.
[0657] Step 6:
[0658] The server continuously monitors emotional data and, if it detects an emotional anomaly, it notifies the designated emergency contact.
[0659] This process analyzes patterns of emotional changes and sends an alert using the notification system if an abnormality is detected.
[0660] The input is continuously collected sentiment data, and the output is the sending of emergency notifications.
[0661] (Application Example 2)
[0662] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0663] In the daily lives of the elderly, there is a need for effective systems that reduce feelings of isolation and can respond quickly to emotional changes and abnormal situations. In particular, a challenge is the inability to detect situations where the elderly require emotional support or where abnormalities are not detected early, potentially jeopardizing their safety. To solve these problems, a system is needed that accurately grasps the user's emotions and enables adaptive responses and rapid action in the event of an abnormality.
[0664] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0665] In this invention, the server includes generation means for generating visual and auditory information of a virtual entity using image data and audio data; analysis means for analyzing the content of conversations with the user and detecting emotions and anomalies; notification means for notifying pre-set emergency contacts and specific locations when an anomaly is detected or when emotions change; and response generation means for providing adaptive responses according to pre-set emotional states. This enables accurate understanding of the user's emotional state, reduces feelings of isolation, and allows for rapid response to anomalies.
[0666] "Image data" refers to data that represents visual information in a digital format.
[0667] "Audio data" refers to data that represents auditory information in a digital format.
[0668] A "virtual entity" is an artificial character or person created using digital technology and represented as visual and auditory information on a computer screen or similar device.
[0669] "Visual information" refers to information obtained through human vision, and is expressed in the form of images, videos, and other visual media.
[0670] "Auditory information" refers to information obtained through human hearing, expressed in the form of sounds, music, and other similar media.
[0671] "Generative means" refers to a technical configuration or device for producing a specific function or result.
[0672] "Analysis means" refers to a technical configuration or device for analyzing input data and extracting specific patterns or information.
[0673] "Detecting emotions" involves analyzing information such as the user's facial expressions and voice to identify the emotions that are expressed.
[0674] "Detecting an anomaly" means detecting a situation or condition that is different from the normal state.
[0675] "Emotional change" refers to the change in the user's emotional state over time.
[0676] "Notification means" refers to a technical configuration or function for informing a third party of specific information.
[0677] "Response generation means" refers to a technical configuration or function for automatically creating an appropriate response based on the user's input and circumstances.
[0678] "Emergency contact information" refers to information indicating how to contact individuals or organizations in the event of an emergency.
[0679] An "adaptive response" is a response that is adjusted according to the user's condition and circumstances.
[0680] To realize this invention, the system consists of a client terminal and a server. The terminal is equipped with a camera and microphone, which can sense the user's facial expressions and voice in real time. The sensed data is sent to the server, which analyzes the data using a generative AI model to identify the user's emotional state.
[0681] Specifically, the server uses OpenCV and dlib as face recognition libraries, and leverages machine learning frameworks such as TensorFlow for speech analysis. The server combines these tools to perform data calculations that extract emotions from facial expressions and speech. Based on the results of the emotion analysis, it generates adaptive responses from pre-configured response patterns. These responses are returned to the user in either voice or text format.
[0682] Furthermore, if a significant change or abnormality is detected in the user's emotional state, the server automatically sends a notification to pre-registered emergency contacts. This helps to ensure the user's safety while also reducing feelings of isolation.
[0683] For example, if a user says, "I'm not feeling very cheerful today," the system will detect "sadness" from the content of their statement and the tone of their voice. The emotion engine will analyze this data and generate a response intended to comfort or support the user. For example, it might say something like, "Why don't you take a short break?"
[0684] An example of a prompt using a generative AI model is, "Please give us some ideas on what to say to comfort a user who appears sad." This prompt can be used to improve the accuracy of response generation.
[0685] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0686] Step 1:
[0687] The device collects user facial expressions and voice data through its camera and microphone. It senses what the user says to the device and their facial expressions in real time and prepares to transmit that data. The input consists of facial image data and voice data, and the output is data to be transmitted to the server.
[0688] Step 2:
[0689] The server receives facial image data and audio data transmitted from the terminal. The server prepares this data for analysis and configures it as a dataset for facial expression analysis and audio analysis. The input is the raw data received from the terminal, and the output is in a data format ready for analysis.
[0690] Step 3:
[0691] The server analyzes facial expression data using OpenCV and dlib, and simultaneously analyzes audio data using TensorFlow to detect emotions. A generative AI model for emotion detection identifies emotions from facial characteristics and voice tone, and obtains analysis results based on this. The input is pre-prepared data for analysis, and the output is the detected emotion information.
[0692] Step 4:
[0693] The server generates responses based on detected emotion information. It uses an AI model to generate appropriate responses according to the emotional state, determining the content of the response to the user. It generates response candidates using prompt text and selects the optimal one. The output is the generated voice or text response.
[0694] Step 5:
[0695] The server sends the generated response to the terminal. The terminal then relays this response to the user, either by providing an audio response through the speaker or by displaying it as text on the screen. The input is the generated response data, and the output is the specific response presented to the user.
[0696] Step 6:
[0697] The server immediately sends a notification to emergency contacts if the detected emotions indicate a significant change or anomaly. Here, it extracts the necessary information based on the anomaly detection and sends an alert to the registered contacts. The output is an emergency contact message.
[0698] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0699] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0700] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0701] [Fourth Embodiment]
[0702] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0703] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0704] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0705] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0706] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0707] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0708] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0709] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0710] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0711] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0712] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0713] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0714] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0715] This invention provides an AI appliance system that alleviates feelings of loneliness among the elderly and enables rapid response in the event of an emergency. This system primarily consists of the interaction between a server, a terminal, and a user. The roles of each component and specific examples are explained below.
[0716] Server Role
[0717] The server is the core of the entire system and is primarily responsible for updating AI models and analyzing data. The server maintains models that generate virtual individuals resembling the user's family members using the latest DeepFake technology, and regularly updates these models with training data. It also analyzes user interactions in real time and uses language processing techniques to detect unnatural conversation patterns and anomalies.
[0718] Terminal role
[0719] The device is a digital photo frame that displays a virtual person using model data transmitted from a server. It recognizes voice input and actions from the user and responds appropriately based on that. By providing voice and visual interaction with the virtual person, the user can feel as if they are always connected to their family. If an abnormal situation is detected, the device quickly sends data to the server, and an emergency contact is made promptly.
[0720] User roles
[0721] Users are the primary users of the system and interact with the device mainly through voice. Users can engage in natural conversations with virtual characters and receive lifestyle advice and reminders from the system. If there are any abnormalities regarding the user's health or activity, the system will make a judgment based on the user's statements and actions and notify them as necessary.
[0722] Examples
[0723] As a concrete example, a user can send a request to their device saying, "I want to talk to my grandchild today." The device analyzes this request, uses a generative model received from the server to generate a virtual person resembling the grandchild, and begins a conversation with the user. While the conversation is ongoing, the server monitors the content and sends an alert to an emergency contact if it detects any unusual patterns.
[0724] This system will improve the quality of life for the elderly and enable a swift and appropriate response in emergencies.
[0725] The following describes the processing flow.
[0726] Step 1:
[0727] The server generates the latest AI model and updates the data to create virtual people resembling the user's family using Deep Fake technology. The updated model is sent to the device in real time.
[0728] Step 2:
[0729] The terminal installs the generative model received from the server and prepares to display the virtual character. It waits for the user to input what they want to control the terminal by voice.
[0730] Step 3:
[0731] The user speaks to the device and gives a voice command saying, "I want to talk to my grandchild." This voice command is received and recognized by the device.
[0732] Step 4:
[0733] The terminal uses a generative model downloaded from the server based on user commands to begin generating a virtual person resembling the grandchild. During this process, both audio and video are synthesized in real time to provide the user with a conversational environment.
[0734] Step 5:
[0735] The server monitors the interaction between the user and the virtual character in real time. It uses natural language processing to analyze the conversation content and determine whether any anomalies have been detected.
[0736] Step 6:
[0737] If the server detects any unnatural patterns or anomalies in the interaction, it will send an alert to a pre-configured emergency contact. The notification will include a statement that an anomaly may have occurred.
[0738] Step 7:
[0739] After the conversation ends, the device records the conversation history and usage, and sends the data to the server as needed. This data will be used for future model updates and feature improvements.
[0740] (Example 1)
[0741] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0742] There is a need for technologies that can alleviate feelings of loneliness among the elderly and enable rapid responses in emergencies. However, conventional technologies struggle to achieve both natural conversation and early detection of abnormalities, and in particular, the ability to balance flexible conversation through voice interfaces with abnormality detection has not been sufficiently achieved. Solving this problem is necessary to provide a safer and more fulfilling living environment.
[0743] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0744] In this invention, the server includes an input means for inputting instructions using voice, a generation means for generating video and audio of a virtual person using generated information, an analysis means for analyzing dialogue information and detecting anomalies, and a notification means for notifying pre-set contacts when an anomaly is detected. This allows elderly people to gain a sense of security through natural dialogue in their daily lives, and enables a quick and appropriate response in emergencies.
[0745] "Sound" is a form of sound, a means by which people communicate information through language.
[0746] "Instructions" refer to commands or requests given to obtain a specific action or result.
[0747] "Input means" refers to devices or methods for importing data or information into a system.
[0748] "Generated information" refers to a collection of data and models used to create the video and audio of a virtual character.
[0749] A "virtual character" refers to a person who does not actually exist but is simulated by a computer.
[0750] "Generation means" refers to devices and methods for creating new images and sounds using data and models.
[0751] "Dialogue information" refers to the content of conversations and communications exchanged between the user and the system.
[0752] "Analytical means" refers to devices and methods for analyzing data and information and extracting useful patterns and insights from them.
[0753] An "abnormality" refers to a state or pattern that is different from the norm and requires attention or action.
[0754] "Notification means" refers to devices or methods for transmitting specific information or messages to other devices or people.
[0755] "Contact information" refers to the contact details of an individual or organization used to communicate in emergencies or for conveying specific information.
[0756] This invention is a virtual dialogue system designed to alleviate feelings of loneliness among the elderly and enable rapid response in emergency situations. It primarily functions by utilizing the interaction between a server, a terminal, and a user.
[0757] Server Role
[0758] The server plays a central role in the entire system, managing AI models and analyzing data. Specifically, it maintains and updates models that generate virtual individuals resembling the user's relatives using DeepFake technology. To this end, it continuously incorporates training data using generative AI models and analyzes user interactions in real time. The server skillfully detects unnatural patterns and anomalies from conversations using natural language processing techniques.
[0759] Terminal role
[0760] The device is a digital photo frame that projects a virtual person onto its screen using model data provided by a server. Users can interact with the virtual person using voice through the device. The device is equipped with voice recognition capabilities, allowing it to recognize and process voice input from the user appropriately. Through conversations with the virtual person, users can experience the feeling of interacting with their family. In addition, if an anomaly is detected, data is quickly sent to the server, and an emergency notification is promptly issued.
[0761] User roles
[0762] Users are the primary users of this system and interact with their devices mainly using voice. Users can enjoy conversations with virtual characters and receive lifestyle advice and reminders from the system. Based on their speech and actions, the system will notify users as needed if it detects any abnormalities in their health or activity.
[0763] Specific example
[0764] For example, a user could request, "I want to talk to my grandchild today." In this case, the device analyzes the request, generates a virtual person from the server, and displays someone who resembles the grandchild. Then, the conversation begins. The server continuously monitors the conversation and immediately sends an alert to an emergency contact if any abnormal patterns are detected.
[0765] Example of a prompt
[0766] A concrete example of a prompt message is input such as, "Please tell me the appropriate procedure for generating a virtual character when a user requests to speak with their grandchild."
[0767] This system reduces feelings of loneliness for the elderly in their daily lives and ensures a swift and appropriate response in emergencies.
[0768] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0769] Step 1:
[0770] The user inputs a request by voice into the device. The user tells the device a specific request in natural language, such as "I want to talk to my grandchild today." This voice input triggers the start of processing.
[0771] Step 2:
[0772] The device receives voice input from the user. Speech recognition software is used to convert the voice data into text. During this process, the device captures the user's voice through the microphone, analyzes it using a speech recognition engine, and outputs it as text data.
[0773] Step 3:
[0774] The terminal parses the transcribed user request and sends a request to the server to generate a virtual person. The terminal extracts key phrases from the text, constructs a prompt, and then sends it to the server as a data packet. This data contains details about the user's request.
[0775] Step 4:
[0776] The server generates a virtual person using a generative AI model based on the request data received from the terminal. The server references stored family data and uses Deep Fake technology to synthesize the virtual person's video and audio. In this process, the model dynamically calculates the data and generates the virtual person's video and audio in real time.
[0777] Step 5:
[0778] The server sends the generated virtual character data to the terminal. The server outputs data containing the virtual character's video and audio to the terminal. Upon receiving this data, the terminal displays it to the user using a display device.
[0779] Step 6:
[0780] The device displays a virtual person and initiates a conversation with the user. The device displays the generated image on the screen and outputs audio through its speaker. This interaction gives the user the feeling of interacting with a family member.
[0781] Step 7:
[0782] The server monitors the conversation between the user and the virtual character and detects anomalies. The server utilizes natural language processing technology to analyze the text data during the conversation, checking for abnormal patterns and changes in emotion.
[0783] Step 8:
[0784] If an anomaly is detected, the server sends an alert to pre-configured contacts. For example, it can send an email or SMS to family members registered as emergency contacts, enabling a quick response. This allows users to receive immediate assistance in emergencies.
[0785] (Application Example 1)
[0786] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0787] Modern seniors face problems such as social isolation and a lack of support in daily life. Furthermore, there is a demand for improved customer experiences and personalized services in physical stores. To address these challenges, an interactive system tailored to individual needs is necessary, but the technology to realize this is still insufficient. This invention aims to provide a system that reduces feelings of loneliness among seniors, improves their safety and quality of life, and simultaneously offers personalized customer experiences in physical stores.
[0788] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0789] In this invention, the server includes a generation means for generating video and audio of a virtual person using image data, an analysis means for analyzing the content of conversations with the user and detecting anomalies, a notification means for notifying pre-set emergency contacts when an anomaly is detected, a recommendation means for making personalized product recommendations based on the user's attribute information, and a display means for presenting virtual objects through visual information. This makes it possible to reduce feelings of loneliness and ensure safety for the elderly, while simultaneously providing a personalized customer experience in physical stores.
[0790] "Image data" refers to a collection of electronically stored visual information, which is used to generate images of virtual characters.
[0791] A "virtual character" is a character that looks like a human being, created using computer generation technology, and is intended to engage in voice and visual interaction with the user.
[0792] "Generation means" refers to a technical device for creating video and audio of a virtual person using image data and other information.
[0793] "Analysis means" refers to a technical device used to analyze the content of conversations with users and detect anomalies or specific patterns.
[0794] "Anomaly" refers to a specific event or situation that is judged to deviate from normal conversation or behavioral patterns.
[0795] A "notification device" is a technical device for transmitting information to a pre-configured emergency contact or other receiving device when an anomaly is detected.
[0796] A "recommendation tool" is a technical device that individually presents appropriate products and information based on the user's attribute information and past behavioral history.
[0797] A "display means" is a device that visually presents virtual objects and related information, and facilitates interaction with the user.
[0798] This invention is a system aimed at improving the quality of life for the elderly and providing personalized customer experiences in physical stores. The configuration of this system is described in detail below.
[0799] The server is responsible for the core functions of the entire system. Using generative AI models, it analyzes pre-stored image data to generate video and audio of virtual characters. For example, if a user prompts "Tell me today's weather," the server uses natural language processing to analyze the request and construct a virtual character to respond. The server performs processing using cloud services such as Google Cloud Platform and Microsoft Azure.
[0800] The terminal serves as the interface closest to the user. This terminal consists of devices such as smart glasses and head-mounted displays, and displays a virtual person based on a generative model sent from the server. It also analyzes the user's voice commands and gestures using Google Cloud Speech-to-Text and Microsoft Cognitive Services, and provides appropriate feedback.
[0801] For example, if a customer using smart glasses in a physical store says, "Tell me what products would suit these," the device analyzes the voice in real time, and a virtual person recommends products based on information from the server. At this time, the recommendations are personalized based on the customer's past purchase information and pre-registered preferences.
[0802] Users are central to the interaction with this system in their daily lives. Elderly individuals can alleviate feelings of loneliness through daily conversations with virtual characters. For example, if a user asks, "Tell me my exercise record for today," the system will provide the relevant information and offer health advice based on it.
[0803] This system is operated using a high-performance cloud infrastructure and advanced AI algorithms to maintain response speed and reliability. Examples of prompts include: "What are this month's recommended products?" and "I want to check my weekend plans."
[0804] The embodiments of this invention aim to improve the user experience by leveraging personalized and interactive features.
[0805] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0806] Step 1:
[0807] The user enters a voice command into the device. The device captures this audio and uses the Google Cloud Speech-to-Text API to convert the audio data into text. The input is the user's voice command, and the output is the command in text format.
[0808] Step 2:
[0809] The server receives text commands sent from the terminal and performs natural language processing. Specifically, it uses GPT-3.5 or similar AI models to analyze the user's intent and generate a response from a virtual character. The input is a user command in text format, and the output is the generated virtual character's response text.
[0810] Step 3:
[0811] The server generates video and audio of a virtual character based on the generated response. It uses Unity or a similar platform with a GPU to render the video and synthesize the speech. The input is the virtual character's response text, and the output is the virtual character's video and synthesized speech.
[0812] Step 4:
[0813] The terminal receives video and audio of a virtual person transmitted from the server and presents it to the user using smart glasses or a display device. The input is video and audio data of the virtual person, and the output is a visual and auditory presentation to the user.
[0814] Step 5:
[0815] The user interacts with a virtual person through the device and receives information and product recommendations in response to prompts. In this step, personalized recommendations are presented based on the user's past purchase history and attribute information. Input is the user's prompts and feedback, and output is customized information and product recommendations.
[0816] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0817] This invention provides an AI appliance system that incorporates an emotion engine to recognize the user's emotions, enabling more responsive conversations, thereby reducing loneliness in the daily lives of the elderly and providing notifications in case of abnormalities.
[0818] Server Role
[0819] The server regularly updates its AI model, including the emotion engine algorithm. It utilizes the latest speech recognition and facial expression analysis technologies to extract and analyze emotions from the user's voice and facial expressions in real time. It also has the functionality to detect anomalies based on the analyzed data and notify emergency contacts as needed.
[0820] Terminal role
[0821] The device is equipped with a camera and microphone, allowing it to detect the user's facial expressions and voice. The device sends this data to a server, where an emotion engine analyzes it and adjusts the conversation based on the results. When the user interacts with the system, the device displays a generated virtual persona and changes the tone and content of its responses according to the user's emotional state, providing a more natural and less stressful conversational experience.
[0822] User roles
[0823] The user faces a digital photo frame and begins interacting with it as usual. In addition to daily reminders and questions, the user talks about how they feel that day, allowing the system to recognize their emotional state. For example, if the user sounds tired, the system will respond by suggesting they rest. If an anomaly is detected, the system will automatically contact the necessary people to alleviate the user's burden.
[0824] Examples
[0825] As a concrete example, suppose a user says to their device, "I'm feeling a little down." The device sends their voice and facial expression to the server, where an emotion engine recognizes the emotion of "sadness." The server analyzes this emotional state and generates a response that includes a more comforting and encouraging tone than a normal conversation. Furthermore, if the emotion of "sadness" is recognized frequently, the server can perceive this persistent emotional change as an anomaly and send an alert to emergency contacts.
[0826] This invention enables faster problem resolution while providing a more interactive and responsive user experience.
[0827] The following describes the processing flow.
[0828] Step 1:
[0829] The user speaks into the device, saying, "I'm feeling a little down today." This input is received by the device via the microphone.
[0830] Step 2:
[0831] The device analyzes the acquired audio data and captures the user's facial expressions with its camera. The audio and facial expression data are transmitted to the server in real time.
[0832] Step 3:
[0833] The server uses an emotion engine to analyze the user's emotional state from the transmitted data. It recognizes emotions such as "sadness" from voice tone and facial expressions.
[0834] Step 4:
[0835] The server generates an appropriate response based on the recognized emotional state. In this example, it prepares dialogue that encourages the user or suggests they take a rest.
[0836] Step 5:
[0837] The terminal uses the response sent from the server to generate a virtual character and proceed with the interaction with the user. This includes speech synthesis and video display.
[0838] Step 6:
[0839] The server continuously monitors the user's emotional state even during interactions. If the emotion of "sadness" is detected frequently over a certain period, it can be identified as an anomaly.
[0840] Step 7:
[0841] If an anomaly is detected, the server will notify pre-configured emergency contacts. This notification will include information about changes in the user's emotional state.
[0842] Step 8:
[0843] After the conversation ends, the terminal saves all conversation data and sentiment analysis results as logs and backs them up to the server for use in future analyses and improvements to the AI model.
[0844] (Example 2)
[0845] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0846] This system aims to alleviate feelings of loneliness among the elderly, provide emotional support in daily life, quickly detect emotional abnormalities, and ensure they receive necessary assistance. In particular, it strives to improve the quality of life for the elderly by recognizing emotional changes in real time and responding appropriately through dialogue.
[0847] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0848] In this invention, the server includes generation means for generating visual and auditory information of a virtual person using voice data and image data; analysis means for analyzing the user's voice and facial expressions and identifying their emotional state; dialogue generation means for generating responses based on the emotion analysis; and notification means for notifying pre-configured emergency contacts when emotional changes or persistent emotional abnormalities are detected. This enables real-time understanding of the user's emotions, the provision of dialogues that alleviate feelings of loneliness, and rapid support in the event of an abnormality.
[0849] "Voice data" refers to information recorded in digital format from the user's speech, and is used for analysis.
[0850] "Image data" refers to visual information that digitally records the user's facial expressions and posture, and is data used for analysis.
[0851] A "virtual character" is an artificial human figure generated by a computer, which interacts with the user through visual and auditory information.
[0852] "Generation means" refers to a device or method for generating visual and auditory information of a virtual person using audio data and image data.
[0853] "Analysis means" refers to a device or method that identifies an emotional state by analyzing the user's voice and facial expressions.
[0854] "Dialogue generation means" refers to a device or method that generates an appropriate response based on the results of emotion analysis and conveys it to the user through a virtual character.
[0855] "Notification means" refers to a device or method that notifies pre-set emergency contacts when it detects changes in emotions or abnormal emotions.
[0856] "Emotional state" refers to the state of the user's internal emotional expression, including the type and intensity of emotions identified through analysis.
[0857] An "emergency contact" refers to an external contact designated to receive notifications in the event of an emergency involving the user.
[0858] This invention utilizes an AI system equipped with an emotion analysis engine to alleviate feelings of loneliness in the user's daily life, detect emotional abnormalities, and provide necessary support.
[0859] The server is a computer device that runs the emotion analysis engine and plays a key role. This server receives audio and image data and uses speech recognition software and facial expression analysis software to analyze them, respectively. Specifically, a general speech processing API is used for speech recognition, and an image processing library is used for facial expression analysis. This identifies the user's emotional state. By utilizing an AI model, a natural language response that matches the emotional situation is generated.
[0860] The terminal is a device equipped with a camera and microphone, and its primary role is to capture the user's voice and facial expressions and transmit them to the server. Furthermore, based on the analysis results received from the server, the terminal displays a virtual person's image and interacts with the user visually and audibly. The terminal provides a natural conversational experience by changing the tone and content of the virtual person's responses according to the user's emotional state.
[0861] Users engage in everyday conversations with their devices. They can talk about their feelings for the day when using reminders or seeking advice. For example, by uttering a prompt such as "I'm feeling a little down," the system instantly analyzes their emotions and offers words of comfort and encouragement. If the change in emotions is deemed abnormal, the server automatically notifies emergency contacts.
[0862] This invention aims to reduce feelings of loneliness in the user's living environment and enable prompt support responses through emotional monitoring.
[0863] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0864] Step 1:
[0865] The device acquires the user's voice and facial expression data.
[0866] Specifically, the system uses a camera to capture an image of the user's face and a microphone to record their voice.
[0867] The audio and image data used as input are sent to the server in real time for the next processing step.
[0868] Step 2:
[0869] The server converts the received audio data into text data.
[0870] This process uses speech recognition software to analyze the audio signal and extract it as text information.
[0871] The input is audio data, and the output is text data.
[0872] Step 3:
[0873] The server analyzes image data to recognize the user's emotions.
[0874] This process uses a facial expression analysis algorithm to analyze facial features in an image and identify the emotional state.
[0875] The input is image data, and the output is data indicating emotional state.
[0876] Step 4:
[0877] The server generates an appropriate response based on the analyzed text data and emotional state.
[0878] We utilize generative AI models to create natural language responses based on textual and emotional information.
[0879] The input consists of text data and sentiment data, and the output is a natural language response.
[0880] Step 5:
[0881] The terminal receives a response from the server and communicates the response to the user using a virtual persona.
[0882] Specifically, the system displays a virtual character on the screen and plays back a response generated using speech synthesis software.
[0883] The input is a natural language response sentence, and the output is the response as visual and auditory information.
[0884] Step 6:
[0885] The server continuously monitors emotional data and, if it detects an emotional anomaly, it notifies the designated emergency contact.
[0886] This process analyzes patterns of emotional changes and sends an alert using the notification system if an abnormality is detected.
[0887] The input is continuously collected sentiment data, and the output is the sending of emergency notifications.
[0888] (Application Example 2)
[0889] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0890] In the daily lives of the elderly, there is a need for effective systems that reduce feelings of isolation and can respond quickly to emotional changes and abnormal situations. In particular, a challenge is the inability to detect situations where the elderly require emotional support or where abnormalities are not detected early, potentially jeopardizing their safety. To solve these problems, a system is needed that accurately grasps the user's emotions and enables adaptive responses and rapid action in the event of an abnormality.
[0891] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0892] In this invention, the server includes generation means for generating visual and auditory information of a virtual entity using image data and audio data; analysis means for analyzing the content of conversations with the user and detecting emotions and anomalies; notification means for notifying pre-set emergency contacts and specific locations when an anomaly is detected or when emotions change; and response generation means for providing adaptive responses according to pre-set emotional states. This enables accurate understanding of the user's emotional state, reduces feelings of isolation, and allows for rapid response to anomalies.
[0893] "Image data" refers to data that represents visual information in a digital format.
[0894] "Audio data" refers to data that represents auditory information in a digital format.
[0895] A "virtual entity" is an artificial character or person created using digital technology and represented as visual and auditory information on a computer screen or similar device.
[0896] "Visual information" refers to information obtained through human vision, and is expressed in the form of images, videos, and other visual media.
[0897] "Auditory information" refers to information obtained through human hearing, expressed in the form of sounds, music, and other similar media.
[0898] "Generative means" refers to a technical configuration or device for producing a specific function or result.
[0899] "Analysis means" refers to a technical configuration or device for analyzing input data and extracting specific patterns or information.
[0900] "Detecting emotions" involves analyzing information such as the user's facial expressions and voice to identify the emotions that are expressed.
[0901] "Detecting an anomaly" means detecting a situation or condition that is different from the normal state.
[0902] "Emotional change" refers to the change in the user's emotional state over time.
[0903] "Notification means" refers to a technical configuration or function for informing a third party of specific information.
[0904] "Response generation means" refers to a technical configuration or function for automatically creating an appropriate response based on the user's input and circumstances.
[0905] "Emergency contact information" refers to information indicating how to contact individuals or organizations in the event of an emergency.
[0906] An "adaptive response" is a response that is adjusted according to the user's condition and circumstances.
[0907] To realize this invention, the system consists of a client terminal and a server. The terminal is equipped with a camera and microphone, which can sense the user's facial expressions and voice in real time. The sensed data is sent to the server, which analyzes the data using a generative AI model to identify the user's emotional state.
[0908] Specifically, the server uses OpenCV and dlib as face recognition libraries, and leverages machine learning frameworks such as TensorFlow for speech analysis. The server combines these tools to perform data calculations that extract emotions from facial expressions and speech. Based on the results of the emotion analysis, it generates adaptive responses from pre-configured response patterns. These responses are returned to the user in either voice or text format.
[0909] Furthermore, if a significant change or abnormality is detected in the user's emotional state, the server automatically sends a notification to pre-registered emergency contacts. This helps to ensure the user's safety while also reducing feelings of isolation.
[0910] For example, if a user says, "I'm not feeling very cheerful today," the system will detect "sadness" from the content of their statement and the tone of their voice. The emotion engine will analyze this data and generate a response intended to comfort or support the user. For example, it might say something like, "Why don't you take a short break?"
[0911] An example of a prompt using a generative AI model is, "Please give us some ideas on what to say to comfort a user who appears sad." This prompt can be used to improve the accuracy of response generation.
[0912] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0913] Step 1:
[0914] The device collects user facial expressions and voice data through its camera and microphone. It senses what the user says to the device and their facial expressions in real time and prepares to transmit that data. The input consists of facial image data and voice data, and the output is data to be transmitted to the server.
[0915] Step 2:
[0916] The server receives facial image data and audio data transmitted from the terminal. The server prepares this data for analysis and configures it as a dataset for facial expression analysis and audio analysis. The input is the raw data received from the terminal, and the output is in a data format ready for analysis.
[0917] Step 3:
[0918] The server analyzes facial expression data using OpenCV and dlib, and simultaneously analyzes audio data using TensorFlow to detect emotions. A generative AI model for emotion detection identifies emotions from facial characteristics and voice tone, and obtains analysis results based on this. The input is pre-prepared data for analysis, and the output is the detected emotion information.
[0919] Step 4:
[0920] The server generates responses based on detected emotion information. It uses an AI model to generate appropriate responses according to the emotional state, determining the content of the response to the user. It generates response candidates using prompt text and selects the optimal one. The output is the generated voice or text response.
[0921] Step 5:
[0922] The server sends the generated response to the terminal. The terminal then relays this response to the user, either by providing an audio response through the speaker or by displaying it as text on the screen. The input is the generated response data, and the output is the specific response presented to the user.
[0923] Step 6:
[0924] The server immediately sends a notification to emergency contacts if the detected emotions indicate a significant change or anomaly. Here, it extracts the necessary information based on the anomaly detection and sends an alert to the registered contacts. The output is an emergency contact message.
[0925] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0926] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0927] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0928] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0929] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0930] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0931] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0932] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0933] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0934] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0935] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0936] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0937] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0938] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0939] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0940] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0941] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0942] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0943] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0944] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0945] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0946] The following is further disclosed regarding the embodiments described above.
[0947] (Claim 1)
[0948] A generation means for generating video and audio of a virtual person using image data,
[0949] An analysis means that analyzes the content of conversations with the user and detects anomalies,
[0950] A notification method that notifies pre-configured emergency contacts when an anomaly is detected,
[0951] A system that includes this.
[0952] (Claim 2)
[0953] The system according to claim 1, which uses natural language processing to analyze the flow of a conversation when detecting anomalies.
[0954] (Claim 3)
[0955] The system according to claim 1, which monitors the interval between utterances, tone of voice, or length of conversation in a dialogue with a generated virtual character, and detects anomalies.
[0956] "Example 1"
[0957] (Claim 1)
[0958] An input method that uses voice to input instructions,
[0959] A generation means for generating video and audio of a virtual person using generated information,
[0960] An analysis means for analyzing dialogue information and detecting anomalies,
[0961] A notification method that notifies pre-configured contacts when an anomaly is detected,
[0962] A system that includes this.
[0963] (Claim 2)
[0964] The system according to claim 1, which analyzes dialogue information using natural language processing and detects anomalies.
[0965] (Claim 3)
[0966] The system according to claim 1, which monitors the interval between utterances, tone of voice, or duration of conversation in a dialogue with a generated virtual character and detects anomalies.
[0967] "Application Example 1"
[0968] (Claim 1)
[0969] A generation means for generating video and audio of a virtual person using image data,
[0970] An analysis means that analyzes the content of conversations with the user and detects anomalies,
[0971] A notification method that notifies pre-configured emergency contacts when an anomaly is detected,
[0972] A recommendation method that provides personalized product recommendations based on user attribute information,
[0973] A display means that presents a virtual object through visual information,
[0974] A system that includes this.
[0975] (Claim 2)
[0976] When detecting anomalies, natural language processing is used to analyze the flow of conversation.
[0977] The system according to claim 1, further adjusting the recommendations based on the user's past behavioral history.
[0978] (Claim 3)
[0979] Monitor the interval between utterances, tone of voice, or length of conversation in interactions with the generated virtual character.
[0980] Furthermore, virtual objects are provided to the user's visual terminal in real time.
[0981] The system according to claim 1 for detecting abnormalities based on selective behavior.
[0982] "Example 2 of combining an emotion engine"
[0983] (Claim 1)
[0984] A generation means for generating visual and auditory information of a virtual person using audio data and image data,
[0985] An analytical means for analyzing the user's voice and facial expressions to identify their emotional state,
[0986] A dialogue generation means that generates responses based on emotion analysis,
[0987] A notification system that notifies pre-configured emergency contacts when it detects changes in emotions or persistent emotional abnormalities,
[0988] A system that includes this.
[0989] (Claim 2)
[0990] The system according to claim 1, which analyzes conversation and the emotional state of the user using natural language processing and facial expression analysis.
[0991] (Claim 3)
[0992] The system according to claim 1, which monitors the interval between utterances, tone of voice, or duration of conversation in a dialogue with a generated virtual character, and detects abnormalities in emotion.
[0993] "Application example 2 when combining with an emotional engine"
[0994] (Claim 1)
[0995] A generation means for generating visual and auditory information of a virtual entity using image data and audio data,
[0996] An analytical means that analyzes the content of conversations with the user and detects emotions and abnormalities,
[0997] A notification system that notifies pre-configured emergency contacts or specific locations when an anomaly is detected or when there is a change in emotion,
[0998] A response generation means that provides an adaptive response in accordance with a pre-set emotional state,
[0999] A system that includes this.
[1000] (Claim 2)
[1001] The system according to claim 1, which uses natural language processing and emotion analysis techniques to analyze the flow of a conversation and changes in emotion when detecting emotions and anomalies.
[1002] (Claim 3)
[1003] The system according to claim 1, which monitors the interval between utterances, voice characteristics, or length of conversation in interaction with a generated virtual entity, and detects anomalies while taking into account the emotional state. [Explanation of Symbols]
[1004] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A generation means for generating video and audio of a virtual person using image data, An analysis means that analyzes the content of conversations with the user and detects anomalies, A notification method that notifies pre-configured emergency contacts when an anomaly is detected, A system that includes this.
2. The system according to claim 1, which uses natural language processing to analyze the flow of a conversation when detecting an anomaly.
3. The system according to claim 1, which monitors the interval between utterances, tone of voice, or length of conversation in a dialogue with a generated virtual character, and detects abnormalities.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A