system
A system with a generative model and virtual person on a display device addresses loneliness and health management issues for the elderly and disabled, offering real-time health support and emotional interaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
The elderly and disabled living alone face significant issues with loneliness, lack of daily conversations, and inadequate health condition management, leading to cognitive decline and delayed emergency responses.
A system utilizing a generative model to render a virtual person on a display device for real-time, two-way conversation, analyzing user health status, and providing personalized health management and emotional support through speech recognition and external health support services.
The system alleviates loneliness and supports effective health management by enabling natural conversational exchanges, detecting abnormalities, and providing timely health support.
Smart Images

Figure 2026070248000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including: receiving a user utterance; adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot; encoding the prompt; and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the situation where the elderly and the disabled live alone, the sense of loneliness and the lack of health condition management in daily life have become serious problems. In addition, since these people have limited opportunities to participate in the local community, the lack of daily conversations and exchanges may lead to a risk of cognitive decline. Furthermore, there are cases where early detection and appropriate support for the deterioration of health conditions cannot be received, and the response to emergencies is often delayed.
Means for Solving the Problems
[0005] This invention provides a system that uses a generative model to render a virtual person on a display device and enables real-time, two-way conversation with the user. This system periodically acquires, stores, and analyzes the user's health status information, and has a function to notify external health support services if an abnormality is detected. Furthermore, by using speech recognition technology to convert the user's input into text and inputting it into the generative model, natural conversational exchange is made possible. Based on the user's health information, the generative model displays appropriate suggestions and alerts, supporting feelings of loneliness and health management in daily life.
[0006] A "generative model" is an algorithm that uses artificial intelligence technology to generate natural language responses.
[0007] A "display device" is a part of an electronic device, including a display, that provides visual information to the user.
[0008] A "virtual character" is a digital avatar generated by a computer and used to simulate conversations with users.
[0009] "Speech recognition technology" is a technology that converts speech into digital data and converts speech input into text information.
[0010] "External health support services" refer to organizations or platforms that provide medical and nursing care support tailored to the user's health condition.
[0011] "Health status information" refers to data about the user's physical and mental condition, including physical condition, mood, and lifestyle.
[0012] A "digital avatar" refers to a virtual character or person used to visually represent interactions with users. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention provides a system that enables two-way communication with a user through a display device. This system mainly consists of a server, a terminal (a display device such as a television), and a user.
[0035] System Configuration
[0036] The server includes a processing unit for running the generative model. The generative model can simulate natural conversations with users. This model generates responses based on text data sent by the user, enabling two-way communication.
[0037] The device uses speech recognition technology to convert the user's voice input into text and sends it to the server. Furthermore, it receives responses from the server and displays a virtual person avatar to provide visual feedback to the user. The device may also ask daily questions about the user's health status, triggered by user input.
[0038] Users interact with the system in the same way as in a normal conversation. During this interaction, they can input information about their health and daily life events via voice. This information is sent to the server through the device and used to build the user profile.
[0039] Program processing
[0040] The server analyzes the text data received from the user and uses a generative model to generate appropriate conversational responses. For example, if a user says, "I'm not feeling well today," the server refers to the user's daily data and, if necessary, generates a more humane response such as, "I'm worried about your health lately. Please get plenty of rest," and sends it to the terminal.
[0041] The device displays the generated response using a virtual person avatar. This avatar is designed to create a sense of familiarity with the user by moving its mouth and changing its facial expressions. In addition, the device uses speech recognition technology to transcribe the user's speech into text and sends it to the server, making it easy to use naturally in everyday situations without burdening the user.
[0042] Through this system, users can not only alleviate feelings of loneliness in their daily lives but also receive support in understanding their own health status. The system evolves on its own, providing communication optimized for each user. For example, if a user says to the system, "I'm worried about not getting enough exercise lately," the system can respond, "Okay, let's work together to come up with an easy exercise plan to get you moving."
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The user turns on the TV, and the device displays the initial setup screen. The device starts a session with the server and establishes a communication channel.
[0046] Step 2:
[0047] The device activates its voice recognition function and waits for the user to speak. When the user says "Hello," the device converts the voice input into text data.
[0048] Step 3:
[0049] The terminal sends the user's text data to the server. The server uses a generative model to generate an appropriate response to the user's "Hello," creating a reply such as "How are you doing today?"
[0050] Step 4:
[0051] The server generates a text response and sends it to the terminal. The terminal uses a virtual person avatar to play the generated response aloud, with mouth movements.
[0052] Step 5:
[0053] The user replies, "I'm a little tired today." The device converts the speech back into text and sends it to the server.
[0054] Step 6:
[0055] The server references past user data and analyzes changes in health status. Using a generative model, it generates a suggestion such as, "That sounds tough, why don't you take a break?"
[0056] Step 7:
[0057] The server sends a suggestion to the terminal. The terminal communicates the suggestion to the user by controlling an avatar. This prompts the user to take appropriate action.
[0058] Step 8:
[0059] The user indicates their intention to end the conversation. If they say, "Thank you, I'm fine now," the device notifies the server that the session has ended.
[0060] Step 9:
[0061] The terminal displays the termination screen and closes the application. The server safely terminates the communication session and saves the log data to the database.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] In modern society, while interest in personal health management is increasing, systems for monitoring health status at home and reducing feelings of isolation are still not adequately developed. In particular, there is a growing need for interactive systems that allow users to understand their own health status through natural conversations at home and receive necessary advice and warnings. Such systems need to support users' daily health management and have features to ensure that urgent health conditions are not overlooked.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for receiving voice input and converting it into a string using speech recognition technology; and means for an analysis device to process the string, generate prompt sentences, and input them into the generative model. This enables the user to monitor their health status through natural conversation and receive quick and appropriate feedback.
[0067] A "generative model" is an artificial intelligence algorithm that generates an appropriate response based on a given input text.
[0068] A "display device" is a device used to present visual information, including virtual character avatars, to a user.
[0069] "Speech recognition technology" is a technology that converts a user's voice into text data.
[0070] A "prompt" is a text-based question or instruction that is input into a generative model.
[0071] An "analysis device" is a device that processes text data and generates prompts adapted to a generative model.
[0072] A "user profile" is an aggregate of personal information built based on a user's attributes and past conversation history.
[0073] "External health-related services" refer to third-party services that utilize users' health information to provide support and advice.
[0074] An "avatar" is a virtual character displayed to visually represent a two-way conversation with the user.
[0075] This invention constitutes a system that provides users with a natural and interactive conversational experience. The system is primarily operated by a server, a terminal, and the user.
[0076] The server is a computing device for running the generative AI model and receives data transcribed into text using speech recognition technology. This data is processed by an analysis device, which generates prompt sentences to be input into the generative model. The server uses these prompt sentences to cause the generative AI model to generate a response. The generated response is sent to the terminal and presented to the user.
[0077] The terminal includes a display device equipped with a microphone and speaker, which receives the user's voice and converts it to text using speech recognition technology. This text data is sent to a server using a secure protocol. The response received from the server is displayed visually through a virtual avatar. The avatar is designed to change its facial expressions and movements in response to the conversation, creating a sense of familiarity with the user.
[0078] Users can interact with this system in a normal conversational manner. They provide information about their health and daily life during the conversation, and this information is reflected in their user profile. This allows them to receive personalized responses based on their past data.
[0079] As a concrete example, if a user says to the device, "I'm worried because I haven't been getting enough exercise lately," speech recognition technology will convert this statement into text and send it to the server. The server will use a generative AI model to generate a response such as, "Okay, let's think together about an easy exercise plan to get you moving," and send it to the device. The device will then present this response to the user through a virtual avatar.
[0080] As an example of a prompt, the input to a generative AI model would look like this:
[0081] "User comment: I'm worried because I haven't been getting enough exercise lately. System response: Well, let's work together to come up with an easy exercise plan to get you moving."
[0082] In this way, the system monitors the user's health while functioning as a conversational partner in their daily life.
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The user emits a voice message through the device's microphone. This voice serves as input, conveying the user's intentions and state.
[0086] Step 2:
[0087] The device processes the audio received from the user using a speech recognition engine and converts it into text format. As a result, the audio data is output in a structured form as text data.
[0088] Step 3:
[0089] The terminal sends the converted text data to the server. The transmitted text serves as input for parsing and response generation.
[0090] Step 4:
[0091] The server analyzes the received text data and uses natural language processing algorithms to understand the appropriate context. The analysis results are output as prompts to the generative AI model.
[0092] Step 5:
[0093] The server inputs a prompt sentence into the generative AI model. The model generates a conversational response based on this input and outputs that response as text.
[0094] Step 6:
[0095] The server sends the generated response text to the terminal. This response is the output necessary to complete the interaction with the user.
[0096] Step 7:
[0097] The device presents received responses to the user visually and audibly using a virtual avatar. The avatar displays facial expressions and movements in accordance with the conversation, providing the user with a natural dialogue experience.
[0098] Step 8:
[0099] The server updates the user profile based on the user's utterances and conversation history with the system. This process prepares the system for more personalized conversations in the future.
[0100] (Application Example 1)
[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0102] In today's real-world retail environments, there is a demand for appropriate customer service and improved customer satisfaction. However, due to limitations in human resources and the difficulty of responding to individual needs, the current situation is generally limited to providing a uniform service. Furthermore, since there is no system in place that can respond immediately on the field based on individual health conditions and needs, there is an urgent need to realize flexible and effective services that meet diverse customer needs.
[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0104] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health status; means for notifying an external health support service when the user's health status meets certain conditions; means for acquiring information via a portable terminal worn by a person in real space and performing voice analysis; and means for suggesting appropriate products and services based on the person's statements. This makes it possible to analyze information emitted by customers in real time and provide services tailored to individual needs.
[0105] A "generative model" is a machine learning-based software program designed to generate natural, two-way conversations with users.
[0106] A "display device" is an electronic device that displays a virtual person to provide visual feedback to the user.
[0107] "Two-way conversation" refers to a form of communication where both the user and the system are interactive and respond to each other.
[0108] "Health status" refers to information that indicates the user's physical and mental condition.
[0109] "External health support services" refer to third-party organizations or systems that support users' health management.
[0110] A "portable terminal" is a device that is portable and designed for communication and information processing.
[0111] "Speech analysis" is a technology that analyzes speech data based on textual information and content.
[0112] "Proposing products or services" refers to the act of recommending appropriate products or services based on the user's needs.
[0113] To realize the system of this invention, a configuration based on the division of roles between the server, terminal, and user is necessary. The server functions as a central processing unit for generating natural, two-way conversations with the user using a generative AI model. The server receives voice data containing text information from the user, analyzes its content, and generates a response message. The software used includes a machine learning library for operating the AI model.
[0114] The device functions as a portable display device, recognizing the user's voice and converting it into text data. It incorporates voice recognition technology and transmits the collected data to a server. It also visually displays the generated messages as a virtual person, interacting with the user. In this configuration, the device can also function as smart glasses or other wearable devices.
[0115] Users, acting as the subjects who make statements and ask questions through this system, input information via voice. The user's voice input may include information related to everyday in-store conversations or purchase preferences. For example, if a user says, "My skin has been in bad condition lately...", the system receives and analyzes that statement and suggests products and services.
[0116] As a concrete example, in a physical store, if a customer says, "I'm looking for a new mystery novel," the system will guide them by saying, "We have the latest mystery novels on this shelf." This allows for real-time analysis of customer information and the provision of services tailored to individual needs. An example of a prompt used in this case would be: "Customer: I'm looking for a new mystery novel\nEmployee: ">
[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0118] Step 1:
[0119] The terminal acquires the user's voice input. Using speech recognition technology, it converts the acquired voice data into text and sends it to the server. In this process, the input is voice data, and the output is string data. The process involves analyzing the voice using speech recognition software and converting it into a string.
[0120] Step 2:
[0121] The server receives string data sent from the terminal. The received text is input into a generative AI model to generate a response based on the user's context and intent. In this step, the input is string data, and the output is a response message. The generative AI model parses the input text and uses prompts to generate an appropriate response.
[0122] Step 3:
[0123] The terminal receives a response message from the server. The received message is displayed using a virtual human avatar. The visualized information is used to adjust the avatar's mouth and facial expressions, and a conversation takes place with the user. The input for this step is the response message, and the output is visual feedback to the user. Virtual human software is used to render the received data for avatar display.
[0124] Step 4:
[0125] The user can ask further questions or make comments based on the displayed information. These comments trigger the next dialogue, returning to step 1. The user's input is voice, and the output generates the next voice input, which the system receives and starts the next cycle.
[0126] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0127] This invention embodies a system that provides a deeper, more personalized, interactive conversational experience with users by using a generative model that combines an emotion engine. This system consists of three parties: a server, a terminal (including a display device), and the user.
[0128] System Configuration
[0129] The server is the core unit that operates the generative model and the emotion engine. In addition to generating regular text responses, the generative model has the ability to generate more emotionally sensitive responses by utilizing user emotion data obtained from the emotion engine. This enables support and suggestions tailored to the user's emotional state.
[0130] The device has voice recognition and emotion analysis sensors, and it converts the user's voice and facial expressions into text and data, which it sends to the server. It also displays the server's response using a virtual character, engaging in visual and auditory interaction with the user. This avatar reflects the user's emotional state and changes its facial expressions appropriately.
[0131] Users can interact with the system in a way that feels similar to everyday communication. Because it provides support not only for the user's health but also for their emotions and mood, it helps them lead a more fulfilling life. The system accumulates daily input from users and analyzes it as emotional trends to generate personalized suggestions and alerts tailored to each user.
[0132] Program processing
[0133] The server processes the voice and facial expression data received from the user through speech recognition and sentiment analysis. The voice data is converted into text data, and the sentiment engine infers the emotional state from the text and facial expressions. For example, if the user says, "I'm feeling a little down today," the server generates an emotion tag such as "sadness" through the sentiment engine.
[0134] The server then uses a generative model to generate emotionally sensitive responses. For example, it might offer the user a comforting suggestion such as, "I see, is there anything I can do to help? Shall we look for something to help you relax?"
[0135] The device expresses the received response through the avatar's facial expressions and voice, providing feedback to the user. The avatar displays a gentle expression that matches the user's emotions, creating a friendly and approachable atmosphere.
[0136] This system allows users to receive not only health management but also psychological care, thus alleviating daily feelings of loneliness and providing support for leading a fulfilling life.
[0137] The following describes the processing flow.
[0138] Step 1:
[0139] The user turns on the TV, and the device starts up. The device activates its voice recognition and emotion analysis sensors and enters standby mode.
[0140] Step 2:
[0141] The user speaks to the device, saying, "I feel kind of tired today." The device captures this audio and converts it into text data using a speech recognition engine.
[0142] Step 3:
[0143] The device acquires user facial expression data via an emotion analysis sensor along with voice input, and sends both sets of data to the server.
[0144] Step 4:
[0145] The server inputs the received text data into the generative model and the facial expression data into the emotion engine. The emotion engine evaluates the user's emotional state as "fatigue."
[0146] Step 5:
[0147] The server integrates the emotion evaluation results from the emotion engine into a generation model to generate responses that take the user's emotions into consideration. For example, it might generate a response such as, "You've been working hard lately, it's important to take a break."
[0148] Step 6:
[0149] The server sends the generated response to the terminal. The terminal controls a virtual character avatar and displays the generated response and a calm expression to the user.
[0150] Step 7:
[0151] The user responds, "Thank you, I think I'll take a short rest." This audio is also converted to text by the device and sent back to the server.
[0152] Step 8:
[0153] The server stores a history of user emotions and responses in a database, which serves as foundational data for analyzing long-term emotional trends.
[0154] Step 9:
[0155] Once the conversation ends and the user turns off the TV, the device stops all processes and returns to a standby state for the next use.
[0156] (Example 2)
[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0158] Conventional two-way conversation systems have struggled to accurately analyze users' emotional states and provide appropriate dialogue and psychological support accordingly. Furthermore, they lacked personalized suggestions based on users' health and psychological conditions, as well as notifications to external support services.
[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0160] In this invention, the server includes means for enabling two-way conversations that respond to the user's emotional state using a generative model and emotion analysis means, means for converting the user's voice and facial expressions into text and emotion data, storing and analyzing them, and means for generating a response that includes psychological support when the user's emotional state meets certain conditions, and notifying an external support service. This makes it possible to provide a conversational experience that is attentive to the user's emotions and individually optimized health and psychological support.
[0161] A "generative model" is an algorithm or program used to generate responses or content based on input data.
[0162] "Emotional analysis methods" refer to technologies and systems that analyze user voice and facial expression data to understand their emotional state.
[0163] A "virtual character" is a character or avatar that is displayed on a display device and used to interact with the user.
[0164] "Two-way conversation" is a method of communication in which the user and the system take turns speaking and responding.
[0165] "Psychological support" refers to the act or system of providing support and care that takes into account the user's emotions and psychological state.
[0166] "External support services" refer to organizations or platforms that exist outside the system and provide additional support or assistance to users.
[0167] A "condition" is an element or premise necessary for a particular situation or state to occur.
[0168] An "alert" is a notification or warning that draws attention to the user.
[0169] The system of the present invention is composed of three main components: a server, a terminal, and a user, and provides an individualized conversational experience based on the user's emotional state.
[0170] The server plays a role in enabling two-way conversations with users using a generative AI model and sentiment analysis tools. The generative AI model generates responses based on data sent by the user, while the sentiment analysis tools infer emotions from voice and facial expression data. Specifically, voice is converted into text data by speech recognition, and sentiment tags are generated from this text and facial expression data. For example, if a user shares something like "I'm sad today," the server receives this, generates the sentiment tag "sadness," and returns a thoughtful response accordingly.
[0171] The device is equipped with voice recognition and emotion sensors to detect the user's voice input and facial expressions. This allows data to be transmitted to the server in real time. The device also displays responses sent from the server as avatars and voices, enabling friendly, two-way dialogue with the user. The avatar can change its expression according to the user's emotional state, for example, by expressing an emotionally sensitive offer in a quiet voice, such as, "Is there anything I can help you with?"
[0172] Users can interact with this system in the same way they interact with their daily lives. Data on their daily physical and mental health is accumulated, and the server can use this data to provide personalized advice and warnings. As a result, users can receive psychological care and support to improve their quality of life.
[0173] For example, if a user says, "I'm very tired today," the server performs sentiment analysis and assigns the emotion tag "fatigue." The generative AI model then generates a response such as, "Why don't you try taking a break early? Let's find a way to relax." An example of a prompt might be, "What approach would you take to generate a response when a user is feeling tired?"
[0174] This system will allow users to enjoy a more fulfilling life while receiving emotionally sensitive support.
[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0176] Step 1:
[0177] The user begins a conversation into the device. The user's voice and facial expressions are captured in real time by the device's voice recognition and emotion sensing sensors. This input includes voice data and facial expression data. The device then prepares this data to send to the server.
[0178] Step 2:
[0179] The terminal converts captured audio data into text data using speech recognition technology. During this process, digital signal processing of the audio waveform generates text in string format. Simultaneously, it packages the facial expression data obtained from the emotion sensor, formatting it appropriately for transmission to the server. The output consists of the converted text string and the prepared facial expression data.
[0180] Step 3:
[0181] The server receives text data and facial expression data sent from the terminal. First, the text data is passed to a natural language processing engine to understand its meaning and context. This process lays the foundation for sentiment analysis. Next, the facial expression data is passed through the sentiment engine to infer the user's emotional state. For example, if the user says "I'm tired," the server generates the sentiment tag "fatigue" from the audio and facial expression. This output is used as input to a generative AI model.
[0182] Step 4:
[0183] The server uses a generative AI model to generate an emotionally sensitive response based on the previously obtained emotion tags and input data. This prompt is set to "How can we provide support tailored to the user's level of fatigue?" The generative model creates an appropriate message that matches the user's emotional state by constructing a suggestion such as "Why don't you take a break today?" The resulting output is a response message.
[0184] Step 5:
[0185] The terminal receives a response message sent from the server and prepares to display it using a virtual person (avatar). The avatar uses this message to play back voice in an expressive and persuasive manner as feedback to the user. Here, the avatar's facial expression is set to be gentle, reflecting the user's fatigue, providing a pleasant response both visually and aurally.
[0186] (Application Example 2)
[0187] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0188] In modern society, there is a demand for support that takes into account individual emotions and health conditions. However, conventional technology makes it difficult to accurately grasp users' emotions and provide appropriate responses immediately. Therefore, there is a need for new means to provide customer-specific services and suggestions in physical stores and other face-to-face service settings, thereby improving customer satisfaction.
[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0190] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health and emotional state; means for providing external support or suggestion services when the user's health or emotional state meets certain conditions; and means for rendering a person according to the emotion using one or more representation means and generating a conversation that matches the user's emotions. This makes it possible to provide personalized responses according to the user's individual emotions and state, and to immediately provide appropriate support and suggestions.
[0191] A "generative model" is an algorithm that generates text and responses based on user input data, enabling natural conversations with users.
[0192] A "display device" is a device that includes digital displays and screens used to visually present virtual people or information to a user.
[0193] A "virtual character" is a digital avatar that has a human-like appearance and movements, enabling visual and auditory interaction with the user.
[0194] "Two-way conversation" is a form of interactive communication in which the user provides input to the system, and the system returns a corresponding response.
[0195] "Health status" refers to information indicating the user's physical and mental well-being, and includes various data used for health management and support.
[0196] "Emotional state" refers to the user's psychological state, including emotions such as joy, anger, sadness, and happiness, inferred from their facial expressions and voice.
[0197] "External support or suggested services" refer to third-party services or suggestions provided according to the user's health and emotional state, intended to assist in improving the user's quality of life and solving problems.
[0198] In the system that implements this application example, the server uses a generative model and an emotion engine to create a two-way conversational experience with the user. The terminal is equipped with speech recognition and emotion analysis sensors, and acquires voice input and facial expression data from the user and sends it to the server. The server uses a speech recognition library to convert the voice into text data and processes that text with a generative model. It also uses an emotion analysis module to analyze the user's emotional state and uses the results as prompts incorporated into the generative AI model.
[0199] The server uses an avatar display module to render a virtual person with facial expressions corresponding to the user's emotional state on the display device, and provides the generated response to the user visually and audibly. This allows the user to receive personalized support through the system that takes into account their psychological well-being and health status.
[0200] For example, if a customer in a store says, "I'm feeling a little down today," the server performs sentiment analysis and infers the emotional state as "sadness." The generative model then generates a response that takes that emotion into consideration, suggesting something like, "Is there anything I can do to help?" through a virtual character. The avatar interacts with the user with a gentle expression, creating a comfortable communication experience.
[0201] An example of a prompt for a generative AI model is, "If the user is complaining of being a little tired, generate appropriate rest suggestions." Based on this prompt, the generative model will provide the user with the most suitable suggestions.
[0202] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0203] Step 1:
[0204] The terminal acquires voice input and facial expression data from the user. Using a voice recognition sensor and an emotion analysis sensor, the user's speech and facial expressions are converted into digital data and prepared as voice data and image data. This becomes the input data to the server.
[0205] Step 2:
[0206] The server converts the audio data received from the terminal into text data using a speech recognition library. This process analyzes the audio waveform, converts it into appropriate strings, and extracts the user's actual spoken content. The accuracy of speech recognition is crucial, as it provides the foundational data for the conversation.
[0207] Step 3:
[0208] The server inputs the acquired text and image data into an emotion analysis module to identify the user's emotional state. The emotion analysis processes the data based on the vocabulary in the text and the facial features in the images, generating emotion tags such as "joy," "sadness," and "anger." This emotional information then serves as a prompt for the generating AI model.
[0209] Step 4:
[0210] The server inputs text data and sentiment tags into a generative model to generate personalized responses that take the user's emotions into consideration. In this process, the generative AI model creates natural language responses based on prompt sentences and prepares suggestions and questions appropriate to the user's state.
[0211] Step 5:
[0212] The server sends the generated response to the terminal, which uses an avatar display module to render a virtual person on the screen. The avatar displays facial expressions corresponding to the user's emotions and plays the generated response aloud. This allows the user to receive responses both visually and aurally, resulting in a more interactive experience.
[0213] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0214] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0215] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0216] [Second Embodiment]
[0217] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0218] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0219] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0220] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0221] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0222] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0223] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0224] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0225] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0226] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0227] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0228] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0229] This invention provides a system that enables two-way communication with a user through a display device. This system mainly consists of a server, a terminal (a display device such as a television), and a user.
[0230] System Configuration
[0231] The server includes a processing unit for running the generative model. The generative model can simulate natural conversations with users. This model generates responses based on text data sent by the user, enabling two-way communication.
[0232] The device uses speech recognition technology to convert the user's voice input into text and sends it to the server. Furthermore, it receives responses from the server and displays a virtual person avatar to provide visual feedback to the user. The device may also ask daily questions about the user's health status, triggered by user input.
[0233] Users interact with the system in the same way as in a normal conversation. During this interaction, they can input information about their health and daily life events via voice. This information is sent to the server through the device and used to build the user profile.
[0234] Program processing
[0235] The server analyzes the text data received from the user and uses a generative model to generate appropriate conversational responses. For example, if a user says, "I'm not feeling well today," the server refers to the user's daily data and, if necessary, generates a more humane response such as, "I'm worried about your health lately. Please get plenty of rest," and sends it to the terminal.
[0236] The device displays the generated response using a virtual person avatar. This avatar is designed to create a sense of familiarity with the user by moving its mouth and changing its facial expressions. In addition, the device uses speech recognition technology to transcribe the user's speech into text and sends it to the server, making it easy to use naturally in everyday situations without burdening the user.
[0237] Through this system, users can not only alleviate feelings of loneliness in their daily lives but also receive support in understanding their own health status. The system evolves on its own, providing communication optimized for each user. For example, if a user says to the system, "I'm worried about not getting enough exercise lately," the system can respond, "Okay, let's work together to come up with an easy exercise plan to get you moving."
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] The user turns on the TV, and the device displays the initial setup screen. The device starts a session with the server and establishes a communication channel.
[0241] Step 2:
[0242] The device activates its voice recognition function and waits for the user to speak. When the user says "Hello," the device converts the voice input into text data.
[0243] Step 3:
[0244] The terminal sends the user's text data to the server. The server uses a generative model to generate an appropriate response to the user's "Hello," creating a reply such as "How are you doing today?"
[0245] Step 4:
[0246] The server generates a text response and sends it to the terminal. The terminal uses a virtual person avatar to play the generated response aloud, with mouth movements.
[0247] Step 5:
[0248] The user replies, "I'm a little tired today." The device converts the speech back into text and sends it to the server.
[0249] Step 6:
[0250] The server references past user data and analyzes changes in health status. Using a generative model, it generates a suggestion such as, "That sounds tough, why don't you take a break?"
[0251] Step 7:
[0252] The server sends a suggestion to the terminal. The terminal communicates the suggestion to the user by controlling an avatar. This prompts the user to take appropriate action.
[0253] Step 8:
[0254] The user indicates their intention to end the conversation. If they say, "Thank you, I'm fine now," the device notifies the server that the session has ended.
[0255] Step 9:
[0256] The terminal displays the termination screen and closes the application. The server safely terminates the communication session and saves the log data to the database.
[0257] (Example 1)
[0258] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0259] In modern society, while interest in personal health management is increasing, systems for monitoring health status at home and reducing feelings of isolation are still not adequately developed. In particular, there is a growing need for interactive systems that allow users to understand their own health status through natural conversations at home and receive necessary advice and warnings. Such systems need to support users' daily health management and have features to ensure that urgent health conditions are not overlooked.
[0260] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0261] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for receiving voice input and converting it into a string using speech recognition technology; and means for an analysis device to process the string, generate prompt sentences, and input them into the generative model. This enables the user to monitor their health status through natural conversation and receive quick and appropriate feedback.
[0262] A "generative model" is an artificial intelligence algorithm that generates an appropriate response based on a given input text.
[0263] A "display device" is a device used to present visual information, including virtual character avatars, to a user.
[0264] "Speech recognition technology" is a technology that converts a user's voice into text data.
[0265] A "prompt" is a text-based question or instruction that is input into a generative model.
[0266] An "analysis device" is a device that processes text data and generates prompts adapted to a generative model.
[0267] A "user profile" is an aggregate of personal information built based on a user's attributes and past conversation history.
[0268] "External health-related services" refer to third-party services that utilize users' health information to provide support and advice.
[0269] An "avatar" is a virtual character displayed to visually represent a two-way conversation with the user.
[0270] This invention constitutes a system that provides users with a natural and interactive conversational experience. The system is primarily operated by a server, a terminal, and the user.
[0271] The server is a computing device for running the generative AI model and receives data transcribed into text using speech recognition technology. This data is processed by an analysis device, which generates prompt sentences to be input into the generative model. The server uses these prompt sentences to cause the generative AI model to generate a response. The generated response is sent to the terminal and presented to the user.
[0272] The terminal includes a display device equipped with a microphone and speaker, which receives the user's voice and converts it to text using speech recognition technology. This text data is sent to a server using a secure protocol. The response received from the server is displayed visually through a virtual avatar. The avatar is designed to change its facial expressions and movements in response to the conversation, creating a sense of familiarity with the user.
[0273] Users can interact with this system in a normal conversational manner. They provide information about their health and daily life during the conversation, and this information is reflected in their user profile. This allows them to receive personalized responses based on their past data.
[0274] As a concrete example, if a user says to the device, "I'm worried because I haven't been getting enough exercise lately," speech recognition technology will convert this statement into text and send it to the server. The server will use a generative AI model to generate a response such as, "Okay, let's think together about an easy exercise plan to get you moving," and send it to the device. The device will then present this response to the user through a virtual avatar.
[0275] As an example of a prompt, the input to a generative AI model would look like this:
[0276] "User comment: I'm worried because I haven't been getting enough exercise lately. System response: Well, let's work together to come up with an easy exercise plan to get you moving."
[0277] In this way, while monitoring the user's health condition, the system functions as a conversation partner in daily life.
[0278] The flow of the specific process in Example 1 will be described with reference to FIG. 11.
[0279] Step 1:
[0280] The user issues a voice message through the microphone of the terminal. This voice serves as an input to convey the user's intention and status.
[0281] Step 2:
[0282] The terminal processes the voice received from the user using a voice recognition engine and converts it into text format. As a result, the voice data is output in a structured form as text data.
[0283] Step 3: [[ID=X]]
[0284] The terminal transmits the converted text data to the server. The transmitted text serves as an input for analysis and response generation.
[0285] Step 4:
[0286] The server analyzes the received text data and uses natural language processing algorithms to understand the appropriate context. The result of the analysis is output as a prompt sentence to the generation AI model.
[0287] Step 5:
[0288] The server inputs the prompt sentence into the generation AI model. The model generates a response to the conversation based on this input and outputs the response as text. <000091X> Step 6:
[0290] <~ It should be noted that there seems to be an error in the original text where the tag is misspelled as <000089X> and
[0289] is misspelled as <000091X> in the provided text. The above translation is based on the corrected version for the purpose of maintaining consistency in the translation process. If the original text has specific requirements regarding these misspellings, the translation may need to be adjusted accordingly.The server sends the generated response text to the terminal. This response is the output necessary to complete the interaction with the user.
[0291] Step 7:
[0292] The device presents received responses to the user visually and audibly using a virtual avatar. The avatar displays facial expressions and movements in accordance with the conversation, providing the user with a natural dialogue experience.
[0293] Step 8:
[0294] The server updates the user profile based on the user's utterances and conversation history with the system. This process prepares the system for more personalized conversations in the future.
[0295] (Application Example 1)
[0296] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0297] In today's real-world retail environments, there is a demand for appropriate customer service and improved customer satisfaction. However, due to limitations in human resources and the difficulty of responding to individual needs, the current situation is generally limited to providing a uniform service. Furthermore, since there is no system in place that can respond immediately on the field based on individual health conditions and needs, there is an urgent need to realize flexible and effective services that meet diverse customer needs.
[0298] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0299] In this invention, the server includes means for drawing a virtual person on a display device using a generation model to enable two-way conversation, means for acquiring, storing, and analyzing information regarding the user's health status, means for notifying an external health support service when the user's health status meets certain conditions, means for acquiring information via a portable terminal worn by a person in the real space and performing voice analysis, and means for proposing appropriate products and services based on the person's speech. Thereby, it becomes possible to analyze in real time the information issued by the customer and provide services according to individual needs.
[0300] The "generation model" is a software program based on machine learning for generating natural two-way conversation with the user.
[0301] The "display device" is an electronic device for drawing a virtual person and providing visual feedback to the user.
[0302] "Two-way conversation" is an interactive communication form in which both the user and the system respond to each other.
[0303] The "health status" is information indicating the physical and mental state of the user.
[0304] The "external health support service" is a third-party organization or system for supporting the user's health management.
[0305] The "portable terminal" is a device with a shape that can be carried by a person and is used for communication and information processing.
[0306] "Voice analysis" is a technology for analyzing voice data based on character information and content.
[0307] "Proposing products and services" is an act of recommending appropriate products or services based on the needs of the user.
[0308] To realize the system of this invention, a configuration based on the division of roles between the server, terminal, and user is necessary. The server functions as a central processing unit for generating natural, two-way conversations with the user using a generative AI model. The server receives voice data containing text information from the user, analyzes its content, and generates a response message. The software used includes a machine learning library for operating the AI model.
[0309] The device functions as a portable display device, recognizing the user's voice and converting it into text data. It incorporates voice recognition technology and transmits the collected data to a server. It also visually displays the generated messages as a virtual person, interacting with the user. In this configuration, the device can also function as smart glasses or other wearable devices.
[0310] Users, acting as the subjects who make statements and ask questions through this system, input information via voice. The user's voice input may include information related to everyday in-store conversations or purchase preferences. For example, if a user says, "My skin has been in bad condition lately...", the system receives and analyzes that statement and suggests products and services.
[0311] As a concrete example, in a physical store, if a customer says, "I'm looking for a new mystery novel," the system will guide them by saying, "We have the latest mystery novels on this shelf." This allows for real-time analysis of customer information and the provision of services tailored to individual needs. An example of a prompt used in this case would be: "Customer: I'm looking for a new mystery novel\nEmployee: ">
[0312] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0313] Step 1:
[0314] The terminal acquires the user's voice input. Using speech recognition technology, it converts the acquired voice data into text and sends it to the server. In this process, the input is voice data, and the output is string data. The process involves analyzing the voice using speech recognition software and converting it into a string.
[0315] Step 2:
[0316] The server receives string data sent from the terminal. The received text is input into a generative AI model to generate a response based on the user's context and intent. In this step, the input is string data, and the output is a response message. The generative AI model parses the input text and uses prompts to generate an appropriate response.
[0317] Step 3:
[0318] The terminal receives a response message from the server. The received message is displayed using a virtual human avatar. The visualized information is used to adjust the avatar's mouth and facial expressions, and a conversation takes place with the user. The input for this step is the response message, and the output is visual feedback to the user. Virtual human software is used to render the received data for avatar display.
[0319] Step 4:
[0320] The user can ask further questions or make comments based on the displayed information. These comments trigger the next dialogue, returning to step 1. The user's input is voice, and the output generates the next voice input, which the system receives and starts the next cycle.
[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0322] This invention embodies a system that provides a deeper, more personalized, interactive conversational experience with users by using a generative model that combines an emotion engine. This system consists of three parties: a server, a terminal (including a display device), and the user.
[0323] System Configuration
[0324] The server is the core unit that operates the generative model and the emotion engine. In addition to generating regular text responses, the generative model has the ability to generate more emotionally sensitive responses by utilizing user emotion data obtained from the emotion engine. This enables support and suggestions tailored to the user's emotional state.
[0325] The device has voice recognition and emotion analysis sensors, and it converts the user's voice and facial expressions into text and data, which it sends to the server. It also displays the server's response using a virtual character, engaging in visual and auditory interaction with the user. This avatar reflects the user's emotional state and changes its facial expressions appropriately.
[0326] Users can interact with the system in a way that feels similar to everyday communication. Because it provides support not only for the user's health but also for their emotions and mood, it helps them lead a more fulfilling life. The system accumulates daily input from users and analyzes it as emotional trends to generate personalized suggestions and alerts tailored to each user.
[0327] Program processing
[0328] The server processes the voice and facial expression data received from the user through speech recognition and sentiment analysis. The voice data is converted into text data, and the sentiment engine infers the emotional state from the text and facial expressions. For example, if the user says, "I'm feeling a little down today," the server generates an emotion tag such as "sadness" through the sentiment engine.
[0329] The server then uses a generative model to generate emotionally sensitive responses. For example, it might offer the user a comforting suggestion such as, "I see, is there anything I can do to help? Shall we look for something to help you relax?"
[0330] The device expresses the received response through the avatar's facial expressions and voice, providing feedback to the user. The avatar displays a gentle expression that matches the user's emotions, creating a friendly and approachable atmosphere.
[0331] This system allows users to receive not only health management but also psychological care, thus alleviating daily feelings of loneliness and providing support for leading a fulfilling life.
[0332] The following describes the processing flow.
[0333] Step 1:
[0334] The user turns on the TV, and the device starts up. The device activates its voice recognition and emotion analysis sensors and enters standby mode.
[0335] Step 2:
[0336] The user speaks to the device, saying, "I feel kind of tired today." The device captures this audio and converts it into text data using a speech recognition engine.
[0337] Step 3:
[0338] The device acquires user facial expression data via an emotion analysis sensor along with voice input, and sends both sets of data to the server.
[0339] Step 4:
[0340] The server inputs the received text data into the generative model and the facial expression data into the emotion engine. The emotion engine evaluates the user's emotional state as "fatigue."
[0341] Step 5:
[0342] The server integrates the emotion evaluation results from the emotion engine into a generation model to generate responses that take the user's emotions into consideration. For example, it might generate a response such as, "You've been working hard lately, it's important to take a break."
[0343] Step 6:
[0344] The server sends the generated response to the terminal. The terminal controls a virtual character avatar and displays the generated response and a calm expression to the user.
[0345] Step 7:
[0346] The user responds, "Thank you, I think I'll take a short rest." This audio is also converted to text by the device and sent back to the server.
[0347] Step 8:
[0348] The server stores a history of user emotions and responses in a database, which serves as foundational data for analyzing long-term emotional trends.
[0349] Step 9:
[0350] Once the conversation ends and the user turns off the TV, the device stops all processes and returns to a standby state for the next use.
[0351] (Example 2)
[0352] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0353] Conventional two-way conversation systems have struggled to accurately analyze users' emotional states and provide appropriate dialogue and psychological support accordingly. Furthermore, they lacked personalized suggestions based on users' health and psychological conditions, as well as notifications to external support services.
[0354] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0355] In this invention, the server includes means for enabling two-way conversations that respond to the user's emotional state using a generative model and emotion analysis means, means for converting the user's voice and facial expressions into text and emotion data, storing and analyzing them, and means for generating a response that includes psychological support when the user's emotional state meets certain conditions, and notifying an external support service. This makes it possible to provide a conversational experience that is attentive to the user's emotions and individually optimized health and psychological support.
[0356] A "generative model" is an algorithm or program used to generate responses or content based on input data.
[0357] "Emotional analysis methods" refer to technologies and systems that analyze user voice and facial expression data to understand their emotional state.
[0358] A "virtual character" is a character or avatar that is displayed on a display device and used to interact with the user.
[0359] "Two-way conversation" is a method of communication in which the user and the system take turns speaking and responding.
[0360] "Psychological support" refers to the act or system of providing support and care that takes into account the user's emotions and psychological state.
[0361] "External support services" refer to organizations or platforms that exist outside the system and provide additional support or assistance to users.
[0362] A "condition" is an element or premise necessary for a particular situation or state to occur.
[0363] An "alert" is a notification or warning that draws attention to the user.
[0364] The system of the present invention is composed of three main components: a server, a terminal, and a user, and provides an individualized conversational experience based on the user's emotional state.
[0365] The server plays a role in enabling two-way conversations with users using a generative AI model and sentiment analysis tools. The generative AI model generates responses based on data sent by the user, while the sentiment analysis tools infer emotions from voice and facial expression data. Specifically, voice is converted into text data by speech recognition, and sentiment tags are generated from this text and facial expression data. For example, if a user shares something like "I'm sad today," the server receives this, generates the sentiment tag "sadness," and returns a thoughtful response accordingly.
[0366] The device is equipped with voice recognition and emotion sensors to detect the user's voice input and facial expressions. This allows data to be transmitted to the server in real time. The device also displays responses sent from the server as avatars and voices, enabling friendly, two-way dialogue with the user. The avatar can change its expression according to the user's emotional state, for example, by expressing an emotionally sensitive offer in a quiet voice, such as, "Is there anything I can help you with?"
[0367] Users can interact with this system in the same way they interact with their daily lives. Data on their daily physical and mental health is accumulated, and the server can use this data to provide personalized advice and warnings. As a result, users can receive psychological care and support to improve their quality of life.
[0368] For example, if a user says, "I'm very tired today," the server performs sentiment analysis and assigns the emotion tag "fatigue." The generative AI model then generates a response such as, "Why don't you try taking a break early? Let's find a way to relax." An example of a prompt might be, "What approach would you take to generate a response when a user is feeling tired?"
[0369] This system will allow users to enjoy a more fulfilling life while receiving emotionally sensitive support.
[0370] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0371] Step 1:
[0372] The user begins a conversation into the device. The user's voice and facial expressions are captured in real time by the device's voice recognition and emotion sensing sensors. This input includes voice data and facial expression data. The device then prepares this data to send to the server.
[0373] Step 2:
[0374] The terminal converts captured audio data into text data using speech recognition technology. During this process, digital signal processing of the audio waveform generates text in string format. Simultaneously, it packages the facial expression data obtained from the emotion sensor, formatting it appropriately for transmission to the server. The output consists of the converted text string and the prepared facial expression data.
[0375] Step 3:
[0376] The server receives text data and facial expression data sent from the terminal. First, the text data is passed to a natural language processing engine to understand its meaning and context. This process lays the foundation for sentiment analysis. Next, the facial expression data is passed through the sentiment engine to infer the user's emotional state. For example, if the user says "I'm tired," the server generates the sentiment tag "fatigue" from the audio and facial expression. This output is used as input to a generative AI model.
[0377] Step 4:
[0378] The server uses a generative AI model to generate an emotionally sensitive response based on the previously obtained emotion tags and input data. This prompt is set to "How can we provide support tailored to the user's level of fatigue?" The generative model creates an appropriate message that matches the user's emotional state by constructing a suggestion such as "Why don't you take a break today?" The resulting output is a response message.
[0379] Step 5:
[0380] The terminal receives a response message sent from the server and prepares to display it using a virtual person (avatar). The avatar uses this message to play back voice in an expressive and persuasive manner as feedback to the user. Here, the avatar's facial expression is set to be gentle, reflecting the user's fatigue, providing a pleasant response both visually and aurally.
[0381] (Application Example 2)
[0382] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0383] In modern society, there is a demand for support that takes into account individual emotions and health conditions. However, conventional technology makes it difficult to accurately grasp users' emotions and provide appropriate responses immediately. Therefore, there is a need for new means to provide customer-specific services and suggestions in physical stores and other face-to-face service settings, thereby improving customer satisfaction.
[0384] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0385] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health and emotional state; means for providing external support or suggestion services when the user's health or emotional state meets certain conditions; and means for rendering a person according to the emotion using one or more representation means and generating a conversation that matches the user's emotions. This makes it possible to provide personalized responses according to the user's individual emotions and state, and to immediately provide appropriate support and suggestions.
[0386] A "generative model" is an algorithm that generates text and responses based on user input data, enabling natural conversations with users.
[0387] A "display device" is a device that includes digital displays and screens used to visually present virtual people or information to a user.
[0388] A "virtual character" is a digital avatar that has a human-like appearance and movements, enabling visual and auditory interaction with the user.
[0389] "Two-way conversation" is a form of interactive communication in which the user provides input to the system, and the system returns a corresponding response.
[0390] "Health status" refers to information indicating the user's physical and mental well-being, and includes various data used for health management and support.
[0391] "Emotional state" refers to the user's psychological state, including emotions such as joy, anger, sadness, and happiness, inferred from their facial expressions and voice.
[0392] "External support or suggested services" refer to third-party services or suggestions provided according to the user's health and emotional state, intended to assist in improving the user's quality of life and solving problems.
[0393] In the system that implements this application example, the server uses a generative model and an emotion engine to create a two-way conversational experience with the user. The terminal is equipped with speech recognition and emotion analysis sensors, and acquires voice input and facial expression data from the user and sends it to the server. The server uses a speech recognition library to convert the voice into text data and processes that text with a generative model. It also uses an emotion analysis module to analyze the user's emotional state and uses the results as prompts incorporated into the generative AI model.
[0394] The server uses an avatar display module to render a virtual person with facial expressions corresponding to the user's emotional state on the display device, and provides the generated response to the user visually and audibly. This allows the user to receive personalized support through the system that takes into account their psychological well-being and health status.
[0395] For example, if a customer in a store says, "I'm feeling a little down today," the server performs sentiment analysis and infers the emotional state as "sadness." The generative model then generates a response that takes that emotion into consideration, suggesting something like, "Is there anything I can do to help?" through a virtual character. The avatar interacts with the user with a gentle expression, creating a comfortable communication experience.
[0396] An example of a prompt for a generative AI model is, "If the user is complaining of being a little tired, generate appropriate rest suggestions." Based on this prompt, the generative model will provide the user with the most suitable suggestions.
[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0398] Step 1:
[0399] The terminal acquires voice input and facial expression data from the user. Using a voice recognition sensor and an emotion analysis sensor, the user's speech and facial expressions are converted into digital data and prepared as voice data and image data. This becomes the input data to the server.
[0400] Step 2:
[0401] The server converts the audio data received from the terminal into text data using a speech recognition library. This process analyzes the audio waveform, converts it into appropriate strings, and extracts the user's actual spoken content. The accuracy of speech recognition is crucial, as it provides the foundational data for the conversation.
[0402] Step 3:
[0403] The server inputs the acquired text and image data into an emotion analysis module to identify the user's emotional state. The emotion analysis processes the data based on the vocabulary in the text and the facial features in the images, generating emotion tags such as "joy," "sadness," and "anger." This emotional information then serves as a prompt for the generating AI model.
[0404] Step 4:
[0405] The server inputs text data and sentiment tags into a generative model to generate personalized responses that take the user's emotions into consideration. In this process, the generative AI model creates natural language responses based on prompt sentences and prepares suggestions and questions appropriate to the user's state.
[0406] Step 5:
[0407] The server sends the generated response to the terminal, which uses an avatar display module to render a virtual person on the screen. The avatar displays facial expressions corresponding to the user's emotions and plays the generated response aloud. This allows the user to receive responses both visually and aurally, resulting in a more interactive experience.
[0408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0409] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0411] [Third Embodiment]
[0412] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0413] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0415] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0419] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0422] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0423] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0424] This invention provides a system that enables two-way communication with a user through a display device. This system mainly consists of a server, a terminal (a display device such as a television), and a user.
[0425] System Configuration
[0426] The server includes a processing unit for running the generative model. The generative model can simulate natural conversations with users. This model generates responses based on text data sent by the user, enabling two-way communication.
[0427] The device uses speech recognition technology to convert the user's voice input into text and sends it to the server. Furthermore, it receives responses from the server and displays a virtual person avatar to provide visual feedback to the user. The device may also ask daily questions about the user's health status, triggered by user input.
[0428] Users interact with the system in the same way as in a normal conversation. During this interaction, they can input information about their health and daily life events via voice. This information is sent to the server through the device and used to build the user profile.
[0429] Program processing
[0430] The server analyzes the text data received from the user and uses a generative model to generate appropriate conversational responses. For example, if a user says, "I'm not feeling well today," the server refers to the user's daily data and, if necessary, generates a more humane response such as, "I'm worried about your health lately. Please get plenty of rest," and sends it to the terminal.
[0431] The device displays the generated response using a virtual person avatar. This avatar is designed to create a sense of familiarity with the user by moving its mouth and changing its facial expressions. In addition, the device uses speech recognition technology to transcribe the user's speech into text and sends it to the server, making it easy to use naturally in everyday situations without burdening the user.
[0432] Through this system, users can not only alleviate feelings of loneliness in their daily lives but also receive support in understanding their own health status. The system evolves on its own, providing communication optimized for each user. For example, if a user says to the system, "I'm worried about not getting enough exercise lately," the system can respond, "Okay, let's work together to come up with an easy exercise plan to get you moving."
[0433] The following describes the processing flow.
[0434] Step 1:
[0435] The user turns on the TV, and the device displays the initial setup screen. The device starts a session with the server and establishes a communication channel.
[0436] Step 2:
[0437] The device activates its voice recognition function and waits for the user to speak. When the user says "Hello," the device converts the voice input into text data.
[0438] Step 3:
[0439] The terminal sends the user's text data to the server. The server uses a generative model to generate an appropriate response to the user's "Hello," creating a reply such as "How are you doing today?"
[0440] Step 4:
[0441] The server generates a text response and sends it to the terminal. The terminal uses a virtual person avatar to play the generated response aloud, with mouth movements.
[0442] Step 5:
[0443] The user replies, "I'm a little tired today." The device converts the speech back into text and sends it to the server.
[0444] Step 6:
[0445] The server references past user data and analyzes changes in health status. Using a generative model, it generates a suggestion such as, "That sounds tough, why don't you take a break?"
[0446] Step 7:
[0447] The server sends a suggestion to the terminal. The terminal communicates the suggestion to the user by controlling an avatar. This prompts the user to take appropriate action.
[0448] Step 8:
[0449] The user indicates their intention to end the conversation. If they say, "Thank you, I'm fine now," the device notifies the server that the session has ended.
[0450] Step 9:
[0451] The terminal displays the termination screen and closes the application. The server safely terminates the communication session and saves the log data to the database.
[0452] (Example 1)
[0453] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0454] In modern society, while interest in personal health management is increasing, systems for monitoring health status at home and reducing feelings of isolation are still not adequately developed. In particular, there is a growing need for interactive systems that allow users to understand their own health status through natural conversations at home and receive necessary advice and warnings. Such systems need to support users' daily health management and have features to ensure that urgent health conditions are not overlooked.
[0455] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0456] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for receiving voice input and converting it into a string using speech recognition technology; and means for an analysis device to process the string, generate prompt sentences, and input them into the generative model. This enables the user to monitor their health status through natural conversation and receive quick and appropriate feedback.
[0457] A "generative model" is an artificial intelligence algorithm that generates an appropriate response based on a given input text.
[0458] A "display device" is a device used to present visual information, including virtual character avatars, to a user.
[0459] "Speech recognition technology" is a technology that converts a user's voice into text data.
[0460] A "prompt" is a text-based question or instruction that is input into a generative model.
[0461] An "analysis device" is a device that processes text data and generates prompts adapted to a generative model.
[0462] A "user profile" is an aggregate of personal information built based on a user's attributes and past conversation history.
[0463] "External health-related services" refer to third-party services that utilize users' health information to provide support and advice.
[0464] An "avatar" is a virtual character displayed to visually represent a two-way conversation with the user.
[0465] This invention constitutes a system that provides users with a natural and interactive conversational experience. The system is primarily operated by a server, a terminal, and the user.
[0466] The server is a computing device for running the generative AI model and receives data transcribed into text using speech recognition technology. This data is processed by an analysis device, which generates prompt sentences to be input into the generative model. The server uses these prompt sentences to cause the generative AI model to generate a response. The generated response is sent to the terminal and presented to the user.
[0467] The terminal includes a display device equipped with a microphone and speaker, which receives the user's voice and converts it to text using speech recognition technology. This text data is sent to a server using a secure protocol. The response received from the server is displayed visually through a virtual avatar. The avatar is designed to change its facial expressions and movements in response to the conversation, creating a sense of familiarity with the user.
[0468] Users can interact with this system in a normal conversational manner. They provide information about their health and daily life during the conversation, and this information is reflected in their user profile. This allows them to receive personalized responses based on their past data.
[0469] As a concrete example, if a user says to the device, "I'm worried because I haven't been getting enough exercise lately," speech recognition technology will convert this statement into text and send it to the server. The server will use a generative AI model to generate a response such as, "Okay, let's think together about an easy exercise plan to get you moving," and send it to the device. The device will then present this response to the user through a virtual avatar.
[0470] As an example of a prompt, the input to a generative AI model would look like this:
[0471] "User comment: I'm worried because I haven't been getting enough exercise lately. System response: Well, let's work together to come up with an easy exercise plan to get you moving."
[0472] In this way, the system monitors the user's health while functioning as a conversational partner in their daily life.
[0473] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0474] Step 1:
[0475] The user emits a voice message through the device's microphone. This voice serves as input, conveying the user's intentions and state.
[0476] Step 2:
[0477] The device processes the audio received from the user using a speech recognition engine and converts it into text format. As a result, the audio data is output in a structured form as text data.
[0478] Step 3:
[0479] The terminal sends the converted text data to the server. The transmitted text serves as input for parsing and response generation.
[0480] Step 4:
[0481] The server analyzes the received text data and uses natural language processing algorithms to understand the appropriate context. The analysis results are output as prompts to the generative AI model.
[0482] Step 5:
[0483] The server inputs a prompt sentence into the generative AI model. The model generates a conversational response based on this input and outputs that response as text.
[0484] Step 6:
[0485] The server sends the generated response text to the terminal. This response is the output necessary to complete the interaction with the user.
[0486] Step 7:
[0487] The device presents received responses to the user visually and audibly using a virtual avatar. The avatar displays facial expressions and movements in accordance with the conversation, providing the user with a natural dialogue experience.
[0488] Step 8:
[0489] The server updates the user profile based on the user's utterances and conversation history with the system. This process prepares the system for more personalized conversations in the future.
[0490] (Application Example 1)
[0491] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0492] In today's real-world retail environments, there is a demand for appropriate customer service and improved customer satisfaction. However, due to limitations in human resources and the difficulty of responding to individual needs, the current situation is generally limited to providing a uniform service. Furthermore, since there is no system in place that can respond immediately on the field based on individual health conditions and needs, there is an urgent need to realize flexible and effective services that meet diverse customer needs.
[0493] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0494] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health status; means for notifying an external health support service when the user's health status meets certain conditions; means for acquiring information via a portable terminal worn by a person in real space and performing voice analysis; and means for suggesting appropriate products and services based on the person's statements. This makes it possible to analyze information emitted by customers in real time and provide services tailored to individual needs.
[0495] A "generative model" is a machine learning-based software program designed to generate natural, two-way conversations with users.
[0496] A "display device" is an electronic device that displays a virtual person to provide visual feedback to the user.
[0497] "Two-way conversation" refers to a form of communication where both the user and the system are interactive and respond to each other.
[0498] "Health status" refers to information that indicates the user's physical and mental condition.
[0499] "External health support services" refer to third-party organizations or systems that support users' health management.
[0500] A "portable terminal" is a device that is portable and designed for communication and information processing.
[0501] "Speech analysis" is a technology that analyzes speech data based on textual information and content.
[0502] "Proposing products or services" refers to the act of recommending appropriate products or services based on the user's needs.
[0503] To realize the system of this invention, a configuration based on the division of roles between the server, terminal, and user is necessary. The server functions as a central processing unit for generating natural, two-way conversations with the user using a generative AI model. The server receives voice data containing text information from the user, analyzes its content, and generates a response message. The software used includes a machine learning library for operating the AI model.
[0504] The device functions as a portable display device, recognizing the user's voice and converting it into text data. It incorporates voice recognition technology and transmits the collected data to a server. It also visually displays the generated messages as a virtual person, interacting with the user. In this configuration, the device can also function as smart glasses or other wearable devices.
[0505] Users, acting as the subjects who make statements and ask questions through this system, input information via voice. The user's voice input may include information related to everyday in-store conversations or purchase preferences. For example, if a user says, "My skin has been in bad condition lately...", the system receives and analyzes that statement and suggests products and services.
[0506] As a concrete example, in a physical store, if a customer says, "I'm looking for a new mystery novel," the system will guide them by saying, "We have the latest mystery novels on this shelf." This allows for real-time analysis of customer information and the provision of services tailored to individual needs. An example of a prompt used in this case would be: "Customer: I'm looking for a new mystery novel\nEmployee: ">
[0507] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0508] Step 1:
[0509] The terminal acquires the user's voice input. Using speech recognition technology, it converts the acquired voice data into text and sends it to the server. In this process, the input is voice data, and the output is string data. The process involves analyzing the voice using speech recognition software and converting it into a string.
[0510] Step 2:
[0511] The server receives string data sent from the terminal. The received text is input into a generative AI model to generate a response based on the user's context and intent. In this step, the input is string data, and the output is a response message. The generative AI model parses the input text and uses prompts to generate an appropriate response.
[0512] Step 3:
[0513] The terminal receives a response message from the server. The received message is displayed using a virtual human avatar. The visualized information is used to adjust the avatar's mouth and facial expressions, and a conversation takes place with the user. The input for this step is the response message, and the output is visual feedback to the user. Virtual human software is used to render the received data for avatar display.
[0514] Step 4:
[0515] The user can ask further questions or make comments based on the displayed information. These comments trigger the next dialogue, returning to step 1. The user's input is voice, and the output generates the next voice input, which the system receives and starts the next cycle.
[0516] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0517] This invention embodies a system that provides a deeper, more personalized, interactive conversational experience with users by using a generative model that combines an emotion engine. This system consists of three parties: a server, a terminal (including a display device), and the user.
[0518] System Configuration
[0519] The server is the core unit that operates the generative model and the emotion engine. In addition to generating regular text responses, the generative model has the ability to generate more emotionally sensitive responses by utilizing user emotion data obtained from the emotion engine. This enables support and suggestions tailored to the user's emotional state.
[0520] The device has voice recognition and emotion analysis sensors, and it converts the user's voice and facial expressions into text and data, which it sends to the server. It also displays the server's response using a virtual character, engaging in visual and auditory interaction with the user. This avatar reflects the user's emotional state and changes its facial expressions appropriately.
[0521] Users can interact with the system in a way that feels similar to everyday communication. Because it provides support not only for the user's health but also for their emotions and mood, it helps them lead a more fulfilling life. The system accumulates daily input from users and analyzes it as emotional trends to generate personalized suggestions and alerts tailored to each user.
[0522] Program processing
[0523] The server processes the voice and facial expression data received from the user through speech recognition and sentiment analysis. The voice data is converted into text data, and the sentiment engine infers the emotional state from the text and facial expressions. For example, if the user says, "I'm feeling a little down today," the server generates an emotion tag such as "sadness" through the sentiment engine.
[0524] The server then uses a generative model to generate emotionally sensitive responses. For example, it might offer the user a comforting suggestion such as, "I see, is there anything I can do to help? Shall we look for something to help you relax?"
[0525] The device expresses the received response through the avatar's facial expressions and voice, providing feedback to the user. The avatar displays a gentle expression that matches the user's emotions, creating a friendly and approachable atmosphere.
[0526] This system allows users to receive not only health management but also psychological care, thus alleviating daily feelings of loneliness and providing support for leading a fulfilling life.
[0527] The following describes the processing flow.
[0528] Step 1:
[0529] The user turns on the TV, and the device starts up. The device activates its voice recognition and emotion analysis sensors and enters standby mode.
[0530] Step 2:
[0531] The user speaks to the device, saying, "I feel kind of tired today." The device captures this audio and converts it into text data using a speech recognition engine.
[0532] Step 3:
[0533] The device acquires user facial expression data via an emotion analysis sensor along with voice input, and sends both sets of data to the server.
[0534] Step 4:
[0535] The server inputs the received text data into the generative model and the facial expression data into the emotion engine. The emotion engine evaluates the user's emotional state as "fatigue."
[0536] Step 5:
[0537] The server integrates the emotion evaluation results from the emotion engine into a generation model to generate responses that take the user's emotions into consideration. For example, it might generate a response such as, "You've been working hard lately, it's important to take a break."
[0538] Step 6:
[0539] The server sends the generated response to the terminal. The terminal controls a virtual character avatar and displays the generated response and a calm expression to the user.
[0540] Step 7:
[0541] The user responds, "Thank you, I think I'll take a short rest." This audio is also converted to text by the device and sent back to the server.
[0542] Step 8:
[0543] The server stores a history of user emotions and responses in a database, which serves as foundational data for analyzing long-term emotional trends.
[0544] Step 9:
[0545] Once the conversation ends and the user turns off the TV, the device stops all processes and returns to a standby state for the next use.
[0546] (Example 2)
[0547] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0548] Conventional two-way conversation systems have struggled to accurately analyze users' emotional states and provide appropriate dialogue and psychological support accordingly. Furthermore, they lacked personalized suggestions based on users' health and psychological conditions, as well as notifications to external support services.
[0549] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0550] In this invention, the server includes means for enabling two-way conversations that respond to the user's emotional state using a generative model and emotion analysis means, means for converting the user's voice and facial expressions into text and emotion data, storing and analyzing them, and means for generating a response that includes psychological support when the user's emotional state meets certain conditions, and notifying an external support service. This makes it possible to provide a conversational experience that is attentive to the user's emotions and individually optimized health and psychological support.
[0551] A "generative model" is an algorithm or program used to generate responses or content based on input data.
[0552] "Emotional analysis methods" refer to technologies and systems that analyze user voice and facial expression data to understand their emotional state.
[0553] A "virtual character" is a character or avatar that is displayed on a display device and used to interact with the user.
[0554] "Two-way conversation" is a method of communication in which the user and the system take turns speaking and responding.
[0555] "Psychological support" refers to the act or system of providing support and care that takes into account the user's emotions and psychological state.
[0556] "External support services" refer to organizations or platforms that exist outside the system and provide additional support or assistance to users.
[0557] A "condition" is an element or premise necessary for a particular situation or state to occur.
[0558] An "alert" is a notification or warning that draws attention to the user.
[0559] The system of the present invention is composed of three main components: a server, a terminal, and a user, and provides an individualized conversational experience based on the user's emotional state.
[0560] The server plays a role in enabling two-way conversations with users using a generative AI model and sentiment analysis tools. The generative AI model generates responses based on data sent by the user, while the sentiment analysis tools infer emotions from voice and facial expression data. Specifically, voice is converted into text data by speech recognition, and sentiment tags are generated from this text and facial expression data. For example, if a user shares something like "I'm sad today," the server receives this, generates the sentiment tag "sadness," and returns a thoughtful response accordingly.
[0561] The device is equipped with voice recognition and emotion sensors to detect the user's voice input and facial expressions. This allows data to be transmitted to the server in real time. The device also displays responses sent from the server as avatars and voices, enabling friendly, two-way dialogue with the user. The avatar can change its expression according to the user's emotional state, for example, by expressing an emotionally sensitive offer in a quiet voice, such as, "Is there anything I can help you with?"
[0562] Users can interact with this system in the same way they interact with their daily lives. Data on their daily physical and mental health is accumulated, and the server can use this data to provide personalized advice and warnings. As a result, users can receive psychological care and support to improve their quality of life.
[0563] For example, if a user says, "I'm very tired today," the server performs sentiment analysis and assigns the emotion tag "fatigue." The generative AI model then generates a response such as, "Why don't you try taking a break early? Let's find a way to relax." An example of a prompt might be, "What approach would you take to generate a response when a user is feeling tired?"
[0564] This system will allow users to enjoy a more fulfilling life while receiving emotionally sensitive support.
[0565] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0566] Step 1:
[0567] The user begins a conversation into the device. The user's voice and facial expressions are captured in real time by the device's voice recognition and emotion sensing sensors. This input includes voice data and facial expression data. The device then prepares this data to send to the server.
[0568] Step 2:
[0569] The terminal converts captured audio data into text data using speech recognition technology. During this process, digital signal processing of the audio waveform generates text in string format. Simultaneously, it packages the facial expression data obtained from the emotion sensor, formatting it appropriately for transmission to the server. The output consists of the converted text string and the prepared facial expression data.
[0570] Step 3:
[0571] The server receives text data and facial expression data sent from the terminal. First, the text data is passed to a natural language processing engine to understand its meaning and context. This process lays the foundation for sentiment analysis. Next, the facial expression data is passed through the sentiment engine to infer the user's emotional state. For example, if the user says "I'm tired," the server generates the sentiment tag "fatigue" from the audio and facial expression. This output is used as input to a generative AI model.
[0572] Step 4:
[0573] The server uses a generative AI model to generate an emotionally sensitive response based on the previously obtained emotion tags and input data. This prompt is set to "How can we provide support tailored to the user's level of fatigue?" The generative model creates an appropriate message that matches the user's emotional state by constructing a suggestion such as "Why don't you take a break today?" The resulting output is a response message.
[0574] Step 5:
[0575] The terminal receives a response message sent from the server and prepares to display it using a virtual person (avatar). The avatar uses this message to play back voice in an expressive and persuasive manner as feedback to the user. Here, the avatar's facial expression is set to be gentle, reflecting the user's fatigue, providing a pleasant response both visually and aurally.
[0576] (Application Example 2)
[0577] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0578] In modern society, there is a demand for support that takes into account individual emotions and health conditions. However, conventional technology makes it difficult to accurately grasp users' emotions and provide appropriate responses immediately. Therefore, there is a need for new means to provide customer-specific services and suggestions in physical stores and other face-to-face service settings, thereby improving customer satisfaction.
[0579] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0580] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health and emotional state; means for providing external support or suggestion services when the user's health or emotional state meets certain conditions; and means for rendering a person according to the emotion using one or more representation means and generating a conversation that matches the user's emotions. This makes it possible to provide personalized responses according to the user's individual emotions and state, and to immediately provide appropriate support and suggestions.
[0581] A "generative model" is an algorithm that generates text and responses based on user input data, enabling natural conversations with users.
[0582] A "display device" is a device that includes digital displays and screens used to visually present virtual people or information to a user.
[0583] A "virtual character" is a digital avatar that has a human-like appearance and movements, enabling visual and auditory interaction with the user.
[0584] "Two-way conversation" is a form of interactive communication in which the user provides input to the system, and the system returns a corresponding response.
[0585] "Health status" refers to information indicating the user's physical and mental well-being, and includes various data used for health management and support.
[0586] "Emotional state" refers to the user's psychological state, including emotions such as joy, anger, sadness, and happiness, inferred from their facial expressions and voice.
[0587] "External support or suggested services" refer to third-party services or suggestions provided according to the user's health and emotional state, intended to assist in improving the user's quality of life and solving problems.
[0588] In the system that implements this application example, the server uses a generative model and an emotion engine to create a two-way conversational experience with the user. The terminal is equipped with speech recognition and emotion analysis sensors, and acquires voice input and facial expression data from the user and sends it to the server. The server uses a speech recognition library to convert the voice into text data and processes that text with a generative model. It also uses an emotion analysis module to analyze the user's emotional state and uses the results as prompts incorporated into the generative AI model.
[0589] The server uses an avatar display module to render a virtual person with facial expressions corresponding to the user's emotional state on the display device, and provides the generated response to the user visually and audibly. This allows the user to receive personalized support through the system that takes into account their psychological well-being and health status.
[0590] For example, if a customer in a store says, "I'm feeling a little down today," the server performs sentiment analysis and infers the emotional state as "sadness." The generative model then generates a response that takes that emotion into consideration, suggesting something like, "Is there anything I can do to help?" through a virtual character. The avatar interacts with the user with a gentle expression, creating a comfortable communication experience.
[0591] An example of a prompt for a generative AI model is, "If the user is complaining of being a little tired, generate appropriate rest suggestions." Based on this prompt, the generative model will provide the user with the most suitable suggestions.
[0592] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0593] Step 1:
[0594] The terminal acquires voice input and facial expression data from the user. Using a voice recognition sensor and an emotion analysis sensor, the user's speech and facial expressions are converted into digital data and prepared as voice data and image data. This becomes the input data to the server.
[0595] Step 2:
[0596] The server converts the audio data received from the terminal into text data using a speech recognition library. This process analyzes the audio waveform, converts it into appropriate strings, and extracts the user's actual spoken content. The accuracy of speech recognition is crucial, as it provides the foundational data for the conversation.
[0597] Step 3:
[0598] The server inputs the acquired text and image data into an emotion analysis module to identify the user's emotional state. The emotion analysis processes the data based on the vocabulary in the text and the facial features in the images, generating emotion tags such as "joy," "sadness," and "anger." This emotional information then serves as a prompt for the generating AI model.
[0599] Step 4:
[0600] The server inputs text data and sentiment tags into a generative model to generate personalized responses that take the user's emotions into consideration. In this process, the generative AI model creates natural language responses based on prompt sentences and prepares suggestions and questions appropriate to the user's state.
[0601] Step 5:
[0602] The server sends the generated response to the terminal, which uses an avatar display module to render a virtual person on the screen. The avatar displays facial expressions corresponding to the user's emotions and plays the generated response aloud. This allows the user to receive responses both visually and aurally, resulting in a more interactive experience.
[0603] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0604] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0605] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0606] [Fourth Embodiment]
[0607] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0608] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0609] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0610] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0611] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0612] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0613] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0614] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0615] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0616] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0617] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0618] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0619] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0620] This invention provides a system that enables two-way communication with a user through a display device. This system mainly consists of a server, a terminal (a display device such as a television), and a user.
[0621] System Configuration
[0622] The server includes a processing unit for running the generative model. The generative model can simulate natural conversations with users. This model generates responses based on text data sent by the user, enabling two-way communication.
[0623] The device uses speech recognition technology to convert the user's voice input into text and sends it to the server. Furthermore, it receives responses from the server and displays a virtual person avatar to provide visual feedback to the user. The device may also ask daily questions about the user's health status, triggered by user input.
[0624] Users interact with the system in the same way as in a normal conversation. During this interaction, they can input information about their health and daily life events via voice. This information is sent to the server through the device and used to build the user profile.
[0625] Program processing
[0626] The server analyzes the text data received from the user and uses a generative model to generate appropriate conversational responses. For example, if a user says, "I'm not feeling well today," the server refers to the user's daily data and, if necessary, generates a more humane response such as, "I'm worried about your health lately. Please get plenty of rest," and sends it to the terminal.
[0627] The device displays the generated response using a virtual person avatar. This avatar is designed to create a sense of familiarity with the user by moving its mouth and changing its facial expressions. In addition, the device uses speech recognition technology to transcribe the user's speech into text and sends it to the server, making it easy to use naturally in everyday situations without burdening the user.
[0628] Through this system, users can not only alleviate feelings of loneliness in their daily lives but also receive support in understanding their own health status. The system evolves on its own, providing communication optimized for each user. For example, if a user says to the system, "I'm worried about not getting enough exercise lately," the system can respond, "Okay, let's work together to come up with an easy exercise plan to get you moving."
[0629] The following describes the processing flow.
[0630] Step 1:
[0631] The user turns on the TV, and the device displays the initial setup screen. The device starts a session with the server and establishes a communication channel.
[0632] Step 2:
[0633] The device activates its voice recognition function and waits for the user to speak. When the user says "Hello," the device converts the voice input into text data.
[0634] Step 3:
[0635] The terminal sends the user's text data to the server. The server uses a generative model to generate an appropriate response to the user's "Hello," creating a reply such as "How are you doing today?"
[0636] Step 4:
[0637] The server generates a text response and sends it to the terminal. The terminal uses a virtual person avatar to play the generated response aloud, with mouth movements.
[0638] Step 5:
[0639] The user replies, "I'm a little tired today." The device converts the speech back into text and sends it to the server.
[0640] Step 6:
[0641] The server references past user data and analyzes changes in health status. Using a generative model, it generates a suggestion such as, "That sounds tough, why don't you take a break?"
[0642] Step 7:
[0643] The server sends a suggestion to the terminal. The terminal communicates the suggestion to the user by controlling an avatar. This prompts the user to take appropriate action.
[0644] Step 8:
[0645] The user indicates their intention to end the conversation. If they say, "Thank you, I'm fine now," the device notifies the server that the session has ended.
[0646] Step 9:
[0647] The terminal displays the termination screen and closes the application. The server safely terminates the communication session and saves the log data to the database.
[0648] (Example 1)
[0649] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0650] In modern society, while interest in personal health management is increasing, systems for monitoring health status at home and reducing feelings of isolation are still not adequately developed. In particular, there is a growing need for interactive systems that allow users to understand their own health status through natural conversations at home and receive necessary advice and warnings. Such systems need to support users' daily health management and have features to ensure that urgent health conditions are not overlooked.
[0651] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0652] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for receiving voice input and converting it into a string using speech recognition technology; and means for an analysis device to process the string, generate prompt sentences, and input them into the generative model. This enables the user to monitor their health status through natural conversation and receive quick and appropriate feedback.
[0653] A "generative model" is an artificial intelligence algorithm that generates an appropriate response based on a given input text.
[0654] A "display device" is a device used to present visual information, including virtual character avatars, to a user.
[0655] "Speech recognition technology" is a technology that converts a user's voice into text data.
[0656] A "prompt" is a text-based question or instruction that is input into a generative model.
[0657] An "analysis device" is a device that processes text data and generates prompts adapted to a generative model.
[0658] A "user profile" is an aggregate of personal information built based on a user's attributes and past conversation history.
[0659] "External health-related services" refer to third-party services that utilize users' health information to provide support and advice.
[0660] An "avatar" is a virtual character displayed to visually represent a two-way conversation with the user.
[0661] This invention constitutes a system that provides users with a natural and interactive conversational experience. The system is primarily operated by a server, a terminal, and the user.
[0662] The server is a computing device for running the generative AI model and receives data transcribed into text using speech recognition technology. This data is processed by an analysis device, which generates prompt sentences to be input into the generative model. The server uses these prompt sentences to cause the generative AI model to generate a response. The generated response is sent to the terminal and presented to the user.
[0663] The terminal includes a display device equipped with a microphone and speaker, which receives the user's voice and converts it to text using speech recognition technology. This text data is sent to a server using a secure protocol. The response received from the server is displayed visually through a virtual avatar. The avatar is designed to change its facial expressions and movements in response to the conversation, creating a sense of familiarity with the user.
[0664] Users can interact with this system in a normal conversational manner. They provide information about their health and daily life during the conversation, and this information is reflected in their user profile. This allows them to receive personalized responses based on their past data.
[0665] As a concrete example, if a user says to the device, "I'm worried because I haven't been getting enough exercise lately," speech recognition technology will convert this statement into text and send it to the server. The server will use a generative AI model to generate a response such as, "Okay, let's think together about an easy exercise plan to get you moving," and send it to the device. The device will then present this response to the user through a virtual avatar.
[0666] As an example of a prompt, the input to a generative AI model would look like this:
[0667] "User comment: I'm worried because I haven't been getting enough exercise lately. System response: Well, let's work together to come up with an easy exercise plan to get you moving."
[0668] In this way, the system monitors the user's health while functioning as a conversational partner in their daily life.
[0669] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0670] Step 1:
[0671] The user emits a voice message through the device's microphone. This voice serves as input, conveying the user's intentions and state.
[0672] Step 2:
[0673] The device processes the audio received from the user using a speech recognition engine and converts it into text format. As a result, the audio data is output in a structured form as text data.
[0674] Step 3:
[0675] The terminal sends the converted text data to the server. The transmitted text serves as input for parsing and response generation.
[0676] Step 4:
[0677] The server analyzes the received text data and uses natural language processing algorithms to understand the appropriate context. The analysis results are output as prompts to the generative AI model.
[0678] Step 5:
[0679] The server inputs a prompt sentence into the generative AI model. The model generates a conversational response based on this input and outputs that response as text.
[0680] Step 6:
[0681] The server sends the generated response text to the terminal. This response is the output necessary to complete the interaction with the user.
[0682] Step 7:
[0683] The device presents received responses to the user visually and audibly using a virtual avatar. The avatar displays facial expressions and movements in accordance with the conversation, providing the user with a natural dialogue experience.
[0684] Step 8:
[0685] The server updates the user profile based on the user's utterances and conversation history with the system. This process prepares the system for more personalized conversations in the future.
[0686] (Application Example 1)
[0687] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0688] In today's real-world retail environments, there is a demand for appropriate customer service and improved customer satisfaction. However, due to limitations in human resources and the difficulty of responding to individual needs, the current situation is generally limited to providing a uniform service. Furthermore, since there is no system in place that can respond immediately on the field based on individual health conditions and needs, there is an urgent need to realize flexible and effective services that meet diverse customer needs.
[0689] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0690] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health status; means for notifying an external health support service when the user's health status meets certain conditions; means for acquiring information via a portable terminal worn by a person in real space and performing voice analysis; and means for suggesting appropriate products and services based on the person's statements. This makes it possible to analyze information emitted by customers in real time and provide services tailored to individual needs.
[0691] A "generative model" is a machine learning-based software program designed to generate natural, two-way conversations with users.
[0692] A "display device" is an electronic device that displays a virtual person to provide visual feedback to the user.
[0693] "Two-way conversation" refers to a form of communication where both the user and the system are interactive and respond to each other.
[0694] "Health status" refers to information that indicates the user's physical and mental condition.
[0695] "External health support services" refer to third-party organizations or systems that support users' health management.
[0696] A "portable terminal" is a device that is portable and designed for communication and information processing.
[0697] "Speech analysis" is a technology that analyzes speech data based on textual information and content.
[0698] "Proposing products or services" refers to the act of recommending appropriate products or services based on the user's needs.
[0699] To realize the system of this invention, a configuration based on the division of roles between the server, terminal, and user is necessary. The server functions as a central processing unit for generating natural, two-way conversations with the user using a generative AI model. The server receives voice data containing text information from the user, analyzes its content, and generates a response message. The software used includes a machine learning library for operating the AI model.
[0700] The device functions as a portable display device, recognizing the user's voice and converting it into text data. It incorporates voice recognition technology and transmits the collected data to a server. It also visually displays the generated messages as a virtual person, interacting with the user. In this configuration, the device can also function as smart glasses or other wearable devices.
[0701] Users, acting as the subjects who make statements and ask questions through this system, input information via voice. The user's voice input may include information related to everyday in-store conversations or purchase preferences. For example, if a user says, "My skin has been in bad condition lately...", the system receives and analyzes that statement and suggests products and services.
[0702] As a concrete example, in a physical store, if a customer says, "I'm looking for a new mystery novel," the system will guide them by saying, "We have the latest mystery novels on this shelf." This allows for real-time analysis of customer information and the provision of services tailored to individual needs. An example of a prompt used in this case would be: "Customer: I'm looking for a new mystery novel\nEmployee: ">
[0703] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0704] Step 1:
[0705] The terminal acquires the user's voice input. Using speech recognition technology, it converts the acquired voice data into text and sends it to the server. In this process, the input is voice data, and the output is string data. The process involves analyzing the voice using speech recognition software and converting it into a string.
[0706] Step 2:
[0707] The server receives string data sent from the terminal. The received text is input into a generative AI model to generate a response based on the user's context and intent. In this step, the input is string data, and the output is a response message. The generative AI model parses the input text and uses prompts to generate an appropriate response.
[0708] Step 3:
[0709] The terminal receives a response message from the server. The received message is displayed using a virtual human avatar. The visualized information is used to adjust the avatar's mouth and facial expressions, and a conversation takes place with the user. The input for this step is the response message, and the output is visual feedback to the user. Virtual human software is used to render the received data for avatar display.
[0710] Step 4:
[0711] The user can ask further questions or make comments based on the displayed information. These comments trigger the next dialogue, returning to step 1. The user's input is voice, and the output generates the next voice input, which the system receives and starts the next cycle.
[0712] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0713] This invention embodies a system that provides a deeper, more personalized, interactive conversational experience with users by using a generative model that combines an emotion engine. This system consists of three parties: a server, a terminal (including a display device), and the user.
[0714] System Configuration
[0715] The server is the core unit that operates the generative model and the emotion engine. In addition to generating regular text responses, the generative model has the ability to generate more emotionally sensitive responses by utilizing user emotion data obtained from the emotion engine. This enables support and suggestions tailored to the user's emotional state.
[0716] The device has voice recognition and emotion analysis sensors, and it converts the user's voice and facial expressions into text and data, which it sends to the server. It also displays the server's response using a virtual character, engaging in visual and auditory interaction with the user. This avatar reflects the user's emotional state and changes its facial expressions appropriately.
[0717] Users can interact with the system in a way that feels similar to everyday communication. Because it provides support not only for the user's health but also for their emotions and mood, it helps them lead a more fulfilling life. The system accumulates daily input from users and analyzes it as emotional trends to generate personalized suggestions and alerts tailored to each user.
[0718] Program processing
[0719] The server processes the voice and facial expression data received from the user through speech recognition and sentiment analysis. The voice data is converted into text data, and the sentiment engine infers the emotional state from the text and facial expressions. For example, if the user says, "I'm feeling a little down today," the server generates an emotion tag such as "sadness" through the sentiment engine.
[0720] The server then uses a generative model to generate emotionally sensitive responses. For example, it might offer the user a comforting suggestion such as, "I see, is there anything I can do to help? Shall we look for something to help you relax?"
[0721] The device expresses the received response through the avatar's facial expressions and voice, providing feedback to the user. The avatar displays a gentle expression that matches the user's emotions, creating a friendly and approachable atmosphere.
[0722] This system allows users to receive not only health management but also psychological care, thus alleviating daily feelings of loneliness and providing support for leading a fulfilling life.
[0723] The following describes the processing flow.
[0724] Step 1:
[0725] The user turns on the TV, and the device starts up. The device activates its voice recognition and emotion analysis sensors and enters standby mode.
[0726] Step 2:
[0727] The user speaks to the device, saying, "I feel kind of tired today." The device captures this audio and converts it into text data using a speech recognition engine.
[0728] Step 3:
[0729] The device acquires user facial expression data via an emotion analysis sensor along with voice input, and sends both sets of data to the server.
[0730] Step 4:
[0731] The server inputs the received text data into the generative model and the facial expression data into the emotion engine. The emotion engine evaluates the user's emotional state as "fatigue."
[0732] Step 5:
[0733] The server integrates the emotion evaluation results from the emotion engine into a generation model to generate responses that take the user's emotions into consideration. For example, it might generate a response such as, "You've been working hard lately, it's important to take a break."
[0734] Step 6:
[0735] The server sends the generated response to the terminal. The terminal controls a virtual character avatar and displays the generated response and a calm expression to the user.
[0736] Step 7:
[0737] The user responds, "Thank you, I think I'll take a short rest." This audio is also converted to text by the device and sent back to the server.
[0738] Step 8:
[0739] The server stores a history of user emotions and responses in a database, which serves as foundational data for analyzing long-term emotional trends.
[0740] Step 9:
[0741] Once the conversation ends and the user turns off the TV, the device stops all processes and returns to a standby state for the next use.
[0742] (Example 2)
[0743] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] Conventional two-way conversation systems have struggled to accurately analyze users' emotional states and provide appropriate dialogue and psychological support accordingly. Furthermore, they lacked personalized suggestions based on users' health and psychological conditions, as well as notifications to external support services.
[0745] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0746] In this invention, the server includes means for enabling two-way conversations that respond to the user's emotional state using a generative model and emotion analysis means, means for converting the user's voice and facial expressions into text and emotion data, storing and analyzing them, and means for generating a response that includes psychological support when the user's emotional state meets certain conditions, and notifying an external support service. This makes it possible to provide a conversational experience that is attentive to the user's emotions and individually optimized health and psychological support.
[0747] A "generative model" is an algorithm or program used to generate responses or content based on input data.
[0748] "Emotional analysis methods" refer to technologies and systems that analyze user voice and facial expression data to understand their emotional state.
[0749] A "virtual character" is a character or avatar that is displayed on a display device and used to interact with the user.
[0750] "Two-way conversation" is a method of communication in which the user and the system take turns speaking and responding.
[0751] "Psychological support" refers to the act or system of providing support and care that takes into account the user's emotions and psychological state.
[0752] "External support services" refer to organizations or platforms that exist outside the system and provide additional support or assistance to users.
[0753] A "condition" is an element or premise necessary for a particular situation or state to occur.
[0754] An "alert" is a notification or warning that draws attention to the user.
[0755] The system of the present invention is composed of three main components: a server, a terminal, and a user, and provides an individualized conversational experience based on the user's emotional state.
[0756] The server plays a role in enabling two-way conversations with users using a generative AI model and sentiment analysis tools. The generative AI model generates responses based on data sent by the user, while the sentiment analysis tools infer emotions from voice and facial expression data. Specifically, voice is converted into text data by speech recognition, and sentiment tags are generated from this text and facial expression data. For example, if a user shares something like "I'm sad today," the server receives this, generates the sentiment tag "sadness," and returns a thoughtful response accordingly.
[0757] The device is equipped with voice recognition and emotion sensors to detect the user's voice input and facial expressions. This allows data to be transmitted to the server in real time. The device also displays responses sent from the server as avatars and voices, enabling friendly, two-way dialogue with the user. The avatar can change its expression according to the user's emotional state, for example, by expressing an emotionally sensitive offer in a quiet voice, such as, "Is there anything I can help you with?"
[0758] Users can interact with this system in the same way they interact with their daily lives. Data on their daily physical and mental health is accumulated, and the server can use this data to provide personalized advice and warnings. As a result, users can receive psychological care and support to improve their quality of life.
[0759] For example, if a user says, "I'm very tired today," the server performs sentiment analysis and assigns the emotion tag "fatigue." The generative AI model then generates a response such as, "Why don't you try taking a break early? Let's find a way to relax." An example of a prompt might be, "What approach would you take to generate a response when a user is feeling tired?"
[0760] This system will allow users to enjoy a more fulfilling life while receiving emotionally sensitive support.
[0761] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0762] Step 1:
[0763] The user begins a conversation into the device. The user's voice and facial expressions are captured in real time by the device's voice recognition and emotion sensing sensors. This input includes voice data and facial expression data. The device then prepares this data to send to the server.
[0764] Step 2:
[0765] The terminal converts captured audio data into text data using speech recognition technology. During this process, digital signal processing of the audio waveform generates text in string format. Simultaneously, it packages the facial expression data obtained from the emotion sensor, formatting it appropriately for transmission to the server. The output consists of the converted text string and the prepared facial expression data.
[0766] Step 3:
[0767] The server receives text data and facial expression data sent from the terminal. First, the text data is passed to a natural language processing engine to understand its meaning and context. This process lays the foundation for sentiment analysis. Next, the facial expression data is passed through the sentiment engine to infer the user's emotional state. For example, if the user says "I'm tired," the server generates the sentiment tag "fatigue" from the audio and facial expression. This output is used as input to a generative AI model.
[0768] Step 4:
[0769] The server uses a generative AI model to generate an emotionally sensitive response based on the previously obtained emotion tags and input data. This prompt is set to "How can we provide support tailored to the user's level of fatigue?" The generative model creates an appropriate message that matches the user's emotional state by constructing a suggestion such as "Why don't you take a break today?" The resulting output is a response message.
[0770] Step 5:
[0771] The terminal receives a response message sent from the server and prepares to display it using a virtual person (avatar). The avatar uses this message to play back voice in an expressive and persuasive manner as feedback to the user. Here, the avatar's facial expression is set to be gentle, reflecting the user's fatigue, providing a pleasant response both visually and aurally.
[0772] (Application Example 2)
[0773] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0774] In modern society, there is a demand for support that takes into account individual emotions and health conditions. However, conventional technology makes it difficult to accurately grasp users' emotions and provide appropriate responses immediately. Therefore, there is a need for new means to provide customer-specific services and suggestions in physical stores and other face-to-face service settings, thereby improving customer satisfaction.
[0775] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0776] In this invention, the server includes means for rendering a virtual person on a display device using a generative model and enabling two-way conversation; means for acquiring, storing, and analyzing information regarding the user's health and emotional state; means for providing external support or suggestion services when the user's health or emotional state meets certain conditions; and means for rendering a person according to the emotion using one or more representation means and generating a conversation that matches the user's emotions. This makes it possible to provide personalized responses according to the user's individual emotions and state, and to immediately provide appropriate support and suggestions.
[0777] A "generative model" is an algorithm that generates text and responses based on user input data, enabling natural conversations with users.
[0778] A "display device" is a device that includes digital displays and screens used to visually present virtual people or information to a user.
[0779] A "virtual character" is a digital avatar that has a human-like appearance and movements, enabling visual and auditory interaction with the user.
[0780] "Two-way conversation" is a form of interactive communication in which the user provides input to the system, and the system returns a corresponding response.
[0781] "Health status" refers to information indicating the user's physical and mental well-being, and includes various data used for health management and support.
[0782] "Emotional state" refers to the user's psychological state, including emotions such as joy, anger, sadness, and happiness, inferred from their facial expressions and voice.
[0783] "External support or suggested services" refer to third-party services or suggestions provided according to the user's health and emotional state, intended to assist in improving the user's quality of life and solving problems.
[0784] In the system that implements this application example, the server uses a generative model and an emotion engine to create a two-way conversational experience with the user. The terminal is equipped with speech recognition and emotion analysis sensors, and acquires voice input and facial expression data from the user and sends it to the server. The server uses a speech recognition library to convert the voice into text data and processes that text with a generative model. It also uses an emotion analysis module to analyze the user's emotional state and uses the results as prompts incorporated into the generative AI model.
[0785] The server uses an avatar display module to render a virtual person with facial expressions corresponding to the user's emotional state on the display device, and provides the generated response to the user visually and audibly. This allows the user to receive personalized support through the system that takes into account their psychological well-being and health status.
[0786] For example, if a customer in a store says, "I'm feeling a little down today," the server performs sentiment analysis and infers the emotional state as "sadness." The generative model then generates a response that takes that emotion into consideration, suggesting something like, "Is there anything I can do to help?" through a virtual character. The avatar interacts with the user with a gentle expression, creating a comfortable communication experience.
[0787] An example of a prompt for a generative AI model is, "If the user is complaining of being a little tired, generate appropriate rest suggestions." Based on this prompt, the generative model will provide the user with the most suitable suggestions.
[0788] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0789] Step 1:
[0790] The terminal acquires voice input and facial expression data from the user. Using a voice recognition sensor and an emotion analysis sensor, the user's speech and facial expressions are converted into digital data and prepared as voice data and image data. This becomes the input data to the server.
[0791] Step 2:
[0792] The server converts the audio data received from the terminal into text data using a speech recognition library. This process analyzes the audio waveform, converts it into appropriate strings, and extracts the user's actual spoken content. The accuracy of speech recognition is crucial, as it provides the foundational data for the conversation.
[0793] Step 3:
[0794] The server inputs the acquired text and image data into an emotion analysis module to identify the user's emotional state. The emotion analysis processes the data based on the vocabulary in the text and the facial features in the images, generating emotion tags such as "joy," "sadness," and "anger." This emotional information then serves as a prompt for the generating AI model.
[0795] Step 4:
[0796] The server inputs text data and sentiment tags into a generative model to generate personalized responses that take the user's emotions into consideration. In this process, the generative AI model creates natural language responses based on prompt sentences and prepares suggestions and questions appropriate to the user's state.
[0797] Step 5:
[0798] The server sends the generated response to the terminal, which uses an avatar display module to render a virtual person on the screen. The avatar displays facial expressions corresponding to the user's emotions and plays the generated response aloud. This allows the user to receive responses both visually and aurally, resulting in a more interactive experience.
[0799] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0800] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0801] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0802] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0803] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0804] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0805] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0806] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0807] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0808] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0809] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0810] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0811] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0812] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0813] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0814] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0815] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0816] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0817] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0818] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0819] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0820] The following is further disclosed regarding the embodiments described above.
[0821] (Claim 1)
[0822] A means for rendering a virtual person on a display device using a generative model and enabling two-way conversation,
[0823] A means of acquiring, storing, and analyzing information about a user's health status,
[0824] A means of notifying external health support services when a user's health status meets certain conditions,
[0825] A system that includes this.
[0826] (Claim 2)
[0827] The system according to claim 1, which converts voice input from a user into a string and inputs the string into a generation model.
[0828] (Claim 3)
[0829] The system according to claim 1, which evaluates user health information based on a generative model and displays suggestions and alerts.
[0830] "Example 1"
[0831] (Claim 1)
[0832] A means for rendering a virtual person on a display device using a generative model and enabling two-way conversation,
[0833] A means for receiving voice input and converting it into a string using speech recognition technology,
[0834] The analysis device processes a string, generates a prompt statement, and inputs it into the generative model.
[0835] A means for a generative model to generate a response based on the user's everyday conversation and transmit that response to a display device,
[0836] A means of acquiring, storing, and analyzing information about a user's health status,
[0837] A means to update user profiles and achieve optimized communication through multiple uses,
[0838] A means of notifying external health-related services when a user's health status meets certain conditions,
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, which converts voice input from a user into a string and inputs the generated string as a prompt sentence into a generation model using an analysis device.
[0842] (Claim 3)
[0843] The system according to claim 1, which evaluates the user's health information using responses generated based on a generative model and displays optimized suggestions and alerts.
[0844] "Application Example 1"
[0845] (Claim 1)
[0846] A means for rendering a virtual person on a display device using a generative model and enabling two-way conversation,
[0847] A means of acquiring, storing, and analyzing information about a user's health status,
[0848] A means of notifying external health support services when a user's health status meets certain conditions,
[0849] A means of acquiring information via a portable device worn by a person in real space and performing voice analysis,
[0850] A means of suggesting appropriate products and services based on a person's statements,
[0851] A system that includes this.
[0852] (Claim 2)
[0853] The system according to claim 1, which converts voice input from a user into a string and inputs the string into a generation model.
[0854] (Claim 3)
[0855] The system according to claim 1, which evaluates user health information based on a generative model and displays suggestions and alerts.
[0856] "Example 2 of combining an emotion engine"
[0857] (Claim 1)
[0858] A means for rendering a virtual person on a display device using a generative model and emotion analysis means, and enabling two-way conversation according to the user's emotional state,
[0859] A means for acquiring a user's voice and facial expressions, converting them into text and emotion data, storing them, and analyzing them,
[0860] A means of generating a response including psychological support when the user's emotional state meets certain conditions, and notifying an external support service,
[0861] A system that includes this.
[0862] (Claim 2)
[0863] The system according to claim 1, which converts voice and facial expression input from a user into string and emotion data and inputs them to a generative model and an emotion analysis engine.
[0864] (Claim 3)
[0865] The system according to claim 1, which evaluates the user's psychological state and health information based on generative models and sentiment data, and displays appropriate suggestions or alerts.
[0866] "Application example 2 when combining with an emotional engine"
[0867] (Claim 1)
[0868] A means for rendering a virtual person on a display device using a generative model and enabling two-way conversation,
[0869] A means for acquiring, storing, and analyzing information regarding the user's health status and emotional state,
[0870] A means of providing external support or suggestion services when the user's health or emotional state meets certain conditions,
[0871] A means of creating character depictions that respond to emotions using one or more means of expression, and generating conversations that match the user's emotions,
[0872] A system that includes this.
[0873] (Claim 2)
[0874] The system according to claim 1, which converts voice input from a user into a string, inputs the string and emotion data into a generative model, and generates an emotion-sensitive response.
[0875] (Claim 3)
[0876] The system according to claim 1, which evaluates user health information and emotional information based on a generative model and displays alerts with suggestions and psychological care. [Explanation of Symbols]
[0877] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for rendering a virtual person on a display device using a generative model and enabling two-way conversation, A means of acquiring, storing, and analyzing information about a user's health status, A means of notifying external health support services when a user's health status meets certain conditions, A system that includes this.
2. The system according to claim 1, which converts voice input from a user into a string and inputs the string into a generation model.
3. The system according to claim 1, which evaluates user health information based on a generative model and displays suggestions and alerts.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A