System

The system addresses the challenge of providing real-time specialized advice by capturing voice input, analyzing intent, selecting avatars, and generating answers, ensuring quick and accurate expert guidance.

JP2026019728APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121476
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing systems struggle to provide quick and accurate specialized advice to users on a wide range of problems in their daily lives, often failing to efficiently analyze voice input, select appropriate avatars, and display relevant answers in real-time.

Method used

A system that captures voice input, converts it into text data, analyzes the intent of the question, selects an avatar with corresponding expertise, generates an answer, and displays it through an avatar, utilizing speech recognition, natural language processing, and generative AI models to provide real-time expert advice.

Benefits of technology

Enables users to receive tailored and specialized advice quickly and accurately through interactive avatars, addressing various daily challenges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019728000001_ABST
    Figure 2026019728000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for capturing a voice input from a user; means for analyzing and converting the captured voice input into text data; means for analyzing an intent of a question based on the text data; means for selecting an avatar having expertise corresponding to the intent of the question; means for generating an answer based on the selected avatar; and means for displaying the generated answer to the user through the avatar.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, people face a wide range of problems in their daily lives, but there are limited means to quickly acquire specialized knowledge and skills. This can result in problems such as difficulty learning a specific language or smoothly completing household tasks. The present invention aims to solve these problems by providing a system that allows users to easily and intuitively obtain specialized advice. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system comprising the following means for providing specialized support tailored to user needs. The system includes a means for capturing voice input from a user, a means for analyzing the captured voice input and converting it into text data, a means for analyzing the intent of the question based on the text data, a means for selecting an avatar with expertise corresponding to the intent of the question, a means for generating an answer based on the selected avatar, and a means for displaying the generated answer to the user via an avatar. This allows a user to easily input a question via voice and receive advice in real time from an avatar with appropriate expertise. Furthermore, by providing avatars with multiple areas of expertise, it is possible to provide information tailored to a wide range of user needs.

[0006] "Voice input" refers to linguistic voice information uttered by the user, and is an input method for the system to capture this information.

[0007] "Capture" refers to obtaining data such as audio input and video in real time.

[0008] "Analysis" refers to the means of processing acquired data and understanding its content and meaning.

[0009] "Text data" refers to data expressed as text information that has been converted from audio or image data.

[0010] "Question intent" refers to analyzing the purpose behind the question a user asks and the information they are seeking.

[0011] An "avatar" is a virtual character form and is a display means within the system for interacting with the user.

[0012] "Selection" refers to choosing the most suitable target based on specific conditions.

[0013] "Generation" refers to automatically creating answers to questions or necessary information.

[0014] "Display" refers to outputting the generated answers and avatars to the screen so that the user can see them.

[0015] "Expertise" refers to information that has a high level of understanding and skill in a particular field. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[0038] System Overview

[0039] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[0040] Program processing

[0041] Audio data capture and analysis

[0042] 1. Capture voice input

[0043] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[0044] 2. Sending audio data

[0045] The captured audio data is transmitted to a server in real time.

[0046] 3. Analysis of audio data

[0047] The server receives the voice data and converts it into text using a speech recognition algorithm.

[0048] 4. Intent Analysis

[0049] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question. For example, the intent of "everyday English conversation" is extracted.

[0050] Selecting specialized avatars and generating answers

[0051] 5. Select your avatar

[0052] Based on the intent of the question, the server selects an avatar with appropriate expertise, for example, an English teacher avatar.

[0053] 6. Answer Generation

[0054] A question is input into the generative AI model, which generates an appropriate answer, such as an English phrase like "Good morning! This is a morning greeting."

[0055] Show Answers

[0056] 7. Submitting avatar and answer data

[0057] The server sends the generated 3D model of the avatar and the response data to the device.

[0058] 8. Displaying Avatars

[0059] The terminal decodes the received avatar and response data and prepares it for display on the mirror device.

[0060] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[0061] Specific examples

[0062] Learning English

[0063] 1. User operations

[0064] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[0065] 2. Terminal Processing

[0066] The device captures the audio and sends it to the server.

[0067] 3. Server Processing

[0068] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[0069] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[0070] 4. Terminal Processing

[0071] The terminal receives the avatar and the answer and displays them on the mirror device.

[0072] The avatar explains, "Good morning! This is a morning greeting."

[0073] Technical Advice

[0074] 1. User operations

[0075] User: Say, "How do I use a wood stove safely?"

[0076] 2. Terminal Processing

[0077] The device captures the audio and sends it to the server.

[0078] 3. Server Processing

[0079] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[0080] An outdoor expert avatar is selected to generate answers regarding safe use.

[0081] 4. Terminal Processing

[0082] The terminal receives the avatar and the answer and displays them on the mirror device.

[0083] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[0084] As described above, the present invention is a system that interactively provides expert advice to users on various problems they may encounter, allowing users to obtain the information and guidance they need in real time.

[0085] The processing flow will be explained below.

[0086] Step 1:

[0087] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[0088] Step 2:

[0089] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[0090] Step 3:

[0091] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[0092] Step 4:

[0093] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[0094] Step 5:

[0095] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[0096] Step 6:

[0097] The server uses a generative AI model to generate an appropriate answer to the question through the selected avatar, for example, "Good morning! This is a morning greeting."

[0098] Step 7:

[0099] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[0100] Step 8:

[0101] The device decodes the avatar and answer data received from the server, loads the decoded 3D model of the avatar and answer data, and prepares them for display on the mirror device's display.

[0102] Step 9:

[0103] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might say, "Good morning! This is a morning greeting."

[0104] Step 10:

[0105] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[0106] Step 11:

[0107] The server receives the new voice data, converts it back to text, analyzes the intent, and generates an appropriate response. This process can be repeated as necessary.

[0108] This allows users to get real-time expert advice and quickly solve everyday problems.

[0109] Example 1

[0110] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0111] Conventional information provision systems have had difficulty responding quickly and accurately to the various problems users face. They also often only provide limited answers to questions that require specialized knowledge. Furthermore, there are issues with real-time performance and security in the process of efficiently analyzing a user's voice input, selecting an appropriate avatar, and displaying the answer.

[0112] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0113] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user through an avatar, means for compressing the captured voice input in real time and securely transmitting it, and means for decoding a three-dimensional model of the selected avatar and the text data and displaying them on a display device, thereby enabling expert advice to be provided quickly and accurately through the avatar for various problems faced by users.

[0114] "User" means any person or end user of the System.

[0115] "Voice input" refers to the words or sounds spoken by the user, and is the data that the system must recognize.

[0116] "Audio capture means" refers to a device or software that detects a user's voice input and captures it as digital data.

[0117] "Means for converting to text data" refers to the process of converting captured audio into text form using speech recognition technology.

[0118] "Means for analyzing the intent of a question" refers to natural language processing technology for understanding the purpose and requirements of a user's question from text data.

[0119] "Avatar selection method" refers to the process of selecting a virtual character with appropriate expertise based on the intent of the analyzed question.

[0120] "Means for generating answers" refers to the ability to use selected avatars and generative AI models to construct appropriate answers to users' questions.

[0121] "Means for displaying to the user through an avatar" refers to the process of providing the generated answer to the user visually and audibly through an avatar.

[0122] "Means for real-time compression and secure transmission" refers to the function of instantly compressing captured audio data and securely transmitting it to a server using encryption technology.

[0123] "Means for decoding the three-dimensional model and text data and displaying it on a display device" refers to the function of decoding the received data and appropriately rendering the three-dimensional avatar and text information for display on a display device.

[0124] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[0125] System Overview

[0126] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[0127] Audio data capture and analysis

[0128] 1. Capture voice input

[0129] The device captures voice input from the user using a microphone. For example, if the user says, "Please teach me everyday English conversation," the device captures the voice using a high-sensitivity microphone and uses noise-canceling technology to remove background noise.

[0130] 2. Sending audio data

[0131] The captured audio data is compressed in real time by the device and securely transmitted to the server using Transport Layer Security (TLS) encryption.

[0132] 3. Analysis of audio data

[0133] The server decodes the received audio data and converts it into text using a speech recognition algorithm (e.g., Google Speech-to-Text API), which is optimized to complete within a few seconds.

[0134] 4. Intent Analysis

[0135] The server then runs the converted text data through a natural language processing (NLP) engine (such as spaCy or the BERT model) to analyze the intent of the user's question. For example, if the intent is extracted as "everyday English conversation," an appropriate answer is prepared based on this.

[0136] Selecting specialized avatars and generating answers

[0137] 5. Select your avatar

[0138] The server selects the avatar with the most appropriate expertise based on the intent of the question, using a database of avatar skills, for example, selecting an "English teacher" avatar.

[0139] 6. Answer Generation

[0140] The server generates answers to questions using a generative AI model (e.g., OpenAI GPT-4). Specifically, in response to the prompt, "Please give me an example of an everyday English conversation," the server generates the answer, "Good morning! This is a morning greeting."

[0141] Show Answers

[0142] 7. Submitting avatar and answer data

[0143] The server combines the generated avatar's 3D model and the text response data into a single data packet and sends it to the terminal. This transmission uses UDP (User Datagram Protocol) to achieve low latency.

[0144] 8. Displaying Avatars

[0145] The device decodes the received data and prepares it for display on the mirror device's display. The device renders a 3D model of the avatar using the device's GPU and uses a voice synthesizer to output the response. The user hears the avatar displayed on the mirror device explain, "Good morning! This is a morning greeting."

[0146] Specific examples

[0147] 1. Example of user operation

[0148] Learning English: The user speaks to the mirror device, saying, "Please teach me everyday English conversation." The device captures the audio and sends it to the server. The server selects an English teacher avatar and generates a response, saying, "Good morning! This is a morning greeting." The device displays the avatar and the response on the mirror device, and the avatar explains the response aloud.

[0149] 2. Examples of technical advice

[0150] The user says, "Please tell me how to use a wood stove safely." The device captures the audio and sends it to the server. The server selects an outdoor expert avatar and generates a response: "First, make sure there are no flammable objects around the wood stove." The device displays the avatar and the response on a mirror-type device, and the avatar explains.

[0151] The system allows users to receive expert advice in real time.

[0152] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0153] Step 1: Capturing Audio Input

[0154] Specific operation: The user stands in front of the mirror-type device and speaks to it, such as "Please teach me some everyday English conversation."

[0155] Input: User's voice

[0156] Processing: The device's built-in high-sensitivity microphone captures the user's voice and uses noise-canceling technology to remove background noise.

[0157] Output: Captured audio data

[0158] Step 2: Sending audio data

[0159] Specific operation: The captured audio data is sent from the device to the server in real time.

[0160] Input: Captured audio data

[0161] Processing: The device compresses the audio data and sends it securely to the server using TLS (Transport Layer Security) encryption.

[0162] Output: Encrypted audio data sent to the server

[0163] Step 3: Analyzing the audio data

[0164] Specific operation: The server analyzes the received voice data using a voice recognition algorithm.

[0165] Input: Encrypted audio data sent to the server

[0166] Processing: The server decodes the audio data and converts it to text using a speech recognition algorithm (e.g., Google Speech-to-Text API).

[0167] Output: Converted text data

[0168] Step 4: Intent Analysis

[0169] Specific operation: The server analyzes the converted text data and understands the intent of the user's question.

[0170] Input: Converted text data

[0171] Processing: The server runs the text data through a natural language processing (NLP) engine (e.g., spaCy or the BERT model) to analyze the intent of the question.

[0172] Output: Analyzed intent information (e.g., "I want to learn everyday English conversation")

[0173] Step 5: Choose your avatar

[0174] Specific operation: The server selects an appropriate avatar.

[0175] Input: Parsed intent information

[0176] Processing: Based on the analyzed intention information, the server refers to the avatar skill database and selects an avatar with appropriate expertise.

[0177] Output: Information about the selected avatar (e.g., an English teacher avatar)

[0178] Step 6: Generate an answer

[0179] Specific behavior: The server generates an answer based on the selected avatar.

[0180] Input: Selected avatar information and intention information

[0181] Processing: The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers to questions.

[0182] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[0183] Step 7: Submit your avatar and answer data

[0184] Specific operation: The server sends the generated avatar's three-dimensional model and text response data to the terminal.

[0185] Input: Generated answers and a 3D avatar model

[0186] Processing: The server collects the generated data into a single data packet and sends it to the terminal using UDP (User Datagram Protocol).

[0187] Output: Data packets sent to the terminal

[0188] Step 8: Displaying the Avatar

[0189] Specific operation: The terminal decodes the received avatar data and response data and displays them on the mirror device.

[0190] Input: Data packets sent to the terminal

[0191] Processing: The device decodes the data, uses the GPU to render a 3D avatar, and uses a voice synthesizer to voice the response.

[0192] Output: Avatar displayed on a mirror device and voice response

[0193] Through these specific actions and data flows at each processing step, real-time expert advice is provided to the user.

[0194] (Application example 1)

[0195] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0196] The objective of this invention is to improve driving safety and the user experience by providing real-time expert advice through voice input in response to various questions and anxieties that users may encounter in an autonomous vehicle. In particular, there is a need to select an appropriate avatar and provide easy-to-understand answers to specific questions such as road traffic information, explanations of vehicle functions, and driving advice.

[0197] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0198] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the converted text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user via the avatar, and means for displaying the generated answer on a vehicle display, thereby enabling the user to receive reliable information and advice in real time within the autonomous vehicle.

[0199] "Voice input" refers to input of voice information generated by a user speaking.

[0200] "Capture" means capturing surrounding sounds using a device such as a microphone.

[0201] "Analysis" is the process of interpreting data and extracting meaning and patterns.

[0202] "Text data" refers to data that has been converted from voice input into text information through analysis.

[0203] The "question intent" indicates the content and purpose of the question that the user is trying to ask through voice input.

[0204] "Specialist knowledge" refers to advanced knowledge or information about a particular area or field.

[0205] An "avatar" is a virtual character that represents a specific role or persona in the digital space.

[0206] "Answer generation" is the process of creating an appropriate answer to a user's question.

[0207] "Display" is the act of visually presenting information to a user.

[0208] "Vehicle display" means a screen device installed inside an autonomous vehicle for displaying information.

[0209] A "generative artificial intelligence model" is a type of artificial intelligence that generates text data in natural language based on given prompts.

[0210] System Overview

[0211] The present invention is a system for enabling users to receive real-time expert advice in an autonomous vehicle. The system captures voice input from the user and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. An avatar with the most appropriate expertise for the question is then selected, and an answer is generated using a generative artificial intelligence (AI) model. The answer is then displayed on the vehicle's display and provided to the user.

[0212] System configuration

[0213] The system includes the following means:

[0214] 1. Capture voice input

[0215] A microphone built into the vehicle captures the user's voice input and is highly sensitive, capturing clear audio while suppressing interior noise.

[0216] 2. Sending audio data

[0217] The captured voice data is sent to a cloud server via the vehicle's communication system, which achieves low latency and high-speed data transfer.

[0218] 3. Speech Recognition and Text Conversion

[0219] The cloud server converts the voice data into text data using a speech recognition algorithm, such as the Google Speech Recognition API.

[0220] 4. Intent Analysis

[0221] The converted text data undergoes intent analysis using an NLP engine (e.g., Hugging Face Transformers), where the "question intent" is extracted.

[0222] 5. Select a professional avatar

[0223] The server selects an avatar with the appropriate expertise based on the intent of the question. For example, if the question is about "tips for driving on rainy days," an avatar of a driving safety expert will be selected.

[0224] 6. Answer Generation

[0225] The question is input as a prompt into a generative AI model (for example, OpenAI's GPT-3), and an appropriate answer is generated. An example of a prompt is "User question: What are some tips for driving on rainy days? Please select an avatar that will provide an appropriate answer."

[0226] 7. Displaying answers and avatars

[0227] The generated answers and a 3D model of the avatar are then displayed on a display inside the vehicle, allowing the user to receive the answer both visually and audibly. The display uses a high-resolution screen, and the audio is output through high-quality speakers.

[0228] Specific examples

[0229] 1. User operations

[0230] User: "What are some tips for driving in the rain?"

[0231] 2. Server Processing

[0232] The system converts speech to text and analyzes the intent of the "Tips for driving on rainy days." It selects a driving safety expert avatar and generates the answer, "Drive at a moderate speed and keep a sufficient distance between vehicles."

[0233] 3. Terminal Processing

[0234] The device receives the avatar and the answer, which it then displays on the display. The avatar explains, "Please drive slowly and keep a safe distance from other vehicles," providing visual support.

[0235] The system of the present invention enables users to receive accurate and expert advice in real time within an autonomous vehicle, improving driving safety and comfort.

[0236] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0237] Step 1:

[0238] The user provides voice input in the autonomous vehicle. Here, the user provides voice information by speaking. For example, the user might say, "Please tell me some tips for driving on rainy days."

[0239] Step 2:

[0240] The device uses the vehicle's built-in microphone to capture voice input from the user, and the captured voice data is temporarily stored in the device as a digital signal.

[0241] Step 3:

[0242] The device transmits the captured audio data to the cloud server in real time. The data is transferred with low latency using a communication system (e.g., LTE or 5G network). The input is the audio data, and the output is the data sent to the cloud server.

[0243] Step 4:

[0244] The server converts the transmitted voice data into text data using a speech recognition algorithm (e.g., Google Speech Recognition API). The input is voice data, and the output is text data.

[0245] Step 5:

[0246] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Hugging Face Transformers) to understand the intent of the question. In this case, it extracts the intent of "advice on safe driving" from the question "Please tell me some tips for driving on rainy days." The input is text data, and the output is the intent of the question.

[0247] Step 6:

[0248] The server selects an avatar with appropriate expertise based on the intent of the question. For example, a driving safety expert avatar is selected. The input is the intent of the question, and the output is the selected avatar.

[0249] Step 7:

[0250] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the selected avatar and the question. An example prompt is "User question: What are some tips for driving on rainy days? Please select the avatar that provides the appropriate answer." The input is the question and prompt, and the output is the generated answer.

[0251] Step 8:

[0252] The server sends the generated 3D model of the avatar and the answer data to the terminal. The input is the 3D model of the avatar and the answer data, and the output is the data sent to the terminal.

[0253] Step 9:

[0254] The device decodes the received avatar and response data and displays them on the vehicle's display. The avatar then provides specific driving advice visually and audibly, such as "Slow down your speed and keep a safe distance between vehicles." The input is the decoded avatar and response data, and the output is the avatar and response displayed on the display.

[0255] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0256] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives, and by combining it with an emotion engine that recognizes the user's emotions, more personalized support is possible. The system aims to analyze voice input from the user, select an avatar with specialized knowledge corresponding to the question, and provide an appropriate answer. In addition, the emotion engine is used to identify the user's emotions, and the answer is adjusted based on this. Specific embodiments of the present invention are described below.

[0257] System Overview

[0258] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. An emotion engine then analyzes the user's emotions, and the avatar's answer is adjusted accordingly. The answer generated through these processes is displayed on the mirror-type device's display and provided to the user.

[0259] Program processing

[0260] Audio data capture and analysis

[0261] 1. Capture voice input

[0262] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[0263] 2. Sending audio data

[0264] The captured audio data is transmitted to a server in real time.

[0265] 3. Analysis of audio data

[0266] The server receives the voice data and converts it into text data using a speech recognition algorithm, for example, "Please teach me everyday English conversation."

[0267] Intention Analysis and Emotion Recognition

[0268] 4. Intent Analysis

[0269] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question, which reveals that the user is seeking information about everyday English conversation.

[0270] 5. Emotion recognition

[0271] The server uses an emotion engine to analyze the user's tone of voice, facial expressions, body language, etc. to recognize the user's emotions, for example, whether the user is feeling stressed.

[0272] Selecting specialized avatars and generating answers

[0273] 6. Avatar Selection

[0274] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[0275] 7. Answer Generation

[0276] A question is input into the generative AI model, which generates an appropriate answer, such as "Good morning! This is a morning greeting."

[0277] 8. Emotional Adjustment

[0278] The server adjusts the generated answers depending on the user's emotions: for example, if the user is tired, the avatar will provide answers in a gentler tone.

[0279] Show Answers

[0280] 9. Submitting avatar and answer data

[0281] The server sends the generated 3D model of the avatar and the response data to the device.

[0282] 10. Display of Avatar

[0283] The terminal decodes the received avatar and response data and prepares it for display on the mirror device's display.

[0284] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[0285] Specific examples

[0286] Learning English

[0287] 1. User operations

[0288] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[0289] 2. Terminal Processing

[0290] The device captures the audio and sends it to the server.

[0291] 3. Server Processing

[0292] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[0293] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[0294] The emotion engine recognizes when a user is feeling anxious and adjusts responses to a gentler tone.

[0295] 4. Terminal Processing

[0296] The terminal receives the avatar and the answer and displays them on the mirror device.

[0297] The avatar gently explains, "Good morning! This is a morning greeting."

[0298] Technical Advice

[0299] 1. User operations

[0300] User: Say, "How do I use a wood stove safely?"

[0301] 2. Terminal Processing

[0302] The device captures the audio and sends it to the server.

[0303] 3. Server Processing

[0304] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[0305] An outdoor expert avatar is selected to generate answers regarding safe use.

[0306] The emotion engine recognizes that the user is relaxed and provides answers in a normal tone.

[0307] 4. Terminal Processing

[0308] The terminal receives the avatar and the answer and displays them on the mirror device.

[0309] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[0310] As described above, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[0311] The processing flow will be explained below.

[0312] Step 1:

[0313] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[0314] Step 2:

[0315] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[0316] Step 3:

[0317] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[0318] Step 4:

[0319] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[0320] Step 5:

[0321] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[0322] Step 6:

[0323] The server inputs a question into the generative AI model and generates an appropriate answer, such as "Good morning! This is a morning greeting."

[0324] Step 7:

[0325] The server uses an emotion engine to recognize the user's emotions, analyzing the user's voice tone and facial expression data to determine, for example, that the user is feeling anxious.

[0326] Step 8:

[0327] The server adjusts responses based on the user's emotions, for example, adjusting the avatar's dialogue to provide responses in a gentler tone if the user is feeling anxious.

[0328] Step 9:

[0329] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[0330] Step 10:

[0331] The device decodes the avatar and answer data received from the server, loads the decoded 3D avatar model, and prepares it for display on the mirror device's display.

[0332] Step 11:

[0333] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might explain in a gentle tone, "Good morning! This is a morning greeting."

[0334] Step 12:

[0335] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[0336] Step 13:

[0337] The server receives the new voice data, converts it back to text, analyzes the intent, generates an appropriate response, adjusts it again based on the user's sentiment, and sends it back to the device. This process can be repeated as needed.

[0338] This allows users to get real-time expert advice and quickly solve everyday problems.

[0339] Example 2

[0340] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0341] Conventional voice input analysis systems can convert user voice input into text, analyze the intent of the question based on the text data, and generate answers. However, these systems cannot adjust the answers taking into account the user's emotions, which leaves the user experience unsatisfied. Furthermore, there is a lack of interactive systems that provide users with appropriate answers visually and audibly through expert avatars.

[0342] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0343] In this invention, the server includes means for capturing a voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for adjusting the generated answer in accordance with emotional data of the user, and means for displaying the generated answer to the user via the avatar, thereby providing a personalized answer in accordance with the emotional state of the user and improving the user experience.

[0344] "Means for capturing voice input from a user" refers to technical means for capturing voice uttered by a user as electronic data using a voice capture device such as a microphone.

[0345] "Means for analyzing captured voice input and converting it into text data" refers to technical means for converting voice data into corresponding text data using a speech recognition algorithm.

[0346] The "means for analyzing the intent of a question based on text data" is a technical means for using a natural language processing engine to understand the intent of a user's question from converted text data.

[0347] "Means for selecting an avatar with the expertise corresponding to the intent of the question" refers to a technical means for selecting a virtual character (avatar) with the most appropriate skills and knowledge based on the analyzed content of the question.

[0348] The "means for generating an answer based on a selected avatar" refers to a technical means for generating an appropriate answer to a question based on the area of ​​expertise of the selected avatar.

[0349] "Means for adjusting the generated response in accordance with the user's emotional data" refers to a technical means for analyzing the user's emotional state and adjusting the tone of voice, expression, etc. of the generated response in accordance with that emotion.

[0350] The "means for displaying the generated answer to the user through an avatar" refers to a technical means for presenting the generated answer to the user through the visual and audio of the avatar.

[0351] This invention is a system that provides expert advice to users on various problems they encounter in their daily lives. Furthermore, by combining it with a function to recognize the user's emotions, it provides more personalized assistance. The system analyzes the user's voice input, selects an avatar with the expertise required for the question, and provides an appropriate answer. It also uses an emotion engine to identify the user's emotions and adjusts the answer accordingly.

[0352] The system operates using the following hardware and software:

[0353] Hardware

[0354] 1. Mirror-type device: This device has a built-in highly sensitive microphone and camera to capture voice input and facial expression data from the user.

[0355] 2. Server: A high-performance computer that analyzes voice data and emotional data, selects avatars, and generates answers.

[0356] software

[0357] 1. Speech recognition algorithms (e.g., Google Speech-to-Text): Used to convert user voice input into text data.

[0358] 2. Natural language processing engine (e.g. GPT-4): Used to analyze the intent of the user's question from text data.

[0359] 3. Emotion engine: Used to recognize emotions by analyzing the user's voice tone and facial expression data.

[0360] 4. Generative AI models (e.g. ChatGPT): Used to generate appropriate answers to user questions.

[0361] How to operate

[0362] The user speaks a question to the mirror device. For example, if the user says, "Please teach me everyday English conversation," the following processing flow is executed.

[0363] 1. Audio input capture: The microphone on the mirror device captures the user's voice.

[0364] 2. Audio data transmission: The captured audio data is transmitted to the server in real time.

[0365] 3. Voice data analysis: The server uses a voice recognition algorithm to convert the voice into text data.

[0366] 4. Intent analysis: The text data is passed through a natural language processing engine to analyze the intent of the question. For example, the intent may be, "I'm looking for information about everyday English conversation."

[0367] 5. Emotion Recognition: The emotion engine analyzes the user's voice tone and facial expression data to recognize the user's emotional state, for example, whether the user is feeling stressed.

[0368] 6. Avatar selection: The AI ​​model selects the avatar with the expertise that best fits the intent of the question, for example, an English teacher avatar.

[0369] 7. Answer generation: A question is input into the generative AI model, and an appropriate answer is generated. For example, the answer generated might be, "Good morning! This is a morning greeting."

[0370] 8. Emotion-based adjustment: The server adjusts the generated answers depending on the user's emotions. For example, if the user is tired, the avatar will provide answers in a gentler tone.

[0371] 9. Answer display: The avatar and answer data are sent from the server to the terminal and displayed on the mirror device's display.

[0372] Specific examples

[0373] Learning English

[0374] User operation: Talk to the mirror device, saying, "Please teach me everyday English conversation."

[0375] Device processing: The device captures the audio and sends it to the server.

[0376] Server processing: The server converts the speech into text and analyzes the intent of the "everyday English conversation." It selects an English teacher avatar and generates a response such as "Good morning! This is a morning greeting." The emotion engine recognizes that the user is feeling anxious and adjusts the response to a gentler tone.

[0377] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar gently explains, "Good morning! This is a morning greeting."

[0378] Technical Advice

[0379] User Action: Say, "How do I use a wood stove safely?"

[0380] Device processing: The device captures the audio and sends it to the server.

[0381] Server processing: The server converts the speech to text and analyzes the intent of "how to use a wood stove." It selects an outdoor expert avatar and generates an answer on how to use it safely. The emotion engine recognizes that the user is relaxed and provides an answer in a normal tone.

[0382] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar explains, "First, make sure there are no flammable objects around the wood stove."

[0383] In this way, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[0384] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0385] Step 1:

[0386] Capture audio input

[0387] The terminal uses a high-sensitivity microphone to capture the voice of the user speaking into the mirror-type device. For example, the user might say, "Please teach me everyday English conversation."

[0388] Input: User's voice

[0389] Output: Digital audio data

[0390] Step 2:

[0391] Sending audio data

[0392] The device compresses the captured audio data and transmits it to the server with low latency, using a protocol to minimize data delays and loss.

[0393] Input: Audio data captured on the device

[0394] Output: Audio data sent to the server

[0395] Step 3:

[0396] Analysis of audio data

[0397] The server converts the received voice data into text data by running it through a speech recognition algorithm (e.g., Google Speech-to-Text). Specifically, it analyzes the frequency patterns of the voice and converts them into corresponding strings of characters.

[0398] Input: Transmitted audio data

[0399] Output: Text data (e.g., "Please teach me everyday English conversations")

[0400] Step 4:

[0401] Intention Analysis

[0402] The server then runs the converted text data through a natural language processing (NLP) engine (e.g., GPT-4) to analyze the intent of the question. During this process, the server identifies the information the user is looking for based on keywords and context within the text data.

[0403] Input: Text data

[0404] Output: Intent data (e.g., "I'm looking for information about everyday English conversations")

[0405] Step 5:

[0406] emotion recognition

[0407] The server uses an emotion engine to analyze the user's voice tone and facial expression data captured by the mirror device's camera to recognize the user's emotions, thereby obtaining information such as whether the user is stressed or relaxed.

[0408] Input: Voice tone data, facial expression data

[0409] Output: Emotion data (e.g., "The user feels anxious")

[0410] Step 6:

[0411] Avatar Selection

[0412] Based on the analyzed intent data, the server selects the avatar with the most appropriate expertise for the question. For example, if the question is about everyday English conversation, an English teacher avatar will be selected.

[0413] Input: Intent data

[0414] Output: Selected avatar

[0415] Step 7:

[0416] Generate answers

[0417] The server inputs the selected avatar's prompt into a generative AI model (e.g., ChatGPT) to generate an appropriate answer. The prompt includes the question and the user's emotional state.

[0418] Input: Prompt sentence (e.g., "The user is asking about everyday English conversation. The user feels anxious.")

[0419] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[0420] Step 8:

[0421] Emotional adjustment

[0422] The server adjusts the generated responses based on the user's emotional data, for example, if the user is tired, it modifies the generated responses to be gentler in tone and easier to understand.

[0423] Input: Generated answers, sentiment data

[0424] Output: A tailored response (e.g., "Good morning! This is a morning greeting" in a gentle tone)

[0425] Step 9:

[0426] Sending avatar and answer data

[0427] The server encodes the generated avatar and answer data and quickly transmits them to the terminal used by the user.

[0428] Input: Avatar data, adjusted answer

[0429] Output: Avatar data and answer data sent to the device

[0430] Step 10:

[0431] Display avatar

[0432] The terminal decodes the received avatar and answer data and displays them on the mirror device's display, allowing the user to receive the answers generated by the avatar visually and audibly.

[0433] Input: Submitted avatar data, response data

[0434] Output: An avatar displayed on the mirror device and a response (e.g., the avatar gently explains, "Good morning! This is a morning greeting.")

[0435] (Application example 2)

[0436] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0437] In current store operations, it is difficult for staff to provide customers with fast and accurate product information and advice, and staff's own emotions and stress can sometimes affect customer service. This can lead to issues such as lower customer satisfaction and reduced work efficiency. The present invention aims to solve these issues and enable store staff to receive expert advice in real time, improving the quality of customer service.

[0438] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with specialized knowledge corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for recognizing the user's emotion, means for adjusting the answer based on the recognized emotion, and means for displaying the generated answer to the user via an avatar. This enables store staff to provide accurate product information and expert advice in real time, improving customer satisfaction.

[0439] "Voice input" refers to voice data generated by a user speaking to the system through a microphone device.

[0440] "Capture" refers to taking in audio data using an input device such as a microphone.

[0441] "Analysis" refers to processing the captured voice data or text data to understand its meaning.

[0442] "Text data" refers to character data converted from voice input using voice recognition technology.

[0443] The "question intent" refers to the purpose or intention that indicates what the user wants to ask or what they are asking the system for.

[0444] "Expertise" refers to deep understanding and technical information in a particular field.

[0445] An "avatar" is a digital character that has specific capabilities or expertise and interacts with a user.

[0446] "Selection" refers to the act of choosing the most suitable candidate from among multiple candidates.

[0447] "Generation" refers to the act of the system creating the information or data it needs.

[0448] "Emotion" refers to a user's psychological state that is recognized based on their tone of voice, facial expressions, body language, etc.

[0449] "Tuning" refers to the act of modifying or optimizing the generated answer for specific conditions or circumstances.

[0450] "Display" refers to the act of outputting the generated information or avatar to a display and visually showing it to the user.

[0451] A system for realizing this invention has a function that allows store staff to wear smart glasses and obtain information about customer service and products in real time. Specific embodiments of the system will be described below.

[0452] Hardware Configuration

[0453] Smart glasses (e.g., Google Glass): Capture voice input and display information on a display.

[0454] Server: Performs voice data analysis, natural language processing, emotion recognition, avatar selection, and answer generation.

[0455] Software Configuration

[0456] Speech recognition engine (e.g., Google Speech-to-Text): Converts voice input into text data.

[0457] Natural language processing engines (e.g., SpaCy): Analyze the intent of text data.

[0458] Emotion engine (e.g. IBM Watson Tone Analyzer): Recognizes the user's emotions.

[0459] Generative AI models (e.g., GPT-4): Generate answers based on prompts.

[0460] Display control software: displays avatar and text information on the smart glasses display.

[0461] Data processing and calculation

[0462] 1. Capture audio input:

[0463] A microphone built into the smart glasses captures the staff's voice input.

[0464] For example: "Where can I find the product this customer is looking for?"

[0465] 2. Sending audio data:

[0466] The audio data is sent to a cloud server in real time.

[0467] 3. Analysis of audio data:

[0468] The server uses Google Speech-to-Text to convert the audio into text data.

[0469] 4. Intention Analysis and Emotion Recognition:

[0470] The converted text data is run through an NLP engine (SpaCy) to analyze the intent of the question.

[0471] The server uses IBM Watson Tone Analyzer to recognize emotions from staff members' tone of voice and facial expressions.

[0472] 5. Avatar Selection and Answer Generation:

[0473] The server selects the most suitable avatar based on the intent of the question.

[0474] A prompt is input into a generative AI model (GPT-4) to generate an answer.

[0475] Example prompt: "The product you are looking for is located on Aisle 3."

[0476] 6. Adjusting and Displaying Answers:

[0477] Tailor the generated answers depending on the sentiment.

[0478] The adjusted answers and avatar are sent to the smart glasses and displayed on the screen.

[0479] Example: The avatar explains in a friendly tone, "The product you are looking for is on Aisle 3."

[0480] Specific examples

[0481] Example 1: How to use it when guiding the location of a product

[0482] 1. User Action:

[0483] A staff member speaks to the smart glasses and asks, "Where is the product this customer is looking for?"

[0484] 2. Terminal processing:

[0485] The smart glasses capture the audio and send it to a cloud server.

[0486] 3. Server processing:

[0487] The server converts the voice into text and analyzes the intent of the question as a "product search."

[0488] The avatar that best suits the question is selected, and the prompt is input into a generative AI model to generate an answer.

[0489] Tailor the generated answer based on sentiment, e.g., "The product you're looking for can be found on Aisle 3."

[0490] 4. Terminal processing:

[0491] The smart glasses display the received avatar and answer on the screen.

[0492] The avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[0493] This system allows store staff to provide information to customers quickly and accurately, improving customer satisfaction.

[0494] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0495] Step 1:

[0496] The user inputs voice into the smart glasses. The input is made by the staff member through a microphone, for example, "Where is the product this customer is looking for?" The smart glasses capture this.

[0497] Step 2:

[0498] The device transmits the captured voice input to the cloud server. The input is the user's voice data, and the output is the transmission of the voice data to the cloud server. For data processing, the voice data is transmitted in real time.

[0499] Step 3:

[0500] The server converts the received voice data into text data using Google Speech-to-Text. The input is voice data, and the output is text data. A speech recognition algorithm is applied to the data.

[0501] Step 4:

[0502] The server runs the converted text data through a natural language processing engine (SpaCy) to analyze the intent of the question. The input is text data, and the output is the intent of the question. The NLP engine analyzes the text data as a data calculation.

[0503] Step 5:

[0504] The server then processes the analyzed text data through an emotion engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The input is data derived from voice tone and facial expressions, and the output is the user's emotional state. An emotion analysis algorithm is applied to process the data.

[0505] Step 6:

[0506] The server selects the avatar with the most appropriate expertise based on the intent of the question and the perceived emotion. The input is the intent of the question and the emotional state, and the output is the selected avatar. The selection algorithm is executed as a data operation.

[0507] Step 7:

[0508] The server inputs a prompt into a generative AI model (GPT-4) to generate an answer. The input is a prompt based on the intent of the question, and the output is a generated answer. As a data calculation, the generative AI model analyzes the prompt and generates an answer.

[0509] Step 8:

[0510] The server adjusts the generated answer according to the user's emotions. The input is the generated answer and the emotional state, and the output is the adjusted answer. The adjustment algorithm is applied as a data operation.

[0511] Step 9:

[0512] The server sends the adjusted answer and avatar to the smart glasses. The input is the adjusted answer and avatar data, and the output is the data transmission to the smart glasses. A transmission protocol is used for data processing.

[0513] Step 10:

[0514] The device displays the received avatar and answer on the smart glasses' display. The input is the received avatar and answer data, and the output is the visuals and audio displayed on the display. Specifically, the avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[0515] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0516] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0517] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0518] [Second embodiment]

[0519] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0520] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0521] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0522] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0523] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0524] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0525] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0526] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0527] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0528] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0529] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0530] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0531] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[0532] System Overview

[0533] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[0534] Program processing

[0535] Audio data capture and analysis

[0536] 1. Capture voice input

[0537] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[0538] 2. Sending audio data

[0539] The captured audio data is transmitted to a server in real time.

[0540] 3. Analysis of audio data

[0541] The server receives the voice data and converts it into text using a speech recognition algorithm.

[0542] 4. Intent Analysis

[0543] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question. For example, the intent of "everyday English conversation" is extracted.

[0544] Selecting specialized avatars and generating answers

[0545] 5. Select your avatar

[0546] Based on the intent of the question, the server selects an avatar with appropriate expertise, for example, an English teacher avatar.

[0547] 6. Answer Generation

[0548] A question is input into the generative AI model, which generates an appropriate answer, such as an English phrase like "Good morning! This is a morning greeting."

[0549] Show Answers

[0550] 7. Submitting avatar and answer data

[0551] The server sends the generated 3D model of the avatar and the response data to the device.

[0552] 8. Displaying Avatars

[0553] The terminal decodes the received avatar and response data and prepares it for display on the mirror device.

[0554] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[0555] Specific examples

[0556] Learning English

[0557] 1. User operations

[0558] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[0559] 2. Terminal Processing

[0560] The device captures the audio and sends it to the server.

[0561] 3. Server Processing

[0562] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[0563] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[0564] 4. Terminal Processing

[0565] The terminal receives the avatar and the answer and displays them on the mirror device.

[0566] The avatar explains, "Good morning! This is a morning greeting."

[0567] Technical Advice

[0568] 1. User operations

[0569] User: Say, "How do I use a wood stove safely?"

[0570] 2. Terminal Processing

[0571] The device captures the audio and sends it to the server.

[0572] 3. Server Processing

[0573] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[0574] An outdoor expert avatar is selected to generate answers regarding safe use.

[0575] 4. Terminal Processing

[0576] The terminal receives the avatar and the answer and displays them on the mirror device.

[0577] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[0578] As described above, the present invention is a system that interactively provides expert advice to users on various problems they may encounter, allowing users to obtain the information and guidance they need in real time.

[0579] The processing flow will be explained below.

[0580] Step 1:

[0581] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[0582] Step 2:

[0583] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[0584] Step 3:

[0585] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[0586] Step 4:

[0587] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[0588] Step 5:

[0589] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[0590] Step 6:

[0591] The server uses a generative AI model to generate an appropriate answer to the question through the selected avatar, for example, "Good morning! This is a morning greeting."

[0592] Step 7:

[0593] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[0594] Step 8:

[0595] The device decodes the avatar and answer data received from the server, loads the decoded 3D model of the avatar and answer data, and prepares them for display on the mirror device's display.

[0596] Step 9:

[0597] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might say, "Good morning! This is a morning greeting."

[0598] Step 10:

[0599] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[0600] Step 11:

[0601] The server receives the new voice data, converts it back to text, analyzes the intent, and generates an appropriate response. This process can be repeated as necessary.

[0602] This allows users to get real-time expert advice and quickly solve everyday problems.

[0603] Example 1

[0604] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0605] Conventional information provision systems have had difficulty responding quickly and accurately to the various problems users face. They also often only provide limited answers to questions that require specialized knowledge. Furthermore, there are issues with real-time performance and security in the process of efficiently analyzing a user's voice input, selecting an appropriate avatar, and displaying the answer.

[0606] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0607] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user through an avatar, means for compressing the captured voice input in real time and securely transmitting it, and means for decoding a three-dimensional model of the selected avatar and the text data and displaying them on a display device, thereby enabling expert advice to be provided quickly and accurately through the avatar for various problems faced by users.

[0608] "User" means any person or end user of the System.

[0609] "Voice input" refers to the words or sounds spoken by the user, and is the data that the system must recognize.

[0610] "Audio capture means" refers to a device or software that detects a user's voice input and captures it as digital data.

[0611] "Means for converting to text data" refers to the process of converting captured audio into text form using speech recognition technology.

[0612] "Means for analyzing the intent of a question" refers to natural language processing technology for understanding the purpose and requirements of a user's question from text data.

[0613] "Avatar selection method" refers to the process of selecting a virtual character with appropriate expertise based on the intent of the analyzed question.

[0614] "Means for generating answers" refers to the ability to use selected avatars and generative AI models to construct appropriate answers to users' questions.

[0615] "Means for displaying to the user through an avatar" refers to the process of providing the generated answer to the user visually and audibly through an avatar.

[0616] "Means for real-time compression and secure transmission" refers to the function of instantly compressing captured audio data and securely transmitting it to a server using encryption technology.

[0617] "Means for decoding the three-dimensional model and text data and displaying it on a display device" refers to the function of decoding the received data and appropriately rendering the three-dimensional avatar and text information for display on a display device.

[0618] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[0619] System Overview

[0620] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[0621] Audio data capture and analysis

[0622] 1. Capture voice input

[0623] The device captures voice input from the user using a microphone. For example, if the user says, "Please teach me everyday English conversation," the device captures the voice using a high-sensitivity microphone and uses noise-canceling technology to remove background noise.

[0624] 2. Sending audio data

[0625] The captured audio data is compressed in real time by the device and securely transmitted to the server using Transport Layer Security (TLS) encryption.

[0626] 3. Analysis of audio data

[0627] The server decodes the received audio data and converts it into text using a speech recognition algorithm (e.g., Google Speech-to-Text API), which is optimized to complete within a few seconds.

[0628] 4. Intent Analysis

[0629] The server then runs the converted text data through a natural language processing (NLP) engine (such as spaCy or the BERT model) to analyze the intent of the user's question. For example, if the intent is extracted as "everyday English conversation," an appropriate answer is prepared based on this.

[0630] Selecting specialized avatars and generating answers

[0631] 5. Select your avatar

[0632] The server selects the avatar with the most appropriate expertise based on the intent of the question, using a database of avatar skills, for example, selecting an "English teacher" avatar.

[0633] 6. Answer Generation

[0634] The server generates answers to questions using a generative AI model (e.g., OpenAI GPT-4). Specifically, in response to the prompt, "Please give me an example of an everyday English conversation," the server generates the answer, "Good morning! This is a morning greeting."

[0635] Show Answers

[0636] 7. Submitting avatar and answer data

[0637] The server combines the generated avatar's 3D model and the text response data into a single data packet and sends it to the terminal. This transmission uses UDP (User Datagram Protocol) to achieve low latency.

[0638] 8. Displaying Avatars

[0639] The device decodes the received data and prepares it for display on the mirror device's display. The device renders a 3D model of the avatar using the device's GPU and uses a voice synthesizer to output the response. The user hears the avatar displayed on the mirror device explain, "Good morning! This is a morning greeting."

[0640] Specific examples

[0641] 1. Example of user operation

[0642] Learning English: The user speaks to the mirror device, saying, "Please teach me everyday English conversation." The device captures the audio and sends it to the server. The server selects an English teacher avatar and generates a response, saying, "Good morning! This is a morning greeting." The device displays the avatar and the response on the mirror device, and the avatar explains the response aloud.

[0643] 2. Examples of technical advice

[0644] The user says, "Please tell me how to use a wood stove safely." The device captures the audio and sends it to the server. The server selects an outdoor expert avatar and generates a response: "First, make sure there are no flammable objects around the wood stove." The device displays the avatar and the response on a mirror-type device, and the avatar explains.

[0645] The system allows users to receive expert advice in real time.

[0646] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0647] Step 1: Capturing Audio Input

[0648] Specific operation: The user stands in front of the mirror-type device and speaks to it, such as "Please teach me some everyday English conversation."

[0649] Input: User's voice

[0650] Processing: The device's built-in high-sensitivity microphone captures the user's voice and uses noise-canceling technology to remove background noise.

[0651] Output: Captured audio data

[0652] Step 2: Sending audio data

[0653] Specific operation: The captured audio data is sent from the device to the server in real time.

[0654] Input: Captured audio data

[0655] Processing: The device compresses the audio data and sends it securely to the server using TLS (Transport Layer Security) encryption.

[0656] Output: Encrypted audio data sent to the server

[0657] Step 3: Analyzing the audio data

[0658] Specific operation: The server analyzes the received voice data using a voice recognition algorithm.

[0659] Input: Encrypted audio data sent to the server

[0660] Processing: The server decodes the audio data and converts it to text using a speech recognition algorithm (e.g., Google Speech-to-Text API).

[0661] Output: Converted text data

[0662] Step 4: Intent Analysis

[0663] Specific operation: The server analyzes the converted text data and understands the intent of the user's question.

[0664] Input: Converted text data

[0665] Processing: The server runs the text data through a natural language processing (NLP) engine (e.g., spaCy or the BERT model) to analyze the intent of the question.

[0666] Output: Analyzed intent information (e.g., "I want to learn everyday English conversation")

[0667] Step 5: Choose your avatar

[0668] Specific operation: The server selects an appropriate avatar.

[0669] Input: Parsed intent information

[0670] Processing: Based on the analyzed intention information, the server refers to the avatar skill database and selects an avatar with appropriate expertise.

[0671] Output: Information about the selected avatar (e.g., an English teacher avatar)

[0672] Step 6: Generate an answer

[0673] Specific behavior: The server generates an answer based on the selected avatar.

[0674] Input: Selected avatar information and intention information

[0675] Processing: The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers to questions.

[0676] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[0677] Step 7: Submit your avatar and answer data

[0678] Specific operation: The server sends the generated avatar's three-dimensional model and text response data to the terminal.

[0679] Input: Generated answers and a 3D avatar model

[0680] Processing: The server collects the generated data into a single data packet and sends it to the terminal using UDP (User Datagram Protocol).

[0681] Output: Data packets sent to the terminal

[0682] Step 8: Displaying the Avatar

[0683] Specific operation: The terminal decodes the received avatar data and response data and displays them on the mirror device.

[0684] Input: Data packets sent to the terminal

[0685] Processing: The device decodes the data, uses the GPU to render a 3D avatar, and uses a voice synthesizer to voice the response.

[0686] Output: Avatar displayed on a mirror device and voice response

[0687] Through these specific actions and data flows at each processing step, real-time expert advice is provided to the user.

[0688] (Application example 1)

[0689] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0690] The objective of this invention is to improve driving safety and the user experience by providing real-time expert advice through voice input in response to various questions and anxieties that users may encounter in an autonomous vehicle. In particular, there is a need to select an appropriate avatar and provide easy-to-understand answers to specific questions such as road traffic information, explanations of vehicle functions, and driving advice.

[0691] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0692] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the converted text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user via the avatar, and means for displaying the generated answer on a vehicle display, thereby enabling the user to receive reliable information and advice in real time within the autonomous vehicle.

[0693] "Voice input" refers to input of voice information generated by a user speaking.

[0694] "Capture" means capturing surrounding sounds using a device such as a microphone.

[0695] "Analysis" is the process of interpreting data and extracting meaning and patterns.

[0696] "Text data" refers to data that has been converted from voice input into text information through analysis.

[0697] The "question intent" indicates the content and purpose of the question that the user is trying to ask through voice input.

[0698] "Specialist knowledge" refers to advanced knowledge or information about a particular area or field.

[0699] An "avatar" is a virtual character that represents a specific role or persona in the digital space.

[0700] "Answer generation" is the process of creating an appropriate answer to a user's question.

[0701] "Display" is the act of visually presenting information to a user.

[0702] "Vehicle display" means a screen device installed inside an autonomous vehicle for displaying information.

[0703] A "generative artificial intelligence model" is a type of artificial intelligence that generates text data in natural language based on given prompts.

[0704] System Overview

[0705] The present invention is a system for enabling users to receive real-time expert advice in an autonomous vehicle. The system captures voice input from the user and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. An avatar with the most appropriate expertise for the question is then selected, and an answer is generated using a generative artificial intelligence (AI) model. The answer is then displayed on the vehicle's display and provided to the user.

[0706] System configuration

[0707] The system includes the following means:

[0708] 1. Capture voice input

[0709] A microphone built into the vehicle captures the user's voice input and is highly sensitive, capturing clear audio while suppressing interior noise.

[0710] 2. Sending audio data

[0711] The captured voice data is sent to a cloud server via the vehicle's communication system, which achieves low latency and high-speed data transfer.

[0712] 3. Speech Recognition and Text Conversion

[0713] The cloud server converts the voice data into text data using a speech recognition algorithm, such as the Google Speech Recognition API.

[0714] 4. Intent Analysis

[0715] The converted text data undergoes intent analysis using an NLP engine (e.g., Hugging Face Transformers), where the "question intent" is extracted.

[0716] 5. Select a professional avatar

[0717] The server selects an avatar with the appropriate expertise based on the intent of the question. For example, if the question is about "tips for driving on rainy days," an avatar of a driving safety expert will be selected.

[0718] 6. Answer Generation

[0719] The question is input as a prompt into a generative AI model (for example, OpenAI's GPT-3), and an appropriate answer is generated. An example of a prompt is "User question: What are some tips for driving on rainy days? Please select an avatar that will provide an appropriate answer."

[0720] 7. Displaying answers and avatars

[0721] The generated answers and a 3D model of the avatar are then displayed on a display inside the vehicle, allowing the user to receive the answer both visually and audibly. The display uses a high-resolution screen, and the audio is output through high-quality speakers.

[0722] Specific examples

[0723] 1. User operations

[0724] User: "What are some tips for driving in the rain?"

[0725] 2. Server Processing

[0726] The system converts speech to text and analyzes the intent of the "Tips for driving on rainy days." It selects a driving safety expert avatar and generates the answer, "Drive at a moderate speed and keep a sufficient distance between vehicles."

[0727] 3. Terminal Processing

[0728] The device receives the avatar and the answer, which it then displays on the display. The avatar explains, "Please drive slowly and keep a safe distance from other vehicles," providing visual support.

[0729] The system of the present invention enables users to receive accurate and expert advice in real time within an autonomous vehicle, improving driving safety and comfort.

[0730] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0731] Step 1:

[0732] The user provides voice input in the autonomous vehicle. Here, the user provides voice information by speaking. For example, the user might say, "Please tell me some tips for driving on rainy days."

[0733] Step 2:

[0734] The device uses the vehicle's built-in microphone to capture voice input from the user, and the captured voice data is temporarily stored in the device as a digital signal.

[0735] Step 3:

[0736] The device transmits the captured audio data to the cloud server in real time. The data is transferred with low latency using a communication system (e.g., LTE or 5G network). The input is the audio data, and the output is the data sent to the cloud server.

[0737] Step 4:

[0738] The server converts the transmitted voice data into text data using a speech recognition algorithm (e.g., Google Speech Recognition API). The input is voice data, and the output is text data.

[0739] Step 5:

[0740] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Hugging Face Transformers) to understand the intent of the question. In this case, it extracts the intent of "advice on safe driving" from the question "Please tell me some tips for driving on rainy days." The input is text data, and the output is the intent of the question.

[0741] Step 6:

[0742] The server selects an avatar with appropriate expertise based on the intent of the question. For example, a driving safety expert avatar is selected. The input is the intent of the question, and the output is the selected avatar.

[0743] Step 7:

[0744] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the selected avatar and the question. An example prompt is "User question: What are some tips for driving on rainy days? Please select the avatar that provides the appropriate answer." The input is the question and prompt, and the output is the generated answer.

[0745] Step 8:

[0746] The server sends the generated 3D model of the avatar and the answer data to the terminal. The input is the 3D model of the avatar and the answer data, and the output is the data sent to the terminal.

[0747] Step 9:

[0748] The device decodes the received avatar and response data and displays them on the vehicle's display. The avatar then provides specific driving advice visually and audibly, such as "Slow down your speed and keep a safe distance between vehicles." The input is the decoded avatar and response data, and the output is the avatar and response displayed on the display.

[0749] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0750] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives, and by combining it with an emotion engine that recognizes the user's emotions, more personalized support is possible. The system aims to analyze voice input from the user, select an avatar with specialized knowledge corresponding to the question, and provide an appropriate answer. In addition, the emotion engine is used to identify the user's emotions, and the answer is adjusted based on this. Specific embodiments of the present invention are described below.

[0751] System Overview

[0752] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. An emotion engine then analyzes the user's emotions, and the avatar's answer is adjusted accordingly. The answer generated through these processes is displayed on the mirror-type device's display and provided to the user.

[0753] Program processing

[0754] Audio data capture and analysis

[0755] 1. Capture voice input

[0756] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[0757] 2. Sending audio data

[0758] The captured audio data is transmitted to a server in real time.

[0759] 3. Analysis of audio data

[0760] The server receives the voice data and converts it into text data using a speech recognition algorithm, for example, "Please teach me everyday English conversation."

[0761] Intention Analysis and Emotion Recognition

[0762] 4. Intent Analysis

[0763] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question, which reveals that the user is seeking information about everyday English conversation.

[0764] 5. Emotion recognition

[0765] The server uses an emotion engine to analyze the user's tone of voice, facial expressions, body language, etc. to recognize the user's emotions, for example, whether the user is feeling stressed.

[0766] Selecting specialized avatars and generating answers

[0767] 6. Avatar Selection

[0768] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[0769] 7. Answer Generation

[0770] A question is input into the generative AI model, which generates an appropriate answer, such as "Good morning! This is a morning greeting."

[0771] 8. Emotional Adjustment

[0772] The server adjusts the generated answers depending on the user's emotions: for example, if the user is tired, the avatar will provide answers in a gentler tone.

[0773] Show Answers

[0774] 9. Submitting avatar and answer data

[0775] The server sends the generated 3D model of the avatar and the response data to the device.

[0776] 10. Display of Avatar

[0777] The terminal decodes the received avatar and response data and prepares it for display on the mirror device's display.

[0778] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[0779] Specific examples

[0780] Learning English

[0781] 1. User operations

[0782] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[0783] 2. Terminal Processing

[0784] The device captures the audio and sends it to the server.

[0785] 3. Server Processing

[0786] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[0787] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[0788] The emotion engine recognizes when a user is feeling anxious and adjusts responses to a gentler tone.

[0789] 4. Terminal Processing

[0790] The terminal receives the avatar and the answer and displays them on the mirror device.

[0791] The avatar gently explains, "Good morning! This is a morning greeting."

[0792] Technical Advice

[0793] 1. User operations

[0794] User: Say, "How do I use a wood stove safely?"

[0795] 2. Terminal Processing

[0796] The device captures the audio and sends it to the server.

[0797] 3. Server Processing

[0798] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[0799] An outdoor expert avatar is selected to generate answers regarding safe use.

[0800] The emotion engine recognizes that the user is relaxed and provides answers in a normal tone.

[0801] 4. Terminal Processing

[0802] The terminal receives the avatar and the answer and displays them on the mirror device.

[0803] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[0804] As described above, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[0805] The processing flow will be explained below.

[0806] Step 1:

[0807] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[0808] Step 2:

[0809] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[0810] Step 3:

[0811] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[0812] Step 4:

[0813] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[0814] Step 5:

[0815] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[0816] Step 6:

[0817] The server inputs a question into the generative AI model and generates an appropriate answer, such as "Good morning! This is a morning greeting."

[0818] Step 7:

[0819] The server uses an emotion engine to recognize the user's emotions, analyzing the user's voice tone and facial expression data to determine, for example, that the user is feeling anxious.

[0820] Step 8:

[0821] The server adjusts responses based on the user's emotions, for example, adjusting the avatar's dialogue to provide responses in a gentler tone if the user is feeling anxious.

[0822] Step 9:

[0823] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[0824] Step 10:

[0825] The device decodes the avatar and answer data received from the server, loads the decoded 3D avatar model, and prepares it for display on the mirror device's display.

[0826] Step 11:

[0827] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might explain in a gentle tone, "Good morning! This is a morning greeting."

[0828] Step 12:

[0829] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[0830] Step 13:

[0831] The server receives the new voice data, converts it back to text, analyzes the intent, generates an appropriate response, adjusts it again based on the user's sentiment, and sends it back to the device. This process can be repeated as needed.

[0832] This allows users to get real-time expert advice and quickly solve everyday problems.

[0833] Example 2

[0834] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0835] Conventional voice input analysis systems can convert user voice input into text, analyze the intent of the question based on the text data, and generate answers. However, these systems cannot adjust the answers taking into account the user's emotions, which leaves the user experience unsatisfied. Furthermore, there is a lack of interactive systems that provide users with appropriate answers visually and audibly through expert avatars.

[0836] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0837] In this invention, the server includes means for capturing a voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for adjusting the generated answer in accordance with emotional data of the user, and means for displaying the generated answer to the user via the avatar, thereby providing a personalized answer in accordance with the emotional state of the user and improving the user experience.

[0838] "Means for capturing voice input from a user" refers to technical means for capturing voice uttered by a user as electronic data using a voice capture device such as a microphone.

[0839] "Means for analyzing captured voice input and converting it into text data" refers to technical means for converting voice data into corresponding text data using a speech recognition algorithm.

[0840] The "means for analyzing the intent of a question based on text data" is a technical means for using a natural language processing engine to understand the intent of a user's question from converted text data.

[0841] "Means for selecting an avatar with the expertise corresponding to the intent of the question" refers to a technical means for selecting a virtual character (avatar) with the most appropriate skills and knowledge based on the analyzed content of the question.

[0842] The "means for generating an answer based on a selected avatar" refers to a technical means for generating an appropriate answer to a question based on the area of ​​expertise of the selected avatar.

[0843] "Means for adjusting the generated response in accordance with the user's emotional data" refers to a technical means for analyzing the user's emotional state and adjusting the tone of voice, expression, etc. of the generated response in accordance with that emotion.

[0844] The "means for displaying the generated answer to the user through an avatar" refers to a technical means for presenting the generated answer to the user through the visual and audio of the avatar.

[0845] This invention is a system that provides expert advice to users on various problems they encounter in their daily lives. Furthermore, by combining it with a function to recognize the user's emotions, it provides more personalized assistance. The system analyzes the user's voice input, selects an avatar with the expertise required for the question, and provides an appropriate answer. It also uses an emotion engine to identify the user's emotions and adjusts the answer accordingly.

[0846] The system operates using the following hardware and software:

[0847] Hardware

[0848] 1. Mirror-type device: This device has a built-in highly sensitive microphone and camera to capture voice input and facial expression data from the user.

[0849] 2. Server: A high-performance computer that analyzes voice data and emotional data, selects avatars, and generates answers.

[0850] software

[0851] 1. Speech recognition algorithms (e.g., Google Speech-to-Text): Used to convert user voice input into text data.

[0852] 2. Natural language processing engine (e.g. GPT-4): Used to analyze the intent of the user's question from text data.

[0853] 3. Emotion engine: Used to recognize emotions by analyzing the user's voice tone and facial expression data.

[0854] 4. Generative AI models (e.g. ChatGPT): Used to generate appropriate answers to user questions.

[0855] How to operate

[0856] The user speaks a question to the mirror device. For example, if the user says, "Please teach me everyday English conversation," the following processing flow is executed.

[0857] 1. Audio input capture: The microphone on the mirror device captures the user's voice.

[0858] 2. Audio data transmission: The captured audio data is transmitted to the server in real time.

[0859] 3. Voice data analysis: The server uses a voice recognition algorithm to convert the voice into text data.

[0860] 4. Intent analysis: The text data is passed through a natural language processing engine to analyze the intent of the question. For example, the intent may be, "I'm looking for information about everyday English conversation."

[0861] 5. Emotion Recognition: The emotion engine analyzes the user's voice tone and facial expression data to recognize the user's emotional state, for example, whether the user is feeling stressed.

[0862] 6. Avatar selection: The AI ​​model selects the avatar with the expertise that best fits the intent of the question, for example, an English teacher avatar.

[0863] 7. Answer generation: A question is input into the generative AI model, and an appropriate answer is generated. For example, the answer generated might be, "Good morning! This is a morning greeting."

[0864] 8. Emotion-based adjustment: The server adjusts the generated answers depending on the user's emotions. For example, if the user is tired, the avatar will provide answers in a gentler tone.

[0865] 9. Answer display: The avatar and answer data are sent from the server to the terminal and displayed on the mirror device's display.

[0866] Specific examples

[0867] Learning English

[0868] User operation: Talk to the mirror device, saying, "Please teach me everyday English conversation."

[0869] Device processing: The device captures the audio and sends it to the server.

[0870] Server processing: The server converts the speech into text and analyzes the intent of the "everyday English conversation." It selects an English teacher avatar and generates a response such as "Good morning! This is a morning greeting." The emotion engine recognizes that the user is feeling anxious and adjusts the response to a gentler tone.

[0871] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar gently explains, "Good morning! This is a morning greeting."

[0872] Technical Advice

[0873] User Action: Say, "How do I use a wood stove safely?"

[0874] Device processing: The device captures the audio and sends it to the server.

[0875] Server processing: The server converts the speech to text and analyzes the intent of "how to use a wood stove." It selects an outdoor expert avatar and generates an answer on how to use it safely. The emotion engine recognizes that the user is relaxed and provides an answer in a normal tone.

[0876] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar explains, "First, make sure there are no flammable objects around the wood stove."

[0877] In this way, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[0878] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0879] Step 1:

[0880] Capture audio input

[0881] The terminal uses a high-sensitivity microphone to capture the voice of the user speaking into the mirror-type device. For example, the user might say, "Please teach me everyday English conversation."

[0882] Input: User's voice

[0883] Output: Digital audio data

[0884] Step 2:

[0885] Sending audio data

[0886] The device compresses the captured audio data and transmits it to the server with low latency, using a protocol to minimize data delays and loss.

[0887] Input: Audio data captured on the device

[0888] Output: Audio data sent to the server

[0889] Step 3:

[0890] Analysis of audio data

[0891] The server converts the received voice data into text data by running it through a speech recognition algorithm (e.g., Google Speech-to-Text). Specifically, it analyzes the frequency patterns of the voice and converts them into corresponding strings of characters.

[0892] Input: Transmitted audio data

[0893] Output: Text data (e.g., "Please teach me everyday English conversations")

[0894] Step 4:

[0895] Intention Analysis

[0896] The server then runs the converted text data through a natural language processing (NLP) engine (e.g., GPT-4) to analyze the intent of the question. During this process, the server identifies the information the user is looking for based on keywords and context within the text data.

[0897] Input: Text data

[0898] Output: Intent data (e.g., "I'm looking for information about everyday English conversations")

[0899] Step 5:

[0900] emotion recognition

[0901] The server uses an emotion engine to analyze the user's voice tone and facial expression data captured by the mirror device's camera to recognize the user's emotions, thereby obtaining information such as whether the user is stressed or relaxed.

[0902] Input: Voice tone data, facial expression data

[0903] Output: Emotion data (e.g., "The user feels anxious")

[0904] Step 6:

[0905] Avatar Selection

[0906] Based on the analyzed intent data, the server selects the avatar with the most appropriate expertise for the question. For example, if the question is about everyday English conversation, an English teacher avatar will be selected.

[0907] Input: Intent data

[0908] Output: Selected avatar

[0909] Step 7:

[0910] Generate answers

[0911] The server inputs the selected avatar's prompt into a generative AI model (e.g., ChatGPT) to generate an appropriate answer. The prompt includes the question and the user's emotional state.

[0912] Input: Prompt sentence (e.g., "The user is asking about everyday English conversation. The user feels anxious.")

[0913] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[0914] Step 8:

[0915] Emotional adjustment

[0916] The server adjusts the generated responses based on the user's emotional data, for example, if the user is tired, it modifies the generated responses to be gentler in tone and easier to understand.

[0917] Input: Generated answers, sentiment data

[0918] Output: A tailored response (e.g., "Good morning! This is a morning greeting" in a gentle tone)

[0919] Step 9:

[0920] Sending avatar and answer data

[0921] The server encodes the generated avatar and answer data and quickly transmits them to the terminal used by the user.

[0922] Input: Avatar data, adjusted answer

[0923] Output: Avatar data and answer data sent to the device

[0924] Step 10:

[0925] Display avatar

[0926] The terminal decodes the received avatar and answer data and displays them on the mirror device's display, allowing the user to receive the answers generated by the avatar visually and audibly.

[0927] Input: Submitted avatar data, response data

[0928] Output: An avatar displayed on the mirror device and a response (e.g., the avatar gently explains, "Good morning! This is a morning greeting.")

[0929] (Application example 2)

[0930] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0931] In current store operations, it is difficult for staff to provide customers with fast and accurate product information and advice, and staff's own emotions and stress can sometimes affect customer service. This can lead to issues such as lower customer satisfaction and reduced work efficiency. The present invention aims to solve these issues and enable store staff to receive expert advice in real time, improving the quality of customer service.

[0932] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with specialized knowledge corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for recognizing the user's emotion, means for adjusting the answer based on the recognized emotion, and means for displaying the generated answer to the user via an avatar. This enables store staff to provide accurate product information and expert advice in real time, improving customer satisfaction.

[0933] "Voice input" refers to voice data generated by a user speaking to the system through a microphone device.

[0934] "Capture" refers to taking in audio data using an input device such as a microphone.

[0935] "Analysis" refers to processing the captured voice data or text data to understand its meaning.

[0936] "Text data" refers to character data converted from voice input using voice recognition technology.

[0937] The "question intent" refers to the purpose or intention that indicates what the user wants to ask or what they are asking the system for.

[0938] "Expertise" refers to deep understanding and technical information in a particular field.

[0939] An "avatar" is a digital character that has specific capabilities or expertise and interacts with a user.

[0940] "Selection" refers to the act of choosing the most suitable candidate from among multiple candidates.

[0941] "Generation" refers to the act of the system creating the information or data it needs.

[0942] "Emotion" refers to a user's psychological state that is recognized based on their tone of voice, facial expressions, body language, etc.

[0943] "Tuning" refers to the act of modifying or optimizing the generated answer for specific conditions or circumstances.

[0944] "Display" refers to the act of outputting the generated information or avatar to a display and visually showing it to the user.

[0945] A system for realizing this invention has a function that allows store staff to wear smart glasses and obtain information about customer service and products in real time. Specific embodiments of the system will be described below.

[0946] Hardware Configuration

[0947] Smart glasses (e.g., Google Glass): Capture voice input and display information on a display.

[0948] Server: Performs voice data analysis, natural language processing, emotion recognition, avatar selection, and answer generation.

[0949] Software Configuration

[0950] Speech recognition engine (e.g., Google Speech-to-Text): Converts voice input into text data.

[0951] Natural language processing engines (e.g., SpaCy): Analyze the intent of text data.

[0952] Emotion engine (e.g. IBM Watson Tone Analyzer): Recognizes the user's emotions.

[0953] Generative AI models (e.g., GPT-4): Generate answers based on prompts.

[0954] Display control software: displays avatar and text information on the smart glasses display.

[0955] Data processing and calculation

[0956] 1. Capture audio input:

[0957] A microphone built into the smart glasses captures the staff's voice input.

[0958] For example: "Where can I find the product this customer is looking for?"

[0959] 2. Sending audio data:

[0960] The audio data is sent to a cloud server in real time.

[0961] 3. Analysis of audio data:

[0962] The server uses Google Speech-to-Text to convert the audio into text data.

[0963] 4. Intention Analysis and Emotion Recognition:

[0964] The converted text data is run through an NLP engine (SpaCy) to analyze the intent of the question.

[0965] The server uses IBM Watson Tone Analyzer to recognize emotions from staff members' tone of voice and facial expressions.

[0966] 5. Avatar Selection and Answer Generation:

[0967] The server selects the most suitable avatar based on the intent of the question.

[0968] A prompt is input into a generative AI model (GPT-4) to generate an answer.

[0969] Example prompt: "The product you are looking for is located on Aisle 3."

[0970] 6. Adjusting and Displaying Answers:

[0971] Tailor the generated answers depending on the sentiment.

[0972] The adjusted answers and avatar are sent to the smart glasses and displayed on the screen.

[0973] Example: The avatar explains in a friendly tone, "The product you are looking for is on Aisle 3."

[0974] Specific examples

[0975] Example 1: How to use it when guiding the location of a product

[0976] 1. User Action:

[0977] A staff member speaks to the smart glasses and asks, "Where is the product this customer is looking for?"

[0978] 2. Terminal processing:

[0979] The smart glasses capture the audio and send it to a cloud server.

[0980] 3. Server processing:

[0981] The server converts the voice into text and analyzes the intent of the question as a "product search."

[0982] The avatar that best suits the question is selected, and the prompt is input into a generative AI model to generate an answer.

[0983] Tailor the generated answer based on sentiment, e.g., "The product you're looking for can be found on Aisle 3."

[0984] 4. Terminal processing:

[0985] The smart glasses display the received avatar and answer on the screen.

[0986] The avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[0987] This system allows store staff to provide information to customers quickly and accurately, improving customer satisfaction.

[0988] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0989] Step 1:

[0990] The user inputs voice into the smart glasses. The input is made by the staff member through a microphone, for example, "Where is the product this customer is looking for?" The smart glasses capture this.

[0991] Step 2:

[0992] The device transmits the captured voice input to the cloud server. The input is the user's voice data, and the output is the transmission of the voice data to the cloud server. For data processing, the voice data is transmitted in real time.

[0993] Step 3:

[0994] The server converts the received voice data into text data using Google Speech-to-Text. The input is voice data, and the output is text data. A speech recognition algorithm is applied to the data.

[0995] Step 4:

[0996] The server runs the converted text data through a natural language processing engine (SpaCy) to analyze the intent of the question. The input is text data, and the output is the intent of the question. The NLP engine analyzes the text data as a data calculation.

[0997] Step 5:

[0998] The server then processes the analyzed text data through an emotion engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The input is data derived from voice tone and facial expressions, and the output is the user's emotional state. An emotion analysis algorithm is applied to process the data.

[0999] Step 6:

[1000] The server selects the avatar with the most appropriate expertise based on the intent of the question and the perceived emotion. The input is the intent of the question and the emotional state, and the output is the selected avatar. The selection algorithm is executed as a data operation.

[1001] Step 7:

[1002] The server inputs a prompt into a generative AI model (GPT-4) to generate an answer. The input is a prompt based on the intent of the question, and the output is a generated answer. As a data calculation, the generative AI model analyzes the prompt and generates an answer.

[1003] Step 8:

[1004] The server adjusts the generated answer according to the user's emotions. The input is the generated answer and the emotional state, and the output is the adjusted answer. The adjustment algorithm is applied as a data operation.

[1005] Step 9:

[1006] The server sends the adjusted answer and avatar to the smart glasses. The input is the adjusted answer and avatar data, and the output is the data transmission to the smart glasses. A transmission protocol is used for data processing.

[1007] Step 10:

[1008] The device displays the received avatar and answer on the smart glasses' display. The input is the received avatar and answer data, and the output is the visuals and audio displayed on the display. Specifically, the avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[1009] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1010] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1011] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1012] [Third embodiment]

[1013] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1014] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1015] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1016] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1017] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1018] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1019] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1020] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1021] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1022] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1023] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1024] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1025] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[1026] System Overview

[1027] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[1028] Program processing

[1029] Audio data capture and analysis

[1030] 1. Capture voice input

[1031] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[1032] 2. Sending audio data

[1033] The captured audio data is transmitted to a server in real time.

[1034] 3. Analysis of audio data

[1035] The server receives the voice data and converts it into text using a speech recognition algorithm.

[1036] 4. Intent Analysis

[1037] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question. For example, the intent of "everyday English conversation" is extracted.

[1038] Selecting specialized avatars and generating answers

[1039] 5. Select your avatar

[1040] Based on the intent of the question, the server selects an avatar with appropriate expertise, for example, an English teacher avatar.

[1041] 6. Answer Generation

[1042] A question is input into the generative AI model, which generates an appropriate answer, such as an English phrase like "Good morning! This is a morning greeting."

[1043] Show Answers

[1044] 7. Submitting avatar and answer data

[1045] The server sends the generated 3D model of the avatar and the response data to the device.

[1046] 8. Displaying Avatars

[1047] The terminal decodes the received avatar and response data and prepares it for display on the mirror device.

[1048] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[1049] Specific examples

[1050] Learning English

[1051] 1. User operations

[1052] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[1053] 2. Terminal Processing

[1054] The device captures the audio and sends it to the server.

[1055] 3. Server Processing

[1056] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[1057] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[1058] 4. Terminal Processing

[1059] The terminal receives the avatar and the answer and displays them on the mirror device.

[1060] The avatar explains, "Good morning! This is a morning greeting."

[1061] Technical Advice

[1062] 1. User operations

[1063] User: Say, "How do I use a wood stove safely?"

[1064] 2. Terminal Processing

[1065] The device captures the audio and sends it to the server.

[1066] 3. Server Processing

[1067] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[1068] An outdoor expert avatar is selected to generate answers regarding safe use.

[1069] 4. Terminal Processing

[1070] The terminal receives the avatar and the answer and displays them on the mirror device.

[1071] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[1072] As described above, the present invention is a system that interactively provides expert advice to users on various problems they may encounter, allowing users to obtain the information and guidance they need in real time.

[1073] The processing flow will be explained below.

[1074] Step 1:

[1075] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[1076] Step 2:

[1077] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[1078] Step 3:

[1079] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[1080] Step 4:

[1081] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[1082] Step 5:

[1083] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[1084] Step 6:

[1085] The server uses a generative AI model to generate an appropriate answer to the question through the selected avatar, for example, "Good morning! This is a morning greeting."

[1086] Step 7:

[1087] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[1088] Step 8:

[1089] The device decodes the avatar and answer data received from the server, loads the decoded 3D model of the avatar and answer data, and prepares them for display on the mirror device's display.

[1090] Step 9:

[1091] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might say, "Good morning! This is a morning greeting."

[1092] Step 10:

[1093] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[1094] Step 11:

[1095] The server receives the new voice data, converts it back to text, analyzes the intent, and generates an appropriate response. This process can be repeated as necessary.

[1096] This allows users to get real-time expert advice and quickly solve everyday problems.

[1097] Example 1

[1098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1099] Conventional information provision systems have had difficulty responding quickly and accurately to the various problems users face. They also often only provide limited answers to questions that require specialized knowledge. Furthermore, there are issues with real-time performance and security in the process of efficiently analyzing a user's voice input, selecting an appropriate avatar, and displaying the answer.

[1100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1101] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user through an avatar, means for compressing the captured voice input in real time and securely transmitting it, and means for decoding a three-dimensional model of the selected avatar and the text data and displaying them on a display device, thereby enabling expert advice to be provided quickly and accurately through the avatar for various problems faced by users.

[1102] "User" means any person or end user of the System.

[1103] "Voice input" refers to the words or sounds spoken by the user, and is the data that the system must recognize.

[1104] "Audio capture means" refers to a device or software that detects a user's voice input and captures it as digital data.

[1105] "Means for converting to text data" refers to the process of converting captured audio into text form using speech recognition technology.

[1106] "Means for analyzing the intent of a question" refers to natural language processing technology for understanding the purpose and requirements of a user's question from text data.

[1107] "Avatar selection method" refers to the process of selecting a virtual character with appropriate expertise based on the intent of the analyzed question.

[1108] "Means for generating answers" refers to the ability to use selected avatars and generative AI models to construct appropriate answers to users' questions.

[1109] "Means for displaying to the user through an avatar" refers to the process of providing the generated answer to the user visually and audibly through an avatar.

[1110] "Means for real-time compression and secure transmission" refers to the function of instantly compressing captured audio data and securely transmitting it to a server using encryption technology.

[1111] "Means for decoding the three-dimensional model and text data and displaying it on a display device" refers to the function of decoding the received data and appropriately rendering the three-dimensional avatar and text information for display on a display device.

[1112] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[1113] System Overview

[1114] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[1115] Audio data capture and analysis

[1116] 1. Capture voice input

[1117] The device captures voice input from the user using a microphone. For example, if the user says, "Please teach me everyday English conversation," the device captures the voice using a high-sensitivity microphone and uses noise-canceling technology to remove background noise.

[1118] 2. Sending audio data

[1119] The captured audio data is compressed in real time by the device and securely transmitted to the server using Transport Layer Security (TLS) encryption.

[1120] 3. Analysis of audio data

[1121] The server decodes the received audio data and converts it into text using a speech recognition algorithm (e.g., Google Speech-to-Text API), which is optimized to complete within a few seconds.

[1122] 4. Intent Analysis

[1123] The server then runs the converted text data through a natural language processing (NLP) engine (such as spaCy or the BERT model) to analyze the intent of the user's question. For example, if the intent is extracted as "everyday English conversation," an appropriate answer is prepared based on this.

[1124] Selecting specialized avatars and generating answers

[1125] 5. Select your avatar

[1126] The server selects the avatar with the most appropriate expertise based on the intent of the question, using a database of avatar skills, for example, selecting an "English teacher" avatar.

[1127] 6. Answer Generation

[1128] The server generates answers to questions using a generative AI model (e.g., OpenAI GPT-4). Specifically, in response to the prompt, "Please give me an example of an everyday English conversation," the server generates the answer, "Good morning! This is a morning greeting."

[1129] Show Answers

[1130] 7. Submitting avatar and answer data

[1131] The server combines the generated avatar's 3D model and the text response data into a single data packet and sends it to the terminal. This transmission uses UDP (User Datagram Protocol) to achieve low latency.

[1132] 8. Displaying Avatars

[1133] The device decodes the received data and prepares it for display on the mirror device's display. The device renders a 3D model of the avatar using the device's GPU and uses a voice synthesizer to output the response. The user hears the avatar displayed on the mirror device explain, "Good morning! This is a morning greeting."

[1134] Specific examples

[1135] 1. Example of user operation

[1136] Learning English: The user speaks to the mirror device, saying, "Please teach me everyday English conversation." The device captures the audio and sends it to the server. The server selects an English teacher avatar and generates a response, saying, "Good morning! This is a morning greeting." The device displays the avatar and the response on the mirror device, and the avatar explains the response aloud.

[1137] 2. Examples of technical advice

[1138] The user says, "Please tell me how to use a wood stove safely." The device captures the audio and sends it to the server. The server selects an outdoor expert avatar and generates a response: "First, make sure there are no flammable objects around the wood stove." The device displays the avatar and the response on a mirror-type device, and the avatar explains.

[1139] The system allows users to receive expert advice in real time.

[1140] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1141] Step 1: Capturing Audio Input

[1142] Specific operation: The user stands in front of the mirror-type device and speaks to it, such as "Please teach me some everyday English conversation."

[1143] Input: User's voice

[1144] Processing: The device's built-in high-sensitivity microphone captures the user's voice and uses noise-canceling technology to remove background noise.

[1145] Output: Captured audio data

[1146] Step 2: Sending audio data

[1147] Specific operation: The captured audio data is sent from the device to the server in real time.

[1148] Input: Captured audio data

[1149] Processing: The device compresses the audio data and sends it securely to the server using TLS (Transport Layer Security) encryption.

[1150] Output: Encrypted audio data sent to the server

[1151] Step 3: Analyzing the audio data

[1152] Specific operation: The server analyzes the received voice data using a voice recognition algorithm.

[1153] Input: Encrypted audio data sent to the server

[1154] Processing: The server decodes the audio data and converts it to text using a speech recognition algorithm (e.g., Google Speech-to-Text API).

[1155] Output: Converted text data

[1156] Step 4: Intent Analysis

[1157] Specific operation: The server analyzes the converted text data and understands the intent of the user's question.

[1158] Input: Converted text data

[1159] Processing: The server runs the text data through a natural language processing (NLP) engine (e.g., spaCy or the BERT model) to analyze the intent of the question.

[1160] Output: Analyzed intent information (e.g., "I want to learn everyday English conversation")

[1161] Step 5: Choose your avatar

[1162] Specific operation: The server selects an appropriate avatar.

[1163] Input: Parsed intent information

[1164] Processing: Based on the analyzed intention information, the server refers to the avatar skill database and selects an avatar with appropriate expertise.

[1165] Output: Information about the selected avatar (e.g., an English teacher avatar)

[1166] Step 6: Generate an answer

[1167] Specific behavior: The server generates an answer based on the selected avatar.

[1168] Input: Selected avatar information and intention information

[1169] Processing: The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers to questions.

[1170] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[1171] Step 7: Submit your avatar and answer data

[1172] Specific operation: The server sends the generated avatar's three-dimensional model and text response data to the terminal.

[1173] Input: Generated answers and a 3D avatar model

[1174] Processing: The server collects the generated data into a single data packet and sends it to the terminal using UDP (User Datagram Protocol).

[1175] Output: Data packets sent to the terminal

[1176] Step 8: Displaying the Avatar

[1177] Specific operation: The terminal decodes the received avatar data and response data and displays them on the mirror device.

[1178] Input: Data packets sent to the terminal

[1179] Processing: The device decodes the data, uses the GPU to render a 3D avatar, and uses a voice synthesizer to voice the response.

[1180] Output: Avatar displayed on a mirror device and voice response

[1181] Through these specific actions and data flows at each processing step, real-time expert advice is provided to the user.

[1182] (Application example 1)

[1183] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1184] The objective of this invention is to improve driving safety and the user experience by providing real-time expert advice through voice input in response to various questions and anxieties that users may encounter in an autonomous vehicle. In particular, there is a need to select an appropriate avatar and provide easy-to-understand answers to specific questions such as road traffic information, explanations of vehicle functions, and driving advice.

[1185] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1186] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the converted text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user via the avatar, and means for displaying the generated answer on a vehicle display, thereby enabling the user to receive reliable information and advice in real time within the autonomous vehicle.

[1187] "Voice input" refers to input of voice information generated by a user speaking.

[1188] "Capture" means capturing surrounding sounds using a device such as a microphone.

[1189] "Analysis" is the process of interpreting data and extracting meaning and patterns.

[1190] "Text data" refers to data that has been converted from voice input into text information through analysis.

[1191] The "question intent" indicates the content and purpose of the question that the user is trying to ask through voice input.

[1192] "Specialist knowledge" refers to advanced knowledge or information about a particular area or field.

[1193] An "avatar" is a virtual character that represents a specific role or persona in the digital space.

[1194] "Answer generation" is the process of creating an appropriate answer to a user's question.

[1195] "Display" is the act of visually presenting information to a user.

[1196] "Vehicle display" means a screen device installed inside an autonomous vehicle for displaying information.

[1197] A "generative artificial intelligence model" is a type of artificial intelligence that generates text data in natural language based on given prompts.

[1198] System Overview

[1199] The present invention is a system for enabling users to receive real-time expert advice in an autonomous vehicle. The system captures voice input from the user and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. An avatar with the most appropriate expertise for the question is then selected, and an answer is generated using a generative artificial intelligence (AI) model. The answer is then displayed on the vehicle's display and provided to the user.

[1200] System configuration

[1201] The system includes the following means:

[1202] 1. Capture voice input

[1203] A microphone built into the vehicle captures the user's voice input and is highly sensitive, capturing clear audio while suppressing interior noise.

[1204] 2. Sending audio data

[1205] The captured voice data is sent to a cloud server via the vehicle's communication system, which achieves low latency and high-speed data transfer.

[1206] 3. Speech Recognition and Text Conversion

[1207] The cloud server converts the voice data into text data using a speech recognition algorithm, such as the Google Speech Recognition API.

[1208] 4. Intent Analysis

[1209] The converted text data undergoes intent analysis using an NLP engine (e.g., Hugging Face Transformers), where the "question intent" is extracted.

[1210] 5. Select a professional avatar

[1211] The server selects an avatar with the appropriate expertise based on the intent of the question. For example, if the question is about "tips for driving on rainy days," an avatar of a driving safety expert will be selected.

[1212] 6. Answer Generation

[1213] The question is input as a prompt into a generative AI model (for example, OpenAI's GPT-3), and an appropriate answer is generated. An example of a prompt is "User question: What are some tips for driving on rainy days? Please select an avatar that will provide an appropriate answer."

[1214] 7. Displaying answers and avatars

[1215] The generated answers and a 3D model of the avatar are then displayed on a display inside the vehicle, allowing the user to receive the answer both visually and audibly. The display uses a high-resolution screen, and the audio is output through high-quality speakers.

[1216] Specific examples

[1217] 1. User operations

[1218] User: "What are some tips for driving in the rain?"

[1219] 2. Server Processing

[1220] The system converts speech to text and analyzes the intent of the "Tips for driving on rainy days." It selects a driving safety expert avatar and generates the answer, "Drive at a moderate speed and keep a sufficient distance between vehicles."

[1221] 3. Terminal Processing

[1222] The device receives the avatar and the answer, which it then displays on the display. The avatar explains, "Please drive slowly and keep a safe distance from other vehicles," providing visual support.

[1223] The system of the present invention enables users to receive accurate and expert advice in real time within an autonomous vehicle, improving driving safety and comfort.

[1224] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1225] Step 1:

[1226] The user provides voice input in the autonomous vehicle. Here, the user provides voice information by speaking. For example, the user might say, "Please tell me some tips for driving on rainy days."

[1227] Step 2:

[1228] The device uses the vehicle's built-in microphone to capture voice input from the user, and the captured voice data is temporarily stored in the device as a digital signal.

[1229] Step 3:

[1230] The device transmits the captured audio data to the cloud server in real time. The data is transferred with low latency using a communication system (e.g., LTE or 5G network). The input is the audio data, and the output is the data sent to the cloud server.

[1231] Step 4:

[1232] The server converts the transmitted voice data into text data using a speech recognition algorithm (e.g., Google Speech Recognition API). The input is voice data, and the output is text data.

[1233] Step 5:

[1234] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Hugging Face Transformers) to understand the intent of the question. In this case, it extracts the intent of "advice on safe driving" from the question "Please tell me some tips for driving on rainy days." The input is text data, and the output is the intent of the question.

[1235] Step 6:

[1236] The server selects an avatar with appropriate expertise based on the intent of the question. For example, a driving safety expert avatar is selected. The input is the intent of the question, and the output is the selected avatar.

[1237] Step 7:

[1238] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the selected avatar and the question. An example prompt is "User question: What are some tips for driving on rainy days? Please select the avatar that provides the appropriate answer." The input is the question and prompt, and the output is the generated answer.

[1239] Step 8:

[1240] The server sends the generated 3D model of the avatar and the answer data to the terminal. The input is the 3D model of the avatar and the answer data, and the output is the data sent to the terminal.

[1241] Step 9:

[1242] The device decodes the received avatar and response data and displays them on the vehicle's display. The avatar then provides specific driving advice visually and audibly, such as "Slow down your speed and keep a safe distance between vehicles." The input is the decoded avatar and response data, and the output is the avatar and response displayed on the display.

[1243] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1244] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives, and by combining it with an emotion engine that recognizes the user's emotions, more personalized support is possible. The system aims to analyze voice input from the user, select an avatar with specialized knowledge corresponding to the question, and provide an appropriate answer. In addition, the emotion engine is used to identify the user's emotions, and the answer is adjusted based on this. Specific embodiments of the present invention are described below.

[1245] System Overview

[1246] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. An emotion engine then analyzes the user's emotions, and the avatar's answer is adjusted accordingly. The answer generated through these processes is displayed on the mirror-type device's display and provided to the user.

[1247] Program processing

[1248] Audio data capture and analysis

[1249] 1. Capture voice input

[1250] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[1251] 2. Sending audio data

[1252] The captured audio data is transmitted to a server in real time.

[1253] 3. Analysis of audio data

[1254] The server receives the voice data and converts it into text data using a speech recognition algorithm, for example, "Please teach me everyday English conversation."

[1255] Intention Analysis and Emotion Recognition

[1256] 4. Intent Analysis

[1257] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question, which reveals that the user is seeking information about everyday English conversation.

[1258] 5. Emotion recognition

[1259] The server uses an emotion engine to analyze the user's tone of voice, facial expressions, body language, etc. to recognize the user's emotions, for example, whether the user is feeling stressed.

[1260] Selecting specialized avatars and generating answers

[1261] 6. Avatar Selection

[1262] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[1263] 7. Answer Generation

[1264] A question is input into the generative AI model, which generates an appropriate answer, such as "Good morning! This is a morning greeting."

[1265] 8. Emotional Adjustment

[1266] The server adjusts the generated answers depending on the user's emotions: for example, if the user is tired, the avatar will provide answers in a gentler tone.

[1267] Show Answers

[1268] 9. Submitting avatar and answer data

[1269] The server sends the generated 3D model of the avatar and the response data to the device.

[1270] 10. Display of Avatar

[1271] The terminal decodes the received avatar and response data and prepares it for display on the mirror device's display.

[1272] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[1273] Specific examples

[1274] Learning English

[1275] 1. User operations

[1276] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[1277] 2. Terminal Processing

[1278] The device captures the audio and sends it to the server.

[1279] 3. Server Processing

[1280] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[1281] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[1282] The emotion engine recognizes when a user is feeling anxious and adjusts responses to a gentler tone.

[1283] 4. Terminal Processing

[1284] The terminal receives the avatar and the answer and displays them on the mirror device.

[1285] The avatar gently explains, "Good morning! This is a morning greeting."

[1286] Technical Advice

[1287] 1. User operations

[1288] User: Say, "How do I use a wood stove safely?"

[1289] 2. Terminal Processing

[1290] The device captures the audio and sends it to the server.

[1291] 3. Server Processing

[1292] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[1293] An outdoor expert avatar is selected to generate answers regarding safe use.

[1294] The emotion engine recognizes that the user is relaxed and provides answers in a normal tone.

[1295] 4. Terminal Processing

[1296] The terminal receives the avatar and the answer and displays them on the mirror device.

[1297] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[1298] As described above, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[1299] The processing flow will be explained below.

[1300] Step 1:

[1301] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[1302] Step 2:

[1303] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[1304] Step 3:

[1305] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[1306] Step 4:

[1307] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[1308] Step 5:

[1309] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[1310] Step 6:

[1311] The server inputs a question into the generative AI model and generates an appropriate answer, such as "Good morning! This is a morning greeting."

[1312] Step 7:

[1313] The server uses an emotion engine to recognize the user's emotions, analyzing the user's voice tone and facial expression data to determine, for example, that the user is feeling anxious.

[1314] Step 8:

[1315] The server adjusts responses based on the user's emotions, for example, adjusting the avatar's dialogue to provide responses in a gentler tone if the user is feeling anxious.

[1316] Step 9:

[1317] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[1318] Step 10:

[1319] The device decodes the avatar and answer data received from the server, loads the decoded 3D avatar model, and prepares it for display on the mirror device's display.

[1320] Step 11:

[1321] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might explain in a gentle tone, "Good morning! This is a morning greeting."

[1322] Step 12:

[1323] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[1324] Step 13:

[1325] The server receives the new voice data, converts it back to text, analyzes the intent, generates an appropriate response, adjusts it again based on the user's sentiment, and sends it back to the device. This process can be repeated as needed.

[1326] This allows users to get real-time expert advice and quickly solve everyday problems.

[1327] Example 2

[1328] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1329] Conventional voice input analysis systems can convert user voice input into text, analyze the intent of the question based on the text data, and generate answers. However, these systems cannot adjust the answers taking into account the user's emotions, which leaves the user experience unsatisfied. Furthermore, there is a lack of interactive systems that provide users with appropriate answers visually and audibly through expert avatars.

[1330] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1331] In this invention, the server includes means for capturing a voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for adjusting the generated answer in accordance with emotional data of the user, and means for displaying the generated answer to the user via the avatar, thereby providing a personalized answer in accordance with the emotional state of the user and improving the user experience.

[1332] "Means for capturing voice input from a user" refers to technical means for capturing voice uttered by a user as electronic data using a voice capture device such as a microphone.

[1333] "Means for analyzing captured voice input and converting it into text data" refers to technical means for converting voice data into corresponding text data using a speech recognition algorithm.

[1334] The "means for analyzing the intent of a question based on text data" is a technical means for using a natural language processing engine to understand the intent of a user's question from converted text data.

[1335] "Means for selecting an avatar with the expertise corresponding to the intent of the question" refers to a technical means for selecting a virtual character (avatar) with the most appropriate skills and knowledge based on the analyzed content of the question.

[1336] The "means for generating an answer based on a selected avatar" refers to a technical means for generating an appropriate answer to a question based on the area of ​​expertise of the selected avatar.

[1337] "Means for adjusting the generated response in accordance with the user's emotional data" refers to a technical means for analyzing the user's emotional state and adjusting the tone of voice, expression, etc. of the generated response in accordance with that emotion.

[1338] The "means for displaying the generated answer to the user through an avatar" refers to a technical means for presenting the generated answer to the user through the visual and audio of the avatar.

[1339] This invention is a system that provides expert advice to users on various problems they encounter in their daily lives. Furthermore, by combining it with a function to recognize the user's emotions, it provides more personalized assistance. The system analyzes the user's voice input, selects an avatar with the expertise required for the question, and provides an appropriate answer. It also uses an emotion engine to identify the user's emotions and adjusts the answer accordingly.

[1340] The system operates using the following hardware and software:

[1341] Hardware

[1342] 1. Mirror-type device: This device has a built-in highly sensitive microphone and camera to capture voice input and facial expression data from the user.

[1343] 2. Server: A high-performance computer that analyzes voice data and emotional data, selects avatars, and generates answers.

[1344] software

[1345] 1. Speech recognition algorithms (e.g., Google Speech-to-Text): Used to convert user voice input into text data.

[1346] 2. Natural language processing engine (e.g. GPT-4): Used to analyze the intent of the user's question from text data.

[1347] 3. Emotion engine: Used to recognize emotions by analyzing the user's voice tone and facial expression data.

[1348] 4. Generative AI models (e.g. ChatGPT): Used to generate appropriate answers to user questions.

[1349] How to operate

[1350] The user speaks a question to the mirror device. For example, if the user says, "Please teach me everyday English conversation," the following processing flow is executed.

[1351] 1. Audio input capture: The microphone on the mirror device captures the user's voice.

[1352] 2. Audio data transmission: The captured audio data is transmitted to the server in real time.

[1353] 3. Voice data analysis: The server uses a voice recognition algorithm to convert the voice into text data.

[1354] 4. Intent analysis: The text data is passed through a natural language processing engine to analyze the intent of the question. For example, the intent may be, "I'm looking for information about everyday English conversation."

[1355] 5. Emotion Recognition: The emotion engine analyzes the user's voice tone and facial expression data to recognize the user's emotional state, for example, whether the user is feeling stressed.

[1356] 6. Avatar selection: The AI ​​model selects the avatar with the expertise that best fits the intent of the question, for example, an English teacher avatar.

[1357] 7. Answer generation: A question is input into the generative AI model, and an appropriate answer is generated. For example, the answer generated might be, "Good morning! This is a morning greeting."

[1358] 8. Emotion-based adjustment: The server adjusts the generated answers depending on the user's emotions. For example, if the user is tired, the avatar will provide answers in a gentler tone.

[1359] 9. Answer display: The avatar and answer data are sent from the server to the terminal and displayed on the mirror device's display.

[1360] Specific examples

[1361] Learning English

[1362] User operation: Talk to the mirror device, saying, "Please teach me everyday English conversation."

[1363] Device processing: The device captures the audio and sends it to the server.

[1364] Server processing: The server converts the speech into text and analyzes the intent of the "everyday English conversation." It selects an English teacher avatar and generates a response such as "Good morning! This is a morning greeting." The emotion engine recognizes that the user is feeling anxious and adjusts the response to a gentler tone.

[1365] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar gently explains, "Good morning! This is a morning greeting."

[1366] Technical Advice

[1367] User Action: Say, "How do I use a wood stove safely?"

[1368] Device processing: The device captures the audio and sends it to the server.

[1369] Server processing: The server converts the speech to text and analyzes the intent of "how to use a wood stove." It selects an outdoor expert avatar and generates an answer on how to use it safely. The emotion engine recognizes that the user is relaxed and provides an answer in a normal tone.

[1370] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar explains, "First, make sure there are no flammable objects around the wood stove."

[1371] In this way, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[1372] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1373] Step 1:

[1374] Capture audio input

[1375] The terminal uses a high-sensitivity microphone to capture the voice of the user speaking into the mirror-type device. For example, the user might say, "Please teach me everyday English conversation."

[1376] Input: User's voice

[1377] Output: Digital audio data

[1378] Step 2:

[1379] Sending audio data

[1380] The device compresses the captured audio data and transmits it to the server with low latency, using a protocol to minimize data delays and loss.

[1381] Input: Audio data captured on the device

[1382] Output: Audio data sent to the server

[1383] Step 3:

[1384] Analysis of audio data

[1385] The server converts the received voice data into text data by running it through a speech recognition algorithm (e.g., Google Speech-to-Text). Specifically, it analyzes the frequency patterns of the voice and converts them into corresponding strings of characters.

[1386] Input: Transmitted audio data

[1387] Output: Text data (e.g., "Please teach me everyday English conversations")

[1388] Step 4:

[1389] Intention Analysis

[1390] The server then runs the converted text data through a natural language processing (NLP) engine (e.g., GPT-4) to analyze the intent of the question. During this process, the server identifies the information the user is looking for based on keywords and context within the text data.

[1391] Input: Text data

[1392] Output: Intent data (e.g., "I'm looking for information about everyday English conversations")

[1393] Step 5:

[1394] emotion recognition

[1395] The server uses an emotion engine to analyze the user's voice tone and facial expression data captured by the mirror device's camera to recognize the user's emotions, thereby obtaining information such as whether the user is stressed or relaxed.

[1396] Input: Voice tone data, facial expression data

[1397] Output: Emotion data (e.g., "The user feels anxious")

[1398] Step 6:

[1399] Avatar Selection

[1400] Based on the analyzed intent data, the server selects the avatar with the most appropriate expertise for the question. For example, if the question is about everyday English conversation, an English teacher avatar will be selected.

[1401] Input: Intent data

[1402] Output: Selected avatar

[1403] Step 7:

[1404] Generate answers

[1405] The server inputs the selected avatar's prompt into a generative AI model (e.g., ChatGPT) to generate an appropriate answer. The prompt includes the question and the user's emotional state.

[1406] Input: Prompt sentence (e.g., "The user is asking about everyday English conversation. The user feels anxious.")

[1407] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[1408] Step 8:

[1409] Emotional adjustment

[1410] The server adjusts the generated responses based on the user's emotional data, for example, if the user is tired, it modifies the generated responses to be gentler in tone and easier to understand.

[1411] Input: Generated answers, sentiment data

[1412] Output: A tailored response (e.g., "Good morning! This is a morning greeting" in a gentle tone)

[1413] Step 9:

[1414] Sending avatar and answer data

[1415] The server encodes the generated avatar and answer data and quickly transmits them to the terminal used by the user.

[1416] Input: Avatar data, adjusted answer

[1417] Output: Avatar data and answer data sent to the device

[1418] Step 10:

[1419] Display avatar

[1420] The terminal decodes the received avatar and answer data and displays them on the mirror device's display, allowing the user to receive the answers generated by the avatar visually and audibly.

[1421] Input: Submitted avatar data, response data

[1422] Output: An avatar displayed on the mirror device and a response (e.g., the avatar gently explains, "Good morning! This is a morning greeting.")

[1423] (Application example 2)

[1424] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1425] In current store operations, it is difficult for staff to provide customers with fast and accurate product information and advice, and staff's own emotions and stress can sometimes affect customer service. This can lead to issues such as lower customer satisfaction and reduced work efficiency. The present invention aims to solve these issues and enable store staff to receive expert advice in real time, improving the quality of customer service.

[1426] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with specialized knowledge corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for recognizing the user's emotion, means for adjusting the answer based on the recognized emotion, and means for displaying the generated answer to the user via an avatar. This enables store staff to provide accurate product information and expert advice in real time, improving customer satisfaction.

[1427] "Voice input" refers to voice data generated by a user speaking to the system through a microphone device.

[1428] "Capture" refers to taking in audio data using an input device such as a microphone.

[1429] "Analysis" refers to processing the captured voice data or text data to understand its meaning.

[1430] "Text data" refers to character data converted from voice input using voice recognition technology.

[1431] The "question intent" refers to the purpose or intention that indicates what the user wants to ask or what they are asking the system for.

[1432] "Expertise" refers to deep understanding and technical information in a particular field.

[1433] An "avatar" is a digital character that has specific capabilities or expertise and interacts with a user.

[1434] "Selection" refers to the act of choosing the most suitable candidate from among multiple candidates.

[1435] "Generation" refers to the act of the system creating the information or data it needs.

[1436] "Emotion" refers to a user's psychological state that is recognized based on their tone of voice, facial expressions, body language, etc.

[1437] "Tuning" refers to the act of modifying or optimizing the generated answer for specific conditions or circumstances.

[1438] "Display" refers to the act of outputting the generated information or avatar to a display and visually showing it to the user.

[1439] A system for realizing this invention has a function that allows store staff to wear smart glasses and obtain information about customer service and products in real time. Specific embodiments of the system will be described below.

[1440] Hardware Configuration

[1441] Smart glasses (e.g., Google Glass): Capture voice input and display information on a display.

[1442] Server: Performs voice data analysis, natural language processing, emotion recognition, avatar selection, and answer generation.

[1443] Software Configuration

[1444] Speech recognition engine (e.g., Google Speech-to-Text): Converts voice input into text data.

[1445] Natural language processing engines (e.g., SpaCy): Analyze the intent of text data.

[1446] Emotion engine (e.g. IBM Watson Tone Analyzer): Recognizes the user's emotions.

[1447] Generative AI models (e.g., GPT-4): Generate answers based on prompts.

[1448] Display control software: displays avatar and text information on the smart glasses display.

[1449] Data processing and calculation

[1450] 1. Capture audio input:

[1451] A microphone built into the smart glasses captures the staff's voice input.

[1452] For example: "Where can I find the product this customer is looking for?"

[1453] 2. Sending audio data:

[1454] The audio data is sent to a cloud server in real time.

[1455] 3. Analysis of audio data:

[1456] The server uses Google Speech-to-Text to convert the audio into text data.

[1457] 4. Intention Analysis and Emotion Recognition:

[1458] The converted text data is run through an NLP engine (SpaCy) to analyze the intent of the question.

[1459] The server uses IBM Watson Tone Analyzer to recognize emotions from staff members' tone of voice and facial expressions.

[1460] 5. Avatar Selection and Answer Generation:

[1461] The server selects the most suitable avatar based on the intent of the question.

[1462] A prompt is input into a generative AI model (GPT-4) to generate an answer.

[1463] Example prompt: "The product you are looking for is located on Aisle 3."

[1464] 6. Adjusting and Displaying Answers:

[1465] Tailor the generated answers depending on the sentiment.

[1466] The adjusted answers and avatar are sent to the smart glasses and displayed on the screen.

[1467] Example: The avatar explains in a friendly tone, "The product you are looking for is on Aisle 3."

[1468] Specific examples

[1469] Example 1: How to use it when guiding the location of a product

[1470] 1. User Action:

[1471] A staff member speaks to the smart glasses and asks, "Where is the product this customer is looking for?"

[1472] 2. Terminal processing:

[1473] The smart glasses capture the audio and send it to a cloud server.

[1474] 3. Server processing:

[1475] The server converts the voice into text and analyzes the intent of the question as a "product search."

[1476] The avatar that best suits the question is selected, and the prompt is input into a generative AI model to generate an answer.

[1477] Tailor the generated answer based on sentiment, e.g., "The product you're looking for can be found on Aisle 3."

[1478] 4. Terminal processing:

[1479] The smart glasses display the received avatar and answer on the screen.

[1480] The avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[1481] This system allows store staff to provide information to customers quickly and accurately, improving customer satisfaction.

[1482] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1483] Step 1:

[1484] The user inputs voice into the smart glasses. The input is made by the staff member through a microphone, for example, "Where is the product this customer is looking for?" The smart glasses capture this.

[1485] Step 2:

[1486] The device transmits the captured voice input to the cloud server. The input is the user's voice data, and the output is the transmission of the voice data to the cloud server. For data processing, the voice data is transmitted in real time.

[1487] Step 3:

[1488] The server converts the received voice data into text data using Google Speech-to-Text. The input is voice data, and the output is text data. A speech recognition algorithm is applied to the data.

[1489] Step 4:

[1490] The server runs the converted text data through a natural language processing engine (SpaCy) to analyze the intent of the question. The input is text data, and the output is the intent of the question. The NLP engine analyzes the text data as a data calculation.

[1491] Step 5:

[1492] The server then processes the analyzed text data through an emotion engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The input is data derived from voice tone and facial expressions, and the output is the user's emotional state. An emotion analysis algorithm is applied to process the data.

[1493] Step 6:

[1494] The server selects the avatar with the most appropriate expertise based on the intent of the question and the perceived emotion. The input is the intent of the question and the emotional state, and the output is the selected avatar. The selection algorithm is executed as a data operation.

[1495] Step 7:

[1496] The server inputs a prompt into a generative AI model (GPT-4) to generate an answer. The input is a prompt based on the intent of the question, and the output is a generated answer. As a data calculation, the generative AI model analyzes the prompt and generates an answer.

[1497] Step 8:

[1498] The server adjusts the generated answer according to the user's emotions. The input is the generated answer and the emotional state, and the output is the adjusted answer. The adjustment algorithm is applied as a data operation.

[1499] Step 9:

[1500] The server sends the adjusted answer and avatar to the smart glasses. The input is the adjusted answer and avatar data, and the output is the data transmission to the smart glasses. A transmission protocol is used for data processing.

[1501] Step 10:

[1502] The device displays the received avatar and answer on the smart glasses' display. The input is the received avatar and answer data, and the output is the visuals and audio displayed on the display. Specifically, the avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[1503] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1504] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1505] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1506] [Fourth embodiment]

[1507] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1508] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1509] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1510] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1511] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1512] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1513] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1514] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1515] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1516] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1517] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1518] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1519] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1520] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[1521] System Overview

[1522] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[1523] Program processing

[1524] Audio data capture and analysis

[1525] 1. Capture voice input

[1526] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[1527] 2. Sending audio data

[1528] The captured audio data is transmitted to a server in real time.

[1529] 3. Analysis of audio data

[1530] The server receives the voice data and converts it into text using a speech recognition algorithm.

[1531] 4. Intent Analysis

[1532] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question. For example, the intent of "everyday English conversation" is extracted.

[1533] Selecting specialized avatars and generating answers

[1534] 5. Select your avatar

[1535] Based on the intent of the question, the server selects an avatar with appropriate expertise, for example, an English teacher avatar.

[1536] 6. Answer Generation

[1537] A question is input into the generative AI model, which generates an appropriate answer, such as an English phrase like "Good morning! This is a morning greeting."

[1538] Show Answers

[1539] 7. Submitting avatar and answer data

[1540] The server sends the generated 3D model of the avatar and the response data to the device.

[1541] 8. Displaying Avatars

[1542] The terminal decodes the received avatar and response data and prepares it for display on the mirror device.

[1543] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[1544] Specific examples

[1545] Learning English

[1546] 1. User operations

[1547] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[1548] 2. Terminal Processing

[1549] The device captures the audio and sends it to the server.

[1550] 3. Server Processing

[1551] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[1552] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[1553] 4. Terminal Processing

[1554] The terminal receives the avatar and the answer and displays them on the mirror device.

[1555] The avatar explains, "Good morning! This is a morning greeting."

[1556] Technical Advice

[1557] 1. User operations

[1558] User: Say, "How do I use a wood stove safely?"

[1559] 2. Terminal Processing

[1560] The device captures the audio and sends it to the server.

[1561] 3. Server Processing

[1562] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[1563] An outdoor expert avatar is selected to generate answers regarding safe use.

[1564] 4. Terminal Processing

[1565] The terminal receives the avatar and the answer and displays them on the mirror device.

[1566] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[1567] As described above, the present invention is a system that interactively provides expert advice to users on various problems they may encounter, allowing users to obtain the information and guidance they need in real time.

[1568] The processing flow will be explained below.

[1569] Step 1:

[1570] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[1571] Step 2:

[1572] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[1573] Step 3:

[1574] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[1575] Step 4:

[1576] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[1577] Step 5:

[1578] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[1579] Step 6:

[1580] The server uses a generative AI model to generate an appropriate answer to the question through the selected avatar, for example, "Good morning! This is a morning greeting."

[1581] Step 7:

[1582] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[1583] Step 8:

[1584] The device decodes the avatar and answer data received from the server, loads the decoded 3D model of the avatar and answer data, and prepares them for display on the mirror device's display.

[1585] Step 9:

[1586] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might say, "Good morning! This is a morning greeting."

[1587] Step 10:

[1588] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[1589] Step 11:

[1590] The server receives the new voice data, converts it back to text, analyzes the intent, and generates an appropriate response. This process can be repeated as necessary.

[1591] This allows users to get real-time expert advice and quickly solve everyday problems.

[1592] Example 1

[1593] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1594] Conventional information provision systems have had difficulty responding quickly and accurately to the various problems users face. They also often only provide limited answers to questions that require specialized knowledge. Furthermore, there are issues with real-time performance and security in the process of efficiently analyzing a user's voice input, selecting an appropriate avatar, and displaying the answer.

[1595] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1596] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user through an avatar, means for compressing the captured voice input in real time and securely transmitting it, and means for decoding a three-dimensional model of the selected avatar and the text data and displaying them on a display device, thereby enabling expert advice to be provided quickly and accurately through the avatar for various problems faced by users.

[1597] "User" means any person or end user of the System.

[1598] "Voice input" refers to the words or sounds spoken by the user, and is the data that the system must recognize.

[1599] "Audio capture means" refers to a device or software that detects a user's voice input and captures it as digital data.

[1600] "Means for converting to text data" refers to the process of converting captured audio into text form using speech recognition technology.

[1601] "Means for analyzing the intent of a question" refers to natural language processing technology for understanding the purpose and requirements of a user's question from text data.

[1602] "Avatar selection method" refers to the process of selecting a virtual character with appropriate expertise based on the intent of the analyzed question.

[1603] "Means for generating answers" refers to the ability to use selected avatars and generative AI models to construct appropriate answers to users' questions.

[1604] "Means for displaying to the user through an avatar" refers to the process of providing the generated answer to the user visually and audibly through an avatar.

[1605] "Means for real-time compression and secure transmission" refers to the function of instantly compressing captured audio data and securely transmitting it to a server using encryption technology.

[1606] "Means for decoding the three-dimensional model and text data and displaying it on a display device" refers to the function of decoding the received data and appropriately rendering the three-dimensional avatar and text information for display on a display device.

[1607] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives. The system analyzes voice input from the user, selects an avatar with specialized knowledge corresponding to the question, and provides an appropriate answer. Specific embodiments of the present invention are described below.

[1608] System Overview

[1609] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. The answer generated through this process is displayed on the mirror-type device's display and provided to the user.

[1610] Audio data capture and analysis

[1611] 1. Capture voice input

[1612] The device captures voice input from the user using a microphone. For example, if the user says, "Please teach me everyday English conversation," the device captures the voice using a high-sensitivity microphone and uses noise-canceling technology to remove background noise.

[1613] 2. Sending audio data

[1614] The captured audio data is compressed in real time by the device and securely transmitted to the server using Transport Layer Security (TLS) encryption.

[1615] 3. Analysis of audio data

[1616] The server decodes the received audio data and converts it into text using a speech recognition algorithm (e.g., Google Speech-to-Text API), which is optimized to complete within a few seconds.

[1617] 4. Intent Analysis

[1618] The server then runs the converted text data through a natural language processing (NLP) engine (such as spaCy or the BERT model) to analyze the intent of the user's question. For example, if the intent is extracted as "everyday English conversation," an appropriate answer is prepared based on this.

[1619] Selecting specialized avatars and generating answers

[1620] 5. Select your avatar

[1621] The server selects the avatar with the most appropriate expertise based on the intent of the question, using a database of avatar skills, for example, selecting an "English teacher" avatar.

[1622] 6. Answer Generation

[1623] The server generates answers to questions using a generative AI model (e.g., OpenAI GPT-4). Specifically, in response to the prompt, "Please give me an example of an everyday English conversation," the server generates the answer, "Good morning! This is a morning greeting."

[1624] Show Answers

[1625] 7. Submitting avatar and answer data

[1626] The server combines the generated avatar's 3D model and the text response data into a single data packet and sends it to the terminal. This transmission uses UDP (User Datagram Protocol) to achieve low latency.

[1627] 8. Displaying Avatars

[1628] The device decodes the received data and prepares it for display on the mirror device's display. The device renders a 3D model of the avatar using the device's GPU and uses a voice synthesizer to output the response. The user hears the avatar displayed on the mirror device explain, "Good morning! This is a morning greeting."

[1629] Specific examples

[1630] 1. Example of user operation

[1631] Learning English: The user speaks to the mirror device, saying, "Please teach me everyday English conversation." The device captures the audio and sends it to the server. The server selects an English teacher avatar and generates a response, saying, "Good morning! This is a morning greeting." The device displays the avatar and the response on the mirror device, and the avatar explains the response aloud.

[1632] 2. Examples of technical advice

[1633] The user says, "Please tell me how to use a wood stove safely." The device captures the audio and sends it to the server. The server selects an outdoor expert avatar and generates a response: "First, make sure there are no flammable objects around the wood stove." The device displays the avatar and the response on a mirror-type device, and the avatar explains.

[1634] The system allows users to receive expert advice in real time.

[1635] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1636] Step 1: Capturing Audio Input

[1637] Specific operation: The user stands in front of the mirror-type device and speaks to it, such as "Please teach me some everyday English conversation."

[1638] Input: User's voice

[1639] Processing: The device's built-in high-sensitivity microphone captures the user's voice and uses noise-canceling technology to remove background noise.

[1640] Output: Captured audio data

[1641] Step 2: Sending audio data

[1642] Specific operation: The captured audio data is sent from the device to the server in real time.

[1643] Input: Captured audio data

[1644] Processing: The device compresses the audio data and sends it securely to the server using TLS (Transport Layer Security) encryption.

[1645] Output: Encrypted audio data sent to the server

[1646] Step 3: Analyzing the audio data

[1647] Specific operation: The server analyzes the received voice data using a voice recognition algorithm.

[1648] Input: Encrypted audio data sent to the server

[1649] Processing: The server decodes the audio data and converts it to text using a speech recognition algorithm (e.g., Google Speech-to-Text API).

[1650] Output: Converted text data

[1651] Step 4: Intent Analysis

[1652] Specific operation: The server analyzes the converted text data and understands the intent of the user's question.

[1653] Input: Converted text data

[1654] Processing: The server runs the text data through a natural language processing (NLP) engine (e.g., spaCy or the BERT model) to analyze the intent of the question.

[1655] Output: Analyzed intent information (e.g., "I want to learn everyday English conversation")

[1656] Step 5: Choose your avatar

[1657] Specific operation: The server selects an appropriate avatar.

[1658] Input: Parsed intent information

[1659] Processing: Based on the analyzed intention information, the server refers to the avatar skill database and selects an avatar with appropriate expertise.

[1660] Output: Information about the selected avatar (e.g., an English teacher avatar)

[1661] Step 6: Generate an answer

[1662] Specific behavior: The server generates an answer based on the selected avatar.

[1663] Input: Selected avatar information and intention information

[1664] Processing: The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers to questions.

[1665] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[1666] Step 7: Submit your avatar and answer data

[1667] Specific operation: The server sends the generated avatar's three-dimensional model and text response data to the terminal.

[1668] Input: Generated answers and a 3D avatar model

[1669] Processing: The server collects the generated data into a single data packet and sends it to the terminal using UDP (User Datagram Protocol).

[1670] Output: Data packets sent to the terminal

[1671] Step 8: Displaying the Avatar

[1672] Specific operation: The terminal decodes the received avatar data and response data and displays them on the mirror device.

[1673] Input: Data packets sent to the terminal

[1674] Processing: The device decodes the data, uses the GPU to render a 3D avatar, and uses a voice synthesizer to voice the response.

[1675] Output: Avatar displayed on a mirror device and voice response

[1676] Through these specific actions and data flows at each processing step, real-time expert advice is provided to the user.

[1677] (Application example 1)

[1678] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1679] The objective of this invention is to improve driving safety and the user experience by providing real-time expert advice through voice input in response to various questions and anxieties that users may encounter in an autonomous vehicle. In particular, there is a need to select an appropriate avatar and provide easy-to-understand answers to specific questions such as road traffic information, explanations of vehicle functions, and driving advice.

[1680] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1681] In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the converted text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for displaying the generated answer to the user via the avatar, and means for displaying the generated answer on a vehicle display, thereby enabling the user to receive reliable information and advice in real time within the autonomous vehicle.

[1682] "Voice input" refers to input of voice information generated by a user speaking.

[1683] "Capture" means capturing surrounding sounds using a device such as a microphone.

[1684] "Analysis" is the process of interpreting data and extracting meaning and patterns.

[1685] "Text data" refers to data that has been converted from voice input into text information through analysis.

[1686] The "question intent" indicates the content and purpose of the question that the user is trying to ask through voice input.

[1687] "Specialist knowledge" refers to advanced knowledge or information about a particular area or field.

[1688] An "avatar" is a virtual character that represents a specific role or persona in the digital space.

[1689] "Answer generation" is the process of creating an appropriate answer to a user's question.

[1690] "Display" is the act of visually presenting information to a user.

[1691] "Vehicle display" means a screen device installed inside an autonomous vehicle for displaying information.

[1692] A "generative artificial intelligence model" is a type of artificial intelligence that generates text data in natural language based on given prompts.

[1693] System Overview

[1694] The present invention is a system for enabling users to receive real-time expert advice in an autonomous vehicle. The system captures voice input from the user and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. An avatar with the most appropriate expertise for the question is then selected, and an answer is generated using a generative artificial intelligence (AI) model. The answer is then displayed on the vehicle's display and provided to the user.

[1695] System configuration

[1696] The system includes the following means:

[1697] 1. Capture voice input

[1698] A microphone built into the vehicle captures the user's voice input and is highly sensitive, capturing clear audio while suppressing interior noise.

[1699] 2. Sending audio data

[1700] The captured voice data is sent to a cloud server via the vehicle's communication system, which achieves low latency and high-speed data transfer.

[1701] 3. Speech Recognition and Text Conversion

[1702] The cloud server converts the voice data into text data using a speech recognition algorithm, such as the Google Speech Recognition API.

[1703] 4. Intent Analysis

[1704] The converted text data undergoes intent analysis using an NLP engine (e.g., Hugging Face Transformers), where the "question intent" is extracted.

[1705] 5. Select a professional avatar

[1706] The server selects an avatar with the appropriate expertise based on the intent of the question. For example, if the question is about "tips for driving on rainy days," an avatar of a driving safety expert will be selected.

[1707] 6. Answer Generation

[1708] The question is input as a prompt into a generative AI model (for example, OpenAI's GPT-3), and an appropriate answer is generated. An example of a prompt is "User question: What are some tips for driving on rainy days? Please select an avatar that will provide an appropriate answer."

[1709] 7. Displaying answers and avatars

[1710] The generated answers and a 3D model of the avatar are then displayed on a display inside the vehicle, allowing the user to receive the answer both visually and audibly. The display uses a high-resolution screen, and the audio is output through high-quality speakers.

[1711] Specific examples

[1712] 1. User operations

[1713] User: "What are some tips for driving in the rain?"

[1714] 2. Server Processing

[1715] The system converts speech to text and analyzes the intent of the "Tips for driving on rainy days." It selects a driving safety expert avatar and generates the answer, "Drive at a moderate speed and keep a sufficient distance between vehicles."

[1716] 3. Terminal Processing

[1717] The device receives the avatar and the answer, which it then displays on the display. The avatar explains, "Please drive slowly and keep a safe distance from other vehicles," providing visual support.

[1718] The system of the present invention enables users to receive accurate and expert advice in real time within an autonomous vehicle, improving driving safety and comfort.

[1719] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1720] Step 1:

[1721] The user provides voice input in the autonomous vehicle. Here, the user provides voice information by speaking. For example, the user might say, "Please tell me some tips for driving on rainy days."

[1722] Step 2:

[1723] The device uses the vehicle's built-in microphone to capture voice input from the user, and the captured voice data is temporarily stored in the device as a digital signal.

[1724] Step 3:

[1725] The device transmits the captured audio data to the cloud server in real time. The data is transferred with low latency using a communication system (e.g., LTE or 5G network). The input is the audio data, and the output is the data sent to the cloud server.

[1726] Step 4:

[1727] The server converts the transmitted voice data into text data using a speech recognition algorithm (e.g., Google Speech Recognition API). The input is voice data, and the output is text data.

[1728] Step 5:

[1729] The server analyzes the converted text data using a natural language processing (NLP) engine (e.g., Hugging Face Transformers) to understand the intent of the question. In this case, it extracts the intent of "advice on safe driving" from the question "Please tell me some tips for driving on rainy days." The input is text data, and the output is the intent of the question.

[1730] Step 6:

[1731] The server selects an avatar with appropriate expertise based on the intent of the question. For example, a driving safety expert avatar is selected. The input is the intent of the question, and the output is the selected avatar.

[1732] Step 7:

[1733] The server uses a generative AI model (e.g., OpenAI's GPT-3) to generate an answer based on the selected avatar and the question. An example prompt is "User question: What are some tips for driving on rainy days? Please select the avatar that provides the appropriate answer." The input is the question and prompt, and the output is the generated answer.

[1734] Step 8:

[1735] The server sends the generated 3D model of the avatar and the answer data to the terminal. The input is the 3D model of the avatar and the answer data, and the output is the data sent to the terminal.

[1736] Step 9:

[1737] The device decodes the received avatar and response data and displays them on the vehicle's display. The avatar then provides specific driving advice visually and audibly, such as "Slow down your speed and keep a safe distance between vehicles." The input is the decoded avatar and response data, and the output is the avatar and response displayed on the display.

[1738] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1739] The present invention is a system for providing expert advice to users on various problems they encounter in their daily lives, and by combining it with an emotion engine that recognizes the user's emotions, more personalized support is possible. The system aims to analyze voice input from the user, select an avatar with specialized knowledge corresponding to the question, and provide an appropriate answer. In addition, the emotion engine is used to identify the user's emotions, and the answer is adjusted based on this. Specific embodiments of the present invention are described below.

[1740] System Overview

[1741] The system begins operation when the user inputs voice through the mirror-type device. The system captures the user's voice input and converts it into text data using speech recognition technology. The converted text data is analyzed by a natural language processing (NLP) engine to understand the intent of the question. The AI ​​model then selects an avatar with the most appropriate expertise for the question and generates an answer. An emotion engine then analyzes the user's emotions, and the avatar's answer is adjusted accordingly. The answer generated through these processes is displayed on the mirror-type device's display and provided to the user.

[1742] Program processing

[1743] Audio data capture and analysis

[1744] 1. Capture voice input

[1745] The device captures voice input from the user using a microphone. For example, the user might say, "Please teach me everyday English conversation."

[1746] 2. Sending audio data

[1747] The captured audio data is transmitted to a server in real time.

[1748] 3. Analysis of audio data

[1749] The server receives the voice data and converts it into text data using a speech recognition algorithm, for example, "Please teach me everyday English conversation."

[1750] Intention Analysis and Emotion Recognition

[1751] 4. Intent Analysis

[1752] The converted text data is then passed through a natural language processing (NLP) engine to analyze the intent of the question, which reveals that the user is seeking information about everyday English conversation.

[1753] 5. Emotion recognition

[1754] The server uses an emotion engine to analyze the user's tone of voice, facial expressions, body language, etc. to recognize the user's emotions, for example, whether the user is feeling stressed.

[1755] Selecting specialized avatars and generating answers

[1756] 6. Avatar Selection

[1757] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[1758] 7. Answer Generation

[1759] A question is input into the generative AI model, which generates an appropriate answer, such as "Good morning! This is a morning greeting."

[1760] 8. Emotional Adjustment

[1761] The server adjusts the generated answers depending on the user's emotions: for example, if the user is tired, the avatar will provide answers in a gentler tone.

[1762] Show Answers

[1763] 9. Submitting avatar and answer data

[1764] The server sends the generated 3D model of the avatar and the response data to the device.

[1765] 10. Display of Avatar

[1766] The terminal decodes the received avatar and response data and prepares it for display on the mirror device's display.

[1767] An avatar is displayed on the mirror-type device and the generated answer is communicated to the user via voice or text.

[1768] Specific examples

[1769] Learning English

[1770] 1. User operations

[1771] User: Talk to the mirror device, "Please teach me some everyday English conversation."

[1772] 2. Terminal Processing

[1773] The device captures the audio and sends it to the server.

[1774] 3. Server Processing

[1775] The server converts the speech into text and analyzes the intent of the "everyday English conversation."

[1776] An English teacher avatar is selected and the answer generated is "Good morning! This is a morning greeting."

[1777] The emotion engine recognizes when a user is feeling anxious and adjusts responses to a gentler tone.

[1778] 4. Terminal Processing

[1779] The terminal receives the avatar and the answer and displays them on the mirror device.

[1780] The avatar gently explains, "Good morning! This is a morning greeting."

[1781] Technical Advice

[1782] 1. User operations

[1783] User: Say, "How do I use a wood stove safely?"

[1784] 2. Terminal Processing

[1785] The device captures the audio and sends it to the server.

[1786] 3. Server Processing

[1787] The server converts the speech into text and analyzes the intent of "how to use a wood stove."

[1788] An outdoor expert avatar is selected to generate answers regarding safe use.

[1789] The emotion engine recognizes that the user is relaxed and provides answers in a normal tone.

[1790] 4. Terminal Processing

[1791] The terminal receives the avatar and the answer and displays them on the mirror device.

[1792] "First, make sure there are no flammable materials around the wood stove," the avatar explains.

[1793] As described above, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[1794] The processing flow will be explained below.

[1795] Step 1:

[1796] The user speaks a question into the mirror-type device, for example, "Please teach me everyday English conversation."

[1797] Step 2:

[1798] The device captures audio input with a microphone, converts the captured audio data into a digital format, compresses it, and transmits it to a server in real time.

[1799] Step 3:

[1800] The server converts the received voice data into text data using a voice recognition algorithm, for example, "Please teach me everyday English conversation."

[1801] Step 4:

[1802] The server then runs the converted text data through a natural language processing (NLP) engine to analyze the intent of the question, extracting the intent that the user is seeking information about everyday English conversation.

[1803] Step 5:

[1804] The server selects an avatar with appropriate expertise based on the intent of the question, for example, an English teacher avatar.

[1805] Step 6:

[1806] The server inputs a question into the generative AI model and generates an appropriate answer, such as "Good morning! This is a morning greeting."

[1807] Step 7:

[1808] The server uses an emotion engine to recognize the user's emotions, analyzing the user's voice tone and facial expression data to determine, for example, that the user is feeling anxious.

[1809] Step 8:

[1810] The server adjusts responses based on the user's emotions, for example, adjusting the avatar's dialogue to provide responses in a gentler tone if the user is feeling anxious.

[1811] Step 9:

[1812] The server encodes the generated avatar and the response data, including the 3D model of the avatar and voice data, and sends it to the device.

[1813] Step 10:

[1814] The device decodes the avatar and answer data received from the server, loads the decoded 3D avatar model, and prepares it for display on the mirror device's display.

[1815] Step 11:

[1816] The device displays an avatar on the mirror and communicates the generated answer to the user via voice or text. For example, an English teacher avatar might explain in a gentle tone, "Good morning! This is a morning greeting."

[1817] Step 12:

[1818] If the user has additional questions or would like more detailed explanations, they can again enter their questions by voice, and the device will again capture the voice data and send it to the server.

[1819] Step 13:

[1820] The server receives the new voice data, converts it back to text, analyzes the intent, generates an appropriate response, adjusts it again based on the user's sentiment, and sends it back to the device. This process can be repeated as needed.

[1821] This allows users to get real-time expert advice and quickly solve everyday problems.

[1822] Example 2

[1823] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1824] Conventional voice input analysis systems can convert user voice input into text, analyze the intent of the question based on the text data, and generate answers. However, these systems cannot adjust the answers taking into account the user's emotions, which leaves the user experience unsatisfied. Furthermore, there is a lack of interactive systems that provide users with appropriate answers visually and audibly through expert avatars.

[1825] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1826] In this invention, the server includes means for capturing a voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with expertise corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for adjusting the generated answer in accordance with emotional data of the user, and means for displaying the generated answer to the user via the avatar, thereby providing a personalized answer in accordance with the emotional state of the user and improving the user experience.

[1827] "Means for capturing voice input from a user" refers to technical means for capturing voice uttered by a user as electronic data using a voice capture device such as a microphone.

[1828] "Means for analyzing captured voice input and converting it into text data" refers to technical means for converting voice data into corresponding text data using a speech recognition algorithm.

[1829] The "means for analyzing the intent of a question based on text data" is a technical means for using a natural language processing engine to understand the intent of a user's question from converted text data.

[1830] "Means for selecting an avatar with the expertise corresponding to the intent of the question" refers to a technical means for selecting a virtual character (avatar) with the most appropriate skills and knowledge based on the analyzed content of the question.

[1831] The "means for generating an answer based on a selected avatar" refers to a technical means for generating an appropriate answer to a question based on the area of ​​expertise of the selected avatar.

[1832] "Means for adjusting the generated response in accordance with the user's emotional data" refers to a technical means for analyzing the user's emotional state and adjusting the tone of voice, expression, etc. of the generated response in accordance with that emotion.

[1833] The "means for displaying the generated answer to the user through an avatar" refers to a technical means for presenting the generated answer to the user through the visual and audio of the avatar.

[1834] This invention is a system that provides expert advice to users on various problems they encounter in their daily lives. Furthermore, by combining it with a function to recognize the user's emotions, it provides more personalized assistance. The system analyzes the user's voice input, selects an avatar with the expertise required for the question, and provides an appropriate answer. It also uses an emotion engine to identify the user's emotions and adjusts the answer accordingly.

[1835] The system operates using the following hardware and software:

[1836] Hardware

[1837] 1. Mirror-type device: This device has a built-in highly sensitive microphone and camera to capture voice input and facial expression data from the user.

[1838] 2. Server: A high-performance computer that analyzes voice data and emotional data, selects avatars, and generates answers.

[1839] software

[1840] 1. Speech recognition algorithms (e.g., Google Speech-to-Text): Used to convert user voice input into text data.

[1841] 2. Natural language processing engine (e.g. GPT-4): Used to analyze the intent of the user's question from text data.

[1842] 3. Emotion engine: Used to recognize emotions by analyzing the user's voice tone and facial expression data.

[1843] 4. Generative AI models (e.g. ChatGPT): Used to generate appropriate answers to user questions.

[1844] How to operate

[1845] The user speaks a question to the mirror device. For example, if the user says, "Please teach me everyday English conversation," the following processing flow is executed.

[1846] 1. Audio input capture: The microphone on the mirror device captures the user's voice.

[1847] 2. Audio data transmission: The captured audio data is transmitted to the server in real time.

[1848] 3. Voice data analysis: The server uses a voice recognition algorithm to convert the voice into text data.

[1849] 4. Intent analysis: The text data is passed through a natural language processing engine to analyze the intent of the question. For example, the intent may be, "I'm looking for information about everyday English conversation."

[1850] 5. Emotion Recognition: The emotion engine analyzes the user's voice tone and facial expression data to recognize the user's emotional state, for example, whether the user is feeling stressed.

[1851] 6. Avatar selection: The AI ​​model selects the avatar with the expertise that best fits the intent of the question, for example, an English teacher avatar.

[1852] 7. Answer generation: A question is input into the generative AI model, and an appropriate answer is generated. For example, the answer generated might be, "Good morning! This is a morning greeting."

[1853] 8. Emotion-based adjustment: The server adjusts the generated answers depending on the user's emotions. For example, if the user is tired, the avatar will provide answers in a gentler tone.

[1854] 9. Answer display: The avatar and answer data are sent from the server to the terminal and displayed on the mirror device's display.

[1855] Specific examples

[1856] Learning English

[1857] User operation: Talk to the mirror device, saying, "Please teach me everyday English conversation."

[1858] Device processing: The device captures the audio and sends it to the server.

[1859] Server processing: The server converts the speech into text and analyzes the intent of the "everyday English conversation." It selects an English teacher avatar and generates a response such as "Good morning! This is a morning greeting." The emotion engine recognizes that the user is feeling anxious and adjusts the response to a gentler tone.

[1860] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar gently explains, "Good morning! This is a morning greeting."

[1861] Technical Advice

[1862] User Action: Say, "How do I use a wood stove safely?"

[1863] Device processing: The device captures the audio and sends it to the server.

[1864] Server processing: The server converts the speech to text and analyzes the intent of "how to use a wood stove." It selects an outdoor expert avatar and generates an answer on how to use it safely. The emotion engine recognizes that the user is relaxed and provides an answer in a normal tone.

[1865] Device processing: The device receives the avatar and the answer and displays them on the mirror device. The avatar explains, "First, make sure there are no flammable objects around the wood stove."

[1866] In this way, the present invention is a system that interactively provides expert advice for various problems faced by users. Furthermore, by combining it with an emotion engine, it is possible to provide personalized answers according to the user's emotions, thereby improving user satisfaction.

[1867] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1868] Step 1:

[1869] Capture audio input

[1870] The terminal uses a high-sensitivity microphone to capture the voice of the user speaking into the mirror-type device. For example, the user might say, "Please teach me everyday English conversation."

[1871] Input: User's voice

[1872] Output: Digital audio data

[1873] Step 2:

[1874] Sending audio data

[1875] The device compresses the captured audio data and transmits it to the server with low latency, using a protocol to minimize data delays and loss.

[1876] Input: Audio data captured on the device

[1877] Output: Audio data sent to the server

[1878] Step 3:

[1879] Analysis of audio data

[1880] The server converts the received voice data into text data by running it through a speech recognition algorithm (e.g., Google Speech-to-Text). Specifically, it analyzes the frequency patterns of the voice and converts them into corresponding strings of characters.

[1881] Input: Transmitted audio data

[1882] Output: Text data (e.g., "Please teach me everyday English conversations")

[1883] Step 4:

[1884] Intention Analysis

[1885] The server then runs the converted text data through a natural language processing (NLP) engine (e.g., GPT-4) to analyze the intent of the question. During this process, the server identifies the information the user is looking for based on keywords and context within the text data.

[1886] Input: Text data

[1887] Output: Intent data (e.g., "I'm looking for information about everyday English conversations")

[1888] Step 5:

[1889] emotion recognition

[1890] The server uses an emotion engine to analyze the user's voice tone and facial expression data captured by the mirror device's camera to recognize the user's emotions, thereby obtaining information such as whether the user is stressed or relaxed.

[1891] Input: Voice tone data, facial expression data

[1892] Output: Emotion data (e.g., "The user feels anxious")

[1893] Step 6:

[1894] Avatar Selection

[1895] Based on the analyzed intent data, the server selects the avatar with the most appropriate expertise for the question. For example, if the question is about everyday English conversation, an English teacher avatar will be selected.

[1896] Input: Intent data

[1897] Output: Selected avatar

[1898] Step 7:

[1899] Generate answers

[1900] The server inputs the selected avatar's prompt into a generative AI model (e.g., ChatGPT) to generate an appropriate answer. The prompt includes the question and the user's emotional state.

[1901] Input: Prompt sentence (e.g., "The user is asking about everyday English conversation. The user feels anxious.")

[1902] Output: Generated answer (e.g. "Good morning! This is a morning greeting")

[1903] Step 8:

[1904] Emotional adjustment

[1905] The server adjusts the generated responses based on the user's emotional data, for example, if the user is tired, it modifies the generated responses to be gentler in tone and easier to understand.

[1906] Input: Generated answers, sentiment data

[1907] Output: A tailored response (e.g., "Good morning! This is a morning greeting" in a gentle tone)

[1908] Step 9:

[1909] Sending avatar and answer data

[1910] The server encodes the generated avatar and answer data and quickly transmits them to the terminal used by the user.

[1911] Input: Avatar data, adjusted answer

[1912] Output: Avatar data and answer data sent to the device

[1913] Step 10:

[1914] Display avatar

[1915] The terminal decodes the received avatar and answer data and displays them on the mirror device's display, allowing the user to receive the answers generated by the avatar visually and audibly.

[1916] Input: Submitted avatar data, response data

[1917] Output: An avatar displayed on the mirror device and a response (e.g., the avatar gently explains, "Good morning! This is a morning greeting.")

[1918] (Application example 2)

[1919] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1920] In current store operations, it is difficult for staff to provide customers with fast and accurate product information and advice, and staff's own emotions and stress can sometimes affect customer service. This can lead to issues such as lower customer satisfaction and reduced work efficiency. The present invention aims to solve these issues and enable store staff to receive expert advice in real time, improving the quality of customer service.

[1921] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice input from a user, means for analyzing the captured voice input and converting it into text data, means for analyzing the intent of the question based on the text data, means for selecting an avatar with specialized knowledge corresponding to the intent of the question, means for generating an answer based on the selected avatar, means for recognizing the user's emotion, means for adjusting the answer based on the recognized emotion, and means for displaying the generated answer to the user via an avatar. This enables store staff to provide accurate product information and expert advice in real time, improving customer satisfaction.

[1922] "Voice input" refers to voice data generated by a user speaking to the system through a microphone device.

[1923] "Capture" refers to taking in audio data using an input device such as a microphone.

[1924] "Analysis" refers to processing the captured voice data or text data to understand its meaning.

[1925] "Text data" refers to character data converted from voice input using voice recognition technology.

[1926] The "question intent" refers to the purpose or intention that indicates what the user wants to ask or what they are asking the system for.

[1927] "Expertise" refers to deep understanding and technical information in a particular field.

[1928] An "avatar" is a digital character that has specific capabilities or expertise and interacts with a user.

[1929] "Selection" refers to the act of choosing the most suitable candidate from among multiple candidates.

[1930] "Generation" refers to the act of the system creating the information or data it needs.

[1931] "Emotion" refers to a user's psychological state that is recognized based on their tone of voice, facial expressions, body language, etc.

[1932] "Tuning" refers to the act of modifying or optimizing the generated answer for specific conditions or circumstances.

[1933] "Display" refers to the act of outputting the generated information or avatar to a display and visually showing it to the user.

[1934] A system for realizing this invention has a function that allows store staff to wear smart glasses and obtain information about customer service and products in real time. Specific embodiments of the system will be described below.

[1935] Hardware Configuration

[1936] Smart glasses (e.g., Google Glass): Capture voice input and display information on a display.

[1937] Server: Performs voice data analysis, natural language processing, emotion recognition, avatar selection, and answer generation.

[1938] Software Configuration

[1939] Speech recognition engine (e.g., Google Speech-to-Text): Converts voice input into text data.

[1940] Natural language processing engines (e.g., SpaCy): Analyze the intent of text data.

[1941] Emotion engine (e.g. IBM Watson Tone Analyzer): Recognizes the user's emotions.

[1942] Generative AI models (e.g., GPT-4): Generate answers based on prompts.

[1943] Display control software: displays avatar and text information on the smart glasses display.

[1944] Data processing and calculation

[1945] 1. Capture audio input:

[1946] A microphone built into the smart glasses captures the staff's voice input.

[1947] For example: "Where can I find the product this customer is looking for?"

[1948] 2. Sending audio data:

[1949] The audio data is sent to a cloud server in real time.

[1950] 3. Analysis of audio data:

[1951] The server uses Google Speech-to-Text to convert the audio into text data.

[1952] 4. Intention Analysis and Emotion Recognition:

[1953] The converted text data is run through an NLP engine (SpaCy) to analyze the intent of the question.

[1954] The server uses IBM Watson Tone Analyzer to recognize emotions from staff members' tone of voice and facial expressions.

[1955] 5. Avatar Selection and Answer Generation:

[1956] The server selects the most suitable avatar based on the intent of the question.

[1957] A prompt is input into a generative AI model (GPT-4) to generate an answer.

[1958] Example prompt: "The product you are looking for is located on Aisle 3."

[1959] 6. Adjusting and Displaying Answers:

[1960] Tailor the generated answers depending on the sentiment.

[1961] The adjusted answers and avatar are sent to the smart glasses and displayed on the screen.

[1962] Example: The avatar explains in a friendly tone, "The product you are looking for is on Aisle 3."

[1963] Specific examples

[1964] Example 1: How to use it when guiding the location of a product

[1965] 1. User Action:

[1966] A staff member speaks to the smart glasses and asks, "Where is the product this customer is looking for?"

[1967] 2. Terminal processing:

[1968] The smart glasses capture the audio and send it to a cloud server.

[1969] 3. Server processing:

[1970] The server converts the voice into text and analyzes the intent of the question as a "product search."

[1971] The avatar that best suits the question is selected, and the prompt is input into a generative AI model to generate an answer.

[1972] Tailor the generated answer based on sentiment, e.g., "The product you're looking for can be found on Aisle 3."

[1973] 4. Terminal processing:

[1974] The smart glasses display the received avatar and answer on the screen.

[1975] The avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[1976] This system allows store staff to provide information to customers quickly and accurately, improving customer satisfaction.

[1977] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1978] Step 1:

[1979] The user inputs voice into the smart glasses. The input is made by the staff member through a microphone, for example, "Where is the product this customer is looking for?" The smart glasses capture this.

[1980] Step 2:

[1981] The device transmits the captured voice input to the cloud server. The input is the user's voice data, and the output is the transmission of the voice data to the cloud server. For data processing, the voice data is transmitted in real time.

[1982] Step 3:

[1983] The server converts the received voice data into text data using Google Speech-to-Text. The input is voice data, and the output is text data. A speech recognition algorithm is applied to the data.

[1984] Step 4:

[1985] The server runs the converted text data through a natural language processing engine (SpaCy) to analyze the intent of the question. The input is text data, and the output is the intent of the question. The NLP engine analyzes the text data as a data calculation.

[1986] Step 5:

[1987] The server then processes the analyzed text data through an emotion engine (IBM Watson Tone Analyzer) to recognize the user's emotions. The input is data derived from voice tone and facial expressions, and the output is the user's emotional state. An emotion analysis algorithm is applied to process the data.

[1988] Step 6:

[1989] The server selects the avatar with the most appropriate expertise based on the intent of the question and the perceived emotion. The input is the intent of the question and the emotional state, and the output is the selected avatar. The selection algorithm is executed as a data operation.

[1990] Step 7:

[1991] The server inputs a prompt into a generative AI model (GPT-4) to generate an answer. The input is a prompt based on the intent of the question, and the output is a generated answer. As a data calculation, the generative AI model analyzes the prompt and generates an answer.

[1992] Step 8:

[1993] The server adjusts the generated answer according to the user's emotions. The input is the generated answer and the emotional state, and the output is the adjusted answer. The adjustment algorithm is applied as a data operation.

[1994] Step 9:

[1995] The server sends the adjusted answer and avatar to the smart glasses. The input is the adjusted answer and avatar data, and the output is the data transmission to the smart glasses. A transmission protocol is used for data processing.

[1996] Step 10:

[1997] The device displays the received avatar and answer on the smart glasses' display. The input is the received avatar and answer data, and the output is the visuals and audio displayed on the display. Specifically, the avatar explains in a gentle tone, "The product you are looking for can be found on Aisle 3."

[1998] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1999] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2000] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2001] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2002] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2003] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2004] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2005] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2006] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2007] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2008] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2009] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2010] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2011] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2012] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2013] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2014] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2015] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2016] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2017] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2018] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2019] The following is further disclosed regarding the above embodiment.

[2020] (Claim 1)

[2021] means for capturing voice input from a user;

[2022] means for analyzing and converting the captured voice input into text data;

[2023] means for analyzing the intent of a question based on the text data;

[2024] a means for selecting an avatar having expertise corresponding to the intent of the question;

[2025] means for generating an answer based on the selected avatar;

[2026] means for displaying the generated answer to a user through an avatar;

[2027] A system including:

[2028] (Claim 2)

[2029] 2. The system according to claim 1, wherein the avatars with specialized knowledge are different avatars across multiple specialized fields.

[2030] (Claim 3)

[2031] 10. The system of claim 1, wherein the avatar's response is provided to the user as audio and visual information.

[2032] "Example 1"

[2033] (Claim 1)

[2034] means for capturing voice input from a user;

[2035] means for analyzing and converting the captured voice input into text data;

[2036] means for analyzing the intent of a question based on the text data;

[2037] a means for selecting an avatar having expertise corresponding to the intent of the question;

[2038] means for generating an answer based on the selected avatar;

[2039] means for displaying the generated answer to a user through an avatar;

[2040] means for compressing and securely transmitting the captured audio input in real time;

[2041] means for decoding the three-dimensional model and text data of the selected avatar and displaying them on a display device;

[2042] A system including:

[2043] (Claim 2)

[2044] 2. The system according to claim 1, wherein the avatars with specialized knowledge are different avatars across multiple specialized fields.

[2045] (Claim 3)

[2046] 10. The system of claim 1, wherein the avatar's responses are provided as audio and visual information.

[2047] "Application Example 1"

[2048] (Claim 1)

[2049] means for capturing voice input from a user;

[2050] means for analyzing and converting the captured voice input into text data;

[2051] means for analyzing the intent of a question based on the text data;

[2052] a means for selecting an avatar having expertise corresponding to the intent of the question;

[2053] means for generating an answer based on the selected avatar;

[2054] means for displaying the generated answer to a user through an avatar;

[2055] means for displaying the generated answer on a vehicle display;

[2056] A system including:

[2057] (Claim 2)

[2058] 2. The system according to claim 1, wherein the avatars with specialized knowledge are different avatars across multiple specialized fields.

[2059] (Claim 3)

[2060] 10. The system of claim 1, wherein the avatar's response is provided to the user as audio and visual information.

[2061] (Claim 4)

[2062] The system according to claim 1, wherein a question is input to the generative artificial intelligence model and an answer is generated.

[2063] "Example 2: Combining Emotion Engines"

[2064] (Claim 1)

[2065] means for capturing voice input from a user;

[2066] means for analyzing and converting the captured voice input into text data;

[2067] means for analyzing the intent of a question based on the text data;

[2068] a means for selecting an avatar having expertise corresponding to the intent of the question;

[2069] means for generating an answer based on the selected avatar;

[2070] means for adjusting the generated answer in accordance with emotion data of the user;

[2071] means for displaying the generated answer to a user through an avatar;

[2072] A system including:

[2073] (Claim 2)

[2074] 2. The system according to claim 1, wherein the avatars with specialized knowledge are different avatars across multiple specialized fields.

[2075] (Claim 3)

[2076] 10. The system of claim 1, wherein the avatar's response is provided to the user as audio and visual information.

[2077] "Application example 2 when combining emotion engines"

[2078] (Claim 1)

[2079] means for capturing voice input from a user;

[2080] means for analyzing and converting the captured voice input into text data;

[2081] means for analyzing the intent of a question based on the text data;

[2082] a means for selecting an avatar having expertise corresponding to the intent of the question;

[2083] means for generating an answer based on the selected avatar;

[2084] means for recognizing a user's emotion;

[2085] means for adjusting a response based on the recognized emotion;

[2086] means for displaying the generated answer to a user through an avatar;

[2087] A system including:

[2088] (Claim 2)

[2089] 2. The system according to claim 1, wherein the avatars with specialized knowledge are different avatars across multiple specialized fields.

[2090] (Claim 3)

[2091] 10. The system of claim 1, wherein the avatar's response is provided to the user as audio and visual information. [Explanation of symbols]

[2092] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for capturing voice input from a user; means for analyzing and converting the captured voice input into text data; means for analyzing the intent of a question based on the text data; a means for selecting an avatar having expertise corresponding to the intent of the question; means for generating an answer based on the selected avatar; means for displaying the generated answer to a user through an avatar; A system including:

2. 2. The system according to claim 1, wherein the avatars with specialized knowledge are different avatars across a plurality of specialized fields.

3. 10. The system of claim 1, wherein the avatar's responses are provided to the user as audio and visual information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A