System
The system addresses the limitations of conventional dialogue systems by converting voice input to text, using generative AI for specialized responses, and providing visual feedback through a virtual character, enhancing interaction experience.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Conventional dialogue systems fail to provide comprehensive and specialized responses to everyday problems, lacking a sense of familiarity and natural interaction experience, and are limited by single generative AI models that cannot meet diverse user needs.
A system that receives voice input, converts it to text, uses a first-stage generation means for general answers, and specialized generation means for detailed answers, providing visual and auditory interaction through a virtual character.
Enables highly accurate and specialized answers with natural visual and auditory interactions, addressing the limitations of conventional systems by combining voice and visual interfaces.
Smart Images

Figure 2026036279000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional dialogue systems have struggled to provide comprehensive and specialized responses to the wide range of everyday problems users face. Furthermore, their limited visual and auditory interfaces lacked a sense of familiarity and natural interaction experience for users. Furthermore, answer generation using a single generative AI model was unable to provide sophisticated answers specialized in specialized fields, failing to meet the diverse needs of users. [Means for solving the problem]
[0005] The present invention is a system that receives voice input from a user, converts it to text, and then uses a first-stage generation means to generate a general answer. The output is then analyzed by specialized generation means for each specialized field to generate a detailed answer. The converted answer is then reconverted into voice data and output to the user through a virtual character, providing a sense of visual familiarity while achieving a specialized and sophisticated answer. Furthermore, by selectively applying multiple generation means for each specialized field, the system can flexibly respond to a wide range of user needs. In this way, the present invention provides a completely new interactive system that combines voice and visual interfaces, unlike conventional systems.
[0006] "Voice input" refers to voice data uttered by a user through a microphone.
[0007] "Speech recognition module" refers to a software and hardware complex for converting received voice data into text data.
[0008] "First stage generation means" refers to the first stage generative AI module for generating a general solution.
[0009] "Generation means for each specialized field" refers to a generative AI module that generates detailed answers specialized for each specialized field, such as law, finance, or medicine.
[0010] "Speech synthesis module" refers to a software and hardware complex for converting text data into natural-sounding speech data.
[0011] "Virtual character" refers to a digital avatar that provides visual and auditory information to a user.
[0012] "Terminal" refers to a hardware device that captures the user's voice and communicates with the system.
[0013] "Server" refers to a remote computer system that performs processes such as speech recognition, natural language processing, generative AI, and voice synthesis. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions.
[0036] Specifically, the following operations are included:
[0037] 1. User voice input
[0038] User: Speaks a question or command. This voice input is received through the device's microphone.
[0039] Device: Captures audio data and sends it to the server.
[0040] 2. Converting voice data to text
[0041] Server: Receives voice data and passes it to the voice recognition module.
[0042] Speech recognition module: Converts voice input into text data.
[0043] Server: Inputs text data into the generative AI.
[0044] 3. Generating a general answer
[0045] Generative AI (first level): Analyzes text data and generates general answers that are appropriate responses to general questions.
[0046] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[0047] 4. Generating detailed answers
[0048] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[0049] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine.
[0050] Server: Passes the output from the specialized AI to the speech synthesis module.
[0051] 5. Generating Audio Data
[0052] Speech synthesis module: Converts text responses into natural-sounding speech data.
[0053] Server: Sends the generated audio data to the device.
[0054] 6. Audio output to the user
[0055] Terminal: Interprets the audio data received from the server, triggers MetaHuman's animation engine, and synchronizes it with the audio playback.
[0056] MetaHuman: Providing visual and auditory voice answers to users.
[0057] Specific examples
[0058] Example 1: Legal advice
[0059] User: Say, "I need help with the contract."
[0060] Device: Captures audio and sends it to the server.
[0061] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[0062] Generative AI (first level): Generates a general answer to the question, "What specifically is the subject of the contract consultation?"
[0063] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[0064] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[0065] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[0066] Example 2: Engineering Support
[0067] User: "I need some advice on defining requirements for my new project."
[0068] Device: Captures audio and sends it to the server.
[0069] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[0070] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[0071] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[0072] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[0073] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[0074] In this way, the present invention processes the user's voice input efficiently and naturally, providing optimal answers, and by using MetaHuman, also provides a visually familiar experience to the user.
[0075] The processing flow will be explained below.
[0076] Step 1:
[0077] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[0078] Step 2:
[0079] The device's microphone captures the user's voice, and the voice data is sent to the server.
[0080] Step 3:
[0081] The server receives the voice data and passes it to a voice recognition module.
[0082] Step 4:
[0083] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[0084] Step 5:
[0085] The server inputs the converted text into the generative AI (first-stage base model).
[0086] Step 6:
[0087] The generative AI (first stage) analyzes the input text and generates a general answer, such as, "What specifically is the content of the contract consultation?"
[0088] Step 7:
[0089] The server receives the output of the first stage generative AI and analyzes the answer, identifying that the user's question is related to law.
[0090] Step 8:
[0091] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[0092] Step 9:
[0093] Specialized AI (legal) generates detailed answers based on the input text, such as "Please tell me about the specific clauses in the contract."
[0094] Step 10:
[0095] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[0096] Step 11:
[0097] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[0098] Step 12:
[0099] The server transmits the generated voice data to the terminal.
[0100] Step 13:
[0101] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0102] Step 14:
[0103] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as, "Tell me about the specific clauses in the contract."
[0104] In this way, the user, terminal, and server work together at each step, enabling this system to achieve advanced voice dialogue.
[0105] Example 1
[0106] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0107] Many conventional dialogue systems convert a user's voice input into text and generate a specific answer based on that text. However, when the user's question is highly specialized, these systems often produce vague and insufficient answers. There is a demand for systems that can provide highly accurate answers, especially for detailed questions related to specialized fields. Furthermore, conventional systems only output voice, making it difficult to provide sufficient visual feedback to the user. Therefore, it is necessary to provide a more natural and intuitive dialogue experience.
[0108] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0109] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the converted text and generating a general answer based on the text, means for generating a detailed answer based on a field of expertise from the general answer generated by the generation means, means for converting the generated detailed answer into voice data, and means for outputting the voice data to the user. This enables highly accurate answers even to specialized and advanced questions, and also makes it possible to provide a natural dialogue experience including visual feedback.
[0110] "Means for receiving voice input from a user" refers to a device or software function that can capture voice uttered by a user and transmit that data to the next processing step.
[0111] "Means for converting speech input to text" refers to a device or software function that analyzes received speech data and converts it into corresponding text data.
[0112] "Generation means for analyzing the converted text and generating a general answer based on the text" refers to the function of a device or software for analyzing the text obtained from the voice data and generating a general answer based on its content.
[0113] "Means for generating detailed answers based on specialized fields from the general answers generated by the generation means" refers to the function of a device or software that converts the initially generated general answers into more detailed answers based on specialized knowledge.
[0114] The "means for converting the generated detailed answer into voice data" refers to a device or software function that analyzes the detailed text answer and converts it into natural, easily understandable voice data.
[0115] The "means for outputting audio data to the user" refers to a device or software function that plays back and provides the generated audio data to the user.
[0116] "Selective application of multiple disciplinary generative tools" is the process of selecting and applying specialized generative tools in response to a specific question or request.
[0117] "Providing visual output using a virtual character" means using a virtual character model and displaying its movements and expressions to the user to provide visual feedback corresponding to the generated audio data.
[0118] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions. The configuration and operation of this system are described in detail below.
[0119] Hardware and Software Overview
[0120] Hardware used
[0121] Device: A user device that includes a microphone for capturing audio data and a speaker for playing audio data.
[0122] Server: A central processing unit for processing voice data and generating answers.
[0123] Software used
[0124] Speech recognition module: Software that converts voice data into text using the Google® Cloud Speech-to-Text API or similar.
[0125] Generative AI model (first level): An AI model for generating general answers using OpenAI's (registered trademark) GPT-3 (registered trademark), etc.
[0126] Specialized AI: Generative AI specialized in a specific field. For example, AI models specialized in fields such as law or medicine.
[0127] Speech synthesis module: Software that converts text data into speech data using the Google Cloud Text-to-Speech API or similar.
[0128] Virtual Character Animation Engine: Software that uses technologies such as MetaHuman to provide visual feedback synchronized with audio data.
[0129] Specific explanation of operation
[0130] The details of the functions are shown below.
[0131] Voice input from the user
[0132] The user speaks questions or commands into the terminal. For example, the user might say, "I'd like some advice on defining the requirements for a new project."
[0133] Capture and transmit audio data
[0134] The device captures the user's voice with a microphone and transmits the data to the server in real time.
[0135] Converting audio data to text
[0136] The server analyzes the received voice data using the Google Cloud Speech-to-Text API and converts it into text data. In this step, the corresponding text is generated from the speech, "I would like some advice on defining the requirements for a new project."
[0137] General answer generation by first-level generative AI
[0138] The text data is analyzed by the server and input into a generative AI model such as OpenAI's GPT-3. An example prompt is "The user is looking for advice on the project requirements definition. Please provide general guidance." Based on this, a general answer is generated: "You're talking about the project requirements definition. Please tell me more specifically."
[0139] Detailed answers generated by specialized AI
[0140] A specialized AI is used to generate a more detailed answer based on a general answer from the first-stage generative AI. An example prompt is, "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of defining the requirements." The specialized AI generates a detailed answer: "Initial research and gathering stakeholder opinions are important for defining requirements."
[0141] Generating and transmitting audio data
[0142] The generated detailed answer is converted into audio data using the Google Cloud Text-to-Speech API, and the converted audio data is sent from the server to the device.
[0143] Audio output and animation playback to the user
[0144] The device receives the audio data and triggers MetaHuman's animation engine to animate the virtual character in sync with the audio playback, with the character's mouth movements and facial expressions matching the audio.
[0145] The above is an embodiment of the present invention. This system allows users to have natural visual and auditory interactions, and provides an advanced interactive system that can also respond to specialized questions.
[0146] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0147] Step 1:
[0148] Voice input from the user
[0149] User: The user speaks questions and commands into the microphone.
[0150] Input: User utterance (e.g., "I need some advice on defining requirements for a new project").
[0151] Output: Audio data
[0152] What it does: When a user speaks, the microphone captures the sound.
[0153] Step 2:
[0154] Capture and transmit audio data
[0155] Device: The device transmits the captured audio data to the server in real time.
[0156] Input: Captured audio data
[0157] Output: Audio data sent to the server
[0158] Specific operation: Audio captured by the microphone is saved in PCM format or similar and sent to the server as an HTTP request.
[0159] Step 3:
[0160] Converting audio data to text
[0161] Server: The server receives the voice data and passes it to the voice recognition module.
[0162] Input: Audio data sent from the device
[0163] Output: Text data
[0164] What it does: The received voice data is sent to the Google Cloud Speech-to-Text API, which converts the voice data into text. The resulting text is, "I'd like some advice on defining the requirements for a new project."
[0165] Step 4:
[0166] General answer generation by first-level generative AI
[0167] Server: The server analyzes the text data and inputs it into the generative AI.
[0168] Input: Text data (e.g., "I would like some advice on defining requirements for a new project.")
[0169] Output: General answer (text)
[0170] Specific operation: The server sends a prompt to OpenAI's GPT-3 or similar (e.g., "The user is looking for advice on the project requirements definition. Please provide general guidance."). The generative AI generates a general answer, such as "You're talking about the project requirements definition. Please be more specific," and replies to the server.
[0171] Step 5:
[0172] Detailed answers generated by specialized AI
[0173] Server: The server analyzes the general answers from the first-level generative AI and sends them to the specialized AI.
[0174] Input: General answer (text)
[0175] Output: Detailed answer (text)
[0176] Specific operation: The server sends a prompt to the specialized AI (e.g., "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of the requirements definition."). The specialized AI generates a detailed answer, such as "Initial research and gathering stakeholder opinions are important for defining requirements," and replies to the server.
[0177] Step 6:
[0178] Generating and transmitting audio data
[0179] Server: The server passes the detailed answer to the speech synthesis module, converts it into voice data, and sends it to the terminal.
[0180] Input: Detailed answer (text)
[0181] Output: Audio data (e.g., WAV or MP3 format)
[0182] Specific operation: The server sends the detailed answer to the Google Cloud Text-to-Speech API, converts the text "Initial research and stakeholder opinion gathering are important for requirements definition" into audio data, and sends the generated audio file to the device.
[0183] Step 7:
[0184] Audio output and animation playback to the user
[0185] Device: The device receives the audio data and triggers MetaHuman's animation engine to synchronize with the audio playback.
[0186] Input: Audio data
[0187] Output: Audio and visual feedback
[0188] How it works: The device plays back the received audio data, and at the same time, MetaHuman's virtual character moves its mouth and changes its facial expression in response to the audio, providing the user with a natural conversational experience.
[0189] (Application example 1)
[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0191] In virtual stores, a system that integrates voice input and visual feedback is necessary to enable more natural and efficient user interaction. However, current technology simply converts the user's voice input into text and provides text-based answers, failing to realize natural dialogue that integrates visual and auditory feedback. Furthermore, there is a lack of effective means for generating detailed answers for each specialized field. This results in low user convenience and poses challenges for improving user satisfaction.
[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0193] In this invention, the server includes means for converting voice input into text, means for analyzing the converted text and generating a general answer, means for generating detailed answers for each specialized field based on the generated general answers, means for converting the generated detailed answers into voice data, means for outputting the voice data to the user, means for providing visual output using a virtual character, and means for providing animated visual feedback synchronized with the voice data, thereby enabling natural dialogue that integrates voice input and visual feedback.
[0194] "Voice input" refers to words or questions spoken by the user, and is the voice data that the system uses to recognize them.
[0195] A "means for converting to text" is a technique or device for analyzing received voice input and converting the content into text data format.
[0196] "First-stage generation means" refers to a first-stage technique or device for generating a general answer based on the converted text.
[0197] The "generating means for each specialized field" is a technology or device that further analyzes the answer obtained by the generating means at the first stage and generates a detailed answer that is specific to a specific specialized field.
[0198] The "means for converting into voice data" refers to a technique or device for synthesizing the generated text-format answers into voice and outputting them as voice data.
[0199] "Visual feedback with animation synchronized with audio data" refers to a technology or device that displays visual actions or animations corresponding to generated audio data and provides them to the user along with the audio.
[0200] A "virtual character" is a visual character created using computer graphics and animation techniques to interact with a user.
[0201] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice form again. This system is particularly effective for realizing natural dialogue in virtual stores, and also provides visual feedback using virtual characters.
[0202] First, the user uses the smart glasses to input voice. The microphone in the smart glasses captures the voice and sends it to the server. The server then converts the received voice data into text using Google Cloud Speech-to-Text. This converted text is then input into GPT-4 (registered trademark), the first stage of the generation process, to generate a general answer.
[0203] The generated general answer is then passed to a specialized generator. In this example, a generator specialized for product information is used to generate a detailed answer. For example, if a user asks, "Tell me about this product," GPT-4 generates a general answer such as, "What category does this product belong to?" Then, a specialized AI generates a detailed answer such as, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[0204] The detailed answer is then converted into speech using Amazon Polly, which is then sent back to the server and played back to the smart glasses, where it uses Unreal Engine's MetaHuman to provide animated visual feedback synchronized with the speech.
[0205] As a concrete example, the following prompt sentence is input to GPT-4 to generate a general answer:
[0206] A user asked the following question:
[0207] Question: "Tell me about this product"
[0208] Generate a general answer to the question.
[0209] You can then create prompts for specialized AI and get detailed answers.
[0210] The following question was asked to the product information specialized AI.
[0211] Ask: "What category does this product belong to?"
[0212] Generate a professional answer to this question.
[0213] In this way, the present invention efficiently and naturally processes voice input from the user, provides optimal answers in the virtual store, and also provides a sense of visual familiarity to the user by using virtual characters.
[0214] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0215] Step 1:
[0216] The user uses the smart glasses to input voice. Specifically, the user asks a question by voice, such as "Tell me about this product." The microphone in the smart glasses captures this voice and inputs it as voice data into the terminal. The terminal then sends this voice data to the server.
[0217] Step 2:
[0218] The server converts the received voice data into text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data. Specifically, the voice recognition algorithm analyzes the voice waveform data and converts it into text such as "Tell me about this product."
[0219] Step 3:
[0220] The server then inputs the converted text data into GPT-4 to generate a general answer. The input is text data, and the output is the general answer text. Based on the prompt, GPT-4 generates a general answer such as, "What category does this product belong to?"
[0221] Step 4:
[0222] The generated general answer is input to a specialized AI on the server, which uses a generation method specialized for product information to generate a detailed answer. The general answer is the input, and the detailed answer text is obtained as the output. Specifically, the detailed answer generated is, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[0223] Step 5:
[0224] The detailed answer text is converted to speech using Amazon Polly. We have the detailed text answer as input and speech as output. A speech synthesis algorithm analyzes the text and generates natural-sounding speech.
[0225] Step 6:
[0226] The generated voice data is sent from the server to the device (smart glasses). The device plays the received voice data. Specifically, a detailed answer is provided to the user through the audio speaker: "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[0227] Step 7:
[0228] At the same time, the server uses Unreal Engine's MetaHuman to generate animated visual feedback corresponding to the generated voice data. The input is voice data, and the output is visual animation synchronized with the voice. Specifically, a virtual character visually presents the answer to the user, moving its mouth in sync with the voice.
[0229] This series of processing steps enables the user to experience natural and detailed interactions in the virtual store.
[0230] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0231] The present invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. This system allows users to have natural visual and auditory interactions, making it an advanced dialogue system that can also handle specialized questions.
[0232] Specifically, the following operations are included:
[0233] 1. User voice input
[0234] User: Speaks a question or instruction. For example, "I'd like to ask about the contents of the contract."
[0235] Device: Captures audio data and sends it to the server.
[0236] 2. Converting voice data to text
[0237] Server: Receives voice data and passes it to the voice recognition module.
[0238] Speech recognition module: Converts voice input into text data. For example, text data such as "I would like to ask for your advice regarding the contents of the contract."
[0239] Server: Inputs text data into the generative AI.
[0240] 3. Emotional Recognition
[0241] Server: Passes the voice input to the emotion engine and analyzes the user's emotions. For example, it analyzes the voice tone, rate, pitch, and volume to recognize that the user is feeling anxious.
[0242] Emotion engine: Recognizes the user's emotions and passes that information to the generative AI.
[0243] 4. Generating a general answer
[0244] Generative AI (first level): Generates general answers based on text data and recognized emotions. For example, it generates answers such as, "What specifically are you discussing about the contract?"
[0245] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[0246] 5. Generating detailed answers
[0247] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[0248] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine. For example, it generates detailed answers such as, "Please tell me about the specific clauses in the contract."
[0249] Server: Passes the output from the specialized AI to the speech synthesis module.
[0250] 6. Generating Audio Data
[0251] Speech synthesis module: Converts text responses into natural-sounding speech, such as "Please tell me about the specific clauses in the contract."
[0252] Server: Sends the generated audio data to the device.
[0253] 7. Audio output to the user
[0254] Terminal: Interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0255] MetaHuman: Provides visual and audible voice responses to the user, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[0256] Specific examples
[0257] Example 1: Legal advice
[0258] User: Say, "I need help with the contract."
[0259] Device: Captures audio and sends it to the server.
[0260] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[0261] Emotion engine: Recognizes that the user's tone of voice indicates anxiety.
[0262] Generative AI (first level): Generates a general answer such as, "You seem anxious. What specifically would you like to discuss regarding the contract?"
[0263] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[0264] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[0265] Device: The voice data is received and MetaHuman plays it back in a gentle tone as if it were a natural conversation.
[0266] Example 2: Engineering Support
[0267] User: "I need some advice on defining requirements for my new project."
[0268] Device: Captures audio and sends it to the server.
[0269] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[0270] Emotion engine: Recognizes that the user's tone of voice is calm.
[0271] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[0272] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[0273] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[0274] Device: The voice data is received and MetaHuman plays it back in a calm, natural-sounding conversation.
[0275] In this way, the present invention processes user voice input efficiently and naturally, providing optimal answers, and by using an emotion engine, it realizes flexible responses that match the user's emotions, providing a more personalized experience.
[0276] The processing flow will be explained below.
[0277] Step 1:
[0278] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[0279] Step 2:
[0280] The device's microphone captures the user's voice, and the voice data is sent to the server.
[0281] Step 3:
[0282] The server receives the voice data and passes it to a voice recognition module.
[0283] Step 4:
[0284] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[0285] Step 5:
[0286] The server receives the converted text data and passes it to the emotion engine.
[0287] Step 6:
[0288] The emotion engine analyzes text and voice data to recognize the user's emotions. For example, it analyzes voice tone, speed, pitch, and volume to determine if the user is feeling anxious.
[0289] Step 7:
[0290] The server inputs the recognized emotional information into the generative AI (first-stage base model).
[0291] Step 8:
[0292] The generative AI (first stage) generates a general answer based on the input text and emotional information. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[0293] Step 9:
[0294] The server analyzes the output of the first-stage generative AI to input it into the specialized AI, identifying the user's question as legal-related.
[0295] Step 10:
[0296] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[0297] Step 11:
[0298] A specialized AI (legal) generates detailed answers based on the input text and emotional information, such as "Please tell me about the specific clauses in the contract."
[0299] Step 12:
[0300] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[0301] Step 13:
[0302] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[0303] Step 14:
[0304] The server transmits the generated voice data to the terminal.
[0305] Step 15:
[0306] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0307] Step 16:
[0308] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[0309] In this way, the user, device, and server work together at each step to realize advanced voice dialogue. In addition, by combining it with an emotion engine, it is possible to provide flexible responses that correspond to the user's emotions.
[0310] Example 2
[0311] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0312] While conventional voice dialogue systems can provide appropriate answers to user voice inputs, they have difficulty generating flexible responses that reflect the user's emotions. Furthermore, for specialized questions, they can only provide general answers, failing to provide the detailed information actually required. Furthermore, there is a need for the answers output from the system to be natural and for the dialogue with the user to be visually acceptable.
[0313] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0314] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer based on the text by analyzing the converted text, means for generating a detailed answer based on a field of expertise from the general answer generated by the first-stage answer generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for recognizing the user's emotion from the voice input, and means for adjusting the generated answer based on the recognized emotion. This makes it possible to provide a flexible and detailed answer that reflects the user's emotion and realize a visually acceptable and natural dialogue.
[0315] The "means for receiving voice input from the user" is a device or function for capturing voice uttered by the user and incorporating it into the system.
[0316] A "means for converting voice input to text" is a device or function that analyzes received voice data and converts the content into text data.
[0317] The "first-stage generator" is an initial generator or algorithm that generates a general answer based on the converted text data.
[0318] The "means for generating a detailed answer based on a field of expertise" is a device or algorithm that generates a detailed answer specialized in a particular field of expertise based on the general answer generated by the first-stage generation means.
[0319] The "means for converting the generated detailed answer into voice data" is a device or function that converts the detailed answer in text format into natural voice data.
[0320] The "means for outputting voice data to the user" refers to a device or function that reproduces and provides the generated voice data to the user.
[0321] A "means for recognizing a user's emotion from speech input" is a device or algorithm that analyzes a user's speech data and determines their emotional state.
[0322] A "means for adjusting the generated answer based on the recognized emotion" is a device or algorithm that appropriately modifies or adjusts the generated answer depending on the user's emotional state.
[0323] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice format. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. Specific implementation methods of the system are described below.
[0324] This system operates primarily using the following hardware and software:
[0325] Hardware: A device equipped with a microphone to capture the user's voice and a speaker to play the output audio data.
[0326] Server: A central server for processing data and performing necessary calculations.
[0327] Software: Speech recognition module, emotion recognition engine, generative AI, specialized AI, voice synthesis module, virtual character animation engine.
[0328] The main processing flow of this system is as follows:
[0329] The user speaks a voice input, which is captured by the device's microphone. Specifically, the user speaks something like "I would like to consult you about the contents of the contract." The device sends the captured voice data to the server. The server uses a voice recognition module (e.g., a voice recognition API) to convert the voice data into text data. For example, the converted text might be something like "I would like to consult you about the contents of the contract."
[0330] The server then passes the text and voice data to an emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, speed, pitch, volume, etc. of the voice to recognize the emotion the user is feeling. Specifically, it may recognize that the user is feeling "anxiety."
[0331] The server inputs the recognized emotions and text data into a generative AI system to generate a general answer. For example, it might generate an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?" A general-purpose generative model is used for this generative AI model.
[0332] The server then passes the general answer to a specialized AI that generates a detailed answer. The specialized AI generates an answer based on knowledge of a specialized field, such as law, finance, or medicine. For example, it might generate a detailed answer such as, "Please tell me about the specific clauses in the contract."
[0333] The generated detailed answer is passed by the server to a speech synthesis module, which converts the text into natural-sounding speech data (e.g., speech synthesis API). The generated speech data will be something like, "Please tell me about the specific clauses of the contract."
[0334] The server sends the final voice data to the terminal, which triggers the virtual character's animation engine to synchronize with the voice playback. The virtual character speaks in a gentle tone, saying, "Please tell me about the specific clauses in the contract."
[0335] Specific examples
[0336] Example 1: Legal advice
[0337] The user says, "I'd like to ask you about the contents of the contract." The device captures the voice and sends it to the server. The server uses a speech recognition module to convert it into text, "I'd like to ask you about the contents of the contract," and uses an emotion engine to recognize the user's anxiety. The generative AI then generates a general answer, "You seem anxious. What specifically do you want to ask you about in the contract?", and the specialized AI generates a detailed answer, "Please tell me about the specific clauses in the contract." The speech synthesis module converts this into voice data, which the device finally provides to the user through a virtual character.
[0338] Prompt Sentence Examples
[0339] "Generate answers that will alleviate the user's concerns when asked about the contents of the contract."
[0340] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0341] Step 1:
[0342] The user speaks the speech input.
[0343] Specific operation: The user says, "I would like to discuss the contents of the contract."
[0344] Input: User's voice.
[0345] Output: The captured audio data.
[0346] Step 2:
[0347] The device captures the audio data and sends it to the server.
[0348] Specific operation: Records audio data using the device's microphone and sends the data to the server.
[0349] Input: The captured audio data.
[0350] Output: The audio data sent to the server.
[0351] Step 3:
[0352] The server receives the voice data and passes it to the voice recognition module.
[0353] Specific operation: The server receives voice data from the terminal and inputs the data into the voice recognition module.
[0354] Input: The audio data sent to the server.
[0355] Output: Audio data as input to the speech recognition module.
[0356] Step 4:
[0357] A voice recognition module converts the voice data into text data.
[0358] Specific operation: The voice recognition module analyzes the voice data and generates text data such as "I would like to consult you about the contents of the contract."
[0359] Input: Audio data as input to the speech recognition module.
[0360] Output: The converted text data.
[0361] Step 5:
[0362] The server passes the text data to an emotion recognition engine to analyze the user's emotions.
[0363] Specific operation: The server passes the text data and voice characteristic information to the emotion recognition engine and begins analysis.
[0364] Input: Translated text data and speech characteristics information.
[0365] Output: Parsed emotion data (e.g., anxiety).
[0366] Step 6:
[0367] The server inputs emotional data and text data into a generative AI to generate a general answer.
[0368] Specific operation: The server inputs emotion data and text data into the generative AI and receives the generated answer. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[0369] Input: Parsed emotion data and converted text data.
[0370] Output: The generated general answer text.
[0371] Step 7:
[0372] The server passes the general answer to a specialized AI, which generates a detailed answer.
[0373] Specific operation: The server passes the general answer text to a specialized AI (e.g., specialized AI for law, finance, medicine, etc.) to generate a detailed answer. For example, it generates a detailed answer such as "Please tell me about the specific clauses in the contract."
[0374] Input: The generated general answer text.
[0375] Output: The generated detailed answer text.
[0376] Step 8:
[0377] The server passes the detailed answer to a speech synthesis module and converts it into speech data.
[0378] Specific operation: The server inputs the detailed answer text into the speech synthesis module to generate speech data, for example, "Please tell me about the specific clauses of the contract."
[0379] Input: The generated long answer text.
[0380] Output: The generated audio data.
[0381] Step 9:
[0382] The server transmits the generated voice data to the terminal.
[0383] Specific operation: The server sends the generated audio data to the device.
[0384] Input: The generated audio data.
[0385] Output: The audio data sent to the device.
[0386] Step 10:
[0387] The device receives the audio data and triggers the animation engine of the virtual character.
[0388] Specific operation: The device interprets the received audio data and triggers the virtual character's animation engine to create a visual representation synchronized with the audio playback.
[0389] Input: Audio data sent to the device.
[0390] Output: Audio and visual answers provided to the user.
[0391] (Application example 2)
[0392] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0393] Conventional food delivery systems have had problems with the user having to make an order, requiring a lot of effort and not being able to respond flexibly to emotions. This can lead to low user satisfaction and delays in the ordering process. Furthermore, advanced technology is required to ensure natural visual and auditory interactions, and there has been a lack of systems that can achieve this.
[0394] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer by analyzing the text, means for generating a detailed answer based on the field of expertise from the general answer generated by the first-stage generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for analyzing the user's emotions, and means for adjusting the answer based on the results of the emotion analysis. This allows the user to place an order through natural dialogue and enables flexible responses according to emotions.
[0395] A "means for receiving voice input from a user" is a device or function for capturing voice data uttered by a user.
[0396] A "means for converting voice input to text" is a device or function that analyzes captured voice data and converts it into corresponding text data.
[0397] "An initial generation means for analyzing the converted text and generating a general answer based on the text" is a device or function that includes an AI module for generating an initial answer based on text data.
[0398] "Means for generating detailed answers based on specialized fields from the general answers generated by the initial generation means" refers to a device or function that includes an AI module for further specialized analysis of the output of the initial generation means and generating detailed answers.
[0399] The "means for converting the generated detailed answer into voice data" is a device or function for converting the generated detailed answer as text into voice data.
[0400] The "means for outputting audio data to the user" refers to a device or function for playing back the generated audio data to the user.
[0401] A "means for analyzing user emotions" is a device or function that analyzes a user's voice input to identify their emotional state.
[0402] The "means for adjusting the answer based on the result of sentiment analysis" is a device or function for taking into account the result of sentiment analysis and adjusting the answer to be generated and the method of presenting it.
[0403] The present invention is a system for efficiently delivering food via voice, allowing users to place orders through natural dialogue and providing flexible responses according to emotions. This system includes the following specific procedures and devices.
[0404] System configuration
[0405] The system takes voice input from the user, converts it to text, uses generative AI to generate a general answer, then generates a more specific answer based on the user's area of expertise (in this case, food delivery), converts the answer back into speech, and delivers it to the user. It can also recognize the user's emotions and adjust the answer accordingly.
[0406] Hardware and software used
[0407] Hardware:
[0408] Smartphone microphone: A device for capturing voice data spoken by a user.
[0409] Smartphone speaker: A device for playing back the generated audio data to the user.
[0410] software:
[0411] speech_recognition: A library for converting speech data into text data.
[0412] Transformers sentiment-analysis: A library for analyzing user sentiment from text data.
[0413] googletrans: A library for translating and converting answers into text.
[0414] gTTS (Google Text-To-Speech): A library for converting text data into audio data and playing it back.
[0415] System Operation
[0416] The server first receives voice input from the user. This voice is captured using a smartphone microphone. It then uses speech recognition software to convert the speech to text. This converted text is passed to a first-stage generative AI to generate a general answer. This general answer is then further processed into a detailed answer by a specialized AI in a specialized field (food delivery).
[0417] Emotion analysis
[0418] At the same time, the user's emotions are also analyzed. This is done using an emotion analysis engine that analyzes the user's voice tone, speed, pitch, and volume. Based on the analysis results, the answer is adjusted. For example, if the user is in a hurry, a quick response is required.
[0419] Speech synthesis and output
[0420] The detailed answer is then converted into audio data by speech synthesis software, which is then delivered to the user through the smartphone speaker, allowing them to order food delivery in a natural, conversational way.
[0421] Specific examples
[0422] Customer: "One pizza please."
[0423] application:
[0424] (Transcription): "One pizza please."
[0425] (emotion recognition): calm
[0426] (Order Processing): "Confirmed. We will process your order immediately."
[0427] (Text-to-Speech): Play "Thank you. We'll get it sorted right away."
[0428] Prompt Sentence Examples
[0429] Prompt: "A customer calmly orders, 'One pizza, please.' Generate an appropriate response for the calm customer."
[0430] Expected response: "Thank you. I'll get it sorted right away."
[0431] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0432] Step 1:
[0433] The server receives the user's voice input from the device. This input is the voice data spoken by the user using the smartphone microphone. The device then sends the captured voice data to the server. This data is transmitted as a voice waveform and is not in a format that can be analyzed as is.
[0434] Step 2:
[0435] The server converts the speech input into text. This is done using speech recognition software, which analyzes the speech waveform data and converts the phoneme sequence into the corresponding text. For example, if the speech input is "One pizza please," it will be converted into text "One pizza please." The output of this step is the converted text data.
[0436] Step 3:
[0437] The server analyzes the converted text and passes it to a first-stage generative AI that generates a general answer. This first-stage generative AI applies a generative AI model based on the text data to generate a general answer. For example, if the input text is "One pizza please," it will generate a general answer of "I understand. We'll get it ready right away." The output of this step is the general answer text.
[0438] Step 4:
[0439] The server passes the answer text obtained from the first-dan generation AI to a specialized AI, which generates a detailed answer. This specialized AI has a knowledge database specialized in a specific field (in this case, food delivery) and complements the first-dan answer in detail. For example, it adds a specific confirmation such as "Is this pizza size okay?" to the general answer "Please confirm." The output of this step is a detailed answer text.
[0440] Step 5:
[0441] The server receives the detailed answer text and analyzes the user's emotions. It uses an emotion analysis engine to identify the user's emotional state (e.g., hurry, calm, happy) by analyzing the tone, rate, pitch, and volume of the speech input. It adjusts the detailed answer based on this analysis data. The output of this step is the final adjusted detailed answer text.
[0442] Step 6:
[0443] The server converts the finalized detailed answer text into speech data. It uses speech synthesis software such as Google Text-To-Speech (gTTS) to convert the text data into speech data. This speech data is provided with natural pronunciation that is easy for the user to understand. The output of this step is synthesized speech data.
[0444] Step 7:
[0445] The server returns the generated voice data to the terminal, which then provides it to the user through the smartphone speaker, allowing the user to receive the voice answer. The final output of this step is the voice answer provided to the user.
[0446] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0447] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0448] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0449] [Second embodiment]
[0450] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0451] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0452] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0453] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0454] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0456] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0457] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0458] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0459] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0460] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0461] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0462] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions.
[0463] Specifically, the following operations are included:
[0464] 1. User voice input
[0465] User: Speaks a question or command. This voice input is received through the device's microphone.
[0466] Device: Captures audio data and sends it to the server.
[0467] 2. Converting voice data to text
[0468] Server: Receives voice data and passes it to the voice recognition module.
[0469] Speech recognition module: Converts voice input into text data.
[0470] Server: Inputs text data into the generative AI.
[0471] 3. Generating a general answer
[0472] Generative AI (first level): Analyzes text data and generates general answers that are appropriate responses to general questions.
[0473] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[0474] 4. Generating detailed answers
[0475] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[0476] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine.
[0477] Server: Passes the output from the specialized AI to the speech synthesis module.
[0478] 5. Generating Audio Data
[0479] Speech synthesis module: Converts text responses into natural-sounding speech data.
[0480] Server: Sends the generated audio data to the device.
[0481] 6. Audio output to the user
[0482] Terminal: Interprets the audio data received from the server, triggers MetaHuman's animation engine, and synchronizes it with the audio playback.
[0483] MetaHuman: Providing visual and auditory voice answers to users.
[0484] Specific examples
[0485] Example 1: Legal advice
[0486] User: Say, "I need help with the contract."
[0487] Device: Captures audio and sends it to the server.
[0488] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[0489] Generative AI (first level): Generates a general answer to the question, "What specifically is the subject of the contract consultation?"
[0490] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[0491] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[0492] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[0493] Example 2: Engineering Support
[0494] User: "I need some advice on defining requirements for my new project."
[0495] Device: Captures audio and sends it to the server.
[0496] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[0497] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[0498] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[0499] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[0500] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[0501] In this way, the present invention processes the user's voice input efficiently and naturally, providing optimal answers, and by using MetaHuman, also provides a visually familiar experience to the user.
[0502] The processing flow will be explained below.
[0503] Step 1:
[0504] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[0505] Step 2:
[0506] The device's microphone captures the user's voice, and the voice data is sent to the server.
[0507] Step 3:
[0508] The server receives the voice data and passes it to a voice recognition module.
[0509] Step 4:
[0510] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[0511] Step 5:
[0512] The server inputs the converted text into the generative AI (first-stage base model).
[0513] Step 6:
[0514] The generative AI (first stage) analyzes the input text and generates a general answer, such as, "What specifically is the content of the contract consultation?"
[0515] Step 7:
[0516] The server receives the output of the first stage generative AI and analyzes the answer, identifying that the user's question is related to law.
[0517] Step 8:
[0518] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[0519] Step 9:
[0520] Specialized AI (legal) generates detailed answers based on the input text, such as "Please tell me about the specific clauses in the contract."
[0521] Step 10:
[0522] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[0523] Step 11:
[0524] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[0525] Step 12:
[0526] The server transmits the generated voice data to the terminal.
[0527] Step 13:
[0528] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0529] Step 14:
[0530] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as, "Tell me about the specific clauses in the contract."
[0531] In this way, the user, terminal, and server work together at each step, enabling this system to achieve advanced voice dialogue.
[0532] Example 1
[0533] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0534] Many conventional dialogue systems convert a user's voice input into text and generate a specific answer based on that text. However, when the user's question is highly specialized, these systems often produce vague and insufficient answers. There is a demand for systems that can provide highly accurate answers, especially for detailed questions related to specialized fields. Furthermore, conventional systems only output voice, making it difficult to provide sufficient visual feedback to the user. Therefore, it is necessary to provide a more natural and intuitive dialogue experience.
[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0536] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the converted text and generating a general answer based on the text, means for generating a detailed answer based on a field of expertise from the general answer generated by the generation means, means for converting the generated detailed answer into voice data, and means for outputting the voice data to the user. This enables highly accurate answers even to specialized and advanced questions, and also makes it possible to provide a natural dialogue experience including visual feedback.
[0537] "Means for receiving voice input from a user" refers to a device or software function that can capture voice uttered by a user and transmit that data to the next processing step.
[0538] "Means for converting speech input to text" refers to a device or software function that analyzes received speech data and converts it into corresponding text data.
[0539] "Generation means for analyzing the converted text and generating a general answer based on the text" refers to the function of a device or software for analyzing the text obtained from the voice data and generating a general answer based on its content.
[0540] "Means for generating detailed answers based on specialized fields from the general answers generated by the generation means" refers to the function of a device or software that converts the initially generated general answers into more detailed answers based on specialized knowledge.
[0541] The "means for converting the generated detailed answer into voice data" refers to a device or software function that analyzes the detailed text answer and converts it into natural, easily understandable voice data.
[0542] The "means for outputting audio data to the user" refers to a device or software function that plays back and provides the generated audio data to the user.
[0543] "Selective application of multiple disciplinary generative tools" is the process of selecting and applying specialized generative tools in response to a specific question or request.
[0544] "Providing visual output using a virtual character" means using a virtual character model and displaying its movements and expressions to the user to provide visual feedback corresponding to the generated audio data.
[0545] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions. The configuration and operation of this system are described in detail below.
[0546] Hardware and Software Overview
[0547] Hardware used
[0548] Device: A user device that includes a microphone for capturing audio data and a speaker for playing audio data.
[0549] Server: A central processing unit for processing voice data and generating answers.
[0550] Software used
[0551] Speech Recognition Module: Software that converts voice data into text using, for example, the Google Cloud Speech-to-Text API.
[0552] Generative AI model (first stage): An AI model that generates general answers using OpenAI's GPT-3, etc.
[0553] Specialized AI: Generative AI specialized in a specific field. For example, AI models specialized in fields such as law or medicine.
[0554] Speech synthesis module: Software that converts text data into speech data using the Google Cloud Text-to-Speech API or similar.
[0555] Virtual Character Animation Engine: Software that uses technologies such as MetaHuman to provide visual feedback synchronized with audio data.
[0556] Specific explanation of operation
[0557] The details of the functions are shown below.
[0558] Voice input from the user
[0559] The user speaks questions or commands into the terminal. For example, the user might say, "I'd like some advice on defining the requirements for a new project."
[0560] Capture and transmit audio data
[0561] The device captures the user's voice with a microphone and transmits the data to the server in real time.
[0562] Converting audio data to text
[0563] The server analyzes the received voice data using the Google Cloud Speech-to-Text API and converts it into text data. In this step, the corresponding text is generated from the speech, "I would like some advice on defining the requirements for a new project."
[0564] General answer generation by first-level generative AI
[0565] The text data is analyzed by the server and input into a generative AI model such as OpenAI's GPT-3. An example prompt is "The user is looking for advice on the project requirements definition. Please provide general guidance." Based on this, a general answer is generated: "You're talking about the project requirements definition. Please tell me more specifically."
[0566] Detailed answers generated by specialized AI
[0567] A specialized AI is used to generate a more detailed answer based on a general answer from the first-stage generative AI. An example prompt is, "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of defining the requirements." The specialized AI generates a detailed answer: "Initial research and gathering stakeholder opinions are important for defining requirements."
[0568] Generating and transmitting audio data
[0569] The generated detailed answer is converted into audio data using the Google Cloud Text-to-Speech API, and the converted audio data is sent from the server to the device.
[0570] Audio output and animation playback to the user
[0571] The device receives the audio data and triggers MetaHuman's animation engine to animate the virtual character in sync with the audio playback, with the character's mouth movements and facial expressions matching the audio.
[0572] The above is an embodiment of the present invention. This system allows users to have natural visual and auditory interactions, and provides an advanced interactive system that can also respond to specialized questions.
[0573] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0574] Step 1:
[0575] Voice input from the user
[0576] User: The user speaks questions and commands into the microphone.
[0577] Input: User utterance (e.g., "I need some advice on defining requirements for a new project").
[0578] Output: Audio data
[0579] What it does: When a user speaks, the microphone captures the sound.
[0580] Step 2:
[0581] Capture and transmit audio data
[0582] Device: The device transmits the captured audio data to the server in real time.
[0583] Input: Captured audio data
[0584] Output: Audio data sent to the server
[0585] Specific operation: Audio captured by the microphone is saved in PCM format or similar and sent to the server as an HTTP request.
[0586] Step 3:
[0587] Converting audio data to text
[0588] Server: The server receives the voice data and passes it to the voice recognition module.
[0589] Input: Audio data sent from the device
[0590] Output: Text data
[0591] What it does: The received voice data is sent to the Google Cloud Speech-to-Text API, which converts the voice data into text. The resulting text is, "I'd like some advice on defining the requirements for a new project."
[0592] Step 4:
[0593] General answer generation by first-level generative AI
[0594] Server: The server analyzes the text data and inputs it into the generative AI.
[0595] Input: Text data (e.g., "I would like some advice on defining requirements for a new project.")
[0596] Output: General answer (text)
[0597] Specific operation: The server sends a prompt to OpenAI's GPT-3 or similar (e.g., "The user is looking for advice on the project requirements definition. Please provide general guidance."). The generative AI generates a general answer, such as "You're talking about the project requirements definition. Please be more specific," and replies to the server.
[0598] Step 5:
[0599] Detailed answers generated by specialized AI
[0600] Server: The server analyzes the general answers from the first-level generative AI and sends them to the specialized AI.
[0601] Input: General answer (text)
[0602] Output: Detailed answer (text)
[0603] Specific operation: The server sends a prompt to the specialized AI (e.g., "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of the requirements definition."). The specialized AI generates a detailed answer, such as "Initial research and gathering stakeholder opinions are important for defining requirements," and replies to the server.
[0604] Step 6:
[0605] Generating and transmitting audio data
[0606] Server: The server passes the detailed answer to the speech synthesis module, converts it into voice data, and sends it to the terminal.
[0607] Input: Detailed answer (text)
[0608] Output: Audio data (e.g., WAV or MP3 format)
[0609] Specific operation: The server sends the detailed answer to the Google Cloud Text-to-Speech API, converts the text "Initial research and stakeholder opinion gathering are important for requirements definition" into audio data, and sends the generated audio file to the device.
[0610] Step 7:
[0611] Audio output and animation playback to the user
[0612] Device: The device receives the audio data and triggers MetaHuman's animation engine to synchronize with the audio playback.
[0613] Input: Audio data
[0614] Output: Audio and visual feedback
[0615] How it works: The device plays back the received audio data, and at the same time, MetaHuman's virtual character moves its mouth and changes its facial expression in response to the audio, providing the user with a natural conversational experience.
[0616] (Application example 1)
[0617] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0618] In virtual stores, a system that integrates voice input and visual feedback is necessary to enable more natural and efficient user interaction. However, current technology simply converts the user's voice input into text and provides text-based answers, failing to realize natural dialogue that integrates visual and auditory feedback. Furthermore, there is a lack of effective means for generating detailed answers for each specialized field. This results in low user convenience and poses challenges for improving user satisfaction.
[0619] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0620] In this invention, the server includes means for converting voice input into text, means for analyzing the converted text and generating a general answer, means for generating detailed answers for each specialized field based on the generated general answers, means for converting the generated detailed answers into voice data, means for outputting the voice data to the user, means for providing visual output using a virtual character, and means for providing animated visual feedback synchronized with the voice data, thereby enabling natural dialogue that integrates voice input and visual feedback.
[0621] "Voice input" refers to words or questions spoken by the user, and is the voice data that the system uses to recognize them.
[0622] A "means for converting to text" is a technique or device for analyzing received voice input and converting the content into text data format.
[0623] "First-stage generation means" refers to a first-stage technique or device for generating a general answer based on the converted text.
[0624] The "generating means for each specialized field" is a technology or device that further analyzes the answer obtained by the generating means at the first stage and generates a detailed answer that is specific to a specific specialized field.
[0625] The "means for converting into voice data" refers to a technique or device for synthesizing the generated text-format answers into voice and outputting them as voice data.
[0626] "Visual feedback with animation synchronized with audio data" refers to a technology or device that displays visual actions or animations corresponding to generated audio data and provides them to the user along with the audio.
[0627] A "virtual character" is a visual character created using computer graphics and animation techniques to interact with a user.
[0628] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice form again. This system is particularly effective for realizing natural dialogue in virtual stores, and also provides visual feedback using virtual characters.
[0629] First, the user uses the smart glasses to input voice. The microphone in the smart glasses captures the voice and sends it to the server. The server then converts the received voice data into text using Google Cloud Speech-to-Text. This converted text is then input into GPT-4, the first stage of the generator, to generate a general answer.
[0630] The generated general answer is then passed to a specialized generator. In this example, a generator specialized for product information is used to generate a detailed answer. For example, if a user asks, "Tell me about this product," GPT-4 generates a general answer such as, "What category does this product belong to?" Then, a specialized AI generates a detailed answer such as, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[0631] The detailed answer is then converted into speech using Amazon Polly, which is then sent back to the server and played back to the smart glasses, where it uses Unreal Engine's MetaHuman to provide animated visual feedback synchronized with the speech.
[0632] As a concrete example, the following prompt sentence is input to GPT-4 to generate a general answer:
[0633] A user asked the following question:
[0634] Question: "Tell me about this product"
[0635] Generate a general answer to the question.
[0636] You can then create prompts for specialized AI and get detailed answers.
[0637] The following question was asked to the product information specialized AI.
[0638] Ask: "What category does this product belong to?"
[0639] Generate a professional answer to this question.
[0640] In this way, the present invention efficiently and naturally processes voice input from the user, provides optimal answers in the virtual store, and also provides a sense of visual familiarity to the user by using virtual characters.
[0641] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0642] Step 1:
[0643] The user uses the smart glasses to input voice. Specifically, the user asks a question by voice, such as "Tell me about this product." The microphone in the smart glasses captures this voice and inputs it as voice data into the terminal. The terminal then sends this voice data to the server.
[0644] Step 2:
[0645] The server converts the received voice data into text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data. Specifically, the voice recognition algorithm analyzes the voice waveform data and converts it into text such as "Tell me about this product."
[0646] Step 3:
[0647] The server then inputs the converted text data into GPT-4 to generate a general answer. The input is text data, and the output is the general answer text. Based on the prompt, GPT-4 generates a general answer such as, "What category does this product belong to?"
[0648] Step 4:
[0649] The generated general answer is input to a specialized AI on the server, which uses a generation method specialized for product information to generate a detailed answer. The general answer is the input, and the detailed answer text is obtained as the output. Specifically, the detailed answer generated is, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[0650] Step 5:
[0651] The detailed answer text is converted to speech using Amazon Polly. We have the detailed text answer as input and speech as output. A speech synthesis algorithm analyzes the text and generates natural-sounding speech.
[0652] Step 6:
[0653] The generated voice data is sent from the server to the device (smart glasses). The device plays the received voice data. Specifically, a detailed answer is provided to the user through the audio speaker: "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[0654] Step 7:
[0655] At the same time, the server uses Unreal Engine's MetaHuman to generate animated visual feedback corresponding to the generated voice data. The input is voice data, and the output is visual animation synchronized with the voice. Specifically, a virtual character visually presents the answer to the user, moving its mouth in sync with the voice.
[0656] This series of processing steps enables the user to experience natural and detailed interactions in the virtual store.
[0657] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0658] The present invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. This system allows users to have natural visual and auditory interactions, making it an advanced dialogue system that can also handle specialized questions.
[0659] Specifically, the following operations are included:
[0660] 1. User voice input
[0661] User: Speaks a question or instruction. For example, "I'd like to ask about the contents of the contract."
[0662] Device: Captures audio data and sends it to the server.
[0663] 2. Converting voice data to text
[0664] Server: Receives voice data and passes it to the voice recognition module.
[0665] Speech recognition module: Converts voice input into text data. For example, text data such as "I would like to ask for your advice regarding the contents of the contract."
[0666] Server: Inputs text data into the generative AI.
[0667] 3. Emotional Recognition
[0668] Server: Passes the voice input to the emotion engine and analyzes the user's emotions. For example, it analyzes the voice tone, rate, pitch, and volume to recognize that the user is feeling anxious.
[0669] Emotion engine: Recognizes the user's emotions and passes that information to the generative AI.
[0670] 4. Generating a general answer
[0671] Generative AI (first level): Generates general answers based on text data and recognized emotions. For example, it generates answers such as, "What specifically are you discussing about the contract?"
[0672] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[0673] 5. Generating detailed answers
[0674] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[0675] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine. For example, it generates detailed answers such as, "Please tell me about the specific clauses in the contract."
[0676] Server: Passes the output from the specialized AI to the speech synthesis module.
[0677] 6. Generating Audio Data
[0678] Speech synthesis module: Converts text responses into natural-sounding speech, such as "Please tell me about the specific clauses in the contract."
[0679] Server: Sends the generated audio data to the device.
[0680] 7. Audio output to the user
[0681] Terminal: Interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0682] MetaHuman: Provides visual and audible voice responses to the user, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[0683] Specific examples
[0684] Example 1: Legal advice
[0685] User: Say, "I need help with the contract."
[0686] Device: Captures audio and sends it to the server.
[0687] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[0688] Emotion engine: Recognizes that the user's tone of voice indicates anxiety.
[0689] Generative AI (first level): Generates a general answer such as, "You seem anxious. What specifically would you like to discuss regarding the contract?"
[0690] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[0691] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[0692] Device: The voice data is received and MetaHuman plays it back in a gentle tone as if it were a natural conversation.
[0693] Example 2: Engineering Support
[0694] User: "I need some advice on defining requirements for my new project."
[0695] Device: Captures audio and sends it to the server.
[0696] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[0697] Emotion engine: Recognizes that the user's tone of voice is calm.
[0698] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[0699] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[0700] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[0701] Device: The voice data is received and MetaHuman plays it back in a calm, natural-sounding conversation.
[0702] In this way, the present invention processes user voice input efficiently and naturally, providing optimal answers, and by using an emotion engine, it realizes flexible responses that match the user's emotions, providing a more personalized experience.
[0703] The processing flow will be explained below.
[0704] Step 1:
[0705] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[0706] Step 2:
[0707] The device's microphone captures the user's voice, and the voice data is sent to the server.
[0708] Step 3:
[0709] The server receives the voice data and passes it to a voice recognition module.
[0710] Step 4:
[0711] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[0712] Step 5:
[0713] The server receives the converted text data and passes it to the emotion engine.
[0714] Step 6:
[0715] The emotion engine analyzes text and voice data to recognize the user's emotions. For example, it analyzes voice tone, speed, pitch, and volume to determine if the user is feeling anxious.
[0716] Step 7:
[0717] The server inputs the recognized emotional information into the generative AI (first-stage base model).
[0718] Step 8:
[0719] The generative AI (first stage) generates a general answer based on the input text and emotional information. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[0720] Step 9:
[0721] The server analyzes the output of the first-stage generative AI to input it into the specialized AI, identifying the user's question as legal-related.
[0722] Step 10:
[0723] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[0724] Step 11:
[0725] A specialized AI (legal) generates detailed answers based on the input text and emotional information, such as "Please tell me about the specific clauses in the contract."
[0726] Step 12:
[0727] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[0728] Step 13:
[0729] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[0730] Step 14:
[0731] The server transmits the generated voice data to the terminal.
[0732] Step 15:
[0733] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0734] Step 16:
[0735] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[0736] In this way, the user, device, and server work together at each step to realize advanced voice dialogue. In addition, by combining it with an emotion engine, it is possible to provide flexible responses that correspond to the user's emotions.
[0737] Example 2
[0738] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0739] While conventional voice dialogue systems can provide appropriate answers to user voice inputs, they have difficulty generating flexible responses that reflect the user's emotions. Furthermore, for specialized questions, they can only provide general answers, failing to provide the detailed information actually required. Furthermore, there is a need for the answers output from the system to be natural and for the dialogue with the user to be visually acceptable.
[0740] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0741] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer based on the text by analyzing the converted text, means for generating a detailed answer based on a field of expertise from the general answer generated by the first-stage answer generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for recognizing the user's emotion from the voice input, and means for adjusting the generated answer based on the recognized emotion. This makes it possible to provide a flexible and detailed answer that reflects the user's emotion and realize a visually acceptable and natural dialogue.
[0742] The "means for receiving voice input from the user" is a device or function for capturing voice uttered by the user and incorporating it into the system.
[0743] A "means for converting voice input to text" is a device or function that analyzes received voice data and converts the content into text data.
[0744] The "first-stage generator" is an initial generator or algorithm that generates a general answer based on the converted text data.
[0745] The "means for generating a detailed answer based on a field of expertise" is a device or algorithm that generates a detailed answer specialized in a particular field of expertise based on the general answer generated by the first-stage generation means.
[0746] The "means for converting the generated detailed answer into voice data" is a device or function that converts the detailed answer in text format into natural voice data.
[0747] The "means for outputting voice data to the user" refers to a device or function that reproduces and provides the generated voice data to the user.
[0748] A "means for recognizing a user's emotion from speech input" is a device or algorithm that analyzes a user's speech data and determines their emotional state.
[0749] A "means for adjusting the generated answer based on the recognized emotion" is a device or algorithm that appropriately modifies or adjusts the generated answer depending on the user's emotional state.
[0750] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice format. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. Specific implementation methods of the system are described below.
[0751] This system operates primarily using the following hardware and software:
[0752] Hardware: A device equipped with a microphone to capture the user's voice and a speaker to play the output audio data.
[0753] Server: A central server for processing data and performing necessary calculations.
[0754] Software: Speech recognition module, emotion recognition engine, generative AI, specialized AI, voice synthesis module, virtual character animation engine.
[0755] The main processing flow of this system is as follows:
[0756] The user speaks a voice input, which is captured by the device's microphone. Specifically, the user speaks something like "I would like to consult you about the contents of the contract." The device sends the captured voice data to the server. The server uses a voice recognition module (e.g., a voice recognition API) to convert the voice data into text data. For example, the converted text might be something like "I would like to consult you about the contents of the contract."
[0757] The server then passes the text and voice data to an emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, speed, pitch, volume, etc. of the voice to recognize the emotion the user is feeling. Specifically, it may recognize that the user is feeling "anxiety."
[0758] The server inputs the recognized emotions and text data into a generative AI system to generate a general answer. For example, it might generate an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?" A general-purpose generative model is used for this generative AI model.
[0759] The server then passes the general answer to a specialized AI that generates a detailed answer. The specialized AI generates an answer based on knowledge of a specialized field, such as law, finance, or medicine. For example, it might generate a detailed answer such as, "Please tell me about the specific clauses in the contract."
[0760] The generated detailed answer is passed by the server to a speech synthesis module, which converts the text into natural-sounding speech data (e.g., speech synthesis API). The generated speech data will be something like, "Please tell me about the specific clauses of the contract."
[0761] The server sends the final voice data to the terminal, which triggers the virtual character's animation engine to synchronize with the voice playback. The virtual character speaks in a gentle tone, saying, "Please tell me about the specific clauses in the contract."
[0762] Specific examples
[0763] Example 1: Legal advice
[0764] The user says, "I'd like to ask you about the contents of the contract." The device captures the voice and sends it to the server. The server uses a speech recognition module to convert it into text, "I'd like to ask you about the contents of the contract," and uses an emotion engine to recognize the user's anxiety. The generative AI then generates a general answer, "You seem anxious. What specifically do you want to ask you about in the contract?", and the specialized AI generates a detailed answer, "Please tell me about the specific clauses in the contract." The speech synthesis module converts this into voice data, which the device finally provides to the user through a virtual character.
[0765] Prompt Sentence Examples
[0766] "Generate answers that will alleviate the user's concerns when asked about the contents of the contract."
[0767] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0768] Step 1:
[0769] The user speaks the speech input.
[0770] Specific operation: The user says, "I would like to discuss the contents of the contract."
[0771] Input: User's voice.
[0772] Output: The captured audio data.
[0773] Step 2:
[0774] The device captures the audio data and sends it to the server.
[0775] Specific operation: Records audio data using the device's microphone and sends the data to the server.
[0776] Input: The captured audio data.
[0777] Output: The audio data sent to the server.
[0778] Step 3:
[0779] The server receives the voice data and passes it to the voice recognition module.
[0780] Specific operation: The server receives voice data from the terminal and inputs the data into the voice recognition module.
[0781] Input: The audio data sent to the server.
[0782] Output: Audio data as input to the speech recognition module.
[0783] Step 4:
[0784] A voice recognition module converts the voice data into text data.
[0785] Specific operation: The voice recognition module analyzes the voice data and generates text data such as "I would like to consult you about the contents of the contract."
[0786] Input: Audio data as input to the speech recognition module.
[0787] Output: The converted text data.
[0788] Step 5:
[0789] The server passes the text data to an emotion recognition engine to analyze the user's emotions.
[0790] Specific operation: The server passes the text data and voice characteristic information to the emotion recognition engine and begins analysis.
[0791] Input: Translated text data and speech characteristics information.
[0792] Output: Parsed emotion data (e.g., anxiety).
[0793] Step 6:
[0794] The server inputs emotional data and text data into a generative AI to generate a general answer.
[0795] Specific operation: The server inputs emotion data and text data into the generative AI and receives the generated answer. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[0796] Input: Parsed emotion data and converted text data.
[0797] Output: The generated general answer text.
[0798] Step 7:
[0799] The server passes the general answer to a specialized AI, which generates a detailed answer.
[0800] Specific operation: The server passes the general answer text to a specialized AI (e.g., specialized AI for law, finance, medicine, etc.) to generate a detailed answer. For example, it generates a detailed answer such as "Please tell me about the specific clauses in the contract."
[0801] Input: The generated general answer text.
[0802] Output: The generated detailed answer text.
[0803] Step 8:
[0804] The server passes the detailed answer to a speech synthesis module and converts it into speech data.
[0805] Specific operation: The server inputs the detailed answer text into the speech synthesis module to generate speech data, for example, "Please tell me about the specific clauses of the contract."
[0806] Input: The generated long answer text.
[0807] Output: The generated audio data.
[0808] Step 9:
[0809] The server transmits the generated voice data to the terminal.
[0810] Specific operation: The server sends the generated audio data to the device.
[0811] Input: The generated audio data.
[0812] Output: The audio data sent to the device.
[0813] Step 10:
[0814] The device receives the audio data and triggers the animation engine of the virtual character.
[0815] Specific operation: The device interprets the received audio data and triggers the virtual character's animation engine to create a visual representation synchronized with the audio playback.
[0816] Input: Audio data sent to the device.
[0817] Output: Audio and visual answers provided to the user.
[0818] (Application example 2)
[0819] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0820] Conventional food delivery systems have had problems with the user having to make an order, requiring a lot of effort and not being able to respond flexibly to emotions. This can lead to low user satisfaction and delays in the ordering process. Furthermore, advanced technology is required to ensure natural visual and auditory interactions, and there has been a lack of systems that can achieve this.
[0821] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer by analyzing the text, means for generating a detailed answer based on the field of expertise from the general answer generated by the first-stage generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for analyzing the user's emotions, and means for adjusting the answer based on the results of the emotion analysis. This allows the user to place an order through natural dialogue and enables flexible responses according to emotions.
[0822] A "means for receiving voice input from a user" is a device or function for capturing voice data uttered by a user.
[0823] A "means for converting voice input to text" is a device or function that analyzes captured voice data and converts it into corresponding text data.
[0824] "An initial generation means for analyzing the converted text and generating a general answer based on the text" is a device or function that includes an AI module for generating an initial answer based on text data.
[0825] "Means for generating detailed answers based on specialized fields from the general answers generated by the initial generation means" refers to a device or function that includes an AI module for further specialized analysis of the output of the initial generation means and generating detailed answers.
[0826] The "means for converting the generated detailed answer into voice data" is a device or function for converting the generated detailed answer as text into voice data.
[0827] The "means for outputting audio data to the user" refers to a device or function for playing back the generated audio data to the user.
[0828] A "means for analyzing user emotions" is a device or function that analyzes a user's voice input to identify their emotional state.
[0829] The "means for adjusting the answer based on the result of sentiment analysis" is a device or function for taking into account the result of sentiment analysis and adjusting the answer to be generated and the method of presenting it.
[0830] The present invention is a system for efficiently delivering food via voice, allowing users to place orders through natural dialogue and providing flexible responses according to emotions. This system includes the following specific procedures and devices.
[0831] System configuration
[0832] The system takes voice input from the user, converts it to text, uses generative AI to generate a general answer, then generates a more specific answer based on the user's area of expertise (in this case, food delivery), converts the answer back into speech, and delivers it to the user. It can also recognize the user's emotions and adjust the answer accordingly.
[0833] Hardware and software used
[0834] Hardware:
[0835] Smartphone microphone: A device for capturing voice data spoken by a user.
[0836] Smartphone speaker: A device for playing back the generated audio data to the user.
[0837] software:
[0838] speech_recognition: A library for converting speech data into text data.
[0839] Transformers sentiment-analysis: A library for analyzing user sentiment from text data.
[0840] googletrans: A library for translating and converting answers into text.
[0841] gTTS (Google Text-To-Speech): A library for converting text data into audio data and playing it back.
[0842] System Operation
[0843] The server first receives voice input from the user. This voice is captured using a smartphone microphone. It then uses speech recognition software to convert the speech to text. This converted text is passed to a first-stage generative AI to generate a general answer. This general answer is then further processed into a detailed answer by a specialized AI in a specialized field (food delivery).
[0844] Emotion analysis
[0845] At the same time, the user's emotions are also analyzed. This is done using an emotion analysis engine that analyzes the user's voice tone, speed, pitch, and volume. Based on the analysis results, the answer is adjusted. For example, if the user is in a hurry, a quick response is required.
[0846] Speech synthesis and output
[0847] The detailed answer is then converted into audio data by speech synthesis software, which is then delivered to the user through the smartphone speaker, allowing them to order food delivery in a natural, conversational way.
[0848] Specific examples
[0849] Customer: "One pizza please."
[0850] application:
[0851] (Transcription): "One pizza please."
[0852] (emotion recognition): calm
[0853] (Order Processing): "Confirmed. We will process your order immediately."
[0854] (Text-to-Speech): Play "Thank you. We'll get it sorted right away."
[0855] Prompt Sentence Examples
[0856] Prompt: "A customer calmly orders, 'One pizza, please.' Generate an appropriate response for the calm customer."
[0857] Expected response: "Thank you. I'll get it sorted right away."
[0858] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0859] Step 1:
[0860] The server receives the user's voice input from the device. This input is the voice data spoken by the user using the smartphone microphone. The device then sends the captured voice data to the server. This data is transmitted as a voice waveform and is not in a format that can be analyzed as is.
[0861] Step 2:
[0862] The server converts the speech input into text. This is done using speech recognition software, which analyzes the speech waveform data and converts the phoneme sequence into the corresponding text. For example, if the speech input is "One pizza please," it will be converted into text "One pizza please." The output of this step is the converted text data.
[0863] Step 3:
[0864] The server analyzes the converted text and passes it to a first-stage generative AI that generates a general answer. This first-stage generative AI applies a generative AI model based on the text data to generate a general answer. For example, if the input text is "One pizza please," it will generate a general answer of "I understand. We'll get it ready right away." The output of this step is the general answer text.
[0865] Step 4:
[0866] The server passes the answer text obtained from the first-dan generation AI to a specialized AI, which generates a detailed answer. This specialized AI has a knowledge database specialized in a specific field (in this case, food delivery) and complements the first-dan answer in detail. For example, it adds a specific confirmation such as "Is this pizza size okay?" to the general answer "Please confirm." The output of this step is a detailed answer text.
[0867] Step 5:
[0868] The server receives the detailed answer text and analyzes the user's emotions. It uses an emotion analysis engine to identify the user's emotional state (e.g., hurry, calm, happy) by analyzing the tone, rate, pitch, and volume of the speech input. It adjusts the detailed answer based on this analysis data. The output of this step is the final adjusted detailed answer text.
[0869] Step 6:
[0870] The server converts the finalized detailed answer text into speech data. It uses speech synthesis software such as Google Text-To-Speech (gTTS) to convert the text data into speech data. This speech data is provided with natural pronunciation that is easy for the user to understand. The output of this step is synthesized speech data.
[0871] Step 7:
[0872] The server returns the generated voice data to the terminal, which then provides it to the user through the smartphone speaker, allowing the user to receive the voice answer. The final output of this step is the voice answer provided to the user.
[0873] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0874] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0875] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0876] [Third embodiment]
[0877] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0878] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0879] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0880] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0881] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0882] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0883] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0884] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0885] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0886] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0887] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0888] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0889] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions.
[0890] Specifically, the following operations are included:
[0891] 1. User voice input
[0892] User: Speaks a question or command. This voice input is received through the device's microphone.
[0893] Device: Captures audio data and sends it to the server.
[0894] 2. Converting voice data to text
[0895] Server: Receives voice data and passes it to the voice recognition module.
[0896] Speech recognition module: Converts voice input into text data.
[0897] Server: Inputs text data into the generative AI.
[0898] 3. Generating a general answer
[0899] Generative AI (first level): Analyzes text data and generates general answers that are appropriate responses to general questions.
[0900] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[0901] 4. Generating detailed answers
[0902] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[0903] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine.
[0904] Server: Passes the output from the specialized AI to the speech synthesis module.
[0905] 5. Generating Audio Data
[0906] Speech synthesis module: Converts text responses into natural-sounding speech data.
[0907] Server: Sends the generated audio data to the device.
[0908] 6. Audio output to the user
[0909] Terminal: Interprets the audio data received from the server, triggers MetaHuman's animation engine, and synchronizes it with the audio playback.
[0910] MetaHuman: Providing visual and auditory voice answers to users.
[0911] Specific examples
[0912] Example 1: Legal advice
[0913] User: Say, "I need help with the contract."
[0914] Device: Captures audio and sends it to the server.
[0915] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[0916] Generative AI (first level): Generates a general answer to the question, "What specifically is the subject of the contract consultation?"
[0917] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[0918] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[0919] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[0920] Example 2: Engineering Support
[0921] User: "I need some advice on defining requirements for my new project."
[0922] Device: Captures audio and sends it to the server.
[0923] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[0924] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[0925] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[0926] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[0927] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[0928] In this way, the present invention processes the user's voice input efficiently and naturally, providing optimal answers, and by using MetaHuman, also provides a visually familiar experience to the user.
[0929] The processing flow will be explained below.
[0930] Step 1:
[0931] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[0932] Step 2:
[0933] The device's microphone captures the user's voice, and the voice data is sent to the server.
[0934] Step 3:
[0935] The server receives the voice data and passes it to a voice recognition module.
[0936] Step 4:
[0937] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[0938] Step 5:
[0939] The server inputs the converted text into the generative AI (first-stage base model).
[0940] Step 6:
[0941] The generative AI (first stage) analyzes the input text and generates a general answer, such as, "What specifically is the content of the contract consultation?"
[0942] Step 7:
[0943] The server receives the output of the first stage generative AI and analyzes the answer, identifying that the user's question is related to law.
[0944] Step 8:
[0945] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[0946] Step 9:
[0947] Specialized AI (legal) generates detailed answers based on the input text, such as "Please tell me about the specific clauses in the contract."
[0948] Step 10:
[0949] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[0950] Step 11:
[0951] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[0952] Step 12:
[0953] The server transmits the generated voice data to the terminal.
[0954] Step 13:
[0955] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[0956] Step 14:
[0957] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as, "Tell me about the specific clauses in the contract."
[0958] In this way, the user, terminal, and server work together at each step, enabling this system to achieve advanced voice dialogue.
[0959] Example 1
[0960] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0961] Many conventional dialogue systems convert a user's voice input into text and generate a specific answer based on that text. However, when the user's question is highly specialized, these systems often produce vague and insufficient answers. There is a demand for systems that can provide highly accurate answers, especially for detailed questions related to specialized fields. Furthermore, conventional systems only output voice, making it difficult to provide sufficient visual feedback to the user. Therefore, it is necessary to provide a more natural and intuitive dialogue experience.
[0962] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0963] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the converted text and generating a general answer based on the text, means for generating a detailed answer based on a field of expertise from the general answer generated by the generation means, means for converting the generated detailed answer into voice data, and means for outputting the voice data to the user. This enables highly accurate answers even to specialized and advanced questions, and also makes it possible to provide a natural dialogue experience including visual feedback.
[0964] "Means for receiving voice input from a user" refers to a device or software function that can capture voice uttered by a user and transmit that data to the next processing step.
[0965] "Means for converting speech input to text" refers to a device or software function that analyzes received speech data and converts it into corresponding text data.
[0966] "Generation means for analyzing the converted text and generating a general answer based on the text" refers to the function of a device or software for analyzing the text obtained from the voice data and generating a general answer based on its content.
[0967] "Means for generating detailed answers based on specialized fields from the general answers generated by the generation means" refers to the function of a device or software that converts the initially generated general answers into more detailed answers based on specialized knowledge.
[0968] The "means for converting the generated detailed answer into voice data" refers to a device or software function that analyzes the detailed text answer and converts it into natural, easily understandable voice data.
[0969] The "means for outputting audio data to the user" refers to a device or software function that plays back and provides the generated audio data to the user.
[0970] "Selective application of multiple disciplinary generative tools" is the process of selecting and applying specialized generative tools in response to a specific question or request.
[0971] "Providing visual output using a virtual character" means using a virtual character model and displaying its movements and expressions to the user to provide visual feedback corresponding to the generated audio data.
[0972] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions. The configuration and operation of this system are described in detail below.
[0973] Hardware and Software Overview
[0974] Hardware used
[0975] Device: A user device that includes a microphone for capturing audio data and a speaker for playing audio data.
[0976] Server: A central processing unit for processing voice data and generating answers.
[0977] Software used
[0978] Speech Recognition Module: Software that converts voice data into text using, for example, the Google Cloud Speech-to-Text API.
[0979] Generative AI model (first stage): An AI model that generates general answers using OpenAI's GPT-3, etc.
[0980] Specialized AI: Generative AI specialized in a specific field. For example, AI models specialized in fields such as law or medicine.
[0981] Speech synthesis module: Software that converts text data into speech data using the Google Cloud Text-to-Speech API or similar.
[0982] Virtual Character Animation Engine: Software that uses technologies such as MetaHuman to provide visual feedback synchronized with audio data.
[0983] Specific explanation of operation
[0984] The details of the functions are shown below.
[0985] Voice input from the user
[0986] The user speaks questions or commands into the terminal. For example, the user might say, "I'd like some advice on defining the requirements for a new project."
[0987] Capture and transmit audio data
[0988] The device captures the user's voice with a microphone and transmits the data to the server in real time.
[0989] Converting audio data to text
[0990] The server analyzes the received voice data using the Google Cloud Speech-to-Text API and converts it into text data. In this step, the corresponding text is generated from the speech, "I would like some advice on defining the requirements for a new project."
[0991] General answer generation by first-level generative AI
[0992] The text data is analyzed by the server and input into a generative AI model such as OpenAI's GPT-3. An example prompt is "The user is looking for advice on the project requirements definition. Please provide general guidance." Based on this, a general answer is generated: "You're talking about the project requirements definition. Please tell me more specifically."
[0993] Detailed answers generated by specialized AI
[0994] A specialized AI is used to generate a more detailed answer based on a general answer from the first-stage generative AI. An example prompt is, "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of defining the requirements." The specialized AI generates a detailed answer: "Initial research and gathering stakeholder opinions are important for defining requirements."
[0995] Generating and transmitting audio data
[0996] The generated detailed answer is converted into audio data using the Google Cloud Text-to-Speech API, and the converted audio data is sent from the server to the device.
[0997] Audio output and animation playback to the user
[0998] The device receives the audio data and triggers MetaHuman's animation engine to animate the virtual character in sync with the audio playback, with the character's mouth movements and facial expressions matching the audio.
[0999] The above is an embodiment of the present invention. This system allows users to have natural visual and auditory interactions, and provides an advanced interactive system that can also respond to specialized questions.
[1000] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1001] Step 1:
[1002] Voice input from the user
[1003] User: The user speaks questions and commands into the microphone.
[1004] Input: User utterance (e.g., "I need some advice on defining requirements for a new project").
[1005] Output: Audio data
[1006] What it does: When a user speaks, the microphone captures the sound.
[1007] Step 2:
[1008] Capture and transmit audio data
[1009] Device: The device transmits the captured audio data to the server in real time.
[1010] Input: Captured audio data
[1011] Output: Audio data sent to the server
[1012] Specific operation: Audio captured by the microphone is saved in PCM format or similar and sent to the server as an HTTP request.
[1013] Step 3:
[1014] Converting audio data to text
[1015] Server: The server receives the voice data and passes it to the voice recognition module.
[1016] Input: Audio data sent from the device
[1017] Output: Text data
[1018] What it does: The received voice data is sent to the Google Cloud Speech-to-Text API, which converts the voice data into text. The resulting text is, "I'd like some advice on defining the requirements for a new project."
[1019] Step 4:
[1020] General answer generation by first-level generative AI
[1021] Server: The server analyzes the text data and inputs it into the generative AI.
[1022] Input: Text data (e.g., "I would like some advice on defining requirements for a new project.")
[1023] Output: General answer (text)
[1024] Specific operation: The server sends a prompt to OpenAI's GPT-3 or similar (e.g., "The user is looking for advice on the project requirements definition. Please provide general guidance."). The generative AI generates a general answer, such as "You're talking about the project requirements definition. Please be more specific," and replies to the server.
[1025] Step 5:
[1026] Detailed answers generated by specialized AI
[1027] Server: The server analyzes the general answers from the first-level generative AI and sends them to the specialized AI.
[1028] Input: General answer (text)
[1029] Output: Detailed answer (text)
[1030] Specific operation: The server sends a prompt to the specialized AI (e.g., "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of the requirements definition."). The specialized AI generates a detailed answer, such as "Initial research and gathering stakeholder opinions are important for defining requirements," and replies to the server.
[1031] Step 6:
[1032] Generating and transmitting audio data
[1033] Server: The server passes the detailed answer to the speech synthesis module, converts it into voice data, and sends it to the terminal.
[1034] Input: Detailed answer (text)
[1035] Output: Audio data (e.g., WAV or MP3 format)
[1036] Specific operation: The server sends the detailed answer to the Google Cloud Text-to-Speech API, converts the text "Initial research and stakeholder opinion gathering are important for requirements definition" into audio data, and sends the generated audio file to the device.
[1037] Step 7:
[1038] Audio output and animation playback to the user
[1039] Device: The device receives the audio data and triggers MetaHuman's animation engine to synchronize with the audio playback.
[1040] Input: Audio data
[1041] Output: Audio and visual feedback
[1042] How it works: The device plays back the received audio data, and at the same time, MetaHuman's virtual character moves its mouth and changes its facial expression in response to the audio, providing the user with a natural conversational experience.
[1043] (Application example 1)
[1044] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1045] In virtual stores, a system that integrates voice input and visual feedback is necessary to enable more natural and efficient user interaction. However, current technology simply converts the user's voice input into text and provides text-based answers, failing to realize natural dialogue that integrates visual and auditory feedback. Furthermore, there is a lack of effective means for generating detailed answers for each specialized field. This results in low user convenience and poses challenges for improving user satisfaction.
[1046] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1047] In this invention, the server includes means for converting voice input into text, means for analyzing the converted text and generating a general answer, means for generating detailed answers for each specialized field based on the generated general answers, means for converting the generated detailed answers into voice data, means for outputting the voice data to the user, means for providing visual output using a virtual character, and means for providing animated visual feedback synchronized with the voice data, thereby enabling natural dialogue that integrates voice input and visual feedback.
[1048] "Voice input" refers to words or questions spoken by the user, and is the voice data that the system uses to recognize them.
[1049] A "means for converting to text" is a technique or device for analyzing received voice input and converting the content into text data format.
[1050] "First-stage generation means" refers to a first-stage technique or device for generating a general answer based on the converted text.
[1051] The "generating means for each specialized field" is a technology or device that further analyzes the answer obtained by the generating means at the first stage and generates a detailed answer that is specific to a specific specialized field.
[1052] The "means for converting into voice data" refers to a technique or device for synthesizing the generated text-format answers into voice and outputting them as voice data.
[1053] "Visual feedback with animation synchronized with audio data" refers to a technology or device that displays visual actions or animations corresponding to generated audio data and provides them to the user along with the audio.
[1054] A "virtual character" is a visual character created using computer graphics and animation techniques to interact with a user.
[1055] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice form again. This system is particularly effective for realizing natural dialogue in virtual stores, and also provides visual feedback using virtual characters.
[1056] First, the user uses the smart glasses to input voice. The microphone in the smart glasses captures the voice and sends it to the server. The server then converts the received voice data into text using Google Cloud Speech-to-Text. This converted text is then input into GPT-4, the first stage of the generator, to generate a general answer.
[1057] The generated general answer is then passed to a specialized generator. In this example, a generator specialized for product information is used to generate a detailed answer. For example, if a user asks, "Tell me about this product," GPT-4 generates a general answer such as, "What category does this product belong to?" Then, a specialized AI generates a detailed answer such as, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[1058] The detailed answer is then converted into speech using Amazon Polly, which is then sent back to the server and played back to the smart glasses, where it uses Unreal Engine's MetaHuman to provide animated visual feedback synchronized with the speech.
[1059] As a concrete example, the following prompt sentence is input to GPT-4 to generate a general answer:
[1060] A user asked the following question:
[1061] Question: "Tell me about this product"
[1062] Generate a general answer to the question.
[1063] You can then create prompts for specialized AI and get detailed answers.
[1064] The following question was asked to the product information specialized AI.
[1065] Ask: "What category does this product belong to?"
[1066] Generate a professional answer to this question.
[1067] In this way, the present invention efficiently and naturally processes voice input from the user, provides optimal answers in the virtual store, and also provides a sense of visual familiarity to the user by using virtual characters.
[1068] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1069] Step 1:
[1070] The user uses the smart glasses to input voice. Specifically, the user asks a question by voice, such as "Tell me about this product." The microphone in the smart glasses captures this voice and inputs it as voice data into the terminal. The terminal then sends this voice data to the server.
[1071] Step 2:
[1072] The server converts the received voice data into text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data. Specifically, the voice recognition algorithm analyzes the voice waveform data and converts it into text such as "Tell me about this product."
[1073] Step 3:
[1074] The server then inputs the converted text data into GPT-4 to generate a general answer. The input is text data, and the output is the general answer text. Based on the prompt, GPT-4 generates a general answer such as, "What category does this product belong to?"
[1075] Step 4:
[1076] The generated general answer is input to a specialized AI on the server, which uses a generation method specialized for product information to generate a detailed answer. The general answer is the input, and the detailed answer text is obtained as the output. Specifically, the detailed answer generated is, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[1077] Step 5:
[1078] The detailed answer text is converted to speech using Amazon Polly. We have the detailed text answer as input and speech as output. A speech synthesis algorithm analyzes the text and generates natural-sounding speech.
[1079] Step 6:
[1080] The generated voice data is sent from the server to the device (smart glasses). The device plays the received voice data. Specifically, a detailed answer is provided to the user through the audio speaker: "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[1081] Step 7:
[1082] At the same time, the server uses Unreal Engine's MetaHuman to generate animated visual feedback corresponding to the generated voice data. The input is voice data, and the output is visual animation synchronized with the voice. Specifically, a virtual character visually presents the answer to the user, moving its mouth in sync with the voice.
[1083] This series of processing steps enables the user to experience natural and detailed interactions in the virtual store.
[1084] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1085] The present invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. This system allows users to have natural visual and auditory interactions, making it an advanced dialogue system that can also handle specialized questions.
[1086] Specifically, the following operations are included:
[1087] 1. User voice input
[1088] User: Speaks a question or instruction. For example, "I'd like to ask about the contents of the contract."
[1089] Device: Captures audio data and sends it to the server.
[1090] 2. Converting voice data to text
[1091] Server: Receives voice data and passes it to the voice recognition module.
[1092] Speech recognition module: Converts voice input into text data. For example, text data such as "I would like to ask for your advice regarding the contents of the contract."
[1093] Server: Inputs text data into the generative AI.
[1094] 3. Emotional Recognition
[1095] Server: Passes the voice input to the emotion engine and analyzes the user's emotions. For example, it analyzes the voice tone, rate, pitch, and volume to recognize that the user is feeling anxious.
[1096] Emotion engine: Recognizes the user's emotions and passes that information to the generative AI.
[1097] 4. Generating a general answer
[1098] Generative AI (first level): Generates general answers based on text data and recognized emotions. For example, it generates answers such as, "What specifically are you discussing about the contract?"
[1099] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[1100] 5. Generating detailed answers
[1101] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[1102] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine. For example, it generates detailed answers such as, "Please tell me about the specific clauses in the contract."
[1103] Server: Passes the output from the specialized AI to the speech synthesis module.
[1104] 6. Generating Audio Data
[1105] Speech synthesis module: Converts text responses into natural-sounding speech, such as "Please tell me about the specific clauses in the contract."
[1106] Server: Sends the generated audio data to the device.
[1107] 7. Audio output to the user
[1108] Terminal: Interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[1109] MetaHuman: Provides visual and audible voice responses to the user, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[1110] Specific examples
[1111] Example 1: Legal advice
[1112] User: Say, "I need help with the contract."
[1113] Device: Captures audio and sends it to the server.
[1114] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[1115] Emotion engine: Recognizes that the user's tone of voice indicates anxiety.
[1116] Generative AI (first level): Generates a general answer such as, "You seem anxious. What specifically would you like to discuss regarding the contract?"
[1117] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[1118] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[1119] Device: The voice data is received and MetaHuman plays it back in a gentle tone as if it were a natural conversation.
[1120] Example 2: Engineering Support
[1121] User: "I need some advice on defining requirements for my new project."
[1122] Device: Captures audio and sends it to the server.
[1123] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[1124] Emotion engine: Recognizes that the user's tone of voice is calm.
[1125] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[1126] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[1127] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[1128] Device: The voice data is received and MetaHuman plays it back in a calm, natural-sounding conversation.
[1129] In this way, the present invention processes user voice input efficiently and naturally, providing optimal answers, and by using an emotion engine, it realizes flexible responses that match the user's emotions, providing a more personalized experience.
[1130] The processing flow will be explained below.
[1131] Step 1:
[1132] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[1133] Step 2:
[1134] The device's microphone captures the user's voice, and the voice data is sent to the server.
[1135] Step 3:
[1136] The server receives the voice data and passes it to a voice recognition module.
[1137] Step 4:
[1138] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[1139] Step 5:
[1140] The server receives the converted text data and passes it to the emotion engine.
[1141] Step 6:
[1142] The emotion engine analyzes text and voice data to recognize the user's emotions. For example, it analyzes voice tone, speed, pitch, and volume to determine if the user is feeling anxious.
[1143] Step 7:
[1144] The server inputs the recognized emotional information into the generative AI (first-stage base model).
[1145] Step 8:
[1146] The generative AI (first stage) generates a general answer based on the input text and emotional information. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[1147] Step 9:
[1148] The server analyzes the output of the first-stage generative AI to input it into the specialized AI, identifying the user's question as legal-related.
[1149] Step 10:
[1150] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[1151] Step 11:
[1152] A specialized AI (legal) generates detailed answers based on the input text and emotional information, such as "Please tell me about the specific clauses in the contract."
[1153] Step 12:
[1154] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[1155] Step 13:
[1156] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[1157] Step 14:
[1158] The server transmits the generated voice data to the terminal.
[1159] Step 15:
[1160] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[1161] Step 16:
[1162] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[1163] In this way, the user, device, and server work together at each step to realize advanced voice dialogue. In addition, by combining it with an emotion engine, it is possible to provide flexible responses that correspond to the user's emotions.
[1164] Example 2
[1165] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1166] While conventional voice dialogue systems can provide appropriate answers to user voice inputs, they have difficulty generating flexible responses that reflect the user's emotions. Furthermore, for specialized questions, they can only provide general answers, failing to provide the detailed information actually required. Furthermore, there is a need for the answers output from the system to be natural and for the dialogue with the user to be visually acceptable.
[1167] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1168] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer based on the text by analyzing the converted text, means for generating a detailed answer based on a field of expertise from the general answer generated by the first-stage answer generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for recognizing the user's emotion from the voice input, and means for adjusting the generated answer based on the recognized emotion. This makes it possible to provide a flexible and detailed answer that reflects the user's emotion and realize a visually acceptable and natural dialogue.
[1169] The "means for receiving voice input from the user" is a device or function for capturing voice uttered by the user and incorporating it into the system.
[1170] A "means for converting voice input to text" is a device or function that analyzes received voice data and converts the content into text data.
[1171] The "first-stage generator" is an initial generator or algorithm that generates a general answer based on the converted text data.
[1172] The "means for generating a detailed answer based on a field of expertise" is a device or algorithm that generates a detailed answer specialized in a particular field of expertise based on the general answer generated by the first-stage generation means.
[1173] The "means for converting the generated detailed answer into voice data" is a device or function that converts the detailed answer in text format into natural voice data.
[1174] The "means for outputting voice data to the user" refers to a device or function that reproduces and provides the generated voice data to the user.
[1175] A "means for recognizing a user's emotion from speech input" is a device or algorithm that analyzes a user's speech data and determines their emotional state.
[1176] A "means for adjusting the generated answer based on the recognized emotion" is a device or algorithm that appropriately modifies or adjusts the generated answer depending on the user's emotional state.
[1177] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice format. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. Specific implementation methods of the system are described below.
[1178] This system operates primarily using the following hardware and software:
[1179] Hardware: A device equipped with a microphone to capture the user's voice and a speaker to play the output audio data.
[1180] Server: A central server for processing data and performing necessary calculations.
[1181] Software: Speech recognition module, emotion recognition engine, generative AI, specialized AI, voice synthesis module, virtual character animation engine.
[1182] The main processing flow of this system is as follows:
[1183] The user speaks a voice input, which is captured by the device's microphone. Specifically, the user speaks something like "I would like to consult you about the contents of the contract." The device sends the captured voice data to the server. The server uses a voice recognition module (e.g., a voice recognition API) to convert the voice data into text data. For example, the converted text might be something like "I would like to consult you about the contents of the contract."
[1184] The server then passes the text and voice data to an emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, speed, pitch, volume, etc. of the voice to recognize the emotion the user is feeling. Specifically, it may recognize that the user is feeling "anxiety."
[1185] The server inputs the recognized emotions and text data into a generative AI system to generate a general answer. For example, it might generate an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?" A general-purpose generative model is used for this generative AI model.
[1186] The server then passes the general answer to a specialized AI that generates a detailed answer. The specialized AI generates an answer based on knowledge of a specialized field, such as law, finance, or medicine. For example, it might generate a detailed answer such as, "Please tell me about the specific clauses in the contract."
[1187] The generated detailed answer is passed by the server to a speech synthesis module, which converts the text into natural-sounding speech data (e.g., speech synthesis API). The generated speech data will be something like, "Please tell me about the specific clauses of the contract."
[1188] The server sends the final voice data to the terminal, which triggers the virtual character's animation engine to synchronize with the voice playback. The virtual character speaks in a gentle tone, saying, "Please tell me about the specific clauses in the contract."
[1189] Specific examples
[1190] Example 1: Legal advice
[1191] The user says, "I'd like to ask you about the contents of the contract." The device captures the voice and sends it to the server. The server uses a speech recognition module to convert it into text, "I'd like to ask you about the contents of the contract," and uses an emotion engine to recognize the user's anxiety. The generative AI then generates a general answer, "You seem anxious. What specifically do you want to ask you about in the contract?", and the specialized AI generates a detailed answer, "Please tell me about the specific clauses in the contract." The speech synthesis module converts this into voice data, which the device finally provides to the user through a virtual character.
[1192] Prompt Sentence Examples
[1193] "Generate answers that will alleviate the user's concerns when asked about the contents of the contract."
[1194] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1195] Step 1:
[1196] The user speaks the speech input.
[1197] Specific operation: The user says, "I would like to discuss the contents of the contract."
[1198] Input: User's voice.
[1199] Output: The captured audio data.
[1200] Step 2:
[1201] The device captures the audio data and sends it to the server.
[1202] Specific operation: Records audio data using the device's microphone and sends the data to the server.
[1203] Input: The captured audio data.
[1204] Output: The audio data sent to the server.
[1205] Step 3:
[1206] The server receives the voice data and passes it to the voice recognition module.
[1207] Specific operation: The server receives voice data from the terminal and inputs the data into the voice recognition module.
[1208] Input: The audio data sent to the server.
[1209] Output: Audio data as input to the speech recognition module.
[1210] Step 4:
[1211] A voice recognition module converts the voice data into text data.
[1212] Specific operation: The voice recognition module analyzes the voice data and generates text data such as "I would like to consult you about the contents of the contract."
[1213] Input: Audio data as input to the speech recognition module.
[1214] Output: The converted text data.
[1215] Step 5:
[1216] The server passes the text data to an emotion recognition engine to analyze the user's emotions.
[1217] Specific operation: The server passes the text data and voice characteristic information to the emotion recognition engine and begins analysis.
[1218] Input: Translated text data and speech characteristics information.
[1219] Output: Parsed emotion data (e.g., anxiety).
[1220] Step 6:
[1221] The server inputs emotional data and text data into a generative AI to generate a general answer.
[1222] Specific operation: The server inputs emotion data and text data into the generative AI and receives the generated answer. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[1223] Input: Parsed emotion data and converted text data.
[1224] Output: The generated general answer text.
[1225] Step 7:
[1226] The server passes the general answer to a specialized AI, which generates a detailed answer.
[1227] Specific operation: The server passes the general answer text to a specialized AI (e.g., specialized AI for law, finance, medicine, etc.) to generate a detailed answer. For example, it generates a detailed answer such as "Please tell me about the specific clauses in the contract."
[1228] Input: The generated general answer text.
[1229] Output: The generated detailed answer text.
[1230] Step 8:
[1231] The server passes the detailed answer to a speech synthesis module and converts it into speech data.
[1232] Specific operation: The server inputs the detailed answer text into the speech synthesis module to generate speech data, for example, "Please tell me about the specific clauses of the contract."
[1233] Input: The generated long answer text.
[1234] Output: The generated audio data.
[1235] Step 9:
[1236] The server transmits the generated voice data to the terminal.
[1237] Specific operation: The server sends the generated audio data to the device.
[1238] Input: The generated audio data.
[1239] Output: The audio data sent to the device.
[1240] Step 10:
[1241] The device receives the audio data and triggers the animation engine of the virtual character.
[1242] Specific operation: The device interprets the received audio data and triggers the virtual character's animation engine to create a visual representation synchronized with the audio playback.
[1243] Input: Audio data sent to the device.
[1244] Output: Audio and visual answers provided to the user.
[1245] (Application example 2)
[1246] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1247] Conventional food delivery systems have had problems with the user having to make an order, requiring a lot of effort and not being able to respond flexibly to emotions. This can lead to low user satisfaction and delays in the ordering process. Furthermore, advanced technology is required to ensure natural visual and auditory interactions, and there has been a lack of systems that can achieve this.
[1248] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer by analyzing the text, means for generating a detailed answer based on the field of expertise from the general answer generated by the first-stage generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for analyzing the user's emotions, and means for adjusting the answer based on the results of the emotion analysis. This allows the user to place an order through natural dialogue and enables flexible responses according to emotions.
[1249] A "means for receiving voice input from a user" is a device or function for capturing voice data uttered by a user.
[1250] A "means for converting voice input to text" is a device or function that analyzes captured voice data and converts it into corresponding text data.
[1251] "An initial generation means for analyzing the converted text and generating a general answer based on the text" is a device or function that includes an AI module for generating an initial answer based on text data.
[1252] "Means for generating detailed answers based on specialized fields from the general answers generated by the initial generation means" refers to a device or function that includes an AI module for further specialized analysis of the output of the initial generation means and generating detailed answers.
[1253] The "means for converting the generated detailed answer into voice data" is a device or function for converting the generated detailed answer as text into voice data.
[1254] The "means for outputting audio data to the user" refers to a device or function for playing back the generated audio data to the user.
[1255] A "means for analyzing user emotions" is a device or function that analyzes a user's voice input to identify their emotional state.
[1256] The "means for adjusting the answer based on the result of sentiment analysis" is a device or function for taking into account the result of sentiment analysis and adjusting the answer to be generated and the method of presenting it.
[1257] The present invention is a system for efficiently delivering food via voice, allowing users to place orders through natural dialogue and providing flexible responses according to emotions. This system includes the following specific procedures and devices.
[1258] System configuration
[1259] The system takes voice input from the user, converts it to text, uses generative AI to generate a general answer, then generates a more specific answer based on the user's area of expertise (in this case, food delivery), converts the answer back into speech, and delivers it to the user. It can also recognize the user's emotions and adjust the answer accordingly.
[1260] Hardware and software used
[1261] Hardware:
[1262] Smartphone microphone: A device for capturing voice data spoken by a user.
[1263] Smartphone speaker: A device for playing back the generated audio data to the user.
[1264] software:
[1265] speech_recognition: A library for converting speech data into text data.
[1266] Transformers sentiment-analysis: A library for analyzing user sentiment from text data.
[1267] googletrans: A library for translating and converting answers into text.
[1268] gTTS (Google Text-To-Speech): A library for converting text data into audio data and playing it back.
[1269] System Operation
[1270] The server first receives voice input from the user. This voice is captured using a smartphone microphone. It then uses speech recognition software to convert the speech to text. This converted text is passed to a first-stage generative AI to generate a general answer. This general answer is then further processed into a detailed answer by a specialized AI in a specialized field (food delivery).
[1271] Emotion analysis
[1272] At the same time, the user's emotions are also analyzed. This is done using an emotion analysis engine that analyzes the user's voice tone, speed, pitch, and volume. Based on the analysis results, the answer is adjusted. For example, if the user is in a hurry, a quick response is required.
[1273] Speech synthesis and output
[1274] The detailed answer is then converted into audio data by speech synthesis software, which is then delivered to the user through the smartphone speaker, allowing them to order food delivery in a natural, conversational way.
[1275] Specific examples
[1276] Customer: "One pizza please."
[1277] application:
[1278] (Transcription): "One pizza please."
[1279] (emotion recognition): calm
[1280] (Order Processing): "Confirmed. We will process your order immediately."
[1281] (Text-to-Speech): Play "Thank you. We'll get it sorted right away."
[1282] Prompt Sentence Examples
[1283] Prompt: "A customer calmly orders, 'One pizza, please.' Generate an appropriate response for the calm customer."
[1284] Expected response: "Thank you. I'll get it sorted right away."
[1285] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1286] Step 1:
[1287] The server receives the user's voice input from the device. This input is the voice data spoken by the user using the smartphone microphone. The device then sends the captured voice data to the server. This data is transmitted as a voice waveform and is not in a format that can be analyzed as is.
[1288] Step 2:
[1289] The server converts the speech input into text. This is done using speech recognition software, which analyzes the speech waveform data and converts the phoneme sequence into the corresponding text. For example, if the speech input is "One pizza please," it will be converted into text "One pizza please." The output of this step is the converted text data.
[1290] Step 3:
[1291] The server analyzes the converted text and passes it to a first-stage generative AI that generates a general answer. This first-stage generative AI applies a generative AI model based on the text data to generate a general answer. For example, if the input text is "One pizza please," it will generate a general answer of "I understand. We'll get it ready right away." The output of this step is the general answer text.
[1292] Step 4:
[1293] The server passes the answer text obtained from the first-dan generation AI to a specialized AI, which generates a detailed answer. This specialized AI has a knowledge database specialized in a specific field (in this case, food delivery) and complements the first-dan answer in detail. For example, it adds a specific confirmation such as "Is this pizza size okay?" to the general answer "Please confirm." The output of this step is a detailed answer text.
[1294] Step 5:
[1295] The server receives the detailed answer text and analyzes the user's emotions. It uses an emotion analysis engine to identify the user's emotional state (e.g., hurry, calm, happy) by analyzing the tone, rate, pitch, and volume of the speech input. It adjusts the detailed answer based on this analysis data. The output of this step is the final adjusted detailed answer text.
[1296] Step 6:
[1297] The server converts the finalized detailed answer text into speech data. It uses speech synthesis software such as Google Text-To-Speech (gTTS) to convert the text data into speech data. This speech data is provided with natural pronunciation that is easy for the user to understand. The output of this step is synthesized speech data.
[1298] Step 7:
[1299] The server returns the generated voice data to the terminal, which then provides it to the user through the smartphone speaker, allowing the user to receive the voice answer. The final output of this step is the voice answer provided to the user.
[1300] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1301] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1302] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1303] [Fourth embodiment]
[1304] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1305] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1306] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1307] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1308] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1309] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1310] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1311] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1312] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1313] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1314] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1315] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1316] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1317] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions.
[1318] Specifically, the following operations are included:
[1319] 1. User voice input
[1320] User: Speaks a question or command. This voice input is received through the device's microphone.
[1321] Device: Captures audio data and sends it to the server.
[1322] 2. Converting voice data to text
[1323] Server: Receives voice data and passes it to the voice recognition module.
[1324] Speech recognition module: Converts voice input into text data.
[1325] Server: Inputs text data into the generative AI.
[1326] 3. Generating a general answer
[1327] Generative AI (first level): Analyzes text data and generates general answers that are appropriate responses to general questions.
[1328] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[1329] 4. Generating detailed answers
[1330] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[1331] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine.
[1332] Server: Passes the output from the specialized AI to the speech synthesis module.
[1333] 5. Generating Audio Data
[1334] Speech synthesis module: Converts text responses into natural-sounding speech data.
[1335] Server: Sends the generated audio data to the device.
[1336] 6. Audio output to the user
[1337] Terminal: Interprets the audio data received from the server, triggers MetaHuman's animation engine, and synchronizes it with the audio playback.
[1338] MetaHuman: Providing visual and auditory voice answers to users.
[1339] Specific examples
[1340] Example 1: Legal advice
[1341] User: Say, "I need help with the contract."
[1342] Device: Captures audio and sends it to the server.
[1343] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[1344] Generative AI (first level): Generates a general answer to the question, "What specifically is the subject of the contract consultation?"
[1345] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[1346] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[1347] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[1348] Example 2: Engineering Support
[1349] User: "I need some advice on defining requirements for my new project."
[1350] Device: Captures audio and sends it to the server.
[1351] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[1352] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[1353] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[1354] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[1355] Terminal: The voice data is received and MetaHuman plays it back as natural conversation.
[1356] In this way, the present invention processes the user's voice input efficiently and naturally, providing optimal answers, and by using MetaHuman, also provides a visually familiar experience to the user.
[1357] The processing flow will be explained below.
[1358] Step 1:
[1359] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[1360] Step 2:
[1361] The device's microphone captures the user's voice, and the voice data is sent to the server.
[1362] Step 3:
[1363] The server receives the voice data and passes it to a voice recognition module.
[1364] Step 4:
[1365] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[1366] Step 5:
[1367] The server inputs the converted text into the generative AI (first-stage base model).
[1368] Step 6:
[1369] The generative AI (first stage) analyzes the input text and generates a general answer, such as, "What specifically is the content of the contract consultation?"
[1370] Step 7:
[1371] The server receives the output of the first stage generative AI and analyzes the answer, identifying that the user's question is related to law.
[1372] Step 8:
[1373] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[1374] Step 9:
[1375] Specialized AI (legal) generates detailed answers based on the input text, such as "Please tell me about the specific clauses in the contract."
[1376] Step 10:
[1377] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[1378] Step 11:
[1379] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[1380] Step 12:
[1381] The server transmits the generated voice data to the terminal.
[1382] Step 13:
[1383] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[1384] Step 14:
[1385] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as, "Tell me about the specific clauses in the contract."
[1386] In this way, the user, terminal, and server work together at each step, enabling this system to achieve advanced voice dialogue.
[1387] Example 1
[1388] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1389] Many conventional dialogue systems convert a user's voice input into text and generate a specific answer based on that text. However, when the user's question is highly specialized, these systems often produce vague and insufficient answers. There is a demand for systems that can provide highly accurate answers, especially for detailed questions related to specialized fields. Furthermore, conventional systems only output voice, making it difficult to provide sufficient visual feedback to the user. Therefore, it is necessary to provide a more natural and intuitive dialogue experience.
[1390] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1391] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for analyzing the converted text and generating a general answer based on the text, means for generating a detailed answer based on a field of expertise from the general answer generated by the generation means, means for converting the generated detailed answer into voice data, and means for outputting the voice data to the user. This enables highly accurate answers even to specialized and advanced questions, and also makes it possible to provide a natural dialogue experience including visual feedback.
[1392] "Means for receiving voice input from a user" refers to a device or software function that can capture voice uttered by a user and transmit that data to the next processing step.
[1393] "Means for converting speech input to text" refers to a device or software function that analyzes received speech data and converts it into corresponding text data.
[1394] "Generation means for analyzing the converted text and generating a general answer based on the text" refers to the function of a device or software for analyzing the text obtained from the voice data and generating a general answer based on its content.
[1395] "Means for generating detailed answers based on specialized fields from the general answers generated by the generation means" refers to the function of a device or software that converts the initially generated general answers into more detailed answers based on specialized knowledge.
[1396] The "means for converting the generated detailed answer into voice data" refers to a device or software function that analyzes the detailed text answer and converts it into natural, easily understandable voice data.
[1397] The "means for outputting audio data to the user" refers to a device or software function that plays back and provides the generated audio data to the user.
[1398] "Selective application of multiple disciplinary generative tools" is the process of selecting and applying specialized generative tools in response to a specific question or request.
[1399] "Providing visual output using a virtual character" means using a virtual character model and displaying its movements and expressions to the user to provide visual feedback corresponding to the generated audio data.
[1400] This invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. This system is an advanced interactive system that allows users to have natural visual and auditory interactions and can also respond to specialized questions. The configuration and operation of this system are described in detail below.
[1401] Hardware and Software Overview
[1402] Hardware used
[1403] Device: A user device that includes a microphone for capturing audio data and a speaker for playing audio data.
[1404] Server: A central processing unit for processing voice data and generating answers.
[1405] Software used
[1406] Speech Recognition Module: Software that converts voice data into text using, for example, the Google Cloud Speech-to-Text API.
[1407] Generative AI model (first stage): An AI model that generates general answers using OpenAI's GPT-3, etc.
[1408] Specialized AI: Generative AI specialized in a specific field. For example, AI models specialized in fields such as law or medicine.
[1409] Speech synthesis module: Software that converts text data into speech data using the Google Cloud Text-to-Speech API or similar.
[1410] Virtual Character Animation Engine: Software that uses technologies such as MetaHuman to provide visual feedback synchronized with audio data.
[1411] Specific explanation of operation
[1412] The details of the functions are shown below.
[1413] Voice input from the user
[1414] The user speaks questions or commands into the terminal. For example, the user might say, "I'd like some advice on defining the requirements for a new project."
[1415] Capture and transmit audio data
[1416] The device captures the user's voice with a microphone and transmits the data to the server in real time.
[1417] Converting audio data to text
[1418] The server analyzes the received voice data using the Google Cloud Speech-to-Text API and converts it into text data. In this step, the corresponding text is generated from the speech, "I would like some advice on defining the requirements for a new project."
[1419] General answer generation by first-level generative AI
[1420] The text data is analyzed by the server and input into a generative AI model such as OpenAI's GPT-3. An example prompt is "The user is looking for advice on the project requirements definition. Please provide general guidance." Based on this, a general answer is generated: "You're talking about the project requirements definition. Please tell me more specifically."
[1421] Detailed answers generated by specialized AI
[1422] A specialized AI is used to generate a more detailed answer based on a general answer from the first-stage generative AI. An example prompt is, "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of defining the requirements." The specialized AI generates a detailed answer: "Initial research and gathering stakeholder opinions are important for defining requirements."
[1423] Generating and transmitting audio data
[1424] The generated detailed answer is converted into audio data using the Google Cloud Text-to-Speech API, and the converted audio data is sent from the server to the device.
[1425] Audio output and animation playback to the user
[1426] The device receives the audio data and triggers MetaHuman's animation engine to animate the virtual character in sync with the audio playback, with the character's mouth movements and facial expressions matching the audio.
[1427] The above is an embodiment of the present invention. This system allows users to have natural visual and auditory interactions, and provides an advanced interactive system that can also respond to specialized questions.
[1428] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1429] Step 1:
[1430] Voice input from the user
[1431] User: The user speaks questions and commands into the microphone.
[1432] Input: User utterance (e.g., "I need some advice on defining requirements for a new project").
[1433] Output: Audio data
[1434] What it does: When a user speaks, the microphone captures the sound.
[1435] Step 2:
[1436] Capture and transmit audio data
[1437] Device: The device transmits the captured audio data to the server in real time.
[1438] Input: Captured audio data
[1439] Output: Audio data sent to the server
[1440] Specific operation: Audio captured by the microphone is saved in PCM format or similar and sent to the server as an HTTP request.
[1441] Step 3:
[1442] Converting audio data to text
[1443] Server: The server receives the voice data and passes it to the voice recognition module.
[1444] Input: Audio data sent from the device
[1445] Output: Text data
[1446] What it does: The received voice data is sent to the Google Cloud Speech-to-Text API, which converts the voice data into text. The resulting text is, "I'd like some advice on defining the requirements for a new project."
[1447] Step 4:
[1448] General answer generation by first-level generative AI
[1449] Server: The server analyzes the text data and inputs it into the generative AI.
[1450] Input: Text data (e.g., "I would like some advice on defining requirements for a new project.")
[1451] Output: General answer (text)
[1452] Specific operation: The server sends a prompt to OpenAI's GPT-3 or similar (e.g., "The user is looking for advice on the project requirements definition. Please provide general guidance."). The generative AI generates a general answer, such as "You're talking about the project requirements definition. Please be more specific," and replies to the server.
[1453] Step 5:
[1454] Detailed answers generated by specialized AI
[1455] Server: The server analyzes the general answers from the first-level generative AI and sends them to the specialized AI.
[1456] Input: General answer (text)
[1457] Output: Detailed answer (text)
[1458] Specific operation: The server sends a prompt to the specialized AI (e.g., "The user is looking for advice on defining the requirements for the project. Please explain the detailed steps of the requirements definition."). The specialized AI generates a detailed answer, such as "Initial research and gathering stakeholder opinions are important for defining requirements," and replies to the server.
[1459] Step 6:
[1460] Generating and transmitting audio data
[1461] Server: The server passes the detailed answer to the speech synthesis module, converts it into voice data, and sends it to the terminal.
[1462] Input: Detailed answer (text)
[1463] Output: Audio data (e.g., WAV or MP3 format)
[1464] Specific operation: The server sends the detailed answer to the Google Cloud Text-to-Speech API, converts the text "Initial research and stakeholder opinion gathering are important for requirements definition" into audio data, and sends the generated audio file to the device.
[1465] Step 7:
[1466] Audio output and animation playback to the user
[1467] Device: The device receives the audio data and triggers MetaHuman's animation engine to synchronize with the audio playback.
[1468] Input: Audio data
[1469] Output: Audio and visual feedback
[1470] How it works: The device plays back the received audio data, and at the same time, MetaHuman's virtual character moves its mouth and changes its facial expression in response to the audio, providing the user with a natural conversational experience.
[1471] (Application example 1)
[1472] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1473] In virtual stores, a system that integrates voice input and visual feedback is necessary to enable more natural and efficient user interaction. However, current technology simply converts the user's voice input into text and provides text-based answers, failing to realize natural dialogue that integrates visual and auditory feedback. Furthermore, there is a lack of effective means for generating detailed answers for each specialized field. This results in low user convenience and poses challenges for improving user satisfaction.
[1474] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1475] In this invention, the server includes means for converting voice input into text, means for analyzing the converted text and generating a general answer, means for generating detailed answers for each specialized field based on the generated general answers, means for converting the generated detailed answers into voice data, means for outputting the voice data to the user, means for providing visual output using a virtual character, and means for providing animated visual feedback synchronized with the voice data, thereby enabling natural dialogue that integrates voice input and visual feedback.
[1476] "Voice input" refers to words or questions spoken by the user, and is the voice data that the system uses to recognize them.
[1477] A "means for converting to text" is a technique or device for analyzing received voice input and converting the content into text data format.
[1478] "First-stage generation means" refers to a first-stage technique or device for generating a general answer based on the converted text.
[1479] The "generating means for each specialized field" is a technology or device that further analyzes the answer obtained by the generating means at the first stage and generates a detailed answer that is specific to a specific specialized field.
[1480] The "means for converting into voice data" refers to a technique or device for synthesizing the generated text-format answers into voice and outputting them as voice data.
[1481] "Visual feedback with animation synchronized with audio data" refers to a technology or device that displays visual actions or animations corresponding to generated audio data and provides them to the user along with the audio.
[1482] A "virtual character" is a visual character created using computer graphics and animation techniques to interact with a user.
[1483] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice form again. This system is particularly effective for realizing natural dialogue in virtual stores, and also provides visual feedback using virtual characters.
[1484] First, the user uses the smart glasses to input voice. The microphone in the smart glasses captures the voice and sends it to the server. The server then converts the received voice data into text using Google Cloud Speech-to-Text. This converted text is then input into GPT-4, the first stage of the generator, to generate a general answer.
[1485] The generated general answer is then passed to a specialized generator. In this example, a generator specialized for product information is used to generate a detailed answer. For example, if a user asks, "Tell me about this product," GPT-4 generates a general answer such as, "What category does this product belong to?" Then, a specialized AI generates a detailed answer such as, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[1486] The detailed answer is then converted into speech using Amazon Polly, which is then sent back to the server and played back to the smart glasses, where it uses Unreal Engine's MetaHuman to provide animated visual feedback synchronized with the speech.
[1487] As a concrete example, the following prompt sentence is input to GPT-4 to generate a general answer:
[1488] A user asked the following question:
[1489] Question: "Tell me about this product"
[1490] Generate a general answer to the question.
[1491] You can then create prompts for specialized AI and get detailed answers.
[1492] The following question was asked to the product information specialized AI.
[1493] Ask: "What category does this product belong to?"
[1494] Generate a professional answer to this question.
[1495] In this way, the present invention efficiently and naturally processes voice input from the user, provides optimal answers in the virtual store, and also provides a sense of visual familiarity to the user by using virtual characters.
[1496] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1497] Step 1:
[1498] The user uses the smart glasses to input voice. Specifically, the user asks a question by voice, such as "Tell me about this product." The microphone in the smart glasses captures this voice and inputs it as voice data into the terminal. The terminal then sends this voice data to the server.
[1499] Step 2:
[1500] The server converts the received voice data into text using Google Cloud Speech-to-Text. The input is voice data, and the output is text data. Specifically, the voice recognition algorithm analyzes the voice waveform data and converts it into text such as "Tell me about this product."
[1501] Step 3:
[1502] The server then inputs the converted text data into GPT-4 to generate a general answer. The input is text data, and the output is the general answer text. Based on the prompt, GPT-4 generates a general answer such as, "What category does this product belong to?"
[1503] Step 4:
[1504] The generated general answer is input to a specialized AI on the server, which uses a generation method specialized for product information to generate a detailed answer. The general answer is the input, and the detailed answer text is obtained as the output. Specifically, the detailed answer generated is, "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[1505] Step 5:
[1506] The detailed answer text is converted to speech using Amazon Polly. We have the detailed text answer as input and speech as output. A speech synthesis algorithm analyzes the text and generates natural-sounding speech.
[1507] Step 6:
[1508] The generated voice data is sent from the server to the device (smart glasses). The device plays the received voice data. Specifically, a detailed answer is provided to the user through the audio speaker: "This product is a high-performance smartwatch with heart rate monitoring and GPS functions."
[1509] Step 7:
[1510] At the same time, the server uses Unreal Engine's MetaHuman to generate animated visual feedback corresponding to the generated voice data. The input is voice data, and the output is visual animation synchronized with the voice. Specifically, a virtual character visually presents the answer to the user, moving its mouth in sync with the voice.
[1511] This series of processing steps enables the user to experience natural and detailed interactions in the virtual store.
[1512] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1513] The present invention relates to a system that receives voice input from a user, converts it into text, generates answers using generative AI, and provides them to the user in voice form again. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. This system allows users to have natural visual and auditory interactions, making it an advanced dialogue system that can also handle specialized questions.
[1514] Specifically, the following operations are included:
[1515] 1. User voice input
[1516] User: Speaks a question or instruction. For example, "I'd like to ask about the contents of the contract."
[1517] Device: Captures audio data and sends it to the server.
[1518] 2. Converting voice data to text
[1519] Server: Receives voice data and passes it to the voice recognition module.
[1520] Speech recognition module: Converts voice input into text data. For example, text data such as "I would like to ask for your advice regarding the contents of the contract."
[1521] Server: Inputs text data into the generative AI.
[1522] 3. Emotional Recognition
[1523] Server: Passes the voice input to the emotion engine and analyzes the user's emotions. For example, it analyzes the voice tone, rate, pitch, and volume to recognize that the user is feeling anxious.
[1524] Emotion engine: Recognizes the user's emotions and passes that information to the generative AI.
[1525] 4. Generating a general answer
[1526] Generative AI (first level): Generates general answers based on text data and recognized emotions. For example, it generates answers such as, "What specifically are you discussing about the contract?"
[1527] Server: Inputs the output of the first-stage generative AI into the specialized AI.
[1528] 5. Generating detailed answers
[1529] Server: Analyzes the answers obtained by the first-stage generation method and, if necessary, inputs them into a generative AI specialized for each field of expertise.
[1530] Specialized AI: Generates detailed answers based on specialized fields such as law, finance, and medicine. For example, it generates detailed answers such as, "Please tell me about the specific clauses in the contract."
[1531] Server: Passes the output from the specialized AI to the speech synthesis module.
[1532] 6. Generating Audio Data
[1533] Speech synthesis module: Converts text responses into natural-sounding speech, such as "Please tell me about the specific clauses in the contract."
[1534] Server: Sends the generated audio data to the device.
[1535] 7. Audio output to the user
[1536] Terminal: Interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[1537] MetaHuman: Provides visual and audible voice responses to the user, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[1538] Specific examples
[1539] Example 1: Legal advice
[1540] User: Say, "I need help with the contract."
[1541] Device: Captures audio and sends it to the server.
[1542] Server: The speech recognition module converts "I would like to consult you about the contents of the contract" into text.
[1543] Emotion engine: Recognizes that the user's tone of voice indicates anxiety.
[1544] Generative AI (first level): Generates a general answer such as, "You seem anxious. What specifically would you like to discuss regarding the contract?"
[1545] Specialized AI (legal): Generates detailed answers to questions such as, "Please tell me the specific clauses in the contract."
[1546] Speech synthesis module: Converts "Please tell me about the specific clauses in the contract" into audio data.
[1547] Device: The voice data is received and MetaHuman plays it back in a gentle tone as if it were a natural conversation.
[1548] Example 2: Engineering Support
[1549] User: "I need some advice on defining requirements for my new project."
[1550] Device: Captures audio and sends it to the server.
[1551] Server: The speech recognition module converts the phrase "I would like some advice on defining the requirements for a new project" into text.
[1552] Emotion engine: Recognizes that the user's tone of voice is calm.
[1553] Generative AI (first level): Generates a general answer such as, "You're talking about the project requirements definition. Please tell me more specifically."
[1554] Specialized AI (Engineering): Generates a detailed answer that says, "Initial research and gathering stakeholder opinions are important for requirements definition."
[1555] Speech synthesis module: Converts the statement "Initial research and stakeholder feedback are important for requirements definition" into speech data.
[1556] Device: The voice data is received and MetaHuman plays it back in a calm, natural-sounding conversation.
[1557] In this way, the present invention processes user voice input efficiently and naturally, providing optimal answers, and by using an emotion engine, it realizes flexible responses that match the user's emotions, providing a more personalized experience.
[1558] The processing flow will be explained below.
[1559] Step 1:
[1560] The user speaks a question. For example, they may say, "I have a question about the contents of the contract."
[1561] Step 2:
[1562] The device's microphone captures the user's voice, and the voice data is sent to the server.
[1563] Step 3:
[1564] The server receives the voice data and passes it to a voice recognition module.
[1565] Step 4:
[1566] The speech recognition module converts the speech data into text, such as "I would like to ask for your advice regarding the contents of the contract."
[1567] Step 5:
[1568] The server receives the converted text data and passes it to the emotion engine.
[1569] Step 6:
[1570] The emotion engine analyzes text and voice data to recognize the user's emotions. For example, it analyzes voice tone, speed, pitch, and volume to determine if the user is feeling anxious.
[1571] Step 7:
[1572] The server inputs the recognized emotional information into the generative AI (first-stage base model).
[1573] Step 8:
[1574] The generative AI (first stage) generates a general answer based on the input text and emotional information. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[1575] Step 9:
[1576] The server analyzes the output of the first-stage generative AI to input it into the specialized AI, identifying the user's question as legal-related.
[1577] Step 10:
[1578] The server selects a specialized field AI (legal specialized model) and inputs the answer from the first-stage generative AI.
[1579] Step 11:
[1580] A specialized AI (legal) generates detailed answers based on the input text and emotional information, such as "Please tell me about the specific clauses in the contract."
[1581] Step 12:
[1582] The server receives the output text from the specialized AI and passes it to the speech synthesis module.
[1583] Step 13:
[1584] The speech synthesis module converts the text data into speech data, for example, generating speech data such as "Please tell me about the specific clauses of the contract."
[1585] Step 14:
[1586] The server transmits the generated voice data to the terminal.
[1587] Step 15:
[1588] The device interprets the audio data received from the server and triggers MetaHuman's animation engine, synchronizing it with the audio playback.
[1589] Step 16:
[1590] MetaHuman provides the user with spoken responses, along with facial expressions and gestures, such as saying in a gentle tone, "Tell me about the specific clauses in the contract."
[1591] In this way, the user, device, and server work together at each step to realize advanced voice dialogue. In addition, by combining it with an emotion engine, it is possible to provide flexible responses that correspond to the user's emotions.
[1592] Example 2
[1593] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1594] While conventional voice dialogue systems can provide appropriate answers to user voice inputs, they have difficulty generating flexible responses that reflect the user's emotions. Furthermore, for specialized questions, they can only provide general answers, failing to provide the detailed information actually required. Furthermore, there is a need for the answers output from the system to be natural and for the dialogue with the user to be visually acceptable.
[1595] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1596] In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer based on the text by analyzing the converted text, means for generating a detailed answer based on a field of expertise from the general answer generated by the first-stage answer generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for recognizing the user's emotion from the voice input, and means for adjusting the generated answer based on the recognized emotion. This makes it possible to provide a flexible and detailed answer that reflects the user's emotion and realize a visually acceptable and natural dialogue.
[1597] The "means for receiving voice input from the user" is a device or function for capturing voice uttered by the user and incorporating it into the system.
[1598] A "means for converting voice input to text" is a device or function that analyzes received voice data and converts the content into text data.
[1599] The "first-stage generator" is an initial generator or algorithm that generates a general answer based on the converted text data.
[1600] The "means for generating a detailed answer based on a field of expertise" is a device or algorithm that generates a detailed answer specialized in a particular field of expertise based on the general answer generated by the first-stage generation means.
[1601] The "means for converting the generated detailed answer into voice data" is a device or function that converts the detailed answer in text format into natural voice data.
[1602] The "means for outputting voice data to the user" refers to a device or function that reproduces and provides the generated voice data to the user.
[1603] A "means for recognizing a user's emotion from speech input" is a device or algorithm that analyzes a user's speech data and determines their emotional state.
[1604] A "means for adjusting the generated answer based on the recognized emotion" is a device or algorithm that appropriately modifies or adjusts the generated answer depending on the user's emotional state.
[1605] The present invention is a system that receives voice input from a user, converts it into text, uses generative AI to generate answers, and provides them to the user in voice format. It also has the ability to recognize the user's emotions and adjust the answers based on those emotions. Specific implementation methods of the system are described below.
[1606] This system operates primarily using the following hardware and software:
[1607] Hardware: A device equipped with a microphone to capture the user's voice and a speaker to play the output audio data.
[1608] Server: A central server for processing data and performing necessary calculations.
[1609] Software: Speech recognition module, emotion recognition engine, generative AI, specialized AI, voice synthesis module, virtual character animation engine.
[1610] The main processing flow of this system is as follows:
[1611] The user speaks a voice input, which is captured by the device's microphone. Specifically, the user speaks something like "I would like to consult you about the contents of the contract." The device sends the captured voice data to the server. The server uses a voice recognition module (e.g., a voice recognition API) to convert the voice data into text data. For example, the converted text might be something like "I would like to consult you about the contents of the contract."
[1612] The server then passes the text and voice data to an emotion engine to analyze the user's emotions. The emotion engine analyzes the tone, speed, pitch, volume, etc. of the voice to recognize the emotion the user is feeling. Specifically, it may recognize that the user is feeling "anxiety."
[1613] The server inputs the recognized emotions and text data into a generative AI system to generate a general answer. For example, it might generate an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?" A general-purpose generative model is used for this generative AI model.
[1614] The server then passes the general answer to a specialized AI that generates a detailed answer. The specialized AI generates an answer based on knowledge of a specialized field, such as law, finance, or medicine. For example, it might generate a detailed answer such as, "Please tell me about the specific clauses in the contract."
[1615] The generated detailed answer is passed by the server to a speech synthesis module, which converts the text into natural-sounding speech data (e.g., speech synthesis API). The generated speech data will be something like, "Please tell me about the specific clauses of the contract."
[1616] The server sends the final voice data to the terminal, which triggers the virtual character's animation engine to synchronize with the voice playback. The virtual character speaks in a gentle tone, saying, "Please tell me about the specific clauses in the contract."
[1617] Specific examples
[1618] Example 1: Legal advice
[1619] The user says, "I'd like to ask you about the contents of the contract." The device captures the voice and sends it to the server. The server uses a speech recognition module to convert it into text, "I'd like to ask you about the contents of the contract," and uses an emotion engine to recognize the user's anxiety. The generative AI then generates a general answer, "You seem anxious. What specifically do you want to ask you about in the contract?", and the specialized AI generates a detailed answer, "Please tell me about the specific clauses in the contract." The speech synthesis module converts this into voice data, which the device finally provides to the user through a virtual character.
[1620] Prompt Sentence Examples
[1621] "Generate answers that will alleviate the user's concerns when asked about the contents of the contract."
[1622] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1623] Step 1:
[1624] The user speaks the speech input.
[1625] Specific operation: The user says, "I would like to discuss the contents of the contract."
[1626] Input: User's voice.
[1627] Output: The captured audio data.
[1628] Step 2:
[1629] The device captures the audio data and sends it to the server.
[1630] Specific operation: Records audio data using the device's microphone and sends the data to the server.
[1631] Input: The captured audio data.
[1632] Output: The audio data sent to the server.
[1633] Step 3:
[1634] The server receives the voice data and passes it to the voice recognition module.
[1635] Specific operation: The server receives voice data from the terminal and inputs the data into the voice recognition module.
[1636] Input: The audio data sent to the server.
[1637] Output: Audio data as input to the speech recognition module.
[1638] Step 4:
[1639] A voice recognition module converts the voice data into text data.
[1640] Specific operation: The voice recognition module analyzes the voice data and generates text data such as "I would like to consult you about the contents of the contract."
[1641] Input: Audio data as input to the speech recognition module.
[1642] Output: The converted text data.
[1643] Step 5:
[1644] The server passes the text data to an emotion recognition engine to analyze the user's emotions.
[1645] Specific operation: The server passes the text data and voice characteristic information to the emotion recognition engine and begins analysis.
[1646] Input: Translated text data and speech characteristics information.
[1647] Output: Parsed emotion data (e.g., anxiety).
[1648] Step 6:
[1649] The server inputs emotional data and text data into a generative AI to generate a general answer.
[1650] Specific operation: The server inputs emotion data and text data into the generative AI and receives the generated answer. For example, it generates an answer such as, "You seem anxious. What specifically do you want to discuss about the contract?"
[1651] Input: Parsed emotion data and converted text data.
[1652] Output: The generated general answer text.
[1653] Step 7:
[1654] The server passes the general answer to a specialized AI, which generates a detailed answer.
[1655] Specific operation: The server passes the general answer text to a specialized AI (e.g., specialized AI for law, finance, medicine, etc.) to generate a detailed answer. For example, it generates a detailed answer such as "Please tell me about the specific clauses in the contract."
[1656] Input: The generated general answer text.
[1657] Output: The generated detailed answer text.
[1658] Step 8:
[1659] The server passes the detailed answer to a speech synthesis module and converts it into speech data.
[1660] Specific operation: The server inputs the detailed answer text into the speech synthesis module to generate speech data, for example, "Please tell me about the specific clauses of the contract."
[1661] Input: The generated long answer text.
[1662] Output: The generated audio data.
[1663] Step 9:
[1664] The server transmits the generated voice data to the terminal.
[1665] Specific operation: The server sends the generated audio data to the device.
[1666] Input: The generated audio data.
[1667] Output: The audio data sent to the device.
[1668] Step 10:
[1669] The device receives the audio data and triggers the animation engine of the virtual character.
[1670] Specific operation: The device interprets the received audio data and triggers the virtual character's animation engine to create a visual representation synchronized with the audio playback.
[1671] Input: Audio data sent to the device.
[1672] Output: Audio and visual answers provided to the user.
[1673] (Application example 2)
[1674] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1675] Conventional food delivery systems have had problems with the user having to make an order, requiring a lot of effort and not being able to respond flexibly to emotions. This can lead to low user satisfaction and delays in the ordering process. Furthermore, advanced technology is required to ensure natural visual and auditory interactions, and there has been a lack of systems that can achieve this.
[1676] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice input from a user, means for converting the voice input into text, means for generating a first-stage answer by analyzing the text, means for generating a detailed answer based on the field of expertise from the general answer generated by the first-stage generation means, means for converting the generated detailed answer into voice data, means for outputting the voice data to the user, means for analyzing the user's emotions, and means for adjusting the answer based on the results of the emotion analysis. This allows the user to place an order through natural dialogue and enables flexible responses according to emotions.
[1677] A "means for receiving voice input from a user" is a device or function for capturing voice data uttered by a user.
[1678] A "means for converting voice input to text" is a device or function that analyzes captured voice data and converts it into corresponding text data.
[1679] "An initial generation means for analyzing the converted text and generating a general answer based on the text" is a device or function that includes an AI module for generating an initial answer based on text data.
[1680] "Means for generating detailed answers based on specialized fields from the general answers generated by the initial generation means" refers to a device or function that includes an AI module for further specialized analysis of the output of the initial generation means and generating detailed answers.
[1681] The "means for converting the generated detailed answer into voice data" is a device or function for converting the generated detailed answer as text into voice data.
[1682] The "means for outputting audio data to the user" refers to a device or function for playing back the generated audio data to the user.
[1683] A "means for analyzing user emotions" is a device or function that analyzes a user's voice input to identify their emotional state.
[1684] The "means for adjusting the answer based on the result of sentiment analysis" is a device or function for taking into account the result of sentiment analysis and adjusting the answer to be generated and the method of presenting it.
[1685] The present invention is a system for efficiently delivering food via voice, allowing users to place orders through natural dialogue and providing flexible responses according to emotions. This system includes the following specific procedures and devices.
[1686] System configuration
[1687] The system takes voice input from the user, converts it to text, uses generative AI to generate a general answer, then generates a more specific answer based on the user's area of expertise (in this case, food delivery), converts the answer back into speech, and delivers it to the user. It can also recognize the user's emotions and adjust the answer accordingly.
[1688] Hardware and software used
[1689] Hardware:
[1690] Smartphone microphone: A device for capturing voice data spoken by a user.
[1691] Smartphone speaker: A device for playing back the generated audio data to the user.
[1692] software:
[1693] speech_recognition: A library for converting speech data into text data.
[1694] Transformers sentiment-analysis: A library for analyzing user sentiment from text data.
[1695] googletrans: A library for translating and converting answers into text.
[1696] gTTS (Google Text-To-Speech): A library for converting text data into audio data and playing it back.
[1697] System Operation
[1698] The server first receives voice input from the user. This voice is captured using a smartphone microphone. It then uses speech recognition software to convert the speech to text. This converted text is passed to a first-stage generative AI to generate a general answer. This general answer is then further processed into a detailed answer by a specialized AI in a specialized field (food delivery).
[1699] Emotion analysis
[1700] At the same time, the user's emotions are also analyzed. This is done using an emotion analysis engine that analyzes the user's voice tone, speed, pitch, and volume. Based on the analysis results, the answer is adjusted. For example, if the user is in a hurry, a quick response is required.
[1701] Speech synthesis and output
[1702] The detailed answer is then converted into audio data by speech synthesis software, which is then delivered to the user through the smartphone speaker, allowing them to order food delivery in a natural, conversational way.
[1703] Specific examples
[1704] Customer: "One pizza please."
[1705] application:
[1706] (Transcription): "One pizza please."
[1707] (emotion recognition): calm
[1708] (Order Processing): "Confirmed. We will process your order immediately."
[1709] (Text-to-Speech): Play "Thank you. We'll get it sorted right away."
[1710] Prompt Sentence Examples
[1711] Prompt: "A customer calmly orders, 'One pizza, please.' Generate an appropriate response for the calm customer."
[1712] Expected response: "Thank you. I'll get it sorted right away."
[1713] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1714] Step 1:
[1715] The server receives the user's voice input from the device. This input is the voice data spoken by the user using the smartphone microphone. The device then sends the captured voice data to the server. This data is transmitted as a voice waveform and is not in a format that can be analyzed as is.
[1716] Step 2:
[1717] The server converts the speech input into text. This is done using speech recognition software, which analyzes the speech waveform data and converts the phoneme sequence into the corresponding text. For example, if the speech input is "One pizza please," it will be converted into text "One pizza please." The output of this step is the converted text data.
[1718] Step 3:
[1719] The server analyzes the converted text and passes it to a first-stage generative AI that generates a general answer. This first-stage generative AI applies a generative AI model based on the text data to generate a general answer. For example, if the input text is "One pizza please," it will generate a general answer of "I understand. We'll get it ready right away." The output of this step is the general answer text.
[1720] Step 4:
[1721] The server passes the answer text obtained from the first-dan generation AI to a specialized AI, which generates a detailed answer. This specialized AI has a knowledge database specialized in a specific field (in this case, food delivery) and complements the first-dan answer in detail. For example, it adds a specific confirmation such as "Is this pizza size okay?" to the general answer "Please confirm." The output of this step is a detailed answer text.
[1722] Step 5:
[1723] The server receives the detailed answer text and analyzes the user's emotions. It uses an emotion analysis engine to identify the user's emotional state (e.g., hurry, calm, happy) by analyzing the tone, rate, pitch, and volume of the speech input. It adjusts the detailed answer based on this analysis data. The output of this step is the final adjusted detailed answer text.
[1724] Step 6:
[1725] The server converts the finalized detailed answer text into speech data. It uses speech synthesis software such as Google Text-To-Speech (gTTS) to convert the text data into speech data. This speech data is provided with natural pronunciation that is easy for the user to understand. The output of this step is synthesized speech data.
[1726] Step 7:
[1727] The server returns the generated voice data to the terminal, which then provides it to the user through the smartphone speaker, allowing the user to receive the voice answer. The final output of this step is the voice answer provided to the user.
[1728] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1729] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1730] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1731] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1732] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1733] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1734] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1735] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1736] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1737] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1738] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1739] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1740] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1741] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1742] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1743] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1744] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1745] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1746] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1747] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1748] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1749] The following is further disclosed regarding the above embodiment.
[1750] (Claim 1)
[1751] means for receiving voice input from a user;
[1752] a means for converting voice input into text;
[1753] a first-stage generating means for analyzing the converted text and generating a general answer based on the text;
[1754] a means for generating a detailed answer based on a specialized field from the general answer generated by the first stage generating means;
[1755] A means for converting the generated detailed answer into audio data;
[1756] means for outputting audio data to a user;
[1757] A system including:
[1758] (Claim 2)
[1759] 2. The system according to claim 1, wherein the means for generating a detailed answer based on a field of expertise selectively applies a generation means for each of a plurality of fields of expertise.
[1760] (Claim 3)
[1761] 10. The system of claim 1, wherein the means for providing an output to the user provides a visual output using a virtual character.
[1762] "Example 1"
[1763] (Claim 1)
[1764] means for receiving voice input from a user;
[1765] a means for converting voice input into text;
[1766] a generating means for analyzing the converted text and generating a general answer based on the text;
[1767] A means for generating a detailed answer based on a specialized field from the general answer generated by the generating means;
[1768] A means for converting the generated detailed answer into audio data;
[1769] means for outputting audio data to a user;
[1770] A system including:
[1771] (Claim 2)
[1772] 2. The system according to claim 1, wherein the means for generating a detailed answer based on a field of expertise selectively applies a generation means for each of a plurality of fields of expertise.
[1773] (Claim 3)
[1774] 10. The system of claim 1, wherein the means for providing an output to the user provides a visual output using a virtual character.
[1775] "Application Example 1"
[1776] (Claim 1)
[1777] means for receiving voice input from a user;
[1778] a means for converting voice input into text;
[1779] a first-stage generating means for analyzing the converted text and generating a general answer based on the text;
[1780] a means for generating a detailed answer based on a specialized field from the general answer generated by the first stage generating means;
[1781] A means for converting the generated detailed answer into audio data;
[1782] means for outputting audio data to a user;
[1783] means for producing a visual output using a virtual character;
[1784] A system including:
[1785] (Claim 2)
[1786] 2. The system according to claim 1, wherein the means for generating a detailed answer based on a field of expertise selectively applies a generation means for each of a plurality of fields of expertise.
[1787] (Claim 3)
[1788] 10. The system of claim 1, wherein the means for outputting to the user provides animated visual feedback synchronized with the audio data.
[1789] "Example 2: Combining Emotion Engines"
[1790] (Claim 1)
[1791] means for receiving voice input from a user;
[1792] a means for converting voice input into text;
[1793] a first-stage generating means for analyzing the converted text and generating a general answer based on the text;
[1794] a means for generating a detailed answer based on a specialized field from the general answer generated by the first stage generating means;
[1795] A means for converting the generated detailed answer into audio data;
[1796] means for outputting audio data to a user;
[1797] means for recognizing a user's emotion from a speech input;
[1798] means for adjusting the generated answers based on the perceived emotions;
[1799] A system including:
[1800] (Claim 2)
[1801] 2. The system according to claim 1, wherein the means for generating a detailed answer based on a field of expertise selectively applies a generation means for each of a plurality of fields of expertise.
[1802] (Claim 3)
[1803] 10. The system of claim 1, wherein the means for providing an output to the user provides a visual output using a virtual character.
[1804] "Application example 2 when combining emotion engines"
[1805] (Claim 1)
[1806] means for receiving voice input from a user;
[1807] a means for converting voice input into text;
[1808] a first-stage generating means for analyzing the converted text and generating a general answer based on the text;
[1809] a means for generating a detailed answer based on a specialized field from the general answer generated by the first stage generating means;
[1810] A means for converting the generated detailed answer into audio data;
[1811] means for outputting audio data to a user;
[1812] means for analyzing user emotions;
[1813] means for adjusting the answer based on the results of the sentiment analysis;
[1814] A system including:
[1815] (Claim 2)
[1816] 2. The system according to claim 1, wherein the means for generating a detailed answer based on a field of expertise selectively applies a generation means for each of a plurality of fields of expertise.
[1817] (Claim 3)
[1818] 2. The system according to claim 1, wherein the means for outputting to the user provides visual output using a virtual character. [Explanation of symbols]
[1819] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving voice input from a user; a means for converting voice input into text; a first-stage generating means for analyzing the converted text and generating a general answer based on the text; A means for generating a detailed answer based on a specialized field from the general answer generated by the first stage generating means; A means for converting the generated detailed answer into audio data; means for outputting audio data to a user; A system including:
2. 2. The system according to claim 1, wherein the means for generating a detailed answer based on a field of expertise selectively applies a generation means for each of a plurality of fields of expertise.
3. 10. The system of claim 1, wherein the means for providing an output to the user provides a visual output using a virtual character.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A