System
The system addresses limited conversational capabilities in communication robots by using a generative model and speech recognition with cloud technology for enhanced dialogue and educational support, achieving natural and effective interactions.
Patent Information
- Application Number
- JP2024121576
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-05
AI Technical Summary
Current communication robots have limited conversational capabilities, especially in educational and medical settings, and the technology for converting user speech into text data and generating appropriate responses is immature, limiting the effectiveness of personalized educational support and conversational assistance.
A system incorporating a generative model for dialogue response generation, speech recognition, and cloud technology to enhance communication capabilities, providing individual educational support and effective dialogue in various fields.
The system enables natural and effective interactions by converting user voice to text, generating appropriate responses, and using cloud technology for advanced dialogue capabilities, improving user experience in diverse scenarios.
Smart Images

Figure 2026019828000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current communication robots have limited conversational capabilities, making it difficult to achieve natural conversations with users, especially in educational, medical, and nursing care settings. Furthermore, technology for converting user speech into text data and using that data to generate appropriate responses is still immature, creating a need for improved user experience. These technological limitations limit the effectiveness of personalized educational support and conversational assistance in medical and nursing care settings. [Means for solving the problem]
[0005] The present invention provides a system including a dialogue response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting the user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Furthermore, by adding a means for communication between the generative model and the user using cloud technology to this system, more advanced dialogue capabilities can be achieved. Furthermore, by including a means for providing individual educational support in the education field and a means for supporting dialogue with patients or elderly people in the medical and nursing care fields, the system can be applied in a wide range of fields and can improve the quality of the user experience.
[0006] A "generative model" is an artificial intelligence model that uses natural language processing to automatically generate appropriate responses from input data.
[0007] "Dialogue response generation means" refers to a function for generating an appropriate response to input data from a user using a generative model.
[0008] "Output means" refers to functions such as audio output and display for providing the response generated by the generative model to the user.
[0009] "Speech recognition means" refers to the technology or software used to convert a user's voice into text data.
[0010] "Transmission means" refers to a communication function for transmitting text data converted by the speech recognition means to the generative model.
[0011] "Cloud technology" refers to technology that uses remote servers and computer resources via the Internet.
[0012] "Educational support" refers to functions that provide individual instruction and assistance to students in the field of education.
[0013] "Dialogue support" refers to functions that enable natural and effective dialogue with patients and elderly people in the medical and nursing care fields. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] Program Overview
[0036] This invention realizes a system including a dialogue response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Furthermore, communication between the generative model and the user is performed using cloud technology.
[0037] Explanation of program processing
[0038] 1. Collect user input:
[0039] User: Talk to the communication robot, for example, "What's the weather like today?"
[0040] Device: Uses a microphone to capture the user's voice.
[0041] Device: Uses speech recognition software to convert speech into text data.
[0042] 2. Sending user input:
[0043] Terminal: Organize text data into a suitable format (e.g. JSON).
[0044] Terminal: Sends text data to the server via an HTTP request.
[0045] 3. Generating dialogue content using a generative model:
[0046] Server: Receives and parses the text data.
[0047] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[0048] Generative Model: Receives requests and generates appropriate responses.
[0049] Generative model: Generates the answer and sends it back to the server as a response.
[0050] 4. Sending and processing the response:
[0051] Server: Receives responses from the generative AI model.
[0052] Server: Sends the received response to the terminal as an HTTP response.
[0053] 5. Response to the user:
[0054] Terminal: Parses the response received from the server.
[0055] Terminal: Passes the parsed response data to speech synthesis software to convert text to speech.
[0056] Terminal: Plays the generated audio using an audio output device (speaker).
[0057] Terminal: If necessary, the response will be displayed as text on the display.
[0058] Specific examples
[0059] For the education sector:
[0060] 1. User (Student): "I don't understand this math problem. How do I solve it?"
[0061] Terminal: Converts student voice into text data and sends it to the server.
[0062] Server: Sends text data to the generative model and generates an appropriate explanation.
[0063] Generative model: Generates explanations such as "First, try adding 2 to both sides of the equation. Then, as a next step..."
[0064] Server: Sends this response to the device.
[0065] Terminal: The robot outputs the explanation aloud and also displays it on the screen.
[0066] For the medical and nursing care sector:
[0067] 1. User (elderly): "I'm not feeling very well today."
[0068] Terminal: Converts the elderly person's voice into text data and sends it to the server.
[0069] Server: Feeds text data into the generative model and generates an appropriate response.
[0070] Generative model: Generate a response such as, "That's worrying. Let's take a few deep breaths together, and then it might be a good idea to talk to your doctor."
[0071] Server: Sends this response to the device.
[0072] Terminal: Have the robot output the response content by voice.
[0073] The system provides users with a smooth and personalized experience. By combining generative models and cloud technology, the communication robot achieves natural and effective interactions in a wide range of scenarios.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] The user speaks to the communication robot (e.g., "What's the weather like today?"). When voice input begins, the robot's terminal uses a microphone to capture the user's voice.
[0077] Step 2:
[0078] The device converts the captured audio into text data using built-in speech recognition software, which uses a speech analysis algorithm to analyze the characteristics of the voice and convert it into text.
[0079] Step 3:
[0080] The terminal organizes the converted text data into an appropriate format, such as JSON, which also includes the user's ID and session information.
[0081] Step 4:
[0082] The device sends the organized text data to the server via an HTTP request, which includes the text data along with user information and related context information.
[0083] Step 5:
[0084] The server parses the received text data, checks the text data, and performs any necessary preprocessing before sending it to the generative model.
[0085] Step 6:
[0086] The server sends the parsed text data as a request to the API endpoint of the generative model, which includes the text data entered by the user.
[0087] Step 7:
[0088] The generative model analyzes the incoming request and generates an appropriate response. The generation process uses natural language processing algorithms to generate an appropriate response based on the context.
[0089] Step 8:
[0090] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[0091] Step 9:
[0092] The server receives the response from the generative model and parses it again to ensure that the response is well formed.
[0093] Step 10:
[0094] The server sends the received response data to the terminal as an HTTP response, which includes the generated response text data.
[0095] Step 11:
[0096] The device parses the response data received from the server and converts the text data into a format suitable for voice output or display.
[0097] Step 12:
[0098] The device passes the parsed response data to speech synthesis software, which converts the text into speech. The speech synthesis process converts the text data into speech waveforms.
[0099] Step 13:
[0100] The terminal plays the generated voice using the audio output device (speaker), and simultaneously displays the response text on the display if necessary.
[0101] This series of processing steps enables the user to have a natural and effective dialogue with the communication robot.
[0102] Example 1
[0103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0104] Conventional interactive response systems have had difficulty efficiently collecting user voice input, converting it into appropriate text data, and generating responses. Furthermore, there are limited means for providing users with generated responses in natural-sounding voices. Furthermore, there is a need to effectively integrate generative AI models with cloud technology to realize personalized educational support and other applications.
[0105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0106] In this invention, the server includes means for capturing voice input from a user, means for converting the voice input into text data, means for organizing the text data into an appropriate data format, means for sending the organized text data to a generative model, means for generating a dialogue response using the generative model, means for providing the generated dialogue response to the user, and means for converting the generated dialogue response into voice data, and includes output means for playing back the voice data. This makes it possible to generate smooth and natural dialogue responses using a generative AI model from the user's voice input and provide them in voice.
[0107] "User" refers to a person who uses the system to provide voice input.
[0108] "Voice input" refers to voice information spoken by a user to the system.
[0109] "Capturing means" refers to a device or software used to collect and record a user's voice input.
[0110] "Means for converting into text data" refers to software or algorithms used to analyze and convert collected voice input into corresponding text data.
[0111] "Means for organizing data" refers to the methods or mechanisms for preparing text data in an appropriate format, such as JSON, for transmission to a generative model.
[0112] "Means for sending" refers to a communication means for sending organized text data to a location where a generative model operates.
[0113] A "generative model" refers to an algorithm or software that generates appropriate dialogue responses based on text data.
[0114] "Means for generating a dialogue response" refers to a process of using a generative model to generate a response corresponding to an input.
[0115] "Means for providing" refers to a method for passing the generated interactive response to the user.
[0116] "Means for converting into voice data" refers to a process for converting the generated textual dialogue response into voice format.
[0117] "Output means" refers to a device or mechanism for playing back the converted audio data.
[0118] "Cloud technology" refers to computing resources and services provided over the internet.
[0119] "Educational support" refers to functions and services that provide individual learning support and guidance in the field of education.
[0120] The present invention is a dialogue response generation system using a generation model, and specifically includes the following means.
[0121] Collecting User Input
[0122] The user speaks to the communication robot. A specific example of a question is, "How's the weather today?" The device uses a microphone to capture the user's voice input. It then uses voice recognition software (e.g., a commonly used voice recognition engine) to convert the voice input into text data. For example, the voice "How's the weather today?" is converted into text "How's the weather today?"
[0123] Sending User Input
[0124] The device organizes the converted text data into an appropriate data format, such as JSON. The organized text data is sent to the server via an HTTP request. This request includes the user input, "What's the weather like today?"
[0125] Generating dialogue content using generative models
[0126] The server processes the HTTP request received from the device and extracts the text data. The extracted text data is sent as a request to the API endpoint of a generative model (for example, a commonly used generative AI model). The generative model receives the request and generates an appropriate response, such as "It's sunny today." This generated response is then sent back to the server.
[0127] Sending and Processing the Response
[0128] The server receives the response from the generative model and sends it to the device as an HTTP response. The received response may be data in the form of, for example, "It's sunny today."
[0129] Responding to the user
[0130] The device analyzes the HTTP response received from the server. It uses speech synthesis software (for example, a commonly used speech synthesis engine) to convert the text data into voice data. It then plays the generated voice using a speaker and responds to the user by saying, "It's a sunny day today." It also displays the response text on the display as needed.
[0131] Specific examples
[0132] Specific examples in the field of education
[0133] User (Student): "I don't understand this math problem. How do I solve it?"
[0134] The terminal converts the student's voice into text data and sends it to the server.
[0135] The server feeds the text data into a generative model to generate an appropriate explanation.
[0136] The generative model generates an explanation such as, "First, try adding 2 to both sides of the equation. Then, as a next step..."
[0137] The server sends this response to the terminal.
[0138] The device will then audibly transmit the generated explanation to the student and also display it on the screen.
[0139] Specific examples in the medical and nursing care fields
[0140] Elderly user: "I'm not feeling too great today."
[0141] The terminal converts the elderly person's voice into text data and sends it to the server.
[0142] The server feeds the text data into the generative model to generate an appropriate response.
[0143] The generative model generates a response such as, "That's worrying. Let's take a deep breath together. Then it might be a good idea to consult a doctor."
[0144] The server sends this response to the terminal.
[0145] The terminal then transmits the generated response to the elderly person by voice.
[0146] This invention leverages the collaboration between generative models and cloud technology to provide users with a smooth and personalized interaction experience. By combining generative AI models with speech recognition and speech synthesis technologies, natural-sounding interactions are possible in a wide range of scenarios.
[0147] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0148] Step 1:
[0149] A user speaks to a communication robot. For example, "How's the weather today?" The device uses a microphone to capture the user's voice input. The input voice information is acquired as an analog signal. The device converts this analog signal into a digital signal and passes it to the voice recognition software. The voice input is "How's the weather today?" and the output is a digital voice signal.
[0150] Step 2:
[0151] The device uses speech recognition software (e.g., a commonly used speech recognition engine) to convert speech into text data. During this process, a digital speech signal is analyzed and converted into corresponding text data. The input is a digital speech signal, and the output is text data such as "What's the weather like today?". Specifically, the speech signal is analyzed based on an acoustic model and a language model.
[0152] Step 3:
[0153] The terminal organizes the converted text data into JSON format. For example, the text data "What's the weather like today?" is converted into the format {"input": "What's the weather like today?"}. The input is text data, and the output is JSON format data. Specifically, the text data is structured with appropriate key-value pairs.
[0154] Step 4:
[0155] The terminal sends organized JSON-formatted text data to the server via an HTTP POST request. The input is JSON-formatted data, and the output is an HTTP request sent to the server. Specifically, an HTTP request containing a URL and header information is generated.
[0156] Step 5:
[0157] The server processes the HTTP request received from the terminal and extracts the text data. The input is the HTTP request, and the output is the extracted text data (for example, "What's the weather like today?"). Specifically, the required data is parsed from the request body.
[0158] Step 6:
[0159] The server sends the extracted text data as a request to the API endpoint of a generative model (e.g., a commonly used generative AI model). The input is the extracted text data, and the output is the request sent to the generative model. The specific operation is to call the appropriate API endpoint.
[0160] Step 7:
[0161] A generative AI model generates an appropriate response based on the request it receives. For example, the input "What's the weather like today?" generates the response "It's sunny today." The input is the request to the generative model, and the output is the generated response (e.g., "It's sunny today"). Specifically, the generative algorithm performs the text analysis and generation process.
[0162] Step 8:
[0163] The generative model returns the generated response to the server as a response. The input is the generated response, and the output is the response returned to the server. As a specific operation, the response data is returned in an appropriate format such as JSON format.
[0164] Step 9:
[0165] The server receives the response from the generative model and sends it to the terminal as an HTTP response. The input is the response from the generative model, and the output is the HTTP response sent to the terminal. In concrete terms, the response data is formatted as an HTTP response.
[0166] Step 10:
[0167] The terminal analyzes the HTTP response received from the server and obtains the text data. The input is the HTTP response, and the output is the text data (for example, "It's a sunny day today"). Specifically, the data is extracted from the response body.
[0168] Step 11:
[0169] The terminal passes the acquired text data to speech synthesis software (for example, a commonly used speech synthesis engine) and converts the text data into speech data. The input is text data and the output is speech data. Specifically, the speech synthesis engine synthesizes the text data to generate a speech signal.
[0170] Step 12:
[0171] The device uses a speaker to play back the generated voice. For example, it may output the voice "It's a sunny day today." The input is voice data, and the output is a voice response to the user. Specifically, the speaker converts the voice signal into physical sound, which the user hears. If necessary, the response is also displayed as text on the display.
[0172] (Application example 1)
[0173] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0174] Conventional food delivery systems have the drawback of requiring users to operate complicated applications when placing an order, and the convenience of voice control is particularly insufficient. Furthermore, there is a lack of systems that utilize advanced dialogue functions to provide optimal suggestions to users, making it difficult to achieve seamless dialogue with users. Therefore, there is a demand for a system that allows users to easily order food delivery using voice commands.
[0175] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0176] In this invention, the server includes an interactive response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, a transmission means for transmitting the text data to the generative model, a speech synthesis means for converting the generated response into a speech output, and an order processing means for processing a user's order. This allows a user to seamlessly order food delivery using only voice input, and the generative model makes optimal suggestions, making it possible to provide a more convenient service.
[0177] "Means for generating dialogue responses using a generative model" refers to a function that executes a generative AI model to generate appropriate responses based on input from a user.
[0178] "Output means" refers to a device or function for providing a response generated by a generative model to a user.
[0179] "Speech recognition means" refers to technology or devices for converting a user's voice input into text data.
[0180] "Transmission means" refers to a communication means for transmitting the text data to the generative model.
[0181] "Speech synthesis means" refers to technology or devices for converting text responses generated by a generative model into speech.
[0182] "Order Processing Means" refers to the systems and functions used to process user orders and facilitate the actual food delivery process.
[0183] "Cloud technology" refers to remote computing resources and services used over the internet to communicate between generative models and users.
[0184] The "food delivery sector" refers to industries and activities related to the delivery of food and beverages to specific locations upon request.
[0185] System Overview
[0186] The present invention provides a system that allows users to place food delivery orders by voice. The system converts the user's voice input into text data, sends the data to a generative model to generate a dialogue response, and provides the response to the user as voice using a speech synthesis means. The system also processes the user's order and uses cloud technology to perform these communications.
[0187] Hardware and software used
[0188] This system uses the following hardware and software:
[0189] Microphone: Used to capture the user's voice.
[0190] Speech Recognition Software: Uses the Google Speech-to-Text API to convert voice input into text data.
[0191] Generative model: Uses the OpenAI GPT-3 API to generate responses based on user text input.
[0192] Speech synthesis software: Uses the Google Text-to-Speech API to convert text responses from the generative model into audio.
[0193] Cloud server: Sends and receives data using cloud technologies such as AWS EC2.
[0194] Data format: Organize and send communication data in JSON format.
[0195] Processing flow
[0196] 1. Collecting user voice input
[0197] The user speaks to the smartphone, for example, "What's your recommended pizza today?"
[0198] The smartphone's microphone captures the user's voice.
[0199] Speech recognition software (Google Speech-to-Text API) converts the speech into text data.
[0200] 2. Sending text data
[0201] The smartphone organizes the text data into the appropriate format (JSON).
[0202] Send text data to the cloud server via an HTTP request.
[0203] 3. Generating dialogue content using a generative model
[0204] The cloud server receives the text data and sends it to the endpoint of the generative model (OpenAI GPT-3 API).
[0205] The generative model receives the request, generates an appropriate response, and sends the response back to the cloud server.
[0206] 4. Response to the user
[0207] The cloud server receives the response from the generative model and sends it to the smartphone as an HTTP response.
[0208] The response received by the smartphone is converted into speech using speech synthesis software (Google Text-to-Speech API) and played through the speaker.
[0209] 5. Order Processing
[0210] The user receives the response and responds audibly, "Yes, I'll order that."
[0211] The order is confirmed through a similar process and the food delivery process is carried out through the order processing means.
[0212] Examples of specific examples and prompts
[0213] Sample prompt 1: "What's your recommended pizza today?"
[0214] Sample prompt 2: "What dessert would you recommend?"
[0215] Prompt 3: "I'll order that."
[0216] Specifically, when a user speaks to their smartphone, "What's the recommended pizza today?", the speech recognition software converts the speech into text and sends it to a cloud server. Based on the text, the generative model generates a response, "Today's recommended pizza is Margherita. Would you like to order?", which is returned to the smartphone, converted into speech by speech synthesis software, and played back to the user. When the user replies, "I'll order," the food delivery procedure is completed by the order processing means.
[0217] In this way, a system is created that allows users to seamlessly order food delivery while engaging in natural interactions via their smartphones.
[0218] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0219] Step 1: Collecting voice input
[0220] User: Talks to their smartphone to order food delivery, for example, "What's your pizza special today?"
[0221] Device: The smartphone's microphone captures the user's voice, which is then fed into voice recognition software.
[0222] On the device: Uses the Google Speech-to-Text API to convert voice data to text data. The input is voice data and the output is text data.
[0223] Terminal: Receives the converted text data, which will be used in the next step.
[0224] Step 2: Send text data
[0225] Terminal: Organizes text data into JSON format. Input is text data, output is JSON format data.
[0226] The device sends the organized JSON data via an HTTP request to the cloud server, which then sends this data to the generative model.
[0227] Step 3: Generating dialogue content using a generative model
[0228] Server: The cloud server receives the text data and sends a request to the API endpoint of the generative model. The input is JSON data, and the output is a request to the generative model.
[0229] Generative Model: The OpenAI GPT-3 API receives requests and generates appropriate responses. The input is the request data and the output is the response data.
[0230] Server: The cloud server receives the generated response data and sends it to the terminal as an HTTP response. The input is the response data, and the output is the HTTP response.
[0231] Step 4: Respond to the user
[0232] Terminal: The smartphone receives the HTTP response from the cloud server. The input is the HTTP response, and the output is the response text data.
[0233] Terminal: Convert the received response data into speech using the Google Text-to-Speech API. The input is the response text data, and the output is the speech data.
[0234] Terminal: Plays voice responses to the user through a speaker and optionally displays the responses on a display. Input is voice data, and output is voice and text display.
[0235] Step 5: Order Processing
[0236] User: In response to the response, responds verbally with "Yes, I'll order that."
[0237] Terminal: The user's response is processed in a similar process (steps 1 to 4). Finally, the order processing means processes the food delivery. Specifically, it stores the order details in a database, notifies the delivery service, and provides the user with an order confirmation.
[0238] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0239] Program Overview
[0240] The present invention aims to make dialogue with a user more natural and effective by combining an emotion engine with a system equipped with a dialogue response generation means using a generative model. The system includes an output means for generating a response based on the generative model and providing the response to the user, a speech recognition means for converting the user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Cloud technology is also used for communication between the generative model and the user. Furthermore, the emotion engine is used to recognize the user's emotional state and adjust the response of the generative model based on the recognition result.
[0241] Explanation of program processing
[0242] 1. Collect user input:
[0243] User: Talks to the communication robot (e.g., "I'm not feeling very good today.").
[0244] Device: Uses a microphone and camera to capture the user's voice and facial expressions.
[0245] Device: Uses speech recognition software to convert speech into text data.
[0246] Device: The user's facial expressions are analyzed using facial expression analysis software from the camera footage, and emotional data is extracted.
[0247] 2. Emotional Data Processing:
[0248] Device: The extracted voice and facial expression data is passed to an emotion recognition engine to estimate the user's emotional state.
[0249] Emotion recognition engine: Analyzes voice characteristics such as tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[0250] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[0251] 3. Sending user input and emotion data:
[0252] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[0253] Terminal: Sends data to the server via an HTTP request.
[0254] 4. Generating dialogue content using a generative model:
[0255] Server: Receives and parses text data and emotion data.
[0256] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[0257] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[0258] Generative model: Generates the answer and sends it back to the server as a response.
[0259] 5. Sending and processing the response:
[0260] Server: Receives the response from the generative model.
[0261] Server: Sends the received response to the terminal as an HTTP response.
[0262] 6. Response to the user:
[0263] Terminal: Parses the response received from the server.
[0264] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[0265] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[0266] Specific examples
[0267] For the education sector:
[0268] Student User: "I don't understand this math problem," says in a confused tone.
[0269] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[0270] Server: Sends text data and confusion emotion data to the generative model.
[0271] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[0272] Device: Provides voice and text responses to students.
[0273] For the medical and nursing care sector:
[0274] Elderly user: "I'm not feeling very well today," says sadly.
[0275] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[0276] Server: Sends text data and sadness emotion data to the generative model.
[0277] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[0278] Device: Provides voice and text responses to seniors.
[0279] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[0280] The processing flow will be explained below.
[0281] Step 1:
[0282] The user speaks to the communication robot (e.g., "I'm not feeling very well today."). When voice input begins, the robot's terminal uses a microphone and camera to capture the user's voice and facial expressions.
[0283] Step 2:
[0284] The device converts the captured voice into text data using built-in voice recognition software, while simultaneously extracting the user's facial expression data from the camera footage using facial expression analysis software.
[0285] Step 3:
[0286] The device passes the text data and facial expression data to an emotion recognition engine to estimate the user's emotional state. The emotion recognition engine analyzes the voice tone, volume, and facial expression data to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[0287] Step 4:
[0288] The device collects the emotion data obtained from the emotion recognition engine along with the text data and organizes it into an appropriate format such as JSON, which also includes the user's ID and session information.
[0289] Step 5:
[0290] The device sends text data and emotion data to the server via an HTTP request, which includes the text data, emotion data, and related user information.
[0291] Step 6:
[0292] The server parses the received text and emotion data, checks them, and performs any necessary preprocessing before sending them to the generative model.
[0293] Step 7:
[0294] The server sends the parsed text data and emotion data as a request to the API endpoint of the generative model. This request includes the text data and emotion data entered by the user.
[0295] Step 8:
[0296] The generative model analyzes the received request and generates an appropriate response, taking into account the emotional data to generate a response that is appropriate for the user's emotional state.
[0297] Step 9:
[0298] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[0299] Step 10:
[0300] The server receives the response from the generative model and then sends the received response data to the terminal as an HTTP response.
[0301] Step 11:
[0302] The device parses (analyzes) the response data received from the server, converting the text data into a format suitable for voice output or display.
[0303] Step 12:
[0304] The device passes the parsed response data to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that takes into account the user's emotional state.
[0305] Step 13:
[0306] The terminal plays back the generated voice using the audio output device (speaker) and, if necessary, displays the response text on the display, so that the user receives appropriate feedback.
[0307] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[0308] Example 2
[0309] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0310] In conventional dialogue systems, the dialogue with the user is one-sided, making it difficult to generate responses that take the user's emotional state into consideration. In particular, in the fields of education, medicine, and nursing care, there are many situations where responses that correspond to the user's emotions are required, and conventional technologies have not always been satisfactory.
[0311] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an output means for providing the user with a response generated by the generative model, a voice recognition means for converting the user's voice input into text data, a transmission means for transmitting the text data to the generative model, an expression analysis means for analyzing the user's facial expression, and an emotion recognition means for estimating the user's emotional state from the user's voice and facial expression. This makes it possible to realize a natural dialogue according to the user's emotional state.
[0312] A "generative model" is an algorithm that generates natural language responses or text based on given input data.
[0313] A "dialogue response generator" is a device or method that uses a generative model to generate a response to an input from a user.
[0314] An "output means" is a device or method for providing the generated response to a user.
[0315] "Speech recognition means" refers to technology or devices for analyzing a user's voice input and converting it into text data.
[0316] "Transmission means" refers to the communication means or protocol for transmitting text data to the generative model.
[0317] The "facial expression analysis means" refers to a technique or device that analyzes the user's facial expression from camera footage and acquires that information.
[0318] "Emotion recognition means" refers to technology or devices for estimating a user's emotional state based on voice and facial expression data.
[0319] "Cloud technology" is a technology that stores data on remote servers via the Internet and performs computational processing.
[0320] "Individual educational support" refers to a method or system that provides educational support tailored to each student's learning situation and level of understanding.
[0321] This invention is a system that combines a generative model and an emotion recognition engine to make interactions with users more natural and effective. The configuration and operation of this system will be described in detail below.
[0322] 1. System Configuration
[0323] The system consists of the following main components:
[0324] Generative model: Uses natural language processing algorithms to generate responses in response to user input.
[0325] Dialogue response generator: Utilizes a generative model to create appropriate responses to user input.
[0326] Output: A device (e.g., speaker, display, etc.) that provides the generated response to the user.
[0327] Speech recognition tool: Software that converts user voice input into text data (e.g., Google Speech-to-Text).
[0328] Transmission medium: The communication protocol (e.g., HTTP) used to send text data to the generative model.
[0329] Facial expression analysis means: Software for analyzing the user's facial expressions from camera footage and extracting emotional data (e.g., Amazon Rekognition).
[0330] Emotion recognizer: An engine for inferring the user's emotional state from their voice and facial expression data (e.g., IBM Watson Tone Analyzer).
[0331] Cloud technology: Technology for performing computational processing on remote servers via the Internet.
[0332] 2. Hardware and Software Configuration
[0333] Hardware:
[0334] Microphone: A device for capturing the user's voice.
[0335] Camera: A device for capturing the user's facial expressions.
[0336] Speaker: A device for providing generated audio responses to a user.
[0337] Display: A device for displaying the generated text response.
[0338] software:
[0339] Speech recognition software: converts speech into text data (e.g., Google Speech-to-Text).
[0340] Facial expression analysis software: Extracting emotional data from camera footage (e.g., Amazon Rekognition).
[0341] Emotion recognition engine: Estimates emotional state from voice tone and facial expressions (e.g., IBM Watson Tone Analyzer).
[0342] Generative model: An algorithm for generating responses based on input text data (e.g., the open-source GPT model).
[0343] 3. System Operation
[0344] When a user speaks to the communication robot, the speech recognition means converts the speech into text data. At the same time, the facial expression analysis means analyzes the user's facial expressions and extracts emotional data. This data is sent to the server via the transmission means, and the generative model generates an appropriate response. The generated response is sent to the user's device via cloud technology and converted into voice using speech synthesis software. The response is finally provided to the user through a speaker.
[0345] Specific examples
[0346] For the education sector:
[0347] User (Student): "I don't understand that math problem," says in a confused tone.
[0348] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[0349] Server: Sends text data and confusion emotion data to the generative model.
[0350] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[0351] Device: Provides voice and text responses to students.
[0352] For the medical and nursing care sector:
[0353] Elderly user: "I'm not feeling very well today," says sadly.
[0354] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[0355] Server: Sends text data and sadness emotion data to the generative model.
[0356] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[0357] Device: Provides voice and text responses to seniors.
[0358] Prompt Sentence Examples
[0359] Education: “When a struggling student says, ‘I don’t understand that math problem,’ how does a generative model respond?”
[0360] Healthcare: "How would you respond to an elderly person who says in a sad tone, 'I'm not feeling very well today'?"
[0361] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[0362] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0363] Step 1:
[0364] Collecting User Input
[0365] User: Talks to the communication robot, "I'm not feeling very good today."
[0366] Input: User's voice
[0367] Output: Raw audio file
[0368] Device: Uses a microphone to capture the user's voice.
[0369] What it does: A microphone records audio and converts it into digital data in real time.
[0370] Device: Uses the camera to capture the user's facial expressions.
[0371] Input: User's facial expression
[0372] Output: Facial expression video data
[0373] Specific operation: The camera captures the user's face and generates video data.
[0374] On your device: Use speech recognition software to convert voice data into text (e.g., Google Speech-to-Text).
[0375] Input: Raw audio file
[0376] Output: Text data
[0377] How it works: Recorded audio data is sent to speech recognition software, which analyzes the audio waveform and converts it into text.
[0378] On the device: Facial expression analysis software is used to analyze the user's facial expressions from the camera footage and extract emotional data (e.g., Amazon Rekognition).
[0379] Input: facial expression video data
[0380] Output: Emotion data
[0381] How it works: The captured image of the face is input into expression analysis software, which extracts facial features and estimates the emotional state based on them.
[0382] Step 2:
[0383] Emotional Data Processing
[0384] Device: The extracted voice and facial expression data is passed to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to estimate the user's emotional state.
[0385] Input: Voice data and facial expression data
[0386] Output: Emotional state
[0387] What it does: Voice and facial expression data are sent to the emotion recognition engine, which then begins the analysis process, identifying the emotional state from changes in voice tone and facial expressions.
[0388] Emotion Recognition Engine: Analyzes voice tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger).
[0389] Input: Voice data and facial expression data
[0390] Output: Emotion label (e.g. sadness, anger)
[0391] What it does: Specific algorithms analyze audio and video data to generate emotion labels.
[0392] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[0393] Input: Emotion recognition engine results, text data
[0394] Output: Organized data (JSON format)
[0395] Step 3:
[0396] Sending user input and emotion data
[0397] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[0398] Input: Text data, emotion data
[0399] Output: JSON format data
[0400] Specific behavior: Text data and emotion data are packaged into a single JSON object.
[0401] Terminal: Sends the organized data to the server via an HTTP request.
[0402] Input: JSON format data
[0403] Output: HTTP request
[0404] What happens: An HTTP request is constructed and data is sent to the server endpoint.
[0405] Step 4:
[0406] Generating dialogue content using generative models
[0407] Server: Receives and parses text and emotion data.
[0408] Input: HTTP request data
[0409] Output: Parsed data
[0410] Specific operation: Parse the received JSON data and extract the necessary fields.
[0411] Server: Sends the parsed data as a request to the API endpoint of the generative model (e.g., the open-source GPT model).
[0412] Input: Parsed data
[0413] Output: API request
[0414] What happens: A request is sent to the generative model's API and the appropriate parameters are set.
[0415] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[0416] Input: API request data
[0417] Output: Response data
[0418] Specific operation: The model analyzes the text data and generates a response that takes into account the emotional state.
[0419] Generative model: Generates the answer and sends it back to the server as a response.
[0420] Input: Response data
[0421] Output: Response data (JSON format)
[0422] Specific behavior: Response data is sent to the server in JSON format.
[0423] Step 5:
[0424] Sending and Processing the Response
[0425] Server: Receives the response from the generative model.
[0426] Input: Response data
[0427] Output: Parsed response data
[0428] Specific operation: The server receives the response from the API and analyzes it.
[0429] Server: Sends the received response to the terminal as an HTTP response.
[0430] Input: Parsed response data
[0431] Output: HTTP response
[0432] Specific operation: The response data is organized and sent as an HTTP response to the device.
[0433] Step 6:
[0434] Responding to the user
[0435] Terminal: Parse the response received from the server.
[0436] Input: HTTP response data
[0437] Output: Parsed response data
[0438] Specific operation: Analyzes the received data and extracts the necessary information.
[0439] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[0440] Input: Response data
[0441] Output: Audio data
[0442] Specific operation: Text is input into the speech synthesis software, and the speech generation process is executed according to the emotional state.
[0443] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[0444] Input: Audio data, text data
[0445] Output: Playback audio, display
[0446] Specific behavior: The speaker plays the generated audio and the response text appears on the display.
[0447] (Application example 2)
[0448] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0449] Dialogue response systems using generative models have the problem of being unable to flexibly respond to the user's emotional state, resulting in mechanical and unnatural dialogue. In particular, in factory environments, dialogue that ignores the worker's emotional state can lead to a deterioration in the working environment and a decrease in efficiency. Therefore, there is a need for technology that can recognize the worker's emotional state and adjust the dialogue content accordingly.
[0450] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an expression recognition means for analyzing the user's expression data, and a means for adjusting the response of the generative model based on the emotion recognition means. This makes it possible to recognize the emotion of the worker and provide a natural and effective dialogue response according to that emotion.
[0451] A "generative model" is a model that generates text data using a machine learning algorithm.
[0452] The "interactive response generation means" is a function that uses a generative model to generate a response required for a dialogue with a user.
[0453] The "output means" is a function for providing the generated response to the user.
[0454] The "voice recognition means" is a function that converts the user's voice input into text data.
[0455] The "transmission means" is a function for transmitting text data to the generative model.
[0456] The "facial expression recognition means" is a function that analyzes the user's facial expression data and estimates the user's emotional state.
[0457] The "emotion recognition means" is a function that estimates the user's emotional state from the tone of voice, facial expression, etc.
[0458] "Cloud technology" is a technology that processes and stores data remotely via the Internet.
[0459] A "work environment" is an environment where a specific task is performed, such as a factory or office.
[0460] "Worker" refers to a person who performs a particular task.
[0461] The present invention provides a dialogue system that recognizes the emotional state of a worker in a factory environment and generates a dialogue response in response to the worker's emotional state. The system includes a dialogue response generation means using a generative model, a speech recognition means for converting a user's voice input into text data, an expression recognition means for analyzing the user's facial expression data, a means for adjusting the response of the generative model based on the emotion recognition means, and a means for processing, transmitting, and receiving data using cloud technology.
[0462] Explanation of program processing
[0463] Hardware and software used
[0464] Hardware:
[0465] Microphone: Used to collect the user's voice.
[0466] Camera: Used to capture the user's facial expressions.
[0467] software:
[0468] Speech Recognition: Uses Google Cloud Speech-to-Text to convert speech to text.
[0469] Facial Expression Recognition: Uses Microsoft Azure Face API to analyze the user's facial expressions.
[0470] Emotion Recognition: Uses IBM Watson Tone Analyzer to infer emotional state from vocal tone and facial expressions.
[0471] Generative model: Uses OpenAI GPT-4 API to generate dialogue responses.
[0472] Communication: Uses HTTP REST APIs to send and receive data.
[0473] Specific examples
[0474] When a factory worker speaks to the device, the device collects the voice using a microphone and converts the voice into text data using Google Cloud Speech-to-Text. At the same time, the device captures the worker's facial expressions using a camera and analyzes the facial expression data using the Microsoft Azure Face API to estimate their emotional state. Next, the device analyzes the voice tone using IBM Watson Tone Analyzer to generate the final emotional data.
[0475] The server receives this data and uses the OpenAI GPT-4 API to generate an appropriate response, taking into account the emotional data and adjusting the content and tone of the response. The generated response is then provided to the user via their device in voice and text format.
[0476] Example prompt sentence:
[0477] User input: "I've been having trouble getting things done lately." (sad tone)
[0478] Sentiment analysis result: "Sad"
[0479] Corresponding response: "That's tough. How about you take a break?"
[0480] This system enables natural dialogue responses that respond to the emotions of workers in a factory environment, which is expected to improve the working environment and increase efficiency. Furthermore, by allowing technicians to provide appropriate support based on their emotions, it is possible to increase worker satisfaction and overall productivity.
[0481] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0482] Step 1:
[0483] The device uses a microphone to collect the user's voice input, which can be expressed as text, such as "I'm having trouble getting my work done lately." The collected voice input is saved as raw audio data.
[0484] Step 2:
[0485] The device converts the collected voice data into text data using Google Cloud Speech-to-Text. The input is voice data, and after analyzing the voice data, it outputs the text data, "I've been having trouble with my work lately."
[0486] Step 3:
[0487] The device captures the user's facial expressions with a camera. The captured image is saved as raw image data, which is used as input for facial recognition.
[0488] Step 4:
[0489] The device uses the Microsoft Azure Face API to analyze the user's facial expressions from the captured image data. The input is image data, and the output is a quantitative emotional state (e.g., sadness, joy, etc.) representing the facial expression data.
[0490] Step 5:
[0491] The device uses IBM Watson Tone Analyzer to analyze the voice tone from the text data obtained by Google Cloud Speech-to-Text. The input is the text data, and the output is the tone analysis result (e.g., sad, wonderful, etc.).
[0492] Step 6:
[0493] The device integrates the results of facial expression analysis and voice tone analysis to generate the final emotion data. The input is quantitative facial expression data and voice tone data, and the output is the integrated emotional state (e.g., sad).
[0494] Step 7:
[0495] The device organizes the generated text data and emotion data into an appropriate format, such as JSON, and sends it to the server via an HTTP request. The input is the text data and emotion data, and the output is an HTTP request sent to the server.
[0496] Step 8:
[0497] The server analyzes the text data and emotion data received from the device and generates a response using the OpenAI GPT-4 API. The input is text data and emotion data, and it outputs an appropriate response text that takes the emotion data into consideration.
[0498] Step 9:
[0499] The server sends the generated response text to the terminal as an HTTP response. The input is the response text obtained from the generative model, and the output is the HTTP response sent to the terminal.
[0500] Step 10:
[0501] The device converts the received response text into audio data using Google Cloud Text-to-Speech. The input is the response text, and the output is audio data.
[0502] Step 11:
[0503] The terminal plays the generated voice using an audio output device, and simultaneously displays the response content as text on a display if necessary. The input is the voice data and the response text, and the output is the voice and text display to the user.
[0504] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0505] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0506] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0507] [Second embodiment]
[0508] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0509] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0510] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0511] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0512] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0513] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0514] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0515] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0516] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0517] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0518] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0519] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0520] Program Overview
[0521] This invention realizes a system including a dialogue response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Furthermore, communication between the generative model and the user is performed using cloud technology.
[0522] Explanation of program processing
[0523] 1. Collect user input:
[0524] User: Talk to the communication robot, for example, "What's the weather like today?"
[0525] Device: Uses a microphone to capture the user's voice.
[0526] Device: Uses speech recognition software to convert speech into text data.
[0527] 2. Sending user input:
[0528] Terminal: Organize text data into a suitable format (e.g. JSON).
[0529] Terminal: Sends text data to the server via an HTTP request.
[0530] 3. Generating dialogue content using a generative model:
[0531] Server: Receives and parses the text data.
[0532] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[0533] Generative Model: Receives requests and generates appropriate responses.
[0534] Generative model: Generates the answer and sends it back to the server as a response.
[0535] 4. Sending and processing the response:
[0536] Server: Receives responses from the generative AI model.
[0537] Server: Sends the received response to the terminal as an HTTP response.
[0538] 5. Response to the user:
[0539] Terminal: Parses the response received from the server.
[0540] Terminal: Passes the parsed response data to speech synthesis software to convert text to speech.
[0541] Terminal: Plays the generated audio using an audio output device (speaker).
[0542] Terminal: If necessary, the response will be displayed as text on the display.
[0543] Specific examples
[0544] For the education sector:
[0545] 1. User (Student): "I don't understand this math problem. How do I solve it?"
[0546] Terminal: Converts student voice into text data and sends it to the server.
[0547] Server: Sends text data to the generative model and generates an appropriate explanation.
[0548] Generative model: Generates explanations such as "First, try adding 2 to both sides of the equation. Then, as a next step..."
[0549] Server: Sends this response to the device.
[0550] Terminal: The robot outputs the explanation aloud and also displays it on the screen.
[0551] For the medical and nursing care sector:
[0552] 1. User (elderly): "I'm not feeling very well today."
[0553] Terminal: Converts the elderly person's voice into text data and sends it to the server.
[0554] Server: Feeds text data into the generative model and generates an appropriate response.
[0555] Generative model: Generate a response such as, "That's worrying. Let's take a few deep breaths together, and then it might be a good idea to talk to your doctor."
[0556] Server: Sends this response to the device.
[0557] Terminal: Have the robot output the response content by voice.
[0558] The system provides users with a smooth and personalized experience. By combining generative models and cloud technology, the communication robot achieves natural and effective interactions in a wide range of scenarios.
[0559] The processing flow will be explained below.
[0560] Step 1:
[0561] The user speaks to the communication robot (e.g., "What's the weather like today?"). When voice input begins, the robot's terminal uses a microphone to capture the user's voice.
[0562] Step 2:
[0563] The device converts the captured audio into text data using built-in speech recognition software, which uses a speech analysis algorithm to analyze the characteristics of the voice and convert it into text.
[0564] Step 3:
[0565] The terminal organizes the converted text data into an appropriate format, such as JSON, which also includes the user's ID and session information.
[0566] Step 4:
[0567] The device sends the organized text data to the server via an HTTP request, which includes the text data along with user information and related context information.
[0568] Step 5:
[0569] The server parses the received text data, checks the text data, and performs any necessary preprocessing before sending it to the generative model.
[0570] Step 6:
[0571] The server sends the parsed text data as a request to the API endpoint of the generative model, which includes the text data entered by the user.
[0572] Step 7:
[0573] The generative model analyzes the incoming request and generates an appropriate response. The generation process uses natural language processing algorithms to generate an appropriate response based on the context.
[0574] Step 8:
[0575] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[0576] Step 9:
[0577] The server receives the response from the generative model and parses it again to ensure that the response is well formed.
[0578] Step 10:
[0579] The server sends the received response data to the terminal as an HTTP response, which includes the generated response text data.
[0580] Step 11:
[0581] The device parses the response data received from the server and converts the text data into a format suitable for voice output or display.
[0582] Step 12:
[0583] The device passes the parsed response data to speech synthesis software, which converts the text into speech. The speech synthesis process converts the text data into speech waveforms.
[0584] Step 13:
[0585] The terminal plays the generated voice using the audio output device (speaker), and simultaneously displays the response text on the display if necessary.
[0586] This series of processing steps enables the user to have a natural and effective dialogue with the communication robot.
[0587] Example 1
[0588] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0589] Conventional interactive response systems have had difficulty efficiently collecting user voice input, converting it into appropriate text data, and generating responses. Furthermore, there are limited means for providing users with generated responses in natural-sounding voices. Furthermore, there is a need to effectively integrate generative AI models with cloud technology to realize personalized educational support and other applications.
[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0591] In this invention, the server includes means for capturing voice input from a user, means for converting the voice input into text data, means for organizing the text data into an appropriate data format, means for sending the organized text data to a generative model, means for generating a dialogue response using the generative model, means for providing the generated dialogue response to the user, and means for converting the generated dialogue response into voice data, and includes output means for playing back the voice data. This makes it possible to generate smooth and natural dialogue responses using a generative AI model from the user's voice input and provide them in voice.
[0592] "User" refers to a person who uses the system to provide voice input.
[0593] "Voice input" refers to voice information spoken by a user to the system.
[0594] "Capturing means" refers to a device or software used to collect and record a user's voice input.
[0595] "Means for converting into text data" refers to software or algorithms used to analyze and convert collected voice input into corresponding text data.
[0596] "Means for organizing data" refers to the methods or mechanisms for preparing text data in an appropriate format, such as JSON, for transmission to a generative model.
[0597] "Means for sending" refers to a communication means for sending organized text data to a location where a generative model operates.
[0598] A "generative model" refers to an algorithm or software that generates appropriate dialogue responses based on text data.
[0599] "Means for generating a dialogue response" refers to a process of using a generative model to generate a response corresponding to an input.
[0600] "Means for providing" refers to a method for passing the generated interactive response to the user.
[0601] "Means for converting into voice data" refers to a process for converting the generated textual dialogue response into voice format.
[0602] "Output means" refers to a device or mechanism for playing back the converted audio data.
[0603] "Cloud technology" refers to computing resources and services provided over the internet.
[0604] "Educational support" refers to functions and services that provide individual learning support and guidance in the field of education.
[0605] The present invention is a dialogue response generation system using a generation model, and specifically includes the following means.
[0606] Collecting User Input
[0607] The user speaks to the communication robot. A specific example of a question is, "How's the weather today?" The device uses a microphone to capture the user's voice input. It then uses voice recognition software (e.g., a commonly used voice recognition engine) to convert the voice input into text data. For example, the voice "How's the weather today?" is converted into text "How's the weather today?"
[0608] Sending User Input
[0609] The device organizes the converted text data into an appropriate data format, such as JSON. The organized text data is sent to the server via an HTTP request. This request includes the user input, "What's the weather like today?"
[0610] Generating dialogue content using generative models
[0611] The server processes the HTTP request received from the device and extracts the text data. The extracted text data is sent as a request to the API endpoint of a generative model (for example, a commonly used generative AI model). The generative model receives the request and generates an appropriate response, such as "It's sunny today." This generated response is then sent back to the server.
[0612] Sending and Processing the Response
[0613] The server receives the response from the generative model and sends it to the device as an HTTP response. The received response may be data in the form of, for example, "It's sunny today."
[0614] Responding to the user
[0615] The device analyzes the HTTP response received from the server. It uses speech synthesis software (for example, a commonly used speech synthesis engine) to convert the text data into voice data. It then plays the generated voice using a speaker and responds to the user by saying, "It's a sunny day today." It also displays the response text on the display as needed.
[0616] Specific examples
[0617] Specific examples in the field of education
[0618] User (Student): "I don't understand this math problem. How do I solve it?"
[0619] The terminal converts the student's voice into text data and sends it to the server.
[0620] The server feeds the text data into a generative model to generate an appropriate explanation.
[0621] The generative model generates an explanation such as, "First, try adding 2 to both sides of the equation. Then, as a next step..."
[0622] The server sends this response to the terminal.
[0623] The device will then audibly transmit the generated explanation to the student and also display it on the screen.
[0624] Specific examples in the medical and nursing care fields
[0625] Elderly user: "I'm not feeling too great today."
[0626] The terminal converts the elderly person's voice into text data and sends it to the server.
[0627] The server feeds the text data into the generative model to generate an appropriate response.
[0628] The generative model generates a response such as, "That's worrying. Let's take a deep breath together. Then it might be a good idea to consult a doctor."
[0629] The server sends this response to the terminal.
[0630] The terminal then transmits the generated response to the elderly person by voice.
[0631] This invention leverages the collaboration between generative models and cloud technology to provide users with a smooth and personalized interaction experience. By combining generative AI models with speech recognition and speech synthesis technologies, natural-sounding interactions are possible in a wide range of scenarios.
[0632] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0633] Step 1:
[0634] A user speaks to a communication robot. For example, "How's the weather today?" The device uses a microphone to capture the user's voice input. The input voice information is acquired as an analog signal. The device converts this analog signal into a digital signal and passes it to the voice recognition software. The voice input is "How's the weather today?" and the output is a digital voice signal.
[0635] Step 2:
[0636] The device uses speech recognition software (e.g., a commonly used speech recognition engine) to convert speech into text data. During this process, a digital speech signal is analyzed and converted into corresponding text data. The input is a digital speech signal, and the output is text data such as "What's the weather like today?". Specifically, the speech signal is analyzed based on an acoustic model and a language model.
[0637] Step 3:
[0638] The terminal organizes the converted text data into JSON format. For example, the text data "What's the weather like today?" is converted into the format {"input": "What's the weather like today?"}. The input is text data, and the output is JSON format data. Specifically, the text data is structured with appropriate key-value pairs.
[0639] Step 4:
[0640] The terminal sends organized JSON-formatted text data to the server via an HTTP POST request. The input is JSON-formatted data, and the output is an HTTP request sent to the server. Specifically, an HTTP request containing a URL and header information is generated.
[0641] Step 5:
[0642] The server processes the HTTP request received from the terminal and extracts the text data. The input is the HTTP request, and the output is the extracted text data (for example, "What's the weather like today?"). Specifically, the required data is parsed from the request body.
[0643] Step 6:
[0644] The server sends the extracted text data as a request to the API endpoint of a generative model (e.g., a commonly used generative AI model). The input is the extracted text data, and the output is the request sent to the generative model. The specific operation is to call the appropriate API endpoint.
[0645] Step 7:
[0646] A generative AI model generates an appropriate response based on the request it receives. For example, the input "What's the weather like today?" generates the response "It's sunny today." The input is the request to the generative model, and the output is the generated response (e.g., "It's sunny today"). Specifically, the generative algorithm performs the text analysis and generation process.
[0647] Step 8:
[0648] The generative model returns the generated response to the server as a response. The input is the generated response, and the output is the response returned to the server. As a specific operation, the response data is returned in an appropriate format such as JSON format.
[0649] Step 9:
[0650] The server receives the response from the generative model and sends it to the terminal as an HTTP response. The input is the response from the generative model, and the output is the HTTP response sent to the terminal. In concrete terms, the response data is formatted as an HTTP response.
[0651] Step 10:
[0652] The terminal analyzes the HTTP response received from the server and obtains the text data. The input is the HTTP response, and the output is the text data (for example, "It's a sunny day today"). Specifically, the data is extracted from the response body.
[0653] Step 11:
[0654] The terminal passes the acquired text data to speech synthesis software (for example, a commonly used speech synthesis engine) and converts the text data into speech data. The input is text data and the output is speech data. Specifically, the speech synthesis engine synthesizes the text data to generate a speech signal.
[0655] Step 12:
[0656] The device uses a speaker to play back the generated voice. For example, it may output the voice "It's a sunny day today." The input is voice data, and the output is a voice response to the user. Specifically, the speaker converts the voice signal into physical sound, which the user hears. If necessary, the response is also displayed as text on the display.
[0657] (Application example 1)
[0658] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0659] Conventional food delivery systems have the drawback of requiring users to operate complicated applications when placing an order, and the convenience of voice control is particularly insufficient. Furthermore, there is a lack of systems that utilize advanced dialogue functions to provide optimal suggestions to users, making it difficult to achieve seamless dialogue with users. Therefore, there is a demand for a system that allows users to easily order food delivery using voice commands.
[0660] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0661] In this invention, the server includes an interactive response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, a transmission means for transmitting the text data to the generative model, a speech synthesis means for converting the generated response into a speech output, and an order processing means for processing a user's order. This allows a user to seamlessly order food delivery using only voice input, and the generative model makes optimal suggestions, making it possible to provide a more convenient service.
[0662] "Means for generating dialogue responses using a generative model" refers to a function that executes a generative AI model to generate appropriate responses based on input from a user.
[0663] "Output means" refers to a device or function for providing a response generated by a generative model to a user.
[0664] "Speech recognition means" refers to technology or devices for converting a user's voice input into text data.
[0665] "Transmission means" refers to a communication means for transmitting the text data to the generative model.
[0666] "Speech synthesis means" refers to technology or devices for converting text responses generated by a generative model into speech.
[0667] "Order Processing Means" refers to the systems and functions used to process user orders and facilitate the actual food delivery process.
[0668] "Cloud technology" refers to remote computing resources and services used over the internet to communicate between generative models and users.
[0669] The "food delivery sector" refers to industries and activities related to the delivery of food and beverages to specific locations upon request.
[0670] System Overview
[0671] The present invention provides a system that allows users to place food delivery orders by voice. The system converts the user's voice input into text data, sends the data to a generative model to generate a dialogue response, and provides the response to the user as voice using a speech synthesis means. The system also processes the user's order and uses cloud technology to perform these communications.
[0672] Hardware and software used
[0673] This system uses the following hardware and software:
[0674] Microphone: Used to capture the user's voice.
[0675] Speech Recognition Software: Uses the Google Speech-to-Text API to convert voice input into text data.
[0676] Generative model: Uses the OpenAI GPT-3 API to generate responses based on user text input.
[0677] Speech synthesis software: Uses the Google Text-to-Speech API to convert text responses from the generative model into audio.
[0678] Cloud server: Sends and receives data using cloud technologies such as AWS EC2.
[0679] Data format: Organize and send communication data in JSON format.
[0680] Processing flow
[0681] 1. Collecting user voice input
[0682] The user speaks to the smartphone, for example, "What's your recommended pizza today?"
[0683] The smartphone's microphone captures the user's voice.
[0684] Speech recognition software (Google Speech-to-Text API) converts the speech into text data.
[0685] 2. Sending text data
[0686] The smartphone organizes the text data into the appropriate format (JSON).
[0687] Send text data to the cloud server via an HTTP request.
[0688] 3. Generating dialogue content using a generative model
[0689] The cloud server receives the text data and sends it to the endpoint of the generative model (OpenAI GPT-3 API).
[0690] The generative model receives the request, generates an appropriate response, and sends the response back to the cloud server.
[0691] 4. Response to the user
[0692] The cloud server receives the response from the generative model and sends it to the smartphone as an HTTP response.
[0693] The response received by the smartphone is converted into speech using speech synthesis software (Google Text-to-Speech API) and played through the speaker.
[0694] 5. Order Processing
[0695] The user receives the response and responds audibly, "Yes, I'll order that."
[0696] The order is confirmed through a similar process and the food delivery process is carried out through the order processing means.
[0697] Examples of specific examples and prompts
[0698] Sample prompt 1: "What's your recommended pizza today?"
[0699] Sample prompt 2: "What dessert would you recommend?"
[0700] Prompt 3: "I'll order that."
[0701] Specifically, when a user speaks to their smartphone, "What's the recommended pizza today?", the speech recognition software converts the speech into text and sends it to a cloud server. Based on the text, the generative model generates a response, "Today's recommended pizza is Margherita. Would you like to order?", which is returned to the smartphone, converted into speech by speech synthesis software, and played back to the user. When the user replies, "I'll order," the food delivery procedure is completed by the order processing means.
[0702] In this way, a system is created that allows users to seamlessly order food delivery while engaging in natural interactions via their smartphones.
[0703] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0704] Step 1: Collecting voice input
[0705] User: Talks to their smartphone to order food delivery, for example, "What's your pizza special today?"
[0706] Device: The smartphone's microphone captures the user's voice, which is then fed into voice recognition software.
[0707] On the device: Uses the Google Speech-to-Text API to convert voice data to text data. The input is voice data and the output is text data.
[0708] Terminal: Receives the converted text data, which will be used in the next step.
[0709] Step 2: Send text data
[0710] Terminal: Organizes text data into JSON format. Input is text data, output is JSON format data.
[0711] The device sends the organized JSON data via an HTTP request to the cloud server, which then sends this data to the generative model.
[0712] Step 3: Generating dialogue content using a generative model
[0713] Server: The cloud server receives the text data and sends a request to the API endpoint of the generative model. The input is JSON data, and the output is a request to the generative model.
[0714] Generative Model: The OpenAI GPT-3 API receives requests and generates appropriate responses. The input is the request data and the output is the response data.
[0715] Server: The cloud server receives the generated response data and sends it to the terminal as an HTTP response. The input is the response data, and the output is the HTTP response.
[0716] Step 4: Respond to the user
[0717] Terminal: The smartphone receives the HTTP response from the cloud server. The input is the HTTP response, and the output is the response text data.
[0718] Terminal: Convert the received response data into speech using the Google Text-to-Speech API. The input is the response text data, and the output is the speech data.
[0719] Terminal: Plays voice responses to the user through a speaker and optionally displays the responses on a display. Input is voice data, and output is voice and text display.
[0720] Step 5: Order Processing
[0721] User: In response to the response, responds verbally with "Yes, I'll order that."
[0722] Terminal: The user's response is processed in a similar process (steps 1 to 4). Finally, the order processing means processes the food delivery. Specifically, it stores the order details in a database, notifies the delivery service, and provides the user with an order confirmation.
[0723] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0724] Program Overview
[0725] The present invention aims to make dialogue with a user more natural and effective by combining an emotion engine with a system equipped with a dialogue response generation means using a generative model. The system includes an output means for generating a response based on the generative model and providing the response to the user, a speech recognition means for converting the user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Cloud technology is also used for communication between the generative model and the user. Furthermore, the emotion engine is used to recognize the user's emotional state and adjust the response of the generative model based on the recognition result.
[0726] Explanation of program processing
[0727] 1. Collect user input:
[0728] User: Talks to the communication robot (e.g., "I'm not feeling very good today.").
[0729] Device: Uses a microphone and camera to capture the user's voice and facial expressions.
[0730] Device: Uses speech recognition software to convert speech into text data.
[0731] Device: The user's facial expressions are analyzed using facial expression analysis software from the camera footage, and emotional data is extracted.
[0732] 2. Emotional Data Processing:
[0733] Device: The extracted voice and facial expression data is passed to an emotion recognition engine to estimate the user's emotional state.
[0734] Emotion recognition engine: Analyzes voice characteristics such as tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[0735] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[0736] 3. Sending user input and emotion data:
[0737] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[0738] Terminal: Sends data to the server via an HTTP request.
[0739] 4. Generating dialogue content using a generative model:
[0740] Server: Receives and parses text data and emotion data.
[0741] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[0742] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[0743] Generative model: Generates the answer and sends it back to the server as a response.
[0744] 5. Sending and processing the response:
[0745] Server: Receives the response from the generative model.
[0746] Server: Sends the received response to the terminal as an HTTP response.
[0747] 6. Response to the user:
[0748] Terminal: Parses the response received from the server.
[0749] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[0750] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[0751] Specific examples
[0752] For the education sector:
[0753] Student User: "I don't understand this math problem," says in a confused tone.
[0754] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[0755] Server: Sends text data and confusion emotion data to the generative model.
[0756] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[0757] Device: Provides voice and text responses to students.
[0758] For the medical and nursing care sector:
[0759] Elderly user: "I'm not feeling very well today," says sadly.
[0760] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[0761] Server: Sends text data and sadness emotion data to the generative model.
[0762] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[0763] Device: Provides voice and text responses to seniors.
[0764] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[0765] The processing flow will be explained below.
[0766] Step 1:
[0767] The user speaks to the communication robot (e.g., "I'm not feeling very well today."). When voice input begins, the robot's terminal uses a microphone and camera to capture the user's voice and facial expressions.
[0768] Step 2:
[0769] The device converts the captured voice into text data using built-in voice recognition software, while simultaneously extracting the user's facial expression data from the camera footage using facial expression analysis software.
[0770] Step 3:
[0771] The device passes the text data and facial expression data to an emotion recognition engine to estimate the user's emotional state. The emotion recognition engine analyzes the voice tone, volume, and facial expression data to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[0772] Step 4:
[0773] The device collects the emotion data obtained from the emotion recognition engine along with the text data and organizes it into an appropriate format such as JSON, which also includes the user's ID and session information.
[0774] Step 5:
[0775] The device sends text data and emotion data to the server via an HTTP request, which includes the text data, emotion data, and related user information.
[0776] Step 6:
[0777] The server parses the received text and emotion data, checks them, and performs any necessary preprocessing before sending them to the generative model.
[0778] Step 7:
[0779] The server sends the parsed text data and emotion data as a request to the API endpoint of the generative model. This request includes the text data and emotion data entered by the user.
[0780] Step 8:
[0781] The generative model analyzes the received request and generates an appropriate response, taking into account the emotional data to generate a response that is appropriate for the user's emotional state.
[0782] Step 9:
[0783] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[0784] Step 10:
[0785] The server receives the response from the generative model and then sends the received response data to the terminal as an HTTP response.
[0786] Step 11:
[0787] The device parses (analyzes) the response data received from the server, converting the text data into a format suitable for voice output or display.
[0788] Step 12:
[0789] The device passes the parsed response data to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that takes into account the user's emotional state.
[0790] Step 13:
[0791] The terminal plays back the generated voice using the audio output device (speaker) and, if necessary, displays the response text on the display, so that the user receives appropriate feedback.
[0792] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[0793] Example 2
[0794] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0795] In conventional dialogue systems, the dialogue with the user is one-sided, making it difficult to generate responses that take the user's emotional state into consideration. In particular, in the fields of education, medicine, and nursing care, there are many situations where responses that correspond to the user's emotions are required, and conventional technologies have not always been satisfactory.
[0796] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an output means for providing the user with a response generated by the generative model, a voice recognition means for converting the user's voice input into text data, a transmission means for transmitting the text data to the generative model, an expression analysis means for analyzing the user's facial expression, and an emotion recognition means for estimating the user's emotional state from the user's voice and facial expression. This makes it possible to realize a natural dialogue according to the user's emotional state.
[0797] A "generative model" is an algorithm that generates natural language responses or text based on given input data.
[0798] A "dialogue response generator" is a device or method that uses a generative model to generate a response to an input from a user.
[0799] An "output means" is a device or method for providing the generated response to a user.
[0800] "Speech recognition means" refers to technology or devices for analyzing a user's voice input and converting it into text data.
[0801] "Transmission means" refers to the communication means or protocol for transmitting text data to the generative model.
[0802] The "facial expression analysis means" refers to a technique or device that analyzes the user's facial expression from camera footage and acquires that information.
[0803] "Emotion recognition means" refers to technology or devices for estimating a user's emotional state based on voice and facial expression data.
[0804] "Cloud technology" is a technology that stores data on remote servers via the Internet and performs computational processing.
[0805] "Individual educational support" refers to a method or system that provides educational support tailored to each student's learning situation and level of understanding.
[0806] This invention is a system that combines a generative model and an emotion recognition engine to make interactions with users more natural and effective. The configuration and operation of this system will be described in detail below.
[0807] 1. System Configuration
[0808] The system consists of the following main components:
[0809] Generative model: Uses natural language processing algorithms to generate responses in response to user input.
[0810] Dialogue response generator: Utilizes a generative model to create appropriate responses to user input.
[0811] Output: A device (e.g., speaker, display, etc.) that provides the generated response to the user.
[0812] Speech recognition tool: Software that converts user voice input into text data (e.g., Google Speech-to-Text).
[0813] Transmission medium: The communication protocol (e.g., HTTP) used to send text data to the generative model.
[0814] Facial expression analysis means: Software for analyzing the user's facial expressions from camera footage and extracting emotional data (e.g., Amazon Rekognition).
[0815] Emotion recognizer: An engine for inferring the user's emotional state from their voice and facial expression data (e.g., IBM Watson Tone Analyzer).
[0816] Cloud technology: Technology for performing computational processing on remote servers via the Internet.
[0817] 2. Hardware and Software Configuration
[0818] Hardware:
[0819] Microphone: A device for capturing the user's voice.
[0820] Camera: A device for capturing the user's facial expressions.
[0821] Speaker: A device for providing generated audio responses to a user.
[0822] Display: A device for displaying the generated text response.
[0823] software:
[0824] Speech recognition software: converts speech into text data (e.g., Google Speech-to-Text).
[0825] Facial expression analysis software: Extracting emotional data from camera footage (e.g., Amazon Rekognition).
[0826] Emotion recognition engine: Estimates emotional state from voice tone and facial expressions (e.g., IBM Watson Tone Analyzer).
[0827] Generative model: An algorithm for generating responses based on input text data (e.g., the open-source GPT model).
[0828] 3. System Operation
[0829] When a user speaks to the communication robot, the speech recognition means converts the speech into text data. At the same time, the facial expression analysis means analyzes the user's facial expressions and extracts emotional data. This data is sent to the server via the transmission means, and the generative model generates an appropriate response. The generated response is sent to the user's device via cloud technology and converted into voice using speech synthesis software. The response is finally provided to the user through a speaker.
[0830] Specific examples
[0831] For the education sector:
[0832] User (Student): "I don't understand that math problem," says in a confused tone.
[0833] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[0834] Server: Sends text data and confusion emotion data to the generative model.
[0835] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[0836] Device: Provides voice and text responses to students.
[0837] For the medical and nursing care sector:
[0838] Elderly user: "I'm not feeling very well today," says sadly.
[0839] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[0840] Server: Sends text data and sadness emotion data to the generative model.
[0841] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[0842] Device: Provides voice and text responses to seniors.
[0843] Prompt Sentence Examples
[0844] Education: “When a struggling student says, ‘I don’t understand that math problem,’ how does a generative model respond?”
[0845] Healthcare: "How would you respond to an elderly person who says in a sad tone, 'I'm not feeling very well today'?"
[0846] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[0847] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0848] Step 1:
[0849] Collecting User Input
[0850] User: Talks to the communication robot, "I'm not feeling very good today."
[0851] Input: User's voice
[0852] Output: Raw audio file
[0853] Device: Uses a microphone to capture the user's voice.
[0854] What it does: A microphone records audio and converts it into digital data in real time.
[0855] Device: Uses the camera to capture the user's facial expressions.
[0856] Input: User's facial expression
[0857] Output: Facial expression video data
[0858] Specific operation: The camera captures the user's face and generates video data.
[0859] On your device: Use speech recognition software to convert voice data into text (e.g., Google Speech-to-Text).
[0860] Input: Raw audio file
[0861] Output: Text data
[0862] How it works: Recorded audio data is sent to speech recognition software, which analyzes the audio waveform and converts it into text.
[0863] On the device: Facial expression analysis software is used to analyze the user's facial expressions from the camera footage and extract emotional data (e.g., Amazon Rekognition).
[0864] Input: facial expression video data
[0865] Output: Emotion data
[0866] How it works: The captured image of the face is input into expression analysis software, which extracts facial features and estimates the emotional state based on them.
[0867] Step 2:
[0868] Emotional Data Processing
[0869] Device: The extracted voice and facial expression data is passed to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to estimate the user's emotional state.
[0870] Input: Voice data and facial expression data
[0871] Output: Emotional state
[0872] What it does: Voice and facial expression data are sent to the emotion recognition engine, which then begins the analysis process, identifying the emotional state from changes in voice tone and facial expressions.
[0873] Emotion Recognition Engine: Analyzes voice tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger).
[0874] Input: Voice data and facial expression data
[0875] Output: Emotion label (e.g. sadness, anger)
[0876] What it does: Specific algorithms analyze audio and video data to generate emotion labels.
[0877] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[0878] Input: Emotion recognition engine results, text data
[0879] Output: Organized data (JSON format)
[0880] Step 3:
[0881] Sending user input and emotion data
[0882] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[0883] Input: Text data, emotion data
[0884] Output: JSON format data
[0885] Specific behavior: Text data and emotion data are packaged into a single JSON object.
[0886] Terminal: Sends the organized data to the server via an HTTP request.
[0887] Input: JSON format data
[0888] Output: HTTP request
[0889] What happens: An HTTP request is constructed and data is sent to the server endpoint.
[0890] Step 4:
[0891] Generating dialogue content using generative models
[0892] Server: Receives and parses text and emotion data.
[0893] Input: HTTP request data
[0894] Output: Parsed data
[0895] Specific operation: Parse the received JSON data and extract the necessary fields.
[0896] Server: Sends the parsed data as a request to the API endpoint of the generative model (e.g., the open-source GPT model).
[0897] Input: Parsed data
[0898] Output: API request
[0899] What happens: A request is sent to the generative model's API and the appropriate parameters are set.
[0900] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[0901] Input: API request data
[0902] Output: Response data
[0903] Specific operation: The model analyzes the text data and generates a response that takes into account the emotional state.
[0904] Generative model: Generates the answer and sends it back to the server as a response.
[0905] Input: Response data
[0906] Output: Response data (JSON format)
[0907] Specific behavior: Response data is sent to the server in JSON format.
[0908] Step 5:
[0909] Sending and Processing the Response
[0910] Server: Receives the response from the generative model.
[0911] Input: Response data
[0912] Output: Parsed response data
[0913] Specific operation: The server receives the response from the API and analyzes it.
[0914] Server: Sends the received response to the terminal as an HTTP response.
[0915] Input: Parsed response data
[0916] Output: HTTP response
[0917] Specific operation: The response data is organized and sent as an HTTP response to the device.
[0918] Step 6:
[0919] Responding to the user
[0920] Terminal: Parse the response received from the server.
[0921] Input: HTTP response data
[0922] Output: Parsed response data
[0923] Specific operation: Analyzes the received data and extracts the necessary information.
[0924] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[0925] Input: Response data
[0926] Output: Audio data
[0927] Specific operation: Text is input into the speech synthesis software, and the speech generation process is executed according to the emotional state.
[0928] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[0929] Input: Audio data, text data
[0930] Output: Playback audio, display
[0931] Specific behavior: The speaker plays the generated audio and the response text appears on the display.
[0932] (Application example 2)
[0933] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0934] Dialogue response systems using generative models have the problem of being unable to flexibly respond to the user's emotional state, resulting in mechanical and unnatural dialogue. In particular, in factory environments, dialogue that ignores the worker's emotional state can lead to a deterioration in the working environment and a decrease in efficiency. Therefore, there is a need for technology that can recognize the worker's emotional state and adjust the dialogue content accordingly.
[0935] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an expression recognition means for analyzing the user's expression data, and a means for adjusting the response of the generative model based on the emotion recognition means. This makes it possible to recognize the emotion of the worker and provide a natural and effective dialogue response according to that emotion.
[0936] A "generative model" is a model that generates text data using a machine learning algorithm.
[0937] The "interactive response generation means" is a function that uses a generative model to generate a response required for a dialogue with a user.
[0938] The "output means" is a function for providing the generated response to the user.
[0939] The "voice recognition means" is a function that converts the user's voice input into text data.
[0940] The "transmission means" is a function for transmitting text data to the generative model.
[0941] The "facial expression recognition means" is a function that analyzes the user's facial expression data and estimates the user's emotional state.
[0942] The "emotion recognition means" is a function that estimates the user's emotional state from the tone of voice, facial expression, etc.
[0943] "Cloud technology" is a technology that processes and stores data remotely via the Internet.
[0944] A "work environment" is an environment where a specific task is performed, such as a factory or office.
[0945] "Worker" refers to a person who performs a particular task.
[0946] The present invention provides a dialogue system that recognizes the emotional state of a worker in a factory environment and generates a dialogue response in response to the worker's emotional state. The system includes a dialogue response generation means using a generative model, a speech recognition means for converting a user's voice input into text data, an expression recognition means for analyzing the user's facial expression data, a means for adjusting the response of the generative model based on the emotion recognition means, and a means for processing, transmitting, and receiving data using cloud technology.
[0947] Explanation of program processing
[0948] Hardware and software used
[0949] Hardware:
[0950] Microphone: Used to collect the user's voice.
[0951] Camera: Used to capture the user's facial expressions.
[0952] software:
[0953] Speech Recognition: Uses Google Cloud Speech-to-Text to convert speech to text.
[0954] Facial Expression Recognition: Uses Microsoft Azure Face API to analyze the user's facial expressions.
[0955] Emotion Recognition: Uses IBM Watson Tone Analyzer to infer emotional state from vocal tone and facial expressions.
[0956] Generative model: Uses OpenAI GPT-4 API to generate dialogue responses.
[0957] Communication: Uses HTTP REST APIs to send and receive data.
[0958] Specific examples
[0959] When a factory worker speaks to the device, the device collects the voice using a microphone and converts the voice into text data using Google Cloud Speech-to-Text. At the same time, the device captures the worker's facial expressions using a camera and analyzes the facial expression data using the Microsoft Azure Face API to estimate their emotional state. Next, the device analyzes the voice tone using IBM Watson Tone Analyzer to generate the final emotional data.
[0960] The server receives this data and uses the OpenAI GPT-4 API to generate an appropriate response, taking into account the emotional data and adjusting the content and tone of the response. The generated response is then provided to the user via their device in voice and text format.
[0961] Example prompt sentence:
[0962] User input: "I've been having trouble getting things done lately." (sad tone)
[0963] Sentiment analysis result: "Sad"
[0964] Corresponding response: "That's tough. How about you take a break?"
[0965] This system enables natural dialogue responses that respond to the emotions of workers in a factory environment, which is expected to improve the working environment and increase efficiency. Furthermore, by allowing technicians to provide appropriate support based on their emotions, it is possible to increase worker satisfaction and overall productivity.
[0966] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0967] Step 1:
[0968] The device uses a microphone to collect the user's voice input, which can be expressed as text, such as "I'm having trouble getting my work done lately." The collected voice input is saved as raw audio data.
[0969] Step 2:
[0970] The device converts the collected voice data into text data using Google Cloud Speech-to-Text. The input is voice data, and after analyzing the voice data, it outputs the text data, "I've been having trouble with my work lately."
[0971] Step 3:
[0972] The device captures the user's facial expressions with a camera. The captured image is saved as raw image data, which is used as input for facial recognition.
[0973] Step 4:
[0974] The device uses the Microsoft Azure Face API to analyze the user's facial expressions from the captured image data. The input is image data, and the output is a quantitative emotional state (e.g., sadness, joy, etc.) representing the facial expression data.
[0975] Step 5:
[0976] The device uses IBM Watson Tone Analyzer to analyze the voice tone from the text data obtained by Google Cloud Speech-to-Text. The input is the text data, and the output is the tone analysis result (e.g., sad, wonderful, etc.).
[0977] Step 6:
[0978] The device integrates the results of facial expression analysis and voice tone analysis to generate the final emotion data. The input is quantitative facial expression data and voice tone data, and the output is the integrated emotional state (e.g., sad).
[0979] Step 7:
[0980] The device organizes the generated text data and emotion data into an appropriate format, such as JSON, and sends it to the server via an HTTP request. The input is the text data and emotion data, and the output is an HTTP request sent to the server.
[0981] Step 8:
[0982] The server analyzes the text data and emotion data received from the device and generates a response using the OpenAI GPT-4 API. The input is text data and emotion data, and it outputs an appropriate response text that takes the emotion data into consideration.
[0983] Step 9:
[0984] The server sends the generated response text to the terminal as an HTTP response. The input is the response text obtained from the generative model, and the output is the HTTP response sent to the terminal.
[0985] Step 10:
[0986] The device converts the received response text into audio data using Google Cloud Text-to-Speech. The input is the response text, and the output is audio data.
[0987] Step 11:
[0988] The terminal plays the generated voice using an audio output device, and simultaneously displays the response content as text on a display if necessary. The input is the voice data and the response text, and the output is the voice and text display to the user.
[0989] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0990] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0991] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0992] [Third embodiment]
[0993] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0994] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0995] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0996] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0997] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0998] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0999] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1000] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1001] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1002] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1003] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1004] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1005] Program Overview
[1006] This invention realizes a system including a dialogue response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Furthermore, communication between the generative model and the user is performed using cloud technology.
[1007] Explanation of program processing
[1008] 1. Collect user input:
[1009] User: Talk to the communication robot, for example, "What's the weather like today?"
[1010] Device: Uses a microphone to capture the user's voice.
[1011] Device: Uses speech recognition software to convert speech into text data.
[1012] 2. Sending user input:
[1013] Terminal: Organize text data into a suitable format (e.g. JSON).
[1014] Terminal: Sends text data to the server via an HTTP request.
[1015] 3. Generating dialogue content using a generative model:
[1016] Server: Receives and parses the text data.
[1017] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[1018] Generative Model: Receives requests and generates appropriate responses.
[1019] Generative model: Generates the answer and sends it back to the server as a response.
[1020] 4. Sending and processing the response:
[1021] Server: Receives responses from the generative AI model.
[1022] Server: Sends the received response to the terminal as an HTTP response.
[1023] 5. Response to the user:
[1024] Terminal: Parses the response received from the server.
[1025] Terminal: Passes the parsed response data to speech synthesis software to convert text to speech.
[1026] Terminal: Plays the generated audio using an audio output device (speaker).
[1027] Terminal: If necessary, the response will be displayed as text on the display.
[1028] Specific examples
[1029] For the education sector:
[1030] 1. User (Student): "I don't understand this math problem. How do I solve it?"
[1031] Terminal: Converts student voice into text data and sends it to the server.
[1032] Server: Sends text data to the generative model and generates an appropriate explanation.
[1033] Generative model: Generates explanations such as "First, try adding 2 to both sides of the equation. Then, as a next step..."
[1034] Server: Sends this response to the device.
[1035] Terminal: The robot outputs the explanation aloud and also displays it on the screen.
[1036] For the medical and nursing care sector:
[1037] 1. User (elderly): "I'm not feeling very well today."
[1038] Terminal: Converts the elderly person's voice into text data and sends it to the server.
[1039] Server: Feeds text data into the generative model and generates an appropriate response.
[1040] Generative model: Generate a response such as, "That's worrying. Let's take a few deep breaths together, and then it might be a good idea to talk to your doctor."
[1041] Server: Sends this response to the device.
[1042] Terminal: Have the robot output the response content by voice.
[1043] The system provides users with a smooth and personalized experience. By combining generative models and cloud technology, the communication robot achieves natural and effective interactions in a wide range of scenarios.
[1044] The processing flow will be explained below.
[1045] Step 1:
[1046] The user speaks to the communication robot (e.g., "What's the weather like today?"). When voice input begins, the robot's terminal uses a microphone to capture the user's voice.
[1047] Step 2:
[1048] The device converts the captured audio into text data using built-in speech recognition software, which uses a speech analysis algorithm to analyze the characteristics of the voice and convert it into text.
[1049] Step 3:
[1050] The terminal organizes the converted text data into an appropriate format, such as JSON, which also includes the user's ID and session information.
[1051] Step 4:
[1052] The device sends the organized text data to the server via an HTTP request, which includes the text data along with user information and related context information.
[1053] Step 5:
[1054] The server parses the received text data, checks the text data, and performs any necessary preprocessing before sending it to the generative model.
[1055] Step 6:
[1056] The server sends the parsed text data as a request to the API endpoint of the generative model, which includes the text data entered by the user.
[1057] Step 7:
[1058] The generative model analyzes the incoming request and generates an appropriate response. The generation process uses natural language processing algorithms to generate an appropriate response based on the context.
[1059] Step 8:
[1060] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[1061] Step 9:
[1062] The server receives the response from the generative model and parses it again to ensure that the response is well formed.
[1063] Step 10:
[1064] The server sends the received response data to the terminal as an HTTP response, which includes the generated response text data.
[1065] Step 11:
[1066] The device parses the response data received from the server and converts the text data into a format suitable for voice output or display.
[1067] Step 12:
[1068] The device passes the parsed response data to speech synthesis software, which converts the text into speech. The speech synthesis process converts the text data into speech waveforms.
[1069] Step 13:
[1070] The terminal plays the generated voice using the audio output device (speaker), and simultaneously displays the response text on the display if necessary.
[1071] This series of processing steps enables the user to have a natural and effective dialogue with the communication robot.
[1072] Example 1
[1073] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1074] Conventional interactive response systems have had difficulty efficiently collecting user voice input, converting it into appropriate text data, and generating responses. Furthermore, there are limited means for providing users with generated responses in natural-sounding voices. Furthermore, there is a need to effectively integrate generative AI models with cloud technology to realize personalized educational support and other applications.
[1075] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1076] In this invention, the server includes means for capturing voice input from a user, means for converting the voice input into text data, means for organizing the text data into an appropriate data format, means for sending the organized text data to a generative model, means for generating a dialogue response using the generative model, means for providing the generated dialogue response to the user, and means for converting the generated dialogue response into voice data, and includes output means for playing back the voice data. This makes it possible to generate smooth and natural dialogue responses using a generative AI model from the user's voice input and provide them in voice.
[1077] "User" refers to a person who uses the system to provide voice input.
[1078] "Voice input" refers to voice information spoken by a user to the system.
[1079] "Capturing means" refers to a device or software used to collect and record a user's voice input.
[1080] "Means for converting into text data" refers to software or algorithms used to analyze and convert collected voice input into corresponding text data.
[1081] "Means for organizing data" refers to the methods or mechanisms for preparing text data in an appropriate format, such as JSON, for transmission to a generative model.
[1082] "Means for sending" refers to a communication means for sending organized text data to a location where a generative model operates.
[1083] A "generative model" refers to an algorithm or software that generates appropriate dialogue responses based on text data.
[1084] "Means for generating a dialogue response" refers to a process of using a generative model to generate a response corresponding to an input.
[1085] "Means for providing" refers to a method for passing the generated interactive response to the user.
[1086] "Means for converting into voice data" refers to a process for converting the generated textual dialogue response into voice format.
[1087] "Output means" refers to a device or mechanism for playing back the converted audio data.
[1088] "Cloud technology" refers to computing resources and services provided over the internet.
[1089] "Educational support" refers to functions and services that provide individual learning support and guidance in the field of education.
[1090] The present invention is a dialogue response generation system using a generation model, and specifically includes the following means.
[1091] Collecting User Input
[1092] The user speaks to the communication robot. A specific example of a question is, "How's the weather today?" The device uses a microphone to capture the user's voice input. It then uses voice recognition software (e.g., a commonly used voice recognition engine) to convert the voice input into text data. For example, the voice "How's the weather today?" is converted into text "How's the weather today?"
[1093] Sending User Input
[1094] The device organizes the converted text data into an appropriate data format, such as JSON. The organized text data is sent to the server via an HTTP request. This request includes the user input, "What's the weather like today?"
[1095] Generating dialogue content using generative models
[1096] The server processes the HTTP request received from the device and extracts the text data. The extracted text data is sent as a request to the API endpoint of a generative model (for example, a commonly used generative AI model). The generative model receives the request and generates an appropriate response, such as "It's sunny today." This generated response is then sent back to the server.
[1097] Sending and Processing the Response
[1098] The server receives the response from the generative model and sends it to the device as an HTTP response. The received response may be data in the form of, for example, "It's sunny today."
[1099] Responding to the user
[1100] The device analyzes the HTTP response received from the server. It uses speech synthesis software (for example, a commonly used speech synthesis engine) to convert the text data into voice data. It then plays the generated voice using a speaker and responds to the user by saying, "It's a sunny day today." It also displays the response text on the display as needed.
[1101] Specific examples
[1102] Specific examples in the field of education
[1103] User (Student): "I don't understand this math problem. How do I solve it?"
[1104] The terminal converts the student's voice into text data and sends it to the server.
[1105] The server feeds the text data into a generative model to generate an appropriate explanation.
[1106] The generative model generates an explanation such as, "First, try adding 2 to both sides of the equation. Then, as a next step..."
[1107] The server sends this response to the terminal.
[1108] The device will then audibly transmit the generated explanation to the student and also display it on the screen.
[1109] Specific examples in the medical and nursing care fields
[1110] Elderly user: "I'm not feeling too great today."
[1111] The terminal converts the elderly person's voice into text data and sends it to the server.
[1112] The server feeds the text data into the generative model to generate an appropriate response.
[1113] The generative model generates a response such as, "That's worrying. Let's take a deep breath together. Then it might be a good idea to consult a doctor."
[1114] The server sends this response to the terminal.
[1115] The terminal then transmits the generated response to the elderly person by voice.
[1116] This invention leverages the collaboration between generative models and cloud technology to provide users with a smooth and personalized interaction experience. By combining generative AI models with speech recognition and speech synthesis technologies, natural-sounding interactions are possible in a wide range of scenarios.
[1117] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1118] Step 1:
[1119] A user speaks to a communication robot. For example, "How's the weather today?" The device uses a microphone to capture the user's voice input. The input voice information is acquired as an analog signal. The device converts this analog signal into a digital signal and passes it to the voice recognition software. The voice input is "How's the weather today?" and the output is a digital voice signal.
[1120] Step 2:
[1121] The device uses speech recognition software (e.g., a commonly used speech recognition engine) to convert speech into text data. During this process, a digital speech signal is analyzed and converted into corresponding text data. The input is a digital speech signal, and the output is text data such as "What's the weather like today?". Specifically, the speech signal is analyzed based on an acoustic model and a language model.
[1122] Step 3:
[1123] The terminal organizes the converted text data into JSON format. For example, the text data "What's the weather like today?" is converted into the format {"input": "What's the weather like today?"}. The input is text data, and the output is JSON format data. Specifically, the text data is structured with appropriate key-value pairs.
[1124] Step 4:
[1125] The terminal sends organized JSON-formatted text data to the server via an HTTP POST request. The input is JSON-formatted data, and the output is an HTTP request sent to the server. Specifically, an HTTP request containing a URL and header information is generated.
[1126] Step 5:
[1127] The server processes the HTTP request received from the terminal and extracts the text data. The input is the HTTP request, and the output is the extracted text data (for example, "What's the weather like today?"). Specifically, the required data is parsed from the request body.
[1128] Step 6:
[1129] The server sends the extracted text data as a request to the API endpoint of a generative model (e.g., a commonly used generative AI model). The input is the extracted text data, and the output is the request sent to the generative model. The specific operation is to call the appropriate API endpoint.
[1130] Step 7:
[1131] A generative AI model generates an appropriate response based on the request it receives. For example, the input "What's the weather like today?" generates the response "It's sunny today." The input is the request to the generative model, and the output is the generated response (e.g., "It's sunny today"). Specifically, the generative algorithm performs the text analysis and generation process.
[1132] Step 8:
[1133] The generative model returns the generated response to the server as a response. The input is the generated response, and the output is the response returned to the server. As a specific operation, the response data is returned in an appropriate format such as JSON format.
[1134] Step 9:
[1135] The server receives the response from the generative model and sends it to the terminal as an HTTP response. The input is the response from the generative model, and the output is the HTTP response sent to the terminal. In concrete terms, the response data is formatted as an HTTP response.
[1136] Step 10:
[1137] The terminal analyzes the HTTP response received from the server and obtains the text data. The input is the HTTP response, and the output is the text data (for example, "It's a sunny day today"). Specifically, the data is extracted from the response body.
[1138] Step 11:
[1139] The terminal passes the acquired text data to speech synthesis software (for example, a commonly used speech synthesis engine) and converts the text data into speech data. The input is text data and the output is speech data. Specifically, the speech synthesis engine synthesizes the text data to generate a speech signal.
[1140] Step 12:
[1141] The device uses a speaker to play back the generated voice. For example, it may output the voice "It's a sunny day today." The input is voice data, and the output is a voice response to the user. Specifically, the speaker converts the voice signal into physical sound, which the user hears. If necessary, the response is also displayed as text on the display.
[1142] (Application example 1)
[1143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1144] Conventional food delivery systems have the drawback of requiring users to operate complicated applications when placing an order, and the convenience of voice control is particularly insufficient. Furthermore, there is a lack of systems that utilize advanced dialogue functions to provide optimal suggestions to users, making it difficult to achieve seamless dialogue with users. Therefore, there is a demand for a system that allows users to easily order food delivery using voice commands.
[1145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1146] In this invention, the server includes an interactive response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, a transmission means for transmitting the text data to the generative model, a speech synthesis means for converting the generated response into a speech output, and an order processing means for processing a user's order. This allows a user to seamlessly order food delivery using only voice input, and the generative model makes optimal suggestions, making it possible to provide a more convenient service.
[1147] "Means for generating dialogue responses using a generative model" refers to a function that executes a generative AI model to generate appropriate responses based on input from a user.
[1148] "Output means" refers to a device or function for providing a response generated by a generative model to a user.
[1149] "Speech recognition means" refers to technology or devices for converting a user's voice input into text data.
[1150] "Transmission means" refers to a communication means for transmitting the text data to the generative model.
[1151] "Speech synthesis means" refers to technology or devices for converting text responses generated by a generative model into speech.
[1152] "Order Processing Means" refers to the systems and functions used to process user orders and facilitate the actual food delivery process.
[1153] "Cloud technology" refers to remote computing resources and services used over the internet to communicate between generative models and users.
[1154] The "food delivery sector" refers to industries and activities related to the delivery of food and beverages to specific locations upon request.
[1155] System Overview
[1156] The present invention provides a system that allows users to place food delivery orders by voice. The system converts the user's voice input into text data, sends the data to a generative model to generate a dialogue response, and provides the response to the user as voice using a speech synthesis means. The system also processes the user's order and uses cloud technology to perform these communications.
[1157] Hardware and software used
[1158] This system uses the following hardware and software:
[1159] Microphone: Used to capture the user's voice.
[1160] Speech Recognition Software: Uses the Google Speech-to-Text API to convert voice input into text data.
[1161] Generative model: Uses the OpenAI GPT-3 API to generate responses based on user text input.
[1162] Speech synthesis software: Uses the Google Text-to-Speech API to convert text responses from the generative model into audio.
[1163] Cloud server: Sends and receives data using cloud technologies such as AWS EC2.
[1164] Data format: Organize and send communication data in JSON format.
[1165] Processing flow
[1166] 1. Collecting user voice input
[1167] The user speaks to the smartphone, for example, "What's your recommended pizza today?"
[1168] The smartphone's microphone captures the user's voice.
[1169] Speech recognition software (Google Speech-to-Text API) converts the speech into text data.
[1170] 2. Sending text data
[1171] The smartphone organizes the text data into the appropriate format (JSON).
[1172] Send text data to the cloud server via an HTTP request.
[1173] 3. Generating dialogue content using a generative model
[1174] The cloud server receives the text data and sends it to the endpoint of the generative model (OpenAI GPT-3 API).
[1175] The generative model receives the request, generates an appropriate response, and sends the response back to the cloud server.
[1176] 4. Response to the user
[1177] The cloud server receives the response from the generative model and sends it to the smartphone as an HTTP response.
[1178] The response received by the smartphone is converted into speech using speech synthesis software (Google Text-to-Speech API) and played through the speaker.
[1179] 5. Order Processing
[1180] The user receives the response and responds audibly, "Yes, I'll order that."
[1181] The order is confirmed through a similar process and the food delivery process is carried out through the order processing means.
[1182] Examples of specific examples and prompts
[1183] Sample prompt 1: "What's your recommended pizza today?"
[1184] Sample prompt 2: "What dessert would you recommend?"
[1185] Prompt 3: "I'll order that."
[1186] Specifically, when a user speaks to their smartphone, "What's the recommended pizza today?", the speech recognition software converts the speech into text and sends it to a cloud server. Based on the text, the generative model generates a response, "Today's recommended pizza is Margherita. Would you like to order?", which is returned to the smartphone, converted into speech by speech synthesis software, and played back to the user. When the user replies, "I'll order," the food delivery procedure is completed by the order processing means.
[1187] In this way, a system is created that allows users to seamlessly order food delivery while engaging in natural interactions via their smartphones.
[1188] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1189] Step 1: Collecting voice input
[1190] User: Talks to their smartphone to order food delivery, for example, "What's your pizza special today?"
[1191] Device: The smartphone's microphone captures the user's voice, which is then fed into voice recognition software.
[1192] On the device: Uses the Google Speech-to-Text API to convert voice data to text data. The input is voice data and the output is text data.
[1193] Terminal: Receives the converted text data, which will be used in the next step.
[1194] Step 2: Send text data
[1195] Terminal: Organizes text data into JSON format. Input is text data, output is JSON format data.
[1196] The device sends the organized JSON data via an HTTP request to the cloud server, which then sends this data to the generative model.
[1197] Step 3: Generating dialogue content using a generative model
[1198] Server: The cloud server receives the text data and sends a request to the API endpoint of the generative model. The input is JSON data, and the output is a request to the generative model.
[1199] Generative Model: The OpenAI GPT-3 API receives requests and generates appropriate responses. The input is the request data and the output is the response data.
[1200] Server: The cloud server receives the generated response data and sends it to the terminal as an HTTP response. The input is the response data, and the output is the HTTP response.
[1201] Step 4: Respond to the user
[1202] Terminal: The smartphone receives the HTTP response from the cloud server. The input is the HTTP response, and the output is the response text data.
[1203] Terminal: Convert the received response data into speech using the Google Text-to-Speech API. The input is the response text data, and the output is the speech data.
[1204] Terminal: Plays voice responses to the user through a speaker and optionally displays the responses on a display. Input is voice data, and output is voice and text display.
[1205] Step 5: Order Processing
[1206] User: In response to the response, responds verbally with "Yes, I'll order that."
[1207] Terminal: The user's response is processed in a similar process (steps 1 to 4). Finally, the order processing means processes the food delivery. Specifically, it stores the order details in a database, notifies the delivery service, and provides the user with an order confirmation.
[1208] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1209] Program Overview
[1210] The present invention aims to make dialogue with a user more natural and effective by combining an emotion engine with a system equipped with a dialogue response generation means using a generative model. The system includes an output means for generating a response based on the generative model and providing the response to the user, a speech recognition means for converting the user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Cloud technology is also used for communication between the generative model and the user. Furthermore, the emotion engine is used to recognize the user's emotional state and adjust the response of the generative model based on the recognition result.
[1211] Explanation of program processing
[1212] 1. Collect user input:
[1213] User: Talks to the communication robot (e.g., "I'm not feeling very good today.").
[1214] Device: Uses a microphone and camera to capture the user's voice and facial expressions.
[1215] Device: Uses speech recognition software to convert speech into text data.
[1216] Device: The user's facial expressions are analyzed using facial expression analysis software from the camera footage, and emotional data is extracted.
[1217] 2. Emotional Data Processing:
[1218] Device: The extracted voice and facial expression data is passed to an emotion recognition engine to estimate the user's emotional state.
[1219] Emotion recognition engine: Analyzes voice characteristics such as tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[1220] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[1221] 3. Sending user input and emotion data:
[1222] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[1223] Terminal: Sends data to the server via an HTTP request.
[1224] 4. Generating dialogue content using a generative model:
[1225] Server: Receives and parses text data and emotion data.
[1226] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[1227] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[1228] Generative model: Generates the answer and sends it back to the server as a response.
[1229] 5. Sending and processing the response:
[1230] Server: Receives the response from the generative model.
[1231] Server: Sends the received response to the terminal as an HTTP response.
[1232] 6. Response to the user:
[1233] Terminal: Parses the response received from the server.
[1234] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[1235] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[1236] Specific examples
[1237] For the education sector:
[1238] Student User: "I don't understand this math problem," says in a confused tone.
[1239] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[1240] Server: Sends text data and confusion emotion data to the generative model.
[1241] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[1242] Device: Provides voice and text responses to students.
[1243] For the medical and nursing care sector:
[1244] Elderly user: "I'm not feeling very well today," says sadly.
[1245] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[1246] Server: Sends text data and sadness emotion data to the generative model.
[1247] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[1248] Device: Provides voice and text responses to seniors.
[1249] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[1250] The processing flow will be explained below.
[1251] Step 1:
[1252] The user speaks to the communication robot (e.g., "I'm not feeling very well today."). When voice input begins, the robot's terminal uses a microphone and camera to capture the user's voice and facial expressions.
[1253] Step 2:
[1254] The device converts the captured voice into text data using built-in voice recognition software, while simultaneously extracting the user's facial expression data from the camera footage using facial expression analysis software.
[1255] Step 3:
[1256] The device passes the text data and facial expression data to an emotion recognition engine to estimate the user's emotional state. The emotion recognition engine analyzes the voice tone, volume, and facial expression data to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[1257] Step 4:
[1258] The device collects the emotion data obtained from the emotion recognition engine along with the text data and organizes it into an appropriate format such as JSON, which also includes the user's ID and session information.
[1259] Step 5:
[1260] The device sends text data and emotion data to the server via an HTTP request, which includes the text data, emotion data, and related user information.
[1261] Step 6:
[1262] The server parses the received text and emotion data, checks them, and performs any necessary preprocessing before sending them to the generative model.
[1263] Step 7:
[1264] The server sends the parsed text data and emotion data as a request to the API endpoint of the generative model. This request includes the text data and emotion data entered by the user.
[1265] Step 8:
[1266] The generative model analyzes the received request and generates an appropriate response, taking into account the emotional data to generate a response that is appropriate for the user's emotional state.
[1267] Step 9:
[1268] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[1269] Step 10:
[1270] The server receives the response from the generative model and then sends the received response data to the terminal as an HTTP response.
[1271] Step 11:
[1272] The device parses (analyzes) the response data received from the server, converting the text data into a format suitable for voice output or display.
[1273] Step 12:
[1274] The device passes the parsed response data to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that takes into account the user's emotional state.
[1275] Step 13:
[1276] The terminal plays back the generated voice using the audio output device (speaker) and, if necessary, displays the response text on the display, so that the user receives appropriate feedback.
[1277] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[1278] Example 2
[1279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1280] In conventional dialogue systems, the dialogue with the user is one-sided, making it difficult to generate responses that take the user's emotional state into consideration. In particular, in the fields of education, medicine, and nursing care, there are many situations where responses that correspond to the user's emotions are required, and conventional technologies have not always been satisfactory.
[1281] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an output means for providing the user with a response generated by the generative model, a voice recognition means for converting the user's voice input into text data, a transmission means for transmitting the text data to the generative model, an expression analysis means for analyzing the user's facial expression, and an emotion recognition means for estimating the user's emotional state from the user's voice and facial expression. This makes it possible to realize a natural dialogue according to the user's emotional state.
[1282] A "generative model" is an algorithm that generates natural language responses or text based on given input data.
[1283] A "dialogue response generator" is a device or method that uses a generative model to generate a response to an input from a user.
[1284] An "output means" is a device or method for providing the generated response to a user.
[1285] "Speech recognition means" refers to technology or devices for analyzing a user's voice input and converting it into text data.
[1286] "Transmission means" refers to the communication means or protocol for transmitting text data to the generative model.
[1287] The "facial expression analysis means" refers to a technique or device that analyzes the user's facial expression from camera footage and acquires that information.
[1288] "Emotion recognition means" refers to technology or devices for estimating a user's emotional state based on voice and facial expression data.
[1289] "Cloud technology" is a technology that stores data on remote servers via the Internet and performs computational processing.
[1290] "Individual educational support" refers to a method or system that provides educational support tailored to each student's learning situation and level of understanding.
[1291] This invention is a system that combines a generative model and an emotion recognition engine to make interactions with users more natural and effective. The configuration and operation of this system will be described in detail below.
[1292] 1. System Configuration
[1293] The system consists of the following main components:
[1294] Generative model: Uses natural language processing algorithms to generate responses in response to user input.
[1295] Dialogue response generator: Utilizes a generative model to create appropriate responses to user input.
[1296] Output: A device (e.g., speaker, display, etc.) that provides the generated response to the user.
[1297] Speech recognition tool: Software that converts user voice input into text data (e.g., Google Speech-to-Text).
[1298] Transmission medium: The communication protocol (e.g., HTTP) used to send text data to the generative model.
[1299] Facial expression analysis means: Software for analyzing the user's facial expressions from camera footage and extracting emotional data (e.g., Amazon Rekognition).
[1300] Emotion recognizer: An engine for inferring the user's emotional state from their voice and facial expression data (e.g., IBM Watson Tone Analyzer).
[1301] Cloud technology: Technology for performing computational processing on remote servers via the Internet.
[1302] 2. Hardware and Software Configuration
[1303] Hardware:
[1304] Microphone: A device for capturing the user's voice.
[1305] Camera: A device for capturing the user's facial expressions.
[1306] Speaker: A device for providing generated audio responses to a user.
[1307] Display: A device for displaying the generated text response.
[1308] software:
[1309] Speech recognition software: converts speech into text data (e.g., Google Speech-to-Text).
[1310] Facial expression analysis software: Extracting emotional data from camera footage (e.g., Amazon Rekognition).
[1311] Emotion recognition engine: Estimates emotional state from voice tone and facial expressions (e.g., IBM Watson Tone Analyzer).
[1312] Generative model: An algorithm for generating responses based on input text data (e.g., the open-source GPT model).
[1313] 3. System Operation
[1314] When a user speaks to the communication robot, the speech recognition means converts the speech into text data. At the same time, the facial expression analysis means analyzes the user's facial expressions and extracts emotional data. This data is sent to the server via the transmission means, and the generative model generates an appropriate response. The generated response is sent to the user's device via cloud technology and converted into voice using speech synthesis software. The response is finally provided to the user through a speaker.
[1315] Specific examples
[1316] For the education sector:
[1317] User (Student): "I don't understand that math problem," says in a confused tone.
[1318] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[1319] Server: Sends text data and confusion emotion data to the generative model.
[1320] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[1321] Device: Provides voice and text responses to students.
[1322] For the medical and nursing care sector:
[1323] Elderly user: "I'm not feeling very well today," says sadly.
[1324] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[1325] Server: Sends text data and sadness emotion data to the generative model.
[1326] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[1327] Device: Provides voice and text responses to seniors.
[1328] Prompt Sentence Examples
[1329] Education: “When a struggling student says, ‘I don’t understand that math problem,’ how does a generative model respond?”
[1330] Healthcare: "How would you respond to an elderly person who says in a sad tone, 'I'm not feeling very well today'?"
[1331] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[1332] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1333] Step 1:
[1334] Collecting User Input
[1335] User: Talks to the communication robot, "I'm not feeling very good today."
[1336] Input: User's voice
[1337] Output: Raw audio file
[1338] Device: Uses a microphone to capture the user's voice.
[1339] What it does: A microphone records audio and converts it into digital data in real time.
[1340] Device: Uses the camera to capture the user's facial expressions.
[1341] Input: User's facial expression
[1342] Output: Facial expression video data
[1343] Specific operation: The camera captures the user's face and generates video data.
[1344] On your device: Use speech recognition software to convert voice data into text (e.g., Google Speech-to-Text).
[1345] Input: Raw audio file
[1346] Output: Text data
[1347] How it works: Recorded audio data is sent to speech recognition software, which analyzes the audio waveform and converts it into text.
[1348] On the device: Facial expression analysis software is used to analyze the user's facial expressions from the camera footage and extract emotional data (e.g., Amazon Rekognition).
[1349] Input: facial expression video data
[1350] Output: Emotion data
[1351] How it works: The captured image of the face is input into expression analysis software, which extracts facial features and estimates the emotional state based on them.
[1352] Step 2:
[1353] Emotional Data Processing
[1354] Device: The extracted voice and facial expression data is passed to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to estimate the user's emotional state.
[1355] Input: Voice data and facial expression data
[1356] Output: Emotional state
[1357] What it does: Voice and facial expression data are sent to the emotion recognition engine, which then begins the analysis process, identifying the emotional state from changes in voice tone and facial expressions.
[1358] Emotion Recognition Engine: Analyzes voice tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger).
[1359] Input: Voice data and facial expression data
[1360] Output: Emotion label (e.g. sadness, anger)
[1361] What it does: Specific algorithms analyze audio and video data to generate emotion labels.
[1362] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[1363] Input: Emotion recognition engine results, text data
[1364] Output: Organized data (JSON format)
[1365] Step 3:
[1366] Sending user input and emotion data
[1367] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[1368] Input: Text data, emotion data
[1369] Output: JSON format data
[1370] Specific behavior: Text data and emotion data are packaged into a single JSON object.
[1371] Terminal: Sends the organized data to the server via an HTTP request.
[1372] Input: JSON format data
[1373] Output: HTTP request
[1374] What happens: An HTTP request is constructed and data is sent to the server endpoint.
[1375] Step 4:
[1376] Generating dialogue content using generative models
[1377] Server: Receives and parses text and emotion data.
[1378] Input: HTTP request data
[1379] Output: Parsed data
[1380] Specific operation: Parse the received JSON data and extract the necessary fields.
[1381] Server: Sends the parsed data as a request to the API endpoint of the generative model (e.g., the open-source GPT model).
[1382] Input: Parsed data
[1383] Output: API request
[1384] What happens: A request is sent to the generative model's API and the appropriate parameters are set.
[1385] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[1386] Input: API request data
[1387] Output: Response data
[1388] Specific operation: The model analyzes the text data and generates a response that takes into account the emotional state.
[1389] Generative model: Generates the answer and sends it back to the server as a response.
[1390] Input: Response data
[1391] Output: Response data (JSON format)
[1392] Specific behavior: Response data is sent to the server in JSON format.
[1393] Step 5:
[1394] Sending and Processing the Response
[1395] Server: Receives the response from the generative model.
[1396] Input: Response data
[1397] Output: Parsed response data
[1398] Specific operation: The server receives the response from the API and analyzes it.
[1399] Server: Sends the received response to the terminal as an HTTP response.
[1400] Input: Parsed response data
[1401] Output: HTTP response
[1402] Specific operation: The response data is organized and sent as an HTTP response to the device.
[1403] Step 6:
[1404] Responding to the user
[1405] Terminal: Parse the response received from the server.
[1406] Input: HTTP response data
[1407] Output: Parsed response data
[1408] Specific operation: Analyzes the received data and extracts the necessary information.
[1409] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[1410] Input: Response data
[1411] Output: Audio data
[1412] Specific operation: Text is input into the speech synthesis software, and the speech generation process is executed according to the emotional state.
[1413] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[1414] Input: Audio data, text data
[1415] Output: Playback audio, display
[1416] Specific behavior: The speaker plays the generated audio and the response text appears on the display.
[1417] (Application example 2)
[1418] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1419] Dialogue response systems using generative models have the problem of being unable to flexibly respond to the user's emotional state, resulting in mechanical and unnatural dialogue. In particular, in factory environments, dialogue that ignores the worker's emotional state can lead to a deterioration in the working environment and a decrease in efficiency. Therefore, there is a need for technology that can recognize the worker's emotional state and adjust the dialogue content accordingly.
[1420] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an expression recognition means for analyzing the user's expression data, and a means for adjusting the response of the generative model based on the emotion recognition means. This makes it possible to recognize the emotion of the worker and provide a natural and effective dialogue response according to that emotion.
[1421] A "generative model" is a model that generates text data using a machine learning algorithm.
[1422] The "interactive response generation means" is a function that uses a generative model to generate a response required for a dialogue with a user.
[1423] The "output means" is a function for providing the generated response to the user.
[1424] The "voice recognition means" is a function that converts the user's voice input into text data.
[1425] The "transmission means" is a function for transmitting text data to the generative model.
[1426] The "facial expression recognition means" is a function that analyzes the user's facial expression data and estimates the user's emotional state.
[1427] The "emotion recognition means" is a function that estimates the user's emotional state from the tone of voice, facial expression, etc.
[1428] "Cloud technology" is a technology that processes and stores data remotely via the Internet.
[1429] A "work environment" is an environment where a specific task is performed, such as a factory or office.
[1430] "Worker" refers to a person who performs a particular task.
[1431] The present invention provides a dialogue system that recognizes the emotional state of a worker in a factory environment and generates a dialogue response in response to the worker's emotional state. The system includes a dialogue response generation means using a generative model, a speech recognition means for converting a user's voice input into text data, an expression recognition means for analyzing the user's facial expression data, a means for adjusting the response of the generative model based on the emotion recognition means, and a means for processing, transmitting, and receiving data using cloud technology.
[1432] Explanation of program processing
[1433] Hardware and software used
[1434] Hardware:
[1435] Microphone: Used to collect the user's voice.
[1436] Camera: Used to capture the user's facial expressions.
[1437] software:
[1438] Speech Recognition: Uses Google Cloud Speech-to-Text to convert speech to text.
[1439] Facial Expression Recognition: Uses Microsoft Azure Face API to analyze the user's facial expressions.
[1440] Emotion Recognition: Uses IBM Watson Tone Analyzer to infer emotional state from vocal tone and facial expressions.
[1441] Generative model: Uses OpenAI GPT-4 API to generate dialogue responses.
[1442] Communication: Uses HTTP REST APIs to send and receive data.
[1443] Specific examples
[1444] When a factory worker speaks to the device, the device collects the voice using a microphone and converts the voice into text data using Google Cloud Speech-to-Text. At the same time, the device captures the worker's facial expressions using a camera and analyzes the facial expression data using the Microsoft Azure Face API to estimate their emotional state. Next, the device analyzes the voice tone using IBM Watson Tone Analyzer to generate the final emotional data.
[1445] The server receives this data and uses the OpenAI GPT-4 API to generate an appropriate response, taking into account the emotional data and adjusting the content and tone of the response. The generated response is then provided to the user via their device in voice and text format.
[1446] Example prompt sentence:
[1447] User input: "I've been having trouble getting things done lately." (sad tone)
[1448] Sentiment analysis result: "Sad"
[1449] Corresponding response: "That's tough. How about you take a break?"
[1450] This system enables natural dialogue responses that respond to the emotions of workers in a factory environment, which is expected to improve the working environment and increase efficiency. Furthermore, by allowing technicians to provide appropriate support based on their emotions, it is possible to increase worker satisfaction and overall productivity.
[1451] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1452] Step 1:
[1453] The device uses a microphone to collect the user's voice input, which can be expressed as text, such as "I'm having trouble getting my work done lately." The collected voice input is saved as raw audio data.
[1454] Step 2:
[1455] The device converts the collected voice data into text data using Google Cloud Speech-to-Text. The input is voice data, and after analyzing the voice data, it outputs the text data, "I've been having trouble with my work lately."
[1456] Step 3:
[1457] The device captures the user's facial expressions with a camera. The captured image is saved as raw image data, which is used as input for facial recognition.
[1458] Step 4:
[1459] The device uses the Microsoft Azure Face API to analyze the user's facial expressions from the captured image data. The input is image data, and the output is a quantitative emotional state (e.g., sadness, joy, etc.) representing the facial expression data.
[1460] Step 5:
[1461] The device uses IBM Watson Tone Analyzer to analyze the voice tone from the text data obtained by Google Cloud Speech-to-Text. The input is the text data, and the output is the tone analysis result (e.g., sad, wonderful, etc.).
[1462] Step 6:
[1463] The device integrates the results of facial expression analysis and voice tone analysis to generate the final emotion data. The input is quantitative facial expression data and voice tone data, and the output is the integrated emotional state (e.g., sad).
[1464] Step 7:
[1465] The device organizes the generated text data and emotion data into an appropriate format, such as JSON, and sends it to the server via an HTTP request. The input is the text data and emotion data, and the output is an HTTP request sent to the server.
[1466] Step 8:
[1467] The server analyzes the text data and emotion data received from the device and generates a response using the OpenAI GPT-4 API. The input is text data and emotion data, and it outputs an appropriate response text that takes the emotion data into consideration.
[1468] Step 9:
[1469] The server sends the generated response text to the terminal as an HTTP response. The input is the response text obtained from the generative model, and the output is the HTTP response sent to the terminal.
[1470] Step 10:
[1471] The device converts the received response text into audio data using Google Cloud Text-to-Speech. The input is the response text, and the output is audio data.
[1472] Step 11:
[1473] The terminal plays the generated voice using an audio output device, and simultaneously displays the response content as text on a display if necessary. The input is the voice data and the response text, and the output is the voice and text display to the user.
[1474] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1475] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1476] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1477] [Fourth embodiment]
[1478] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1479] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1480] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1481] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1482] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1483] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1484] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1485] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1486] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1487] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1488] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1489] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1490] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1491] Program Overview
[1492] This invention realizes a system including a dialogue response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Furthermore, communication between the generative model and the user is performed using cloud technology.
[1493] Explanation of program processing
[1494] 1. Collect user input:
[1495] User: Talk to the communication robot, for example, "What's the weather like today?"
[1496] Device: Uses a microphone to capture the user's voice.
[1497] Device: Uses speech recognition software to convert speech into text data.
[1498] 2. Sending user input:
[1499] Terminal: Organize text data into a suitable format (e.g. JSON).
[1500] Terminal: Sends text data to the server via an HTTP request.
[1501] 3. Generating dialogue content using a generative model:
[1502] Server: Receives and parses the text data.
[1503] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[1504] Generative Model: Receives requests and generates appropriate responses.
[1505] Generative model: Generates the answer and sends it back to the server as a response.
[1506] 4. Sending and processing the response:
[1507] Server: Receives responses from the generative AI model.
[1508] Server: Sends the received response to the terminal as an HTTP response.
[1509] 5. Response to the user:
[1510] Terminal: Parses the response received from the server.
[1511] Terminal: Passes the parsed response data to speech synthesis software to convert text to speech.
[1512] Terminal: Plays the generated audio using an audio output device (speaker).
[1513] Terminal: If necessary, the response will be displayed as text on the display.
[1514] Specific examples
[1515] For the education sector:
[1516] 1. User (Student): "I don't understand this math problem. How do I solve it?"
[1517] Terminal: Converts student voice into text data and sends it to the server.
[1518] Server: Sends text data to the generative model and generates an appropriate explanation.
[1519] Generative model: Generates explanations such as "First, try adding 2 to both sides of the equation. Then, as a next step..."
[1520] Server: Sends this response to the device.
[1521] Terminal: The robot outputs the explanation aloud and also displays it on the screen.
[1522] For the medical and nursing care sector:
[1523] 1. User (elderly): "I'm not feeling very well today."
[1524] Terminal: Converts the elderly person's voice into text data and sends it to the server.
[1525] Server: Feeds text data into the generative model and generates an appropriate response.
[1526] Generative model: Generate a response such as, "That's worrying. Let's take a few deep breaths together, and then it might be a good idea to talk to your doctor."
[1527] Server: Sends this response to the device.
[1528] Terminal: Have the robot output the response content by voice.
[1529] The system provides users with a smooth and personalized experience. By combining generative models and cloud technology, the communication robot achieves natural and effective interactions in a wide range of scenarios.
[1530] The processing flow will be explained below.
[1531] Step 1:
[1532] The user speaks to the communication robot (e.g., "What's the weather like today?"). When voice input begins, the robot's terminal uses a microphone to capture the user's voice.
[1533] Step 2:
[1534] The device converts the captured audio into text data using built-in speech recognition software, which uses a speech analysis algorithm to analyze the characteristics of the voice and convert it into text.
[1535] Step 3:
[1536] The terminal organizes the converted text data into an appropriate format, such as JSON, which also includes the user's ID and session information.
[1537] Step 4:
[1538] The device sends the organized text data to the server via an HTTP request, which includes the text data along with user information and related context information.
[1539] Step 5:
[1540] The server parses the received text data, checks the text data, and performs any necessary preprocessing before sending it to the generative model.
[1541] Step 6:
[1542] The server sends the parsed text data as a request to the API endpoint of the generative model, which includes the text data entered by the user.
[1543] Step 7:
[1544] The generative model analyzes the incoming request and generates an appropriate response. The generation process uses natural language processing algorithms to generate an appropriate response based on the context.
[1545] Step 8:
[1546] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[1547] Step 9:
[1548] The server receives the response from the generative model and parses it again to ensure that the response is well formed.
[1549] Step 10:
[1550] The server sends the received response data to the terminal as an HTTP response, which includes the generated response text data.
[1551] Step 11:
[1552] The device parses the response data received from the server and converts the text data into a format suitable for voice output or display.
[1553] Step 12:
[1554] The device passes the parsed response data to speech synthesis software, which converts the text into speech. The speech synthesis process converts the text data into speech waveforms.
[1555] Step 13:
[1556] The terminal plays the generated voice using the audio output device (speaker), and simultaneously displays the response text on the display if necessary.
[1557] This series of processing steps enables the user to have a natural and effective dialogue with the communication robot.
[1558] Example 1
[1559] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1560] Conventional interactive response systems have had difficulty efficiently collecting user voice input, converting it into appropriate text data, and generating responses. Furthermore, there are limited means for providing users with generated responses in natural-sounding voices. Furthermore, there is a need to effectively integrate generative AI models with cloud technology to realize personalized educational support and other applications.
[1561] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1562] In this invention, the server includes means for capturing voice input from a user, means for converting the voice input into text data, means for organizing the text data into an appropriate data format, means for sending the organized text data to a generative model, means for generating a dialogue response using the generative model, means for providing the generated dialogue response to the user, and means for converting the generated dialogue response into voice data, and includes output means for playing back the voice data. This makes it possible to generate smooth and natural dialogue responses using a generative AI model from the user's voice input and provide them in voice.
[1563] "User" refers to a person who uses the system to provide voice input.
[1564] "Voice input" refers to voice information spoken by a user to the system.
[1565] "Capturing means" refers to a device or software used to collect and record a user's voice input.
[1566] "Means for converting into text data" refers to software or algorithms used to analyze and convert collected voice input into corresponding text data.
[1567] "Means for organizing data" refers to the methods or mechanisms for preparing text data in an appropriate format, such as JSON, for transmission to a generative model.
[1568] "Means for sending" refers to a communication means for sending organized text data to a location where a generative model operates.
[1569] A "generative model" refers to an algorithm or software that generates appropriate dialogue responses based on text data.
[1570] "Means for generating a dialogue response" refers to a process of using a generative model to generate a response corresponding to an input.
[1571] "Means for providing" refers to a method for passing the generated interactive response to the user.
[1572] "Means for converting into voice data" refers to a process for converting the generated textual dialogue response into voice format.
[1573] "Output means" refers to a device or mechanism for playing back the converted audio data.
[1574] "Cloud technology" refers to computing resources and services provided over the internet.
[1575] "Educational support" refers to functions and services that provide individual learning support and guidance in the field of education.
[1576] The present invention is a dialogue response generation system using a generation model, and specifically includes the following means.
[1577] Collecting User Input
[1578] The user speaks to the communication robot. A specific example of a question is, "How's the weather today?" The device uses a microphone to capture the user's voice input. It then uses voice recognition software (e.g., a commonly used voice recognition engine) to convert the voice input into text data. For example, the voice "How's the weather today?" is converted into text "How's the weather today?"
[1579] Sending User Input
[1580] The device organizes the converted text data into an appropriate data format, such as JSON. The organized text data is sent to the server via an HTTP request. This request includes the user input, "What's the weather like today?"
[1581] Generating dialogue content using generative models
[1582] The server processes the HTTP request received from the device and extracts the text data. The extracted text data is sent as a request to the API endpoint of a generative model (for example, a commonly used generative AI model). The generative model receives the request and generates an appropriate response, such as "It's sunny today." This generated response is then sent back to the server.
[1583] Sending and Processing the Response
[1584] The server receives the response from the generative model and sends it to the device as an HTTP response. The received response may be data in the form of, for example, "It's sunny today."
[1585] Responding to the user
[1586] The device analyzes the HTTP response received from the server. It uses speech synthesis software (for example, a commonly used speech synthesis engine) to convert the text data into voice data. It then plays the generated voice using a speaker and responds to the user by saying, "It's a sunny day today." It also displays the response text on the display as needed.
[1587] Specific examples
[1588] Specific examples in the field of education
[1589] User (Student): "I don't understand this math problem. How do I solve it?"
[1590] The terminal converts the student's voice into text data and sends it to the server.
[1591] The server feeds the text data into a generative model to generate an appropriate explanation.
[1592] The generative model generates an explanation such as, "First, try adding 2 to both sides of the equation. Then, as a next step..."
[1593] The server sends this response to the terminal.
[1594] The device will then audibly transmit the generated explanation to the student and also display it on the screen.
[1595] Specific examples in the medical and nursing care fields
[1596] Elderly user: "I'm not feeling too great today."
[1597] The terminal converts the elderly person's voice into text data and sends it to the server.
[1598] The server feeds the text data into the generative model to generate an appropriate response.
[1599] The generative model generates a response such as, "That's worrying. Let's take a deep breath together. Then it might be a good idea to consult a doctor."
[1600] The server sends this response to the terminal.
[1601] The terminal then transmits the generated response to the elderly person by voice.
[1602] This invention leverages the collaboration between generative models and cloud technology to provide users with a smooth and personalized interaction experience. By combining generative AI models with speech recognition and speech synthesis technologies, natural-sounding interactions are possible in a wide range of scenarios.
[1603] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1604] Step 1:
[1605] A user speaks to a communication robot. For example, "How's the weather today?" The device uses a microphone to capture the user's voice input. The input voice information is acquired as an analog signal. The device converts this analog signal into a digital signal and passes it to the voice recognition software. The voice input is "How's the weather today?" and the output is a digital voice signal.
[1606] Step 2:
[1607] The device uses speech recognition software (e.g., a commonly used speech recognition engine) to convert speech into text data. During this process, a digital speech signal is analyzed and converted into corresponding text data. The input is a digital speech signal, and the output is text data such as "What's the weather like today?". Specifically, the speech signal is analyzed based on an acoustic model and a language model.
[1608] Step 3:
[1609] The terminal organizes the converted text data into JSON format. For example, the text data "What's the weather like today?" is converted into the format {"input": "What's the weather like today?"}. The input is text data, and the output is JSON format data. Specifically, the text data is structured with appropriate key-value pairs.
[1610] Step 4:
[1611] The terminal sends organized JSON-formatted text data to the server via an HTTP POST request. The input is JSON-formatted data, and the output is an HTTP request sent to the server. Specifically, an HTTP request containing a URL and header information is generated.
[1612] Step 5:
[1613] The server processes the HTTP request received from the terminal and extracts the text data. The input is the HTTP request, and the output is the extracted text data (for example, "What's the weather like today?"). Specifically, the required data is parsed from the request body.
[1614] Step 6:
[1615] The server sends the extracted text data as a request to the API endpoint of a generative model (e.g., a commonly used generative AI model). The input is the extracted text data, and the output is the request sent to the generative model. The specific operation is to call the appropriate API endpoint.
[1616] Step 7:
[1617] A generative AI model generates an appropriate response based on the request it receives. For example, the input "What's the weather like today?" generates the response "It's sunny today." The input is the request to the generative model, and the output is the generated response (e.g., "It's sunny today"). Specifically, the generative algorithm performs the text analysis and generation process.
[1618] Step 8:
[1619] The generative model returns the generated response to the server as a response. The input is the generated response, and the output is the response returned to the server. As a specific operation, the response data is returned in an appropriate format such as JSON format.
[1620] Step 9:
[1621] The server receives the response from the generative model and sends it to the terminal as an HTTP response. The input is the response from the generative model, and the output is the HTTP response sent to the terminal. In concrete terms, the response data is formatted as an HTTP response.
[1622] Step 10:
[1623] The terminal analyzes the HTTP response received from the server and obtains the text data. The input is the HTTP response, and the output is the text data (for example, "It's a sunny day today"). Specifically, the data is extracted from the response body.
[1624] Step 11:
[1625] The terminal passes the acquired text data to speech synthesis software (for example, a commonly used speech synthesis engine) and converts the text data into speech data. The input is text data and the output is speech data. Specifically, the speech synthesis engine synthesizes the text data to generate a speech signal.
[1626] Step 12:
[1627] The device uses a speaker to play back the generated voice. For example, it may output the voice "It's a sunny day today." The input is voice data, and the output is a voice response to the user. Specifically, the speaker converts the voice signal into physical sound, which the user hears. If necessary, the response is also displayed as text on the display.
[1628] (Application example 1)
[1629] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1630] Conventional food delivery systems have the drawback of requiring users to operate complicated applications when placing an order, and the convenience of voice control is particularly insufficient. Furthermore, there is a lack of systems that utilize advanced dialogue functions to provide optimal suggestions to users, making it difficult to achieve seamless dialogue with users. Therefore, there is a demand for a system that allows users to easily order food delivery using voice commands.
[1631] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1632] In this invention, the server includes an interactive response generation means using a generative model, an output means for providing a user with a response generated by the generative model, a speech recognition means for converting a user's voice input into text data, a transmission means for transmitting the text data to the generative model, a speech synthesis means for converting the generated response into a speech output, and an order processing means for processing a user's order. This allows a user to seamlessly order food delivery using only voice input, and the generative model makes optimal suggestions, making it possible to provide a more convenient service.
[1633] "Means for generating dialogue responses using a generative model" refers to a function that executes a generative AI model to generate appropriate responses based on input from a user.
[1634] "Output means" refers to a device or function for providing a response generated by a generative model to a user.
[1635] "Speech recognition means" refers to technology or devices for converting a user's voice input into text data.
[1636] "Transmission means" refers to a communication means for transmitting the text data to the generative model.
[1637] "Speech synthesis means" refers to technology or devices for converting text responses generated by a generative model into speech.
[1638] "Order Processing Means" refers to the systems and functions used to process user orders and facilitate the actual food delivery process.
[1639] "Cloud technology" refers to remote computing resources and services used over the internet to communicate between generative models and users.
[1640] The "food delivery sector" refers to industries and activities related to the delivery of food and beverages to specific locations upon request.
[1641] System Overview
[1642] The present invention provides a system that allows users to place food delivery orders by voice. The system converts the user's voice input into text data, sends the data to a generative model to generate a dialogue response, and provides the response to the user as voice using a speech synthesis means. The system also processes the user's order and uses cloud technology to perform these communications.
[1643] Hardware and software used
[1644] This system uses the following hardware and software:
[1645] Microphone: Used to capture the user's voice.
[1646] Speech Recognition Software: Uses the Google Speech-to-Text API to convert voice input into text data.
[1647] Generative model: Uses the OpenAI GPT-3 API to generate responses based on user text input.
[1648] Speech synthesis software: Uses the Google Text-to-Speech API to convert text responses from the generative model into audio.
[1649] Cloud server: Sends and receives data using cloud technologies such as AWS EC2.
[1650] Data format: Organize and send communication data in JSON format.
[1651] Processing flow
[1652] 1. Collecting user voice input
[1653] The user speaks to the smartphone, for example, "What's your recommended pizza today?"
[1654] The smartphone's microphone captures the user's voice.
[1655] Speech recognition software (Google Speech-to-Text API) converts the speech into text data.
[1656] 2. Sending text data
[1657] The smartphone organizes the text data into the appropriate format (JSON).
[1658] Send text data to the cloud server via an HTTP request.
[1659] 3. Generating dialogue content using a generative model
[1660] The cloud server receives the text data and sends it to the endpoint of the generative model (OpenAI GPT-3 API).
[1661] The generative model receives the request, generates an appropriate response, and sends the response back to the cloud server.
[1662] 4. Response to the user
[1663] The cloud server receives the response from the generative model and sends it to the smartphone as an HTTP response.
[1664] The response received by the smartphone is converted into speech using speech synthesis software (Google Text-to-Speech API) and played through the speaker.
[1665] 5. Order Processing
[1666] The user receives the response and responds audibly, "Yes, I'll order that."
[1667] The order is confirmed through a similar process and the food delivery process is carried out through the order processing means.
[1668] Examples of specific examples and prompts
[1669] Sample prompt 1: "What's your recommended pizza today?"
[1670] Sample prompt 2: "What dessert would you recommend?"
[1671] Prompt 3: "I'll order that."
[1672] Specifically, when a user speaks to their smartphone, "What's the recommended pizza today?", the speech recognition software converts the speech into text and sends it to a cloud server. Based on the text, the generative model generates a response, "Today's recommended pizza is Margherita. Would you like to order?", which is returned to the smartphone, converted into speech by speech synthesis software, and played back to the user. When the user replies, "I'll order," the food delivery procedure is completed by the order processing means.
[1673] In this way, a system is created that allows users to seamlessly order food delivery while engaging in natural interactions via their smartphones.
[1674] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1675] Step 1: Collecting voice input
[1676] User: Talks to their smartphone to order food delivery, for example, "What's your pizza special today?"
[1677] Device: The smartphone's microphone captures the user's voice, which is then fed into voice recognition software.
[1678] On the device: Uses the Google Speech-to-Text API to convert voice data to text data. The input is voice data and the output is text data.
[1679] Terminal: Receives the converted text data, which will be used in the next step.
[1680] Step 2: Send text data
[1681] Terminal: Organizes text data into JSON format. Input is text data, output is JSON format data.
[1682] The device sends the organized JSON data via an HTTP request to the cloud server, which then sends this data to the generative model.
[1683] Step 3: Generating dialogue content using a generative model
[1684] Server: The cloud server receives the text data and sends a request to the API endpoint of the generative model. The input is JSON data, and the output is a request to the generative model.
[1685] Generative Model: The OpenAI GPT-3 API receives requests and generates appropriate responses. The input is the request data and the output is the response data.
[1686] Server: The cloud server receives the generated response data and sends it to the terminal as an HTTP response. The input is the response data, and the output is the HTTP response.
[1687] Step 4: Respond to the user
[1688] Terminal: The smartphone receives the HTTP response from the cloud server. The input is the HTTP response, and the output is the response text data.
[1689] Terminal: Convert the received response data into speech using the Google Text-to-Speech API. The input is the response text data, and the output is the speech data.
[1690] Terminal: Plays voice responses to the user through a speaker and optionally displays the responses on a display. Input is voice data, and output is voice and text display.
[1691] Step 5: Order Processing
[1692] User: In response to the response, responds verbally with "Yes, I'll order that."
[1693] Terminal: The user's response is processed in a similar process (steps 1 to 4). Finally, the order processing means processes the food delivery. Specifically, it stores the order details in a database, notifies the delivery service, and provides the user with an order confirmation.
[1694] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1695] Program Overview
[1696] The present invention aims to make dialogue with a user more natural and effective by combining an emotion engine with a system equipped with a dialogue response generation means using a generative model. The system includes an output means for generating a response based on the generative model and providing the response to the user, a speech recognition means for converting the user's voice input into text data, and a transmission means for transmitting the text data to the generative model. Cloud technology is also used for communication between the generative model and the user. Furthermore, the emotion engine is used to recognize the user's emotional state and adjust the response of the generative model based on the recognition result.
[1697] Explanation of program processing
[1698] 1. Collect user input:
[1699] User: Talks to the communication robot (e.g., "I'm not feeling very good today.").
[1700] Device: Uses a microphone and camera to capture the user's voice and facial expressions.
[1701] Device: Uses speech recognition software to convert speech into text data.
[1702] Device: The user's facial expressions are analyzed using facial expression analysis software from the camera footage, and emotional data is extracted.
[1703] 2. Emotional Data Processing:
[1704] Device: The extracted voice and facial expression data is passed to an emotion recognition engine to estimate the user's emotional state.
[1705] Emotion recognition engine: Analyzes voice characteristics such as tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[1706] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[1707] 3. Sending user input and emotion data:
[1708] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[1709] Terminal: Sends data to the server via an HTTP request.
[1710] 4. Generating dialogue content using a generative model:
[1711] Server: Receives and parses text data and emotion data.
[1712] Server: Sends the parsed data as a request to the API endpoint of the generative model.
[1713] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[1714] Generative model: Generates the answer and sends it back to the server as a response.
[1715] 5. Sending and processing the response:
[1716] Server: Receives the response from the generative model.
[1717] Server: Sends the received response to the terminal as an HTTP response.
[1718] 6. Response to the user:
[1719] Terminal: Parses the response received from the server.
[1720] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[1721] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[1722] Specific examples
[1723] For the education sector:
[1724] Student User: "I don't understand this math problem," says in a confused tone.
[1725] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[1726] Server: Sends text data and confusion emotion data to the generative model.
[1727] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[1728] Device: Provides voice and text responses to students.
[1729] For the medical and nursing care sector:
[1730] Elderly user: "I'm not feeling very well today," says sadly.
[1731] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[1732] Server: Sends text data and sadness emotion data to the generative model.
[1733] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[1734] Device: Provides voice and text responses to seniors.
[1735] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[1736] The processing flow will be explained below.
[1737] Step 1:
[1738] The user speaks to the communication robot (e.g., "I'm not feeling very well today."). When voice input begins, the robot's terminal uses a microphone and camera to capture the user's voice and facial expressions.
[1739] Step 2:
[1740] The device converts the captured voice into text data using built-in voice recognition software, while simultaneously extracting the user's facial expression data from the camera footage using facial expression analysis software.
[1741] Step 3:
[1742] The device passes the text data and facial expression data to an emotion recognition engine to estimate the user's emotional state. The emotion recognition engine analyzes the voice tone, volume, and facial expression data to identify the user's emotions (e.g., joy, sadness, anger, etc.).
[1743] Step 4:
[1744] The device collects the emotion data obtained from the emotion recognition engine along with the text data and organizes it into an appropriate format such as JSON, which also includes the user's ID and session information.
[1745] Step 5:
[1746] The device sends text data and emotion data to the server via an HTTP request, which includes the text data, emotion data, and related user information.
[1747] Step 6:
[1748] The server parses the received text and emotion data, checks them, and performs any necessary preprocessing before sending them to the generative model.
[1749] Step 7:
[1750] The server sends the parsed text data and emotion data as a request to the API endpoint of the generative model. This request includes the text data and emotion data entered by the user.
[1751] Step 8:
[1752] The generative model analyzes the received request and generates an appropriate response, taking into account the emotional data to generate a response that is appropriate for the user's emotional state.
[1753] Step 9:
[1754] The generative model returns the generated response to the server as a response, which includes the text data generated by the generative model.
[1755] Step 10:
[1756] The server receives the response from the generative model and then sends the received response data to the terminal as an HTTP response.
[1757] Step 11:
[1758] The device parses (analyzes) the response data received from the server, converting the text data into a format suitable for voice output or display.
[1759] Step 12:
[1760] The device passes the parsed response data to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that takes into account the user's emotional state.
[1761] Step 13:
[1762] The terminal plays back the generated voice using the audio output device (speaker) and, if necessary, displays the response text on the display, so that the user receives appropriate feedback.
[1763] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[1764] Example 2
[1765] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1766] In conventional dialogue systems, the dialogue with the user is one-sided, making it difficult to generate responses that take the user's emotional state into consideration. In particular, in the fields of education, medicine, and nursing care, there are many situations where responses that correspond to the user's emotions are required, and conventional technologies have not always been satisfactory.
[1767] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an output means for providing the user with a response generated by the generative model, a voice recognition means for converting the user's voice input into text data, a transmission means for transmitting the text data to the generative model, an expression analysis means for analyzing the user's facial expression, and an emotion recognition means for estimating the user's emotional state from the user's voice and facial expression. This makes it possible to realize a natural dialogue according to the user's emotional state.
[1768] A "generative model" is an algorithm that generates natural language responses or text based on given input data.
[1769] A "dialogue response generator" is a device or method that uses a generative model to generate a response to an input from a user.
[1770] An "output means" is a device or method for providing the generated response to a user.
[1771] "Speech recognition means" refers to technology or devices for analyzing a user's voice input and converting it into text data.
[1772] "Transmission means" refers to the communication means or protocol for transmitting text data to the generative model.
[1773] The "facial expression analysis means" refers to a technique or device that analyzes the user's facial expression from camera footage and acquires that information.
[1774] "Emotion recognition means" refers to technology or devices for estimating a user's emotional state based on voice and facial expression data.
[1775] "Cloud technology" is a technology that stores data on remote servers via the Internet and performs computational processing.
[1776] "Individual educational support" refers to a method or system that provides educational support tailored to each student's learning situation and level of understanding.
[1777] This invention is a system that combines a generative model and an emotion recognition engine to make interactions with users more natural and effective. The configuration and operation of this system will be described in detail below.
[1778] 1. System Configuration
[1779] The system consists of the following main components:
[1780] Generative model: Uses natural language processing algorithms to generate responses in response to user input.
[1781] Dialogue response generator: Utilizes a generative model to create appropriate responses to user input.
[1782] Output: A device (e.g., speaker, display, etc.) that provides the generated response to the user.
[1783] Speech recognition tool: Software that converts user voice input into text data (e.g., Google Speech-to-Text).
[1784] Transmission medium: The communication protocol (e.g., HTTP) used to send text data to the generative model.
[1785] Facial expression analysis means: Software for analyzing the user's facial expressions from camera footage and extracting emotional data (e.g., Amazon Rekognition).
[1786] Emotion recognizer: An engine for inferring the user's emotional state from their voice and facial expression data (e.g., IBM Watson Tone Analyzer).
[1787] Cloud technology: Technology for performing computational processing on remote servers via the Internet.
[1788] 2. Hardware and Software Configuration
[1789] Hardware:
[1790] Microphone: A device for capturing the user's voice.
[1791] Camera: A device for capturing the user's facial expressions.
[1792] Speaker: A device for providing generated audio responses to a user.
[1793] Display: A device for displaying the generated text response.
[1794] software:
[1795] Speech recognition software: converts speech into text data (e.g., Google Speech-to-Text).
[1796] Facial expression analysis software: Extracting emotional data from camera footage (e.g., Amazon Rekognition).
[1797] Emotion recognition engine: Estimates emotional state from voice tone and facial expressions (e.g., IBM Watson Tone Analyzer).
[1798] Generative model: An algorithm for generating responses based on input text data (e.g., the open-source GPT model).
[1799] 3. System Operation
[1800] When a user speaks to the communication robot, the speech recognition means converts the speech into text data. At the same time, the facial expression analysis means analyzes the user's facial expressions and extracts emotional data. This data is sent to the server via the transmission means, and the generative model generates an appropriate response. The generated response is sent to the user's device via cloud technology and converted into voice using speech synthesis software. The response is finally provided to the user through a speaker.
[1801] Specific examples
[1802] For the education sector:
[1803] User (Student): "I don't understand that math problem," says in a confused tone.
[1804] Device: Captures students' voices and facial expressions, converts the voices into text, and analyzes facial expressions to generate confusion emotion data.
[1805] Server: Sends text data and confusion emotion data to the generative model.
[1806] Generative model: Generates responses such as "Okay, let me explain it again," and delivers them in an emotionally sensitive tone.
[1807] Device: Provides voice and text responses to students.
[1808] For the medical and nursing care sector:
[1809] Elderly user: "I'm not feeling very well today," says sadly.
[1810] Device: Captures voice and facial expressions, converts voice to text, and generates sadness emotion data through facial expression analysis.
[1811] Server: Sends text data and sadness emotion data to the generative model.
[1812] Generative model: Generates a response such as "That's worrying. Let's relax a bit," delivered in a gentle tone.
[1813] Device: Provides voice and text responses to seniors.
[1814] Prompt Sentence Examples
[1815] Education: “When a struggling student says, ‘I don’t understand that math problem,’ how does a generative model respond?”
[1816] Healthcare: "How would you respond to an elderly person who says in a sad tone, 'I'm not feeling very well today'?"
[1817] This system makes it possible to realize natural dialogue based on the user's emotions, and is expected to be particularly applicable in education, medical care, and nursing care.
[1818] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1819] Step 1:
[1820] Collecting User Input
[1821] User: Talks to the communication robot, "I'm not feeling very good today."
[1822] Input: User's voice
[1823] Output: Raw audio file
[1824] Device: Uses a microphone to capture the user's voice.
[1825] What it does: A microphone records audio and converts it into digital data in real time.
[1826] Device: Uses the camera to capture the user's facial expressions.
[1827] Input: User's facial expression
[1828] Output: Facial expression video data
[1829] Specific operation: The camera captures the user's face and generates video data.
[1830] On your device: Use speech recognition software to convert voice data into text (e.g., Google Speech-to-Text).
[1831] Input: Raw audio file
[1832] Output: Text data
[1833] How it works: Recorded audio data is sent to speech recognition software, which analyzes the audio waveform and converts it into text.
[1834] On the device: Facial expression analysis software is used to analyze the user's facial expressions from the camera footage and extract emotional data (e.g., Amazon Rekognition).
[1835] Input: facial expression video data
[1836] Output: Emotion data
[1837] How it works: The captured image of the face is input into expression analysis software, which extracts facial features and estimates the emotional state based on them.
[1838] Step 2:
[1839] Emotional Data Processing
[1840] Device: The extracted voice and facial expression data is passed to an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to estimate the user's emotional state.
[1841] Input: Voice data and facial expression data
[1842] Output: Emotional state
[1843] What it does: Voice and facial expression data are sent to the emotion recognition engine, which then begins the analysis process, identifying the emotional state from changes in voice tone and facial expressions.
[1844] Emotion Recognition Engine: Analyzes voice tone, volume, and facial expressions to identify the user's emotions (e.g., joy, sadness, anger).
[1845] Input: Voice data and facial expression data
[1846] Output: Emotion label (e.g. sadness, anger)
[1847] What it does: Specific algorithms analyze audio and video data to generate emotion labels.
[1848] Terminal: The results of the emotion recognition engine are organized along with the text data and sent to the server.
[1849] Input: Emotion recognition engine results, text data
[1850] Output: Organized data (JSON format)
[1851] Step 3:
[1852] Sending user input and emotion data
[1853] Terminal: Organize the text data and emotion data into an appropriate format, such as JSON.
[1854] Input: Text data, emotion data
[1855] Output: JSON format data
[1856] Specific behavior: Text data and emotion data are packaged into a single JSON object.
[1857] Terminal: Sends the organized data to the server via an HTTP request.
[1858] Input: JSON format data
[1859] Output: HTTP request
[1860] What happens: An HTTP request is constructed and data is sent to the server endpoint.
[1861] Step 4:
[1862] Generating dialogue content using generative models
[1863] Server: Receives and parses text and emotion data.
[1864] Input: HTTP request data
[1865] Output: Parsed data
[1866] Specific operation: Parse the received JSON data and extract the necessary fields.
[1867] Server: Sends the parsed data as a request to the API endpoint of the generative model (e.g., the open-source GPT model).
[1868] Input: Parsed data
[1869] Output: API request
[1870] What happens: A request is sent to the generative model's API and the appropriate parameters are set.
[1871] Generative models analyze text and sentiment data to generate appropriate responses, adjusting the tone and content of responses based on sentiment data.
[1872] Input: API request data
[1873] Output: Response data
[1874] Specific operation: The model analyzes the text data and generates a response that takes into account the emotional state.
[1875] Generative model: Generates the answer and sends it back to the server as a response.
[1876] Input: Response data
[1877] Output: Response data (JSON format)
[1878] Specific behavior: Response data is sent to the server in JSON format.
[1879] Step 5:
[1880] Sending and Processing the Response
[1881] Server: Receives the response from the generative model.
[1882] Input: Response data
[1883] Output: Parsed response data
[1884] Specific operation: The server receives the response from the API and analyzes it.
[1885] Server: Sends the received response to the terminal as an HTTP response.
[1886] Input: Parsed response data
[1887] Output: HTTP response
[1888] Specific operation: The response data is organized and sent as an HTTP response to the device.
[1889] Step 6:
[1890] Responding to the user
[1891] Terminal: Parse the response received from the server.
[1892] Input: HTTP response data
[1893] Output: Parsed response data
[1894] Specific operation: Analyzes the received data and extracts the necessary information.
[1895] On the device: The parsed response data is passed to speech synthesis software, which converts the text into speech, generating speech with a tone and tempo that reflects the user's emotional state.
[1896] Input: Response data
[1897] Output: Audio data
[1898] Specific operation: Text is input into the speech synthesis software, and the speech generation process is executed according to the emotional state.
[1899] Terminal: Plays back the generated audio using an audio output device (speaker), and simultaneously displays the response text on a display if necessary.
[1900] Input: Audio data, text data
[1901] Output: Playback audio, display
[1902] Specific behavior: The speaker plays the generated audio and the response text appears on the display.
[1903] (Application example 2)
[1904] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1905] Dialogue response systems using generative models have the problem of being unable to flexibly respond to the user's emotional state, resulting in mechanical and unnatural dialogue. In particular, in factory environments, dialogue that ignores the worker's emotional state can lead to a deterioration in the working environment and a decrease in efficiency. Therefore, there is a need for technology that can recognize the worker's emotional state and adjust the dialogue content accordingly.
[1906] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a dialogue response generation means using a generative model, an expression recognition means for analyzing the user's expression data, and a means for adjusting the response of the generative model based on the emotion recognition means. This makes it possible to recognize the emotion of the worker and provide a natural and effective dialogue response according to that emotion.
[1907] A "generative model" is a model that generates text data using a machine learning algorithm.
[1908] The "interactive response generation means" is a function that uses a generative model to generate a response required for a dialogue with a user.
[1909] The "output means" is a function for providing the generated response to the user.
[1910] The "voice recognition means" is a function that converts the user's voice input into text data.
[1911] The "transmission means" is a function for transmitting text data to the generative model.
[1912] The "facial expression recognition means" is a function that analyzes the user's facial expression data and estimates the user's emotional state.
[1913] The "emotion recognition means" is a function that estimates the user's emotional state from the tone of voice, facial expression, etc.
[1914] "Cloud technology" is a technology that processes and stores data remotely via the Internet.
[1915] A "work environment" is an environment where a specific task is performed, such as a factory or office.
[1916] "Worker" refers to a person who performs a particular task.
[1917] The present invention provides a dialogue system that recognizes the emotional state of a worker in a factory environment and generates a dialogue response in response to the worker's emotional state. The system includes a dialogue response generation means using a generative model, a speech recognition means for converting a user's voice input into text data, an expression recognition means for analyzing the user's facial expression data, a means for adjusting the response of the generative model based on the emotion recognition means, and a means for processing, transmitting, and receiving data using cloud technology.
[1918] Explanation of program processing
[1919] Hardware and software used
[1920] Hardware:
[1921] Microphone: Used to collect the user's voice.
[1922] Camera: Used to capture the user's facial expressions.
[1923] software:
[1924] Speech Recognition: Uses Google Cloud Speech-to-Text to convert speech to text.
[1925] Facial Expression Recognition: Uses Microsoft Azure Face API to analyze the user's facial expressions.
[1926] Emotion Recognition: Uses IBM Watson Tone Analyzer to infer emotional state from vocal tone and facial expressions.
[1927] Generative model: Uses OpenAI GPT-4 API to generate dialogue responses.
[1928] Communication: Uses HTTP REST APIs to send and receive data.
[1929] Specific examples
[1930] When a factory worker speaks to the device, the device collects the voice using a microphone and converts the voice into text data using Google Cloud Speech-to-Text. At the same time, the device captures the worker's facial expressions using a camera and analyzes the facial expression data using the Microsoft Azure Face API to estimate their emotional state. Next, the device analyzes the voice tone using IBM Watson Tone Analyzer to generate the final emotional data.
[1931] The server receives this data and uses the OpenAI GPT-4 API to generate an appropriate response, taking into account the emotional data and adjusting the content and tone of the response. The generated response is then provided to the user via their device in voice and text format.
[1932] Example prompt sentence:
[1933] User input: "I've been having trouble getting things done lately." (sad tone)
[1934] Sentiment analysis result: "Sad"
[1935] Corresponding response: "That's tough. How about you take a break?"
[1936] This system enables natural dialogue responses that respond to the emotions of workers in a factory environment, which is expected to improve the working environment and increase efficiency. Furthermore, by allowing technicians to provide appropriate support based on their emotions, it is possible to increase worker satisfaction and overall productivity.
[1937] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1938] Step 1:
[1939] The device uses a microphone to collect the user's voice input, which can be expressed as text, such as "I'm having trouble getting my work done lately." The collected voice input is saved as raw audio data.
[1940] Step 2:
[1941] The device converts the collected voice data into text data using Google Cloud Speech-to-Text. The input is voice data, and after analyzing the voice data, it outputs the text data, "I've been having trouble with my work lately."
[1942] Step 3:
[1943] The device captures the user's facial expressions with a camera. The captured image is saved as raw image data, which is used as input for facial recognition.
[1944] Step 4:
[1945] The device uses the Microsoft Azure Face API to analyze the user's facial expressions from the captured image data. The input is image data, and the output is a quantitative emotional state (e.g., sadness, joy, etc.) representing the facial expression data.
[1946] Step 5:
[1947] The device uses IBM Watson Tone Analyzer to analyze the voice tone from the text data obtained by Google Cloud Speech-to-Text. The input is the text data, and the output is the tone analysis result (e.g., sad, wonderful, etc.).
[1948] Step 6:
[1949] The device integrates the results of facial expression analysis and voice tone analysis to generate the final emotion data. The input is quantitative facial expression data and voice tone data, and the output is the integrated emotional state (e.g., sad).
[1950] Step 7:
[1951] The device organizes the generated text data and emotion data into an appropriate format, such as JSON, and sends it to the server via an HTTP request. The input is the text data and emotion data, and the output is an HTTP request sent to the server.
[1952] Step 8:
[1953] The server analyzes the text data and emotion data received from the device and generates a response using the OpenAI GPT-4 API. The input is text data and emotion data, and it outputs an appropriate response text that takes the emotion data into consideration.
[1954] Step 9:
[1955] The server sends the generated response text to the terminal as an HTTP response. The input is the response text obtained from the generative model, and the output is the HTTP response sent to the terminal.
[1956] Step 10:
[1957] The device converts the received response text into audio data using Google Cloud Text-to-Speech. The input is the response text, and the output is audio data.
[1958] Step 11:
[1959] The terminal plays the generated voice using an audio output device, and simultaneously displays the response content as text on a display if necessary. The input is the voice data and the response text, and the output is the voice and text display to the user.
[1960] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1961] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1962] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1963] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1964] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1965] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1966] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1967] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1968] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1969] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1970] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1971] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1972] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1973] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1974] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1975] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1976] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1977] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1978] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1979] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1980] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1981] The following is further disclosed regarding the above embodiment.
[1982] (Claim 1)
[1983] A dialogue response generation means using a generative model;
[1984] an output means for providing a response generated by the generative model to a user;
[1985] a speech recognition means for converting a user's speech input into text data;
[1986] a transmitting means for transmitting the text data to a generative model.
[1987] (Claim 2)
[1988] 10. The system of claim 1, comprising means for communicating between the generative model and the user using cloud technology.
[1989] (Claim 3)
[1990] 10. The system of claim 1, comprising means for providing personalized educational support in the field of education.
[1991] (Claim 4)
[1992] 10. The system according to claim 1, comprising means for supporting dialogue with patients or elderly people in the medical and nursing care field.
[1993] "Example 1"
[1994] (Claim 1)
[1995] means for capturing voice input from a user;
[1996] means for converting the voice input into text data;
[1997] means for organizing the text data into an appropriate data format;
[1998] means for transmitting the prepared text data to a generative model;
[1999] means for generating a dialogue response using the generative model;
[2000] means for providing the generated interactive response to a user;
[2001] means for converting the generated dialogue response into voice data;
[2002] A system including an output means for playing audio data.
[2003] (Claim 2)
[2004] 10. The system of claim 1, comprising means for communicating between the generative model and the user using cloud technology.
[2005] (Claim 3)
[2006] 10. The system of claim 1, comprising means for providing personalized educational support in the field of education.
[2007] "Application Example 1"
[2008] (Claim 1)
[2009] A dialogue response generation means using a generative model;
[2010] an output means for providing a response generated by the generative model to a user;
[2011] a speech recognition means for converting a user's speech input into text data;
[2012] a transmitting means for transmitting the text data to a generative model;
[2013] speech synthesis means for converting the generated response into a speech output;
[2014] The system includes an order processing means for processing user orders.
[2015] (Claim 2)
[2016] 10. The system of claim 1, further comprising means for communicating between the generative model and the user using cloud technology.
[2017] (Claim 3)
[2018] The system according to claim 1, further comprising means for assisting users in placing orders in the food delivery field.
[2019] "Example 2: Combining Emotion Engines"
[2020] (Claim 1)
[2021] A dialogue response generation means using a generative model;
[2022] an output means for providing a response generated by the generative model to a user;
[2023] a speech recognition means for converting a user's speech input into text data;
[2024] a transmitting means for transmitting the text data to a generative model;
[2025] Facial expression analysis means for analyzing a facial expression of a user;
[2026] A system including an emotion recognition means for inferring an emotional state from a user's voice and facial expressions.
[2027] (Claim 2)
[2028] 10. The system of claim 1, comprising means for communicating between the generative model and the user using cloud technology.
[2029] (Claim 3)
[2030] 10. The system of claim 1, comprising means for providing personalized educational support in the field of education.
[2031] "Application example 2 when combining emotion engines"
[2032] (Claim 1)
[2033] A dialogue response generation means using a generative model;
[2034] an output means for providing a response generated by the generative model to a user;
[2035] a speech recognition means for converting a user's speech input into text data;
[2036] a transmitting means for transmitting the text data to a generative model;
[2037] facial expression recognition means for analyzing facial expression data of a user;
[2038] an emotion recognition means for estimating an emotional state by integrating the facial expression data and the text data;
[2039] The system includes a means for adjusting the response of the generative model based on the emotion recognition means.
[2040] (Claim 2)
[2041] 10. The system of claim 1, comprising means for communicating between the generative model and the user using cloud technology.
[2042] (Claim 3)
[2043] 10. The system of claim 1, further comprising means for providing emotion-based assistance to a worker in a work environment. [Explanation of symbols]
[2044] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A dialogue response generation means using a generative model; an output means for providing a response generated by the generative model to a user; a speech recognition means for converting a user's speech input into text data; a transmitting means for transmitting the text data to a generative model.
2. The system of claim 1 , further comprising means for communicating between the Generative Model and the user using cloud technology.
3. 10. The system of claim 1, comprising means for providing personalized educational support in the field of education.
4. The system according to claim 1, further comprising means for supporting dialogue with patients or elderly people in the medical and nursing care field.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A