System
A system using generative AI in physical vending machines addresses the undervaluation of information by allowing users to easily obtain high-quality information, promoting a sustainable information ecosystem.
Patent Information
- Application Number
- JP2024123854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
The culture of valuing and paying for high-quality information is lacking, especially among younger generations, leading to unfair compensation for information providers and limited access to reliable information without internet reliance.
A system that accepts user input, analyzes and converts it into text data, transmits it to a server using generative AI (GPT model) for response generation, and displays the response on a terminal, allowing users to purchase information services from physical vending machines.
Establishes a culture of recognizing information value and provides a sustainable information provision environment where providers receive fair compensation, enabling instant access to high-quality information without internet connection.
Smart Images

Figure 2026022337000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the culture of properly evaluating and paying for high-quality information has not yet taken root, and this tendency is particularly pronounced among younger generations. This situation prevents information providers from receiving fair compensation and hinders the creation of a sustainable information provision environment. Another problem is the limited means of instantly obtaining high-quality information without relying on the Internet. The challenge is to improve this current situation and establish a culture in which people of all ages recognize the value of information and pay for it. [Means for solving the problem]
[0005] The present invention provides a system including means for accepting information input from a user, means for analyzing the information input and converting it into text data, means for transmitting the text data to a server, means for receiving response data from the server, and means for displaying the response data to the user. Furthermore, the server generates response data based on the received text data using generative artificial intelligence (GPT model) and transmits the response data to the terminal, thereby enabling the user to instantly obtain high-quality information. The system also includes means for adjusting parameters to generate optimal responses based on the user's questions, thereby achieving more appropriate information provision. This provides an environment in which generative AI-based information services can be easily purchased from physical vending machines, realizing a sustainable information provision system in which information providers receive fair rewards.
[0006] "User" refers to the person who operates the system and inputs information.
[0007] "Means for accepting information input" refers to an interface or device for receiving input data from a user.
[0008] "Means for analyzing input information and converting it into textual data" refers to a process and device that analyzes input information and converts it into digital textual data.
[0009] The "means for transmitting to the server" refers to a communication means for transmitting the analyzed and converted text data to the server via a network.
[0010] "Server" refers to a computer system that receives and processes data over a network.
[0011] "Response data" refers to information that the server generates as a processing result and sends to the terminal.
[0012] The "means for receiving response data" refers to a communication means and device for receiving data sent from the server.
[0013] "Means for displaying to the user" refers to a device or interface for visually presenting the received response data to the user.
[0014] "Generative artificial intelligence (GPT model)" refers to an artificial intelligence model based on the Generative Pre-trained Transformer, and is particularly capable of generating text through natural language processing.
[0015] "Means for adjusting parameters" refers to the process and devices that change and optimize various values and settings that are set by the Generative AI to generate an optimal response. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. Specifically, it provides a mechanism in which the user inputs a question, the server generates a response to the question, and displays it to the user through the terminal.
[0038] First, a user inputs a question using a user interface such as a touch screen on a vending machine terminal, which can be specific, such as "What are some recommended travel destinations for my next holiday?"
[0039] The terminal analyzes the input question and converts it into text data, which is then sent to a server via a communication method such as the Internet.
[0040] The server generates response data using generative AI (GPT model) based on the received text data. The generated response data, such as "Kyoto is a recommended travel destination," is then sent back to the device via communication means.
[0041] The terminal displays the response data received from the server to the user, allowing the user to easily obtain the necessary information from the vending machine.
[0042] This system is expected to have the effect of establishing a culture among users that recognizes the value of information and pays for it. It will also contribute to the creation of a sustainable information provision environment in which information providers can receive fair compensation.
[0043] For example, consider the following user actions and system behaviors:
[0044] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0045] 2. The device converts the question into text data and sends it to the server.
[0046] 3. The server analyzes the received question using the GPT model and generates a response such as "Kyoto is a recommended travel destination."
[0047] 4. The generated response is sent to the terminal, which displays the response to the user.
[0048] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0049] The unique feature of this system is that the service is provided through a physical vending machine, without the need for an internet connection. This has the advantage that users can easily obtain the information they need wherever they are, and it is thought that it can also be used as part of information literacy education, which is particularly useful for young people.
[0050] The processing flow will be explained below.
[0051] Step 1:
[0052] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0053] Step 2:
[0054] The terminal parses the input from the user and obtains the question data as a string, which is then converted into a text format.
[0055] Step 3:
[0056] The device prepares an HTTP POST request to send text-formatted question data to the server, specifying the destination URL and including the question data in the body of the HTTP request.
[0057] Step 4:
[0058] The device executes an HTTP POST request and sends the query data to the server.
[0059] Step 5:
[0060] The server receives the HTTP request sent from the terminal and extracts the question data from the request body.
[0061] Step 6:
[0062] The server passes the extracted question data to the generation AI (GPT model) and generates a response. At this time, the API key is used to send a request to the API endpoint of the GPT model.
[0063] Step 7:
[0064] The server receives the response data returned from the GPT model. For example, it receives the text "Kyoto is a recommended travel destination."
[0065] Step 8:
[0066] The server formats the response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[0067] Step 9:
[0068] The server sends the formatted response data to the terminal.
[0069] Step 10:
[0070] The terminal receives the HTTP response from the server and analyzes the response data.
[0071] Step 11:
[0072] The terminal displays the analyzed response data on the user interface.
[0073] Step 12:
[0074] The user can check the screen of the device and see the displayed response, "Kyoto is a recommended travel destination."
[0075] Example 1
[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0077] In modern society, much information is digitized, and users want to obtain the information they need quickly and easily. However, conventional information acquisition methods have problems such as the need for an Internet connection and uncertainty about the accuracy of the information. The purpose of this invention is to solve these problems and build a system that quickly provides highly reliable information.
[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0079] In this invention, the server includes means for generating response data using a generative AI model based on received text data, means for transmitting the text data to the server using a communication means, and means for displaying the response data on a user interface, thereby making it possible to provide advanced information in real time by utilizing the generative AI model.
[0080] "Means for accepting information input" refers to an interface for a user to input information and a mechanism for receiving that information.
[0081] "Means for converting into text format data" refers to a process for recognizing information entered by a user as a digital string and converting it into text data.
[0082] "Means for transmitting to a server using a communication means" refers to infrastructure and software for transmitting text-format data to a server via a communication network such as the Internet.
[0083] The "means for receiving response data from the server" refers to a communication protocol and hardware for receiving response data sent from the server.
[0084] "Means for displaying on a user interface" refers to a display and display software for displaying received response data in a form that can be viewed by a user.
[0085] "Generative artificial intelligence model" refers to the machine learning model and associated algorithms used to generate appropriate responses based on received text data.
[0086] "Means for adjusting the algorithm" refers to the process of dynamically changing and tuning input parameters and processing methods so that the generative artificial intelligence model can output the optimal response.
[0087] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. The system uses a terminal, a server, and a generative AI model to provide users with efficient and accurate information.
[0088] First, the user uses the touchscreen of the vending machine terminal to input a question, such as the prompt "Where would you recommend for my next holiday?" The terminal is equipped with a high-precision touchscreen and software to process the input.
[0089] The terminal converts the information entered by the user into text data using character recognition software to convert the input string into digital data. The converted text data is then sent to a server via a communication method (e.g., an internet connection). The communication protocol used is HTTP or HTTPS.
[0090] The server analyzes the received text data and generates response data using a generative AI model (for example, OpenAI's GPT-4). The server has high-performance computing resources and can run Node.js or Python-based application servers. As a concrete example, the response generated is "Kyoto is a recommended travel destination."
[0091] The generated response data is then sent back to the terminal via the communication means. The terminal analyzes the received response data and displays it on the user interface. The display typically uses an LCD panel or an organic EL panel. The user can check this and obtain the necessary information.
[0092] This system allows users to obtain fast and accurate information through a physical vending machine. Its unique feature is that it does not require an internet connection, making it possible to conveniently obtain information anytime, anywhere. Furthermore, by utilizing generative AI models, advanced information is provided in real time, improving user satisfaction.
[0093] The specific operation flow is as follows:
[0094] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0095] 2. The device converts the question into text data and sends it to the server.
[0096] 3. The server analyzes the received question using a generative AI model and generates a response such as, "Kyoto is a recommended travel destination."
[0097] 4. The generated response is sent to the terminal, which displays the response to the user.
[0098] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0099] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0100] Step 1:
[0101] The user operates the touch screen of the vending machine terminal and inputs a question.
[0102] Input: A question that the user types into the touchscreen (e.g., "What are your recommended travel destinations for my next holiday?").
[0103] Output: Text data recognized by the device.
[0104] How it works: The user types "What are some recommended travel destinations for my next holiday?" into the touchscreen, and the input information is captured by the touchscreen's sensors.
[0105] Step 2:
[0106] The terminal converts the question entered by the user into text data.
[0107] Input: Touchscreen input data.
[0108] Output: Data in text format (digital string).
[0109] Specific operation: The device's internal software analyzes the input data obtained from the sensor and converts it into text data in UTF-8 format.
[0110] Step 3:
[0111] The terminal transmits the converted text data to the server using a communication means.
[0112] Input: Data in text format.
[0113] Output: The HTTP POST request sent to the server.
[0114] Specific operation: An HTTP POST request is generated from the terminal to the server and sent to the server via the Internet.
[0115] Step 4:
[0116] The server analyzes the received text data and generates response data using a generative artificial intelligence model.
[0117] Input: Text data to be included in the HTTP POST request.
[0118] Output: The generated response data.
[0119] How it works: The server uses a Node.js or Python-based application server to send text data to the GPT-4 API and generate an appropriate response.
[0120] Step 5:
[0121] The server transmits the generated response data to the terminal using the communication means.
[0122] Input: The generated response data.
[0123] Output: The response data as an HTTP response.
[0124] Specific operation: The server generates response data in JSON format as an HTTP response and sends it to the terminal.
[0125] Step 6:
[0126] The terminal analyzes the response data received from the server and displays it on the user interface.
[0127] Input: The HTTP response data sent by the server.
[0128] Output: The text that appears in the user interface.
[0129] Specific operation: The device's software analyzes the HTTP response and displays "Kyoto is a recommended travel destination" on the touchscreen.
[0130] Step 7:
[0131] The user checks and acquires the information displayed on the screen of the terminal.
[0132] Input: The response data displayed in the user interface.
[0133] Output: The information obtained by the user.
[0134] Specific operation: The user visually recognizes the information displayed on the touch screen, "Kyoto is a recommended travel destination," and obtains the necessary information.
[0135] (Application example 1)
[0136] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0137] In conventional information provision systems using vending machines, users had to use a physical keyboard or touch screen to input questions, which made the input process cumbersome. In addition, there were limitations to the ways in which users could visually obtain information, so a more intuitive and faster way to obtain information was needed.
[0138] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0139] In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text data, means for transmitting the text data to the server, means for receiving response data from the server, means for displaying the response data to the user, speech recognition means for accepting voice input and converting it into text data, and means for visually displaying the response data, thereby enabling a user to ask questions using voice input and obtain responses in an intuitive manner.
[0140] "Means for accepting information input" refers to a device or method for accepting questions or instructions from a user.
[0141] "Means for analyzing and converting into text format data" refers to a device or function that analyzes input information and converts it into text format data that can be processed by a computer.
[0142] The "means for transmitting to a server" refers to a communication device or system for transmitting the converted text format data to a remote server.
[0143] The "means for receiving response data" refers to a device or method for receiving response data sent from a server.
[0144] The term "means for displaying response data to a user" refers to a device or system that visually displays the received response data to a user.
[0145] "Speech recognition means" refers to a device or technology that converts a user's voice input into digital data and then into an analyzable text format.
[0146] "Visual display means" refers to a device or system that displays data on a screen or display in a format that is easy for a user to understand.
[0147] This invention relates to a system that uses a generative AI model to provide information based on user input. The system for implementing the invention mainly includes the following components:
[0148] 1. Information input method
[0149] Users input questions by voice using devices such as smart glasses or smartphones. Using speech recognition technology (e.g., the speech_recognition library), the voice can be converted into text data.
[0150] 2. Analysis and text conversion methods
[0151] The voice input data is automatically converted into text format. For example, Google's speech recognition API is used to convert voice data into text data.
[0152] 3. Means of communication
[0153] The converted text data is sent to a remote server via the Internet using an HTTP request library (e.g., the requests library).
[0154] 4. Response Data Generation Method
[0155] On the server side, appropriate response data is generated using a generative AI model (for example, a GPT model) based on the received text data, and this response data is again sent to the device via the network.
[0156] 5. Response data display method
[0157] The device visually displays the received response data to the user: in the case of smart glasses, this is done by using a display that shows the information in the user's field of view, while in the case of smartphones, this is done by displaying the information on a screen.
[0158] Specific examples
[0159] For example, if a user wears smart glasses and speaks a question such as "What is a recommended travel destination for my next holiday?", the question is converted into text data by a speech recognition means and sent to a server. The server uses the GPT model to generate response data such as "Kyoto is a recommended travel destination" and returns it to the device. In this way, the user can visually obtain the information "Kyoto is a recommended travel destination" on the display of the smart glasses.
[0160] Prompt Sentence Examples
[0161] The user enters the following prompt sentence as an example of a question:
[0162] "What's your recommended destination for your next holiday?"
[0163] The GPT model, upon seeing this prompt, generates the optimal response based on the specified conditions and presents it to the user.
[0164] By combining the above components, the present invention allows users to obtain information intuitively and quickly, contributing to more efficient information provision and improved user experience.
[0165] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0166] Step 1:
[0167] The user uses the smart glasses to input a question by voice.
[0168] (Input) The user's spoken question.
[0169] (Operation) A voice recognition means captures a voice query.
[0170] (Output) The captured audio data.
[0171] Step 2:
[0172] The terminal converts the voice data into text data.
[0173] (Input) The captured audio data.
[0174] (Operation) Analyze the voice data using a voice recognition method (for example, Google's voice recognition API) and convert it into text data.
[0175] (Output) Question data in text format.
[0176] Step 3:
[0177] The terminal transmits question data in text format to the server.
[0178] (Input) Question data in text format.
[0179] (Operation) Use an HTTP request library (for example, the requests library) to send text-formatted question data to the server.
[0180] (Output) The question data sent to the server in text format.
[0181] Step 4:
[0182] The server generates response data using a generative AI model based on the received text data.
[0183] (Input) The textual question data sent to the server.
[0184] (Operation) The server analyzes the text data using a generative AI model (GPT model) and generates optimal response data.
[0185] (Output) The generated response data (e.g., "Kyoto is a recommended travel destination").
[0186] Step 5:
[0187] The server sends the generated response data to the terminal.
[0188] (Input) The generated response data.
[0189] (Operation) The server sends response data to the terminal.
[0190] (Output) Response data sent to the terminal.
[0191] Step 6:
[0192] The response data received by the terminal is visually displayed to the user.
[0193] (Input) Response data sent to the terminal.
[0194] (Operation) The terminal display is used to display the response data within the user's field of view.
[0195] (Output) Response data that the user visually obtains (e.g., "Kyoto is a recommended travel destination").
[0196] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0197] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[0198] First, the user inputs a question or information using the device's touchscreen. For example, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and operation.
[0199] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent.
[0200] The server generates response data using generative AI (GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling depressed, the generated response might be something like, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0201] The generated response data is converted into an appropriate format and sent to the terminal, which then displays the received response data on a user interface in a format that is easy for the user to understand.
[0202] For example, consider the following user actions and system behaviors:
[0203] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[0204] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[0205] 3. The server uses a GPT model to analyze the received question and sentiment data and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0206] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[0207] 5. The terminal displays the received response to the user.
[0208] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[0209] In this way, the system of the present invention, by combining an emotion recognition engine and a generation AI, can provide high-quality information according to the user's emotional state. This allows users to receive more satisfying responses, and realizes a sustainable information provision environment in which information providers can receive fair rewards.
[0210] The processing flow will be explained below.
[0211] Step 1:
[0212] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0213] Step 2:
[0214] The device analyzes the input from the user and sends the content, input speed, touch pattern, voice tone, etc. to the emotion recognition engine, which then analyzes the user's emotions based on this data.
[0215] Step 3:
[0216] The device receives the input question and emotion data from the emotion recognition engine. For example, the emotion data may indicate that the user is depressed.
[0217] Step 4:
[0218] The device converts the question data and emotion data into text format and prepares an HTTP POST request to send to the server. The destination URL is specified, and the question data and emotion data are included in the HTTP request body.
[0219] Step 5:
[0220] The device executes an HTTP POST request and sends the question data and emotion data to the server.
[0221] Step 6:
[0222] The server receives the HTTP request sent from the terminal and extracts the question data and emotion data from the request body.
[0223] Step 7:
[0224] The server uses generative AI (GPT model) to generate response data based on the question data and emotion data extracted by the server. For example, if the user is feeling down, the server generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0225] Step 8:
[0226] The server formats the generated response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[0227] Step 9:
[0228] The server sends the formatted response data to the terminal.
[0229] Step 10:
[0230] The terminal receives the HTTP response from the server and analyzes the response data.
[0231] Step 11:
[0232] The terminal analyzes the response data and displays it on the user interface, which is displayed according to the user's emotions.
[0233] Step 12:
[0234] The user can check the screen of their device and see the response displayed: "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0235] This series of processes allows optimal information to be provided according to the user's emotional state.
[0236] Example 2
[0237] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0238] Conventional vending machine-type information service systems have difficulty providing appropriate responses according to the user's individual emotional state. Because users need different information and responses depending on their emotional state at any given time, standard responses do not provide sufficient satisfaction. It is necessary to solve this problem and improve user satisfaction.
[0239] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0240] In this invention, the server includes means for generating artificial intelligence that generates response data based on received text data and emotional data, means for transmitting the response data to the terminal, and means for adjusting parameters for generating an optimal response based on the user's question and emotional state, thereby making it possible to provide response data that is adjusted according to the user's emotional state.
[0241] The "means for accepting information input from the user" refers to an interface that allows the user to input questions or information into the terminal. Specifically, it includes input devices such as a touch screen and a keyboard.
[0242] The "means for converting the information input into text format data" refers to a process and system for digitizing the input information and converting it into a format that can be processed as text data.
[0243] The "means for transmitting the text format data to the server" is a device that includes a network connection and a communication protocol for communicating text data from the terminal to the server.
[0244] A "server equipped with artificial intelligence that generates response data adjusted according to the user's emotional state based on received text data" is a server device that includes an AI engine for analyzing text data and emotional data and generating appropriate responses.
[0245] The "means for transmitting the response data to the terminal" is a device that includes a network connection and a communication protocol for communicating the generated response data from the server to the terminal.
[0246] The "means for displaying response data to the user on the terminal" refers to a system including a display device and display software for visually presenting the response results to the user on the terminal.
[0247] "Generative artificial intelligence that generates response data based on received text data and emotional data" refers to an artificial intelligence algorithm that generates optimal responses based on input text data and the results of emotional analysis.
[0248] "Means for adjusting parameters to generate optimal responses based on the user's question and emotional state" refers to processes and systems that modify internal settings to optimize the output of an AI model, taking into account input from the user and their emotional state.
[0249] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[0250] First, the user uses the device's touchscreen to input a question or piece of information. For example, they might input a question like, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and how they operate the device. The emotion recognition engine used uses a common machine learning algorithm, for example, and can analyze input speed, touch strength, voice tone, and other factors.
[0251] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent. Communication is performed using an internet connection and communication protocols (e.g., HTTP or HTTPS).
[0252] The server generates response data using generative AI (e.g., a GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling down, a response such as "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is generated. For example, OpenAI's GPT-3 model is used as the generative AI model.
[0253] The generated response data is converted into an appropriate format and sent to the terminal. The terminal displays the received response data on a user interface, providing it in a format that is easy for the user to understand. The display device of the terminal used can be a touch screen or a display.
[0254] For example, consider the following user actions and system behaviors:
[0255] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[0256] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[0257] 3. The server uses a generative AI to analyze the question and emotional data it receives and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0258] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[0259] 5. The terminal displays the received response to the user.
[0260] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[0261] An example of a prompt might be:
[0262] "Please tell me where I should go on my next holiday. I'm feeling low."
[0263] This invention enables the provision of high-quality information tailored to the user's emotional state, thereby increasing user satisfaction. By combining generative AI with an emotion recognition engine, this system can provide more personalized responses.
[0264] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0265] Step 1:
[0266] The user uses the vending machine terminal's touchscreen to input a question or information, such as "What are some recommended travel destinations for my next holiday?" The input is temporarily stored in the terminal's internal memory.
[0267] Input: User question: "What are some recommended travel destinations for my next holiday?"
[0268] Output: Questions as text data
[0269] Specific operation: The user enters a question by touching the touchscreen and typing. The entered information is internally converted to text format.
[0270] Step 2:
[0271] The device uses a built-in emotion recognition engine to analyze the user's emotions based on their input and operation methods, such as typing speed, touch strength, and tone, to determine whether the user is depressed.
[0272] Input: Questions as text data, and user interaction data such as typing speed and touch strength
[0273] Output: Emotion data (e.g., "depressed")
[0274] Specific operation: The emotion recognition engine analyzes the user's input characteristics and estimates their emotional state. As a result of the analysis, emotional data such as "depressed" is generated.
[0275] Step 3:
[0276] The device sends text data and emotion data to the server using communication protocols such as HTTP and HTTPS.
[0277] Input: Questions as text data, emotion data
[0278] Output: Data sent to the server
[0279] Specific operation: The device uses a network connection to send the analyzed data to the server, which includes text data and emotion data.
[0280] Step 4:
[0281] The server generates response data using a generative AI model (e.g., a GPT model) based on the received text data and emotion data. The generated response is adjusted according to the user's emotional state.
[0282] Input: Received text data, emotion data
[0283] Output: Generated response data
[0284] Specific operation: The server's AI model generates the optimal response for the user based on the text data "What is the recommended travel destination for your next holiday?" and the emotion data "I feel depressed." Example: "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0285] Step 5:
[0286] The generated response data is sent from the server to the terminal, again using a communication protocol such as HTTP or HTTPS.
[0287] Input: Generated response data
[0288] Output: Data sent to the terminal
[0289] Specific operation: The server formats the response data appropriately and sends it to the terminal. The sent data includes the response in text format.
[0290] Step 6:
[0291] The device displays the received response data on the user interface. For example, the text "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is displayed on the touch screen.
[0292] Input: Received response data
[0293] Output: Information displayed to the user
[0294] Specific operation: The terminal displays the received response data on the screen and provides information to the user. The user can check the displayed information and use it as reference.
[0295] Through these steps, users can receive information tailored to their emotional state. This system combines a generative AI model and an emotion recognition engine, aiming to increase user satisfaction.
[0296] (Application example 2)
[0297] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0298] Autonomous vehicles are required to recognize passenger emotions in real time and provide optimal services based on those emotions. This will improve passenger comfort and safety and enable more effective service provision. To solve this problem, the present invention provides a system that combines emotion recognition functionality with a generative AI model.
[0299] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text-format data, means for transmitting the text-format data to the server, means for receiving response data from the server, means for displaying the response data to the user, means for recognizing passenger emotions, means for transmitting the emotion information to a generative AI model, and means for adjusting the response data based on the emotion information and providing a service. This makes it possible to provide optimal services according to passenger emotions, thereby improving comfort and safety.
[0300] The "means for accepting information input from the user" is an interface for the user to input information, and may use a touch screen or a voice input device.
[0301] The "means for analyzing the information input and converting it into text format data" refers to a process for analyzing the information input by the user and converting it into digital text data.
[0302] The "means for transmitting the text format data to the server" refers to a communication means for transmitting the converted text data to the server via a network.
[0303] The "means for receiving response data from the server" refers to a communication means for receiving response data sent from the server.
[0304] The "means for displaying the response data to the user" refers to a display device for visually presenting the received response data to the user.
[0305] "Means for recognizing passenger emotions" refers to an emotion recognition engine that analyzes emotions from passengers' facial expressions, tone of voice, etc.
[0306] "Means for transmitting the emotion information to the generative AI model" refers to a communication means for transmitting the recognized emotion information to the generative AI model as input.
[0307] The "means for adjusting response data based on the emotion information and providing a service" refers to a means for providing response data generated based on emotion information in a form suitable for passengers.
[0308] The present invention relates to a system in which a user inputs information through a system for an autonomous vehicle and an optimal response is provided by a generative AI model. The system is implemented in the following configuration.
[0309] First, the user inputs information using the in-car touchscreen or voice input device, and an emotion recognition engine is activated to analyze the passenger's emotions from facial expressions, tone of voice, etc., and collect emotional data.
[0310] Next, the emotion data analyzed by the emotion recognition engine and the information entered by the user are converted into text using natural language processing technology. The converted text data and emotion data are then sent to a server via a network.
[0311] The server uses a generative AI model (e.g., a GPT model) to analyze the received text and emotion data and generate an appropriate response. This response is tailored based on the passenger's emotional state. For example, if the passenger is feeling depressed, the server might generate a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[0312] The generated response data is transmitted in real time to a terminal inside the autonomous vehicle, which then displays the received response data on a user interface in an easy-to-understand format for passengers.
[0313] Specifically, the device is equipped with a camera and microphone, which use OpenCV and the SpeechRecognition library to perform emotion analysis, and the server side uses the Transformers pipeline (the Hugging Face GPT model) as a generative AI model to generate responses.
[0314] Examples of specific prompts for generative AI models include:
[0315] "User sentiment: Depressed. User asks: 'What's the best place to go for my next vacation?'. Answer:"
[0316] Based on this prompt, the generative AI model generates a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[0317] This system allows passengers to receive the most appropriate service according to their mood, improving comfort and safety.
[0318] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0319] Step 1:
[0320] Users input information using the car's touchscreen or voice input device, such as a question like, "What are some recommended travel destinations for my next vacation?" The input information is treated as text data.
[0321] Step 2:
[0322] The device's emotion recognition engine uses the camera and microphone to analyze emotion data from the user's facial expressions and tone of voice. This process utilizes OpenCV and the SpeechRecognition library. The input is the user's voice and video data, and the output is emotion data such as "depression" or "joy."
[0323] Step 3:
[0324] The device collects the user's input information (text data) and emotion data and sends them to the server. The input is text data and emotion data, and the output is data sent to the server.
[0325] Step 4:
[0326] The server generates response data using a generative AI model (GPT model) based on the received text data and emotion data. First, it analyzes the text data, and then adjusts the response taking into account the emotion data. The input is text data and emotion data, and the output is the optimal response data to the user's question.
[0327] Step 5:
[0328] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the transmission of data to the terminal.
[0329] Step 6:
[0330] The terminal displays the received response data on the user interface. Specifically, it uses a touch screen to display a response such as "We recommend a place surrounded by beautiful nature for you to refresh yourself." The input is the response data sent from the server, and the output is visual information provided to the user.
[0331] Through these processing steps, optimal information is provided based on the user's emotions, improving passenger comfort and safety.
[0332] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0333] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0334] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0335] [Second embodiment]
[0336] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0337] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0338] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0339] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0340] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0341] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0342] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0343] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0344] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0345] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0346] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0347] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0348] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. Specifically, it provides a mechanism in which the user inputs a question, the server generates a response to the question, and displays it to the user through the terminal.
[0349] First, a user inputs a question using a user interface such as a touch screen on a vending machine terminal, which can be specific, such as "What are some recommended travel destinations for my next holiday?"
[0350] The terminal analyzes the input question and converts it into text data, which is then sent to a server via a communication method such as the Internet.
[0351] The server generates response data using generative AI (GPT model) based on the received text data. The generated response data, such as "Kyoto is a recommended travel destination," is then sent back to the device via communication means.
[0352] The terminal displays the response data received from the server to the user, allowing the user to easily obtain the necessary information from the vending machine.
[0353] This system is expected to have the effect of establishing a culture among users that recognizes the value of information and pays for it. It will also contribute to the creation of a sustainable information provision environment in which information providers can receive fair compensation.
[0354] For example, consider the following user actions and system behaviors:
[0355] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0356] 2. The device converts the question into text data and sends it to the server.
[0357] 3. The server analyzes the received question using the GPT model and generates a response such as "Kyoto is a recommended travel destination."
[0358] 4. The generated response is sent to the terminal, which displays the response to the user.
[0359] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0360] The unique feature of this system is that the service is provided through a physical vending machine, without the need for an internet connection. This has the advantage that users can easily obtain the information they need wherever they are, and it is thought that it can also be used as part of information literacy education, which is particularly useful for young people.
[0361] The processing flow will be explained below.
[0362] Step 1:
[0363] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0364] Step 2:
[0365] The terminal parses the input from the user and obtains the question data as a string, which is then converted into a text format.
[0366] Step 3:
[0367] The device prepares an HTTP POST request to send text-formatted question data to the server, specifying the destination URL and including the question data in the body of the HTTP request.
[0368] Step 4:
[0369] The device executes an HTTP POST request and sends the query data to the server.
[0370] Step 5:
[0371] The server receives the HTTP request sent from the terminal and extracts the question data from the request body.
[0372] Step 6:
[0373] The server passes the extracted question data to the generation AI (GPT model) and generates a response. At this time, the API key is used to send a request to the API endpoint of the GPT model.
[0374] Step 7:
[0375] The server receives the response data returned from the GPT model. For example, it receives the text "Kyoto is a recommended travel destination."
[0376] Step 8:
[0377] The server formats the response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[0378] Step 9:
[0379] The server sends the formatted response data to the terminal.
[0380] Step 10:
[0381] The terminal receives the HTTP response from the server and analyzes the response data.
[0382] Step 11:
[0383] The terminal displays the analyzed response data on the user interface.
[0384] Step 12:
[0385] The user can check the screen of the device and see the displayed response, "Kyoto is a recommended travel destination."
[0386] Example 1
[0387] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0388] In modern society, much information is digitized, and users want to obtain the information they need quickly and easily. However, conventional information acquisition methods have problems such as the need for an Internet connection and uncertainty about the accuracy of the information. The purpose of this invention is to solve these problems and build a system that quickly provides highly reliable information.
[0389] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0390] In this invention, the server includes means for generating response data using a generative AI model based on received text data, means for transmitting the text data to the server using a communication means, and means for displaying the response data on a user interface, thereby making it possible to provide advanced information in real time by utilizing the generative AI model.
[0391] "Means for accepting information input" refers to an interface for a user to input information and a mechanism for receiving that information.
[0392] "Means for converting into text format data" refers to a process for recognizing information entered by a user as a digital string and converting it into text data.
[0393] "Means for transmitting to a server using a communication means" refers to infrastructure and software for transmitting text-format data to a server via a communication network such as the Internet.
[0394] The "means for receiving response data from the server" refers to a communication protocol and hardware for receiving response data sent from the server.
[0395] "Means for displaying on a user interface" refers to a display and display software for displaying received response data in a form that can be viewed by a user.
[0396] "Generative artificial intelligence model" refers to the machine learning model and associated algorithms used to generate appropriate responses based on received text data.
[0397] "Means for adjusting the algorithm" refers to the process of dynamically changing and tuning input parameters and processing methods so that the generative artificial intelligence model can output the optimal response.
[0398] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. The system uses a terminal, a server, and a generative AI model to provide users with efficient and accurate information.
[0399] First, the user uses the touchscreen of the vending machine terminal to input a question, such as the prompt "Where would you recommend for my next holiday?" The terminal is equipped with a high-precision touchscreen and software to process the input.
[0400] The terminal converts the information entered by the user into text data using character recognition software to convert the input string into digital data. The converted text data is then sent to a server via a communication method (e.g., an internet connection). The communication protocol used is HTTP or HTTPS.
[0401] The server analyzes the received text data and generates response data using a generative AI model (for example, OpenAI's GPT-4). The server has high-performance computing resources and can run Node.js or Python-based application servers. As a concrete example, the response generated is "Kyoto is a recommended travel destination."
[0402] The generated response data is then sent back to the terminal via the communication means. The terminal analyzes the received response data and displays it on the user interface. The display typically uses an LCD panel or an organic EL panel. The user can check this and obtain the necessary information.
[0403] This system allows users to obtain fast and accurate information through a physical vending machine. Its unique feature is that it does not require an internet connection, making it possible to conveniently obtain information anytime, anywhere. Furthermore, by utilizing generative AI models, advanced information is provided in real time, improving user satisfaction.
[0404] The specific operation flow is as follows:
[0405] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0406] 2. The device converts the question into text data and sends it to the server.
[0407] 3. The server analyzes the received question using a generative AI model and generates a response such as, "Kyoto is a recommended travel destination."
[0408] 4. The generated response is sent to the terminal, which displays the response to the user.
[0409] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0410] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0411] Step 1:
[0412] The user operates the touch screen of the vending machine terminal and inputs a question.
[0413] Input: A question that the user types into the touchscreen (e.g., "What are your recommended travel destinations for my next holiday?").
[0414] Output: Text data recognized by the device.
[0415] How it works: The user types "What are some recommended travel destinations for my next holiday?" into the touchscreen, and the input information is captured by the touchscreen's sensors.
[0416] Step 2:
[0417] The terminal converts the question entered by the user into text data.
[0418] Input: Touchscreen input data.
[0419] Output: Data in text format (digital string).
[0420] Specific operation: The device's internal software analyzes the input data obtained from the sensor and converts it into text data in UTF-8 format.
[0421] Step 3:
[0422] The terminal transmits the converted text data to the server using a communication means.
[0423] Input: Data in text format.
[0424] Output: The HTTP POST request sent to the server.
[0425] Specific operation: An HTTP POST request is generated from the terminal to the server and sent to the server via the Internet.
[0426] Step 4:
[0427] The server analyzes the received text data and generates response data using a generative artificial intelligence model.
[0428] Input: Text data to be included in the HTTP POST request.
[0429] Output: The generated response data.
[0430] How it works: The server uses a Node.js or Python-based application server to send text data to the GPT-4 API and generate an appropriate response.
[0431] Step 5:
[0432] The server transmits the generated response data to the terminal using the communication means.
[0433] Input: The generated response data.
[0434] Output: The response data as an HTTP response.
[0435] Specific operation: The server generates response data in JSON format as an HTTP response and sends it to the terminal.
[0436] Step 6:
[0437] The terminal analyzes the response data received from the server and displays it on the user interface.
[0438] Input: The HTTP response data sent by the server.
[0439] Output: The text that appears in the user interface.
[0440] Specific operation: The device's software analyzes the HTTP response and displays "Kyoto is a recommended travel destination" on the touchscreen.
[0441] Step 7:
[0442] The user checks and acquires the information displayed on the screen of the terminal.
[0443] Input: The response data displayed in the user interface.
[0444] Output: The information obtained by the user.
[0445] Specific operation: The user visually recognizes the information displayed on the touch screen, "Kyoto is a recommended travel destination," and obtains the necessary information.
[0446] (Application example 1)
[0447] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0448] In conventional information provision systems using vending machines, users had to use a physical keyboard or touch screen to input questions, which made the input process cumbersome. In addition, there were limitations to the ways in which users could visually obtain information, so a more intuitive and faster way to obtain information was needed.
[0449] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0450] In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text data, means for transmitting the text data to the server, means for receiving response data from the server, means for displaying the response data to the user, speech recognition means for accepting voice input and converting it into text data, and means for visually displaying the response data, thereby enabling a user to ask questions using voice input and obtain responses in an intuitive manner.
[0451] "Means for accepting information input" refers to a device or method for accepting questions or instructions from a user.
[0452] "Means for analyzing and converting into text format data" refers to a device or function that analyzes input information and converts it into text format data that can be processed by a computer.
[0453] The "means for transmitting to a server" refers to a communication device or system for transmitting the converted text format data to a remote server.
[0454] The "means for receiving response data" refers to a device or method for receiving response data sent from a server.
[0455] The term "means for displaying response data to a user" refers to a device or system that visually displays the received response data to a user.
[0456] "Speech recognition means" refers to a device or technology that converts a user's voice input into digital data and then into an analyzable text format.
[0457] "Visual display means" refers to a device or system that displays data on a screen or display in a format that is easy for a user to understand.
[0458] This invention relates to a system that uses a generative AI model to provide information based on user input. The system for implementing the invention mainly includes the following components:
[0459] 1. Information input method
[0460] Users input questions by voice using devices such as smart glasses or smartphones. Using speech recognition technology (e.g., the speech_recognition library), the voice can be converted into text data.
[0461] 2. Analysis and text conversion methods
[0462] The voice input data is automatically converted into text format. For example, Google's speech recognition API is used to convert voice data into text data.
[0463] 3. Means of communication
[0464] The converted text data is sent to a remote server via the Internet using an HTTP request library (e.g., the requests library).
[0465] 4. Response Data Generation Method
[0466] On the server side, appropriate response data is generated using a generative AI model (for example, a GPT model) based on the received text data, and this response data is again sent to the device via the network.
[0467] 5. Response data display method
[0468] The device visually displays the received response data to the user: in the case of smart glasses, this is done by using a display that shows the information in the user's field of view, while in the case of smartphones, this is done by displaying the information on a screen.
[0469] Specific examples
[0470] For example, if a user wears smart glasses and speaks a question such as "What is a recommended travel destination for my next holiday?", the question is converted into text data by a speech recognition means and sent to a server. The server uses the GPT model to generate response data such as "Kyoto is a recommended travel destination" and returns it to the device. In this way, the user can visually obtain the information "Kyoto is a recommended travel destination" on the display of the smart glasses.
[0471] Prompt Sentence Examples
[0472] The user enters the following prompt sentence as an example of a question:
[0473] "What's your recommended destination for your next holiday?"
[0474] The GPT model, upon seeing this prompt, generates the optimal response based on the specified conditions and presents it to the user.
[0475] By combining the above components, the present invention allows users to obtain information intuitively and quickly, contributing to more efficient information provision and improved user experience.
[0476] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0477] Step 1:
[0478] The user uses the smart glasses to input a question by voice.
[0479] (Input) The user's spoken question.
[0480] (Operation) A voice recognition means captures a voice query.
[0481] (Output) The captured audio data.
[0482] Step 2:
[0483] The terminal converts the voice data into text data.
[0484] (Input) The captured audio data.
[0485] (Operation) Analyze the voice data using a voice recognition method (for example, Google's voice recognition API) and convert it into text data.
[0486] (Output) Question data in text format.
[0487] Step 3:
[0488] The terminal transmits question data in text format to the server.
[0489] (Input) Question data in text format.
[0490] (Operation) Use an HTTP request library (for example, the requests library) to send text-formatted question data to the server.
[0491] (Output) The question data sent to the server in text format.
[0492] Step 4:
[0493] The server generates response data using a generative AI model based on the received text data.
[0494] (Input) The textual question data sent to the server.
[0495] (Operation) The server analyzes the text data using a generative AI model (GPT model) and generates optimal response data.
[0496] (Output) The generated response data (e.g., "Kyoto is a recommended travel destination").
[0497] Step 5:
[0498] The server sends the generated response data to the terminal.
[0499] (Input) The generated response data.
[0500] (Operation) The server sends response data to the terminal.
[0501] (Output) Response data sent to the terminal.
[0502] Step 6:
[0503] The response data received by the terminal is visually displayed to the user.
[0504] (Input) Response data sent to the terminal.
[0505] (Operation) The terminal display is used to display the response data within the user's field of view.
[0506] (Output) Response data that the user visually obtains (e.g., "Kyoto is a recommended travel destination").
[0507] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0508] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[0509] First, the user inputs a question or information using the device's touchscreen. For example, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and operation.
[0510] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent.
[0511] The server generates response data using generative AI (GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling depressed, the generated response might be something like, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0512] The generated response data is converted into an appropriate format and sent to the terminal, which then displays the received response data on a user interface in a format that is easy for the user to understand.
[0513] For example, consider the following user actions and system behaviors:
[0514] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[0515] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[0516] 3. The server uses a GPT model to analyze the received question and sentiment data and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0517] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[0518] 5. The terminal displays the received response to the user.
[0519] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[0520] In this way, the system of the present invention, by combining an emotion recognition engine and a generation AI, can provide high-quality information according to the user's emotional state. This allows users to receive more satisfying responses, and realizes a sustainable information provision environment in which information providers can receive fair rewards.
[0521] The processing flow will be explained below.
[0522] Step 1:
[0523] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0524] Step 2:
[0525] The device analyzes the input from the user and sends the content, input speed, touch pattern, voice tone, etc. to the emotion recognition engine, which then analyzes the user's emotions based on this data.
[0526] Step 3:
[0527] The device receives the input question and emotion data from the emotion recognition engine. For example, the emotion data may indicate that the user is depressed.
[0528] Step 4:
[0529] The device converts the question data and emotion data into text format and prepares an HTTP POST request to send to the server. The destination URL is specified, and the question data and emotion data are included in the HTTP request body.
[0530] Step 5:
[0531] The device executes an HTTP POST request and sends the question data and emotion data to the server.
[0532] Step 6:
[0533] The server receives the HTTP request sent from the terminal and extracts the question data and emotion data from the request body.
[0534] Step 7:
[0535] The server uses generative AI (GPT model) to generate response data based on the question data and emotion data extracted by the server. For example, if the user is feeling down, the server generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0536] Step 8:
[0537] The server formats the generated response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[0538] Step 9:
[0539] The server sends the formatted response data to the terminal.
[0540] Step 10:
[0541] The terminal receives the HTTP response from the server and analyzes the response data.
[0542] Step 11:
[0543] The terminal analyzes the response data and displays it on the user interface, which is displayed according to the user's emotions.
[0544] Step 12:
[0545] The user can check the screen of their device and see the response displayed: "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0546] This series of processes allows optimal information to be provided according to the user's emotional state.
[0547] Example 2
[0548] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0549] Conventional vending machine-type information service systems have difficulty providing appropriate responses according to the user's individual emotional state. Because users need different information and responses depending on their emotional state at any given time, standard responses do not provide sufficient satisfaction. It is necessary to solve this problem and improve user satisfaction.
[0550] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0551] In this invention, the server includes means for generating artificial intelligence that generates response data based on received text data and emotional data, means for transmitting the response data to the terminal, and means for adjusting parameters for generating an optimal response based on the user's question and emotional state, thereby making it possible to provide response data that is adjusted according to the user's emotional state.
[0552] The "means for accepting information input from the user" refers to an interface that allows the user to input questions or information into the terminal. Specifically, it includes input devices such as a touch screen and a keyboard.
[0553] The "means for converting the information input into text format data" refers to a process and system for digitizing the input information and converting it into a format that can be processed as text data.
[0554] The "means for transmitting the text format data to the server" is a device that includes a network connection and a communication protocol for communicating text data from the terminal to the server.
[0555] A "server equipped with artificial intelligence that generates response data adjusted according to the user's emotional state based on received text data" is a server device that includes an AI engine for analyzing text data and emotional data and generating appropriate responses.
[0556] The "means for transmitting the response data to the terminal" is a device that includes a network connection and a communication protocol for communicating the generated response data from the server to the terminal.
[0557] The "means for displaying response data to the user on the terminal" refers to a system including a display device and display software for visually presenting the response results to the user on the terminal.
[0558] "Generative artificial intelligence that generates response data based on received text data and emotional data" refers to an artificial intelligence algorithm that generates optimal responses based on input text data and the results of emotional analysis.
[0559] "Means for adjusting parameters to generate optimal responses based on the user's question and emotional state" refers to processes and systems that modify internal settings to optimize the output of an AI model, taking into account input from the user and their emotional state.
[0560] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[0561] First, the user uses the device's touchscreen to input a question or piece of information. For example, they might input a question like, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and how they operate the device. The emotion recognition engine used uses a common machine learning algorithm, for example, and can analyze input speed, touch strength, voice tone, and other factors.
[0562] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent. Communication is performed using an internet connection and communication protocols (e.g., HTTP or HTTPS).
[0563] The server generates response data using generative AI (e.g., a GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling down, a response such as "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is generated. For example, OpenAI's GPT-3 model is used as the generative AI model.
[0564] The generated response data is converted into an appropriate format and sent to the terminal. The terminal displays the received response data on a user interface, providing it in a format that is easy for the user to understand. The display device of the terminal used can be a touch screen or a display.
[0565] For example, consider the following user actions and system behaviors:
[0566] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[0567] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[0568] 3. The server uses a generative AI to analyze the question and emotional data it receives and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0569] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[0570] 5. The terminal displays the received response to the user.
[0571] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[0572] An example of a prompt might be:
[0573] "Please tell me where I should go on my next holiday. I'm feeling low."
[0574] This invention enables the provision of high-quality information tailored to the user's emotional state, thereby increasing user satisfaction. By combining generative AI with an emotion recognition engine, this system can provide more personalized responses.
[0575] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0576] Step 1:
[0577] The user uses the vending machine terminal's touchscreen to input a question or information, such as "What are some recommended travel destinations for my next holiday?" The input is temporarily stored in the terminal's internal memory.
[0578] Input: User question: "What are some recommended travel destinations for my next holiday?"
[0579] Output: Questions as text data
[0580] Specific operation: The user enters a question by touching the touchscreen and typing. The entered information is internally converted to text format.
[0581] Step 2:
[0582] The device uses a built-in emotion recognition engine to analyze the user's emotions based on their input and operation methods, such as typing speed, touch strength, and tone, to determine whether the user is depressed.
[0583] Input: Questions as text data, and user interaction data such as typing speed and touch strength
[0584] Output: Emotion data (e.g., "depressed")
[0585] Specific operation: The emotion recognition engine analyzes the user's input characteristics and estimates their emotional state. As a result of the analysis, emotional data such as "depressed" is generated.
[0586] Step 3:
[0587] The device sends text data and emotion data to the server using communication protocols such as HTTP and HTTPS.
[0588] Input: Questions as text data, emotion data
[0589] Output: Data sent to the server
[0590] Specific operation: The device uses a network connection to send the analyzed data to the server, which includes text data and emotion data.
[0591] Step 4:
[0592] The server generates response data using a generative AI model (e.g., a GPT model) based on the received text data and emotion data. The generated response is adjusted according to the user's emotional state.
[0593] Input: Received text data, emotion data
[0594] Output: Generated response data
[0595] Specific operation: The server's AI model generates the optimal response for the user based on the text data "What is the recommended travel destination for your next holiday?" and the emotion data "I feel depressed." Example: "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0596] Step 5:
[0597] The generated response data is sent from the server to the terminal, again using a communication protocol such as HTTP or HTTPS.
[0598] Input: Generated response data
[0599] Output: Data sent to the terminal
[0600] Specific operation: The server formats the response data appropriately and sends it to the terminal. The sent data includes the response in text format.
[0601] Step 6:
[0602] The device displays the received response data on the user interface. For example, the text "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is displayed on the touch screen.
[0603] Input: Received response data
[0604] Output: Information displayed to the user
[0605] Specific operation: The terminal displays the received response data on the screen and provides information to the user. The user can check the displayed information and use it as reference.
[0606] Through these steps, users can receive information tailored to their emotional state. This system combines a generative AI model and an emotion recognition engine, aiming to increase user satisfaction.
[0607] (Application example 2)
[0608] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0609] Autonomous vehicles are required to recognize passenger emotions in real time and provide optimal services based on those emotions. This will improve passenger comfort and safety and enable more effective service provision. To solve this problem, the present invention provides a system that combines emotion recognition functionality with a generative AI model.
[0610] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text-format data, means for transmitting the text-format data to the server, means for receiving response data from the server, means for displaying the response data to the user, means for recognizing passenger emotions, means for transmitting the emotion information to a generative AI model, and means for adjusting the response data based on the emotion information and providing a service. This makes it possible to provide optimal services according to passenger emotions, thereby improving comfort and safety.
[0611] The "means for accepting information input from the user" is an interface for the user to input information, and may use a touch screen or a voice input device.
[0612] The "means for analyzing the information input and converting it into text format data" refers to a process for analyzing the information input by the user and converting it into digital text data.
[0613] The "means for transmitting the text format data to the server" refers to a communication means for transmitting the converted text data to the server via a network.
[0614] The "means for receiving response data from the server" refers to a communication means for receiving response data sent from the server.
[0615] The "means for displaying the response data to the user" refers to a display device for visually presenting the received response data to the user.
[0616] "Means for recognizing passenger emotions" refers to an emotion recognition engine that analyzes emotions from passengers' facial expressions, tone of voice, etc.
[0617] "Means for transmitting the emotion information to the generative AI model" refers to a communication means for transmitting the recognized emotion information to the generative AI model as input.
[0618] The "means for adjusting response data based on the emotion information and providing a service" refers to a means for providing response data generated based on emotion information in a form suitable for passengers.
[0619] The present invention relates to a system in which a user inputs information through a system for an autonomous vehicle and an optimal response is provided by a generative AI model. The system is implemented in the following configuration.
[0620] First, the user inputs information using the in-car touchscreen or voice input device, and an emotion recognition engine is activated to analyze the passenger's emotions from facial expressions, tone of voice, etc., and collect emotional data.
[0621] Next, the emotion data analyzed by the emotion recognition engine and the information entered by the user are converted into text using natural language processing technology. The converted text data and emotion data are then sent to a server via a network.
[0622] The server uses a generative AI model (e.g., a GPT model) to analyze the received text and emotion data and generate an appropriate response. This response is tailored based on the passenger's emotional state. For example, if the passenger is feeling depressed, the server might generate a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[0623] The generated response data is transmitted in real time to a terminal inside the autonomous vehicle, which then displays the received response data on a user interface in an easy-to-understand format for passengers.
[0624] Specifically, the device is equipped with a camera and microphone, which use OpenCV and the SpeechRecognition library to perform emotion analysis, and the server side uses the Transformers pipeline (the Hugging Face GPT model) as a generative AI model to generate responses.
[0625] Examples of specific prompts for generative AI models include:
[0626] "User sentiment: Depressed. User asks: 'What's the best place to go for my next vacation?'. Answer:"
[0627] Based on this prompt, the generative AI model generates a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[0628] This system allows passengers to receive the most appropriate service according to their mood, improving comfort and safety.
[0629] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0630] Step 1:
[0631] Users input information using the car's touchscreen or voice input device, such as a question like, "What are some recommended travel destinations for my next vacation?" The input information is treated as text data.
[0632] Step 2:
[0633] The device's emotion recognition engine uses the camera and microphone to analyze emotion data from the user's facial expressions and tone of voice. This process utilizes OpenCV and the SpeechRecognition library. The input is the user's voice and video data, and the output is emotion data such as "depression" or "joy."
[0634] Step 3:
[0635] The device collects the user's input information (text data) and emotion data and sends them to the server. The input is text data and emotion data, and the output is data sent to the server.
[0636] Step 4:
[0637] The server generates response data using a generative AI model (GPT model) based on the received text data and emotion data. First, it analyzes the text data, and then adjusts the response taking into account the emotion data. The input is text data and emotion data, and the output is the optimal response data to the user's question.
[0638] Step 5:
[0639] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the transmission of data to the terminal.
[0640] Step 6:
[0641] The terminal displays the received response data on the user interface. Specifically, it uses a touch screen to display a response such as "We recommend a place surrounded by beautiful nature for you to refresh yourself." The input is the response data sent from the server, and the output is visual information provided to the user.
[0642] Through these processing steps, optimal information is provided based on the user's emotions, improving passenger comfort and safety.
[0643] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0644] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0645] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0646] [Third embodiment]
[0647] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0648] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0649] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0650] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0651] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0652] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0653] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0654] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0655] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0656] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0657] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0658] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0659] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. Specifically, it provides a mechanism in which the user inputs a question, the server generates a response to the question, and displays it to the user through the terminal.
[0660] First, a user inputs a question using a user interface such as a touch screen on a vending machine terminal, which can be specific, such as "What are some recommended travel destinations for my next holiday?"
[0661] The terminal analyzes the input question and converts it into text data, which is then sent to a server via a communication method such as the Internet.
[0662] The server generates response data using generative AI (GPT model) based on the received text data. The generated response data, such as "Kyoto is a recommended travel destination," is then sent back to the device via communication means.
[0663] The terminal displays the response data received from the server to the user, allowing the user to easily obtain the necessary information from the vending machine.
[0664] This system is expected to have the effect of establishing a culture among users that recognizes the value of information and pays for it. It will also contribute to the creation of a sustainable information provision environment in which information providers can receive fair compensation.
[0665] For example, consider the following user actions and system behaviors:
[0666] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0667] 2. The device converts the question into text data and sends it to the server.
[0668] 3. The server analyzes the received question using the GPT model and generates a response such as "Kyoto is a recommended travel destination."
[0669] 4. The generated response is sent to the terminal, which displays the response to the user.
[0670] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0671] The unique feature of this system is that the service is provided through a physical vending machine, without the need for an internet connection. This has the advantage that users can easily obtain the information they need wherever they are, and it is thought that it can also be used as part of information literacy education, which is particularly useful for young people.
[0672] The processing flow will be explained below.
[0673] Step 1:
[0674] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0675] Step 2:
[0676] The terminal parses the input from the user and obtains the question data as a string, which is then converted into a text format.
[0677] Step 3:
[0678] The device prepares an HTTP POST request to send text-formatted question data to the server, specifying the destination URL and including the question data in the body of the HTTP request.
[0679] Step 4:
[0680] The device executes an HTTP POST request and sends the query data to the server.
[0681] Step 5:
[0682] The server receives the HTTP request sent from the terminal and extracts the question data from the request body.
[0683] Step 6:
[0684] The server passes the extracted question data to the generation AI (GPT model) and generates a response. At this time, the API key is used to send a request to the API endpoint of the GPT model.
[0685] Step 7:
[0686] The server receives the response data returned from the GPT model. For example, it receives the text "Kyoto is a recommended travel destination."
[0687] Step 8:
[0688] The server formats the response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[0689] Step 9:
[0690] The server sends the formatted response data to the terminal.
[0691] Step 10:
[0692] The terminal receives the HTTP response from the server and analyzes the response data.
[0693] Step 11:
[0694] The terminal displays the analyzed response data on the user interface.
[0695] Step 12:
[0696] The user can check the screen of the device and see the displayed response, "Kyoto is a recommended travel destination."
[0697] Example 1
[0698] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0699] In modern society, much information is digitized, and users want to obtain the information they need quickly and easily. However, conventional information acquisition methods have problems such as the need for an Internet connection and uncertainty about the accuracy of the information. The purpose of this invention is to solve these problems and build a system that quickly provides highly reliable information.
[0700] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0701] In this invention, the server includes means for generating response data using a generative AI model based on received text data, means for transmitting the text data to the server using a communication means, and means for displaying the response data on a user interface, thereby making it possible to provide advanced information in real time by utilizing the generative AI model.
[0702] "Means for accepting information input" refers to an interface for a user to input information and a mechanism for receiving that information.
[0703] "Means for converting into text format data" refers to a process for recognizing information entered by a user as a digital string and converting it into text data.
[0704] "Means for transmitting to a server using a communication means" refers to infrastructure and software for transmitting text-format data to a server via a communication network such as the Internet.
[0705] The "means for receiving response data from the server" refers to a communication protocol and hardware for receiving response data sent from the server.
[0706] "Means for displaying on a user interface" refers to a display and display software for displaying received response data in a form that can be viewed by a user.
[0707] "Generative artificial intelligence model" refers to the machine learning model and associated algorithms used to generate appropriate responses based on received text data.
[0708] "Means for adjusting the algorithm" refers to the process of dynamically changing and tuning input parameters and processing methods so that the generative artificial intelligence model can output the optimal response.
[0709] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. The system uses a terminal, a server, and a generative AI model to provide users with efficient and accurate information.
[0710] First, the user uses the touchscreen of the vending machine terminal to input a question, such as the prompt "Where would you recommend for my next holiday?" The terminal is equipped with a high-precision touchscreen and software to process the input.
[0711] The terminal converts the information entered by the user into text data using character recognition software to convert the input string into digital data. The converted text data is then sent to a server via a communication method (e.g., an internet connection). The communication protocol used is HTTP or HTTPS.
[0712] The server analyzes the received text data and generates response data using a generative AI model (for example, OpenAI's GPT-4). The server has high-performance computing resources and can run Node.js or Python-based application servers. As a concrete example, the response generated is "Kyoto is a recommended travel destination."
[0713] The generated response data is then sent back to the terminal via the communication means. The terminal analyzes the received response data and displays it on the user interface. The display typically uses an LCD panel or an organic EL panel. The user can check this and obtain the necessary information.
[0714] This system allows users to obtain fast and accurate information through a physical vending machine. Its unique feature is that it does not require an internet connection, making it possible to conveniently obtain information anytime, anywhere. Furthermore, by utilizing generative AI models, advanced information is provided in real time, improving user satisfaction.
[0715] The specific operation flow is as follows:
[0716] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0717] 2. The device converts the question into text data and sends it to the server.
[0718] 3. The server analyzes the received question using a generative AI model and generates a response such as, "Kyoto is a recommended travel destination."
[0719] 4. The generated response is sent to the terminal, which displays the response to the user.
[0720] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0721] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0722] Step 1:
[0723] The user operates the touch screen of the vending machine terminal and inputs a question.
[0724] Input: A question that the user types into the touchscreen (e.g., "What are your recommended travel destinations for my next holiday?").
[0725] Output: Text data recognized by the device.
[0726] How it works: The user types "What are some recommended travel destinations for my next holiday?" into the touchscreen, and the input information is captured by the touchscreen's sensors.
[0727] Step 2:
[0728] The terminal converts the question entered by the user into text data.
[0729] Input: Touchscreen input data.
[0730] Output: Data in text format (digital string).
[0731] Specific operation: The device's internal software analyzes the input data obtained from the sensor and converts it into text data in UTF-8 format.
[0732] Step 3:
[0733] The terminal transmits the converted text data to the server using a communication means.
[0734] Input: Data in text format.
[0735] Output: The HTTP POST request sent to the server.
[0736] Specific operation: An HTTP POST request is generated from the terminal to the server and sent to the server via the Internet.
[0737] Step 4:
[0738] The server analyzes the received text data and generates response data using a generative artificial intelligence model.
[0739] Input: Text data to be included in the HTTP POST request.
[0740] Output: The generated response data.
[0741] How it works: The server uses a Node.js or Python-based application server to send text data to the GPT-4 API and generate an appropriate response.
[0742] Step 5:
[0743] The server transmits the generated response data to the terminal using the communication means.
[0744] Input: The generated response data.
[0745] Output: The response data as an HTTP response.
[0746] Specific operation: The server generates response data in JSON format as an HTTP response and sends it to the terminal.
[0747] Step 6:
[0748] The terminal analyzes the response data received from the server and displays it on the user interface.
[0749] Input: The HTTP response data sent by the server.
[0750] Output: The text that appears in the user interface.
[0751] Specific operation: The device's software analyzes the HTTP response and displays "Kyoto is a recommended travel destination" on the touchscreen.
[0752] Step 7:
[0753] The user checks and acquires the information displayed on the screen of the terminal.
[0754] Input: The response data displayed in the user interface.
[0755] Output: The information obtained by the user.
[0756] Specific operation: The user visually recognizes the information displayed on the touch screen, "Kyoto is a recommended travel destination," and obtains the necessary information.
[0757] (Application example 1)
[0758] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0759] In conventional information provision systems using vending machines, users had to use a physical keyboard or touch screen to input questions, which made the input process cumbersome. In addition, there were limitations to the ways in which users could visually obtain information, so a more intuitive and faster way to obtain information was needed.
[0760] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0761] In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text data, means for transmitting the text data to the server, means for receiving response data from the server, means for displaying the response data to the user, speech recognition means for accepting voice input and converting it into text data, and means for visually displaying the response data, thereby enabling a user to ask questions using voice input and obtain responses in an intuitive manner.
[0762] "Means for accepting information input" refers to a device or method for accepting questions or instructions from a user.
[0763] "Means for analyzing and converting into text format data" refers to a device or function that analyzes input information and converts it into text format data that can be processed by a computer.
[0764] The "means for transmitting to a server" refers to a communication device or system for transmitting the converted text format data to a remote server.
[0765] The "means for receiving response data" refers to a device or method for receiving response data sent from a server.
[0766] The term "means for displaying response data to a user" refers to a device or system that visually displays the received response data to a user.
[0767] "Speech recognition means" refers to a device or technology that converts a user's voice input into digital data and then into an analyzable text format.
[0768] "Visual display means" refers to a device or system that displays data on a screen or display in a format that is easy for a user to understand.
[0769] This invention relates to a system that uses a generative AI model to provide information based on user input. The system for implementing the invention mainly includes the following components:
[0770] 1. Information input method
[0771] Users input questions by voice using devices such as smart glasses or smartphones. Using speech recognition technology (e.g., the speech_recognition library), the voice can be converted into text data.
[0772] 2. Analysis and text conversion methods
[0773] The voice input data is automatically converted into text format. For example, Google's speech recognition API is used to convert voice data into text data.
[0774] 3. Means of communication
[0775] The converted text data is sent to a remote server via the Internet using an HTTP request library (e.g., the requests library).
[0776] 4. Response Data Generation Method
[0777] On the server side, appropriate response data is generated using a generative AI model (for example, a GPT model) based on the received text data, and this response data is again sent to the device via the network.
[0778] 5. Response data display method
[0779] The device visually displays the received response data to the user: in the case of smart glasses, this is done by using a display that shows the information in the user's field of view, while in the case of smartphones, this is done by displaying the information on a screen.
[0780] Specific examples
[0781] For example, if a user wears smart glasses and speaks a question such as "What is a recommended travel destination for my next holiday?", the question is converted into text data by a speech recognition means and sent to a server. The server uses the GPT model to generate response data such as "Kyoto is a recommended travel destination" and returns it to the device. In this way, the user can visually obtain the information "Kyoto is a recommended travel destination" on the display of the smart glasses.
[0782] Prompt Sentence Examples
[0783] The user enters the following prompt sentence as an example of a question:
[0784] "What's your recommended destination for your next holiday?"
[0785] The GPT model, upon seeing this prompt, generates the optimal response based on the specified conditions and presents it to the user.
[0786] By combining the above components, the present invention allows users to obtain information intuitively and quickly, contributing to more efficient information provision and improved user experience.
[0787] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0788] Step 1:
[0789] The user uses the smart glasses to input a question by voice.
[0790] (Input) The user's spoken question.
[0791] (Operation) A voice recognition means captures a voice query.
[0792] (Output) The captured audio data.
[0793] Step 2:
[0794] The terminal converts the voice data into text data.
[0795] (Input) The captured audio data.
[0796] (Operation) Analyze the voice data using a voice recognition method (for example, Google's voice recognition API) and convert it into text data.
[0797] (Output) Question data in text format.
[0798] Step 3:
[0799] The terminal transmits question data in text format to the server.
[0800] (Input) Question data in text format.
[0801] (Operation) Use an HTTP request library (for example, the requests library) to send text-formatted question data to the server.
[0802] (Output) The question data sent to the server in text format.
[0803] Step 4:
[0804] The server generates response data using a generative AI model based on the received text data.
[0805] (Input) The textual question data sent to the server.
[0806] (Operation) The server analyzes the text data using a generative AI model (GPT model) and generates optimal response data.
[0807] (Output) The generated response data (e.g., "Kyoto is a recommended travel destination").
[0808] Step 5:
[0809] The server sends the generated response data to the terminal.
[0810] (Input) The generated response data.
[0811] (Operation) The server sends response data to the terminal.
[0812] (Output) Response data sent to the terminal.
[0813] Step 6:
[0814] The response data received by the terminal is visually displayed to the user.
[0815] (Input) Response data sent to the terminal.
[0816] (Operation) The terminal display is used to display the response data within the user's field of view.
[0817] (Output) Response data that the user visually obtains (e.g., "Kyoto is a recommended travel destination").
[0818] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0819] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[0820] First, the user inputs a question or information using the device's touchscreen. For example, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and operation.
[0821] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent.
[0822] The server generates response data using generative AI (GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling depressed, the generated response might be something like, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0823] The generated response data is converted into an appropriate format and sent to the terminal, which then displays the received response data on a user interface in a format that is easy for the user to understand.
[0824] For example, consider the following user actions and system behaviors:
[0825] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[0826] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[0827] 3. The server uses a GPT model to analyze the received question and sentiment data and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0828] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[0829] 5. The terminal displays the received response to the user.
[0830] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[0831] In this way, the system of the present invention, by combining an emotion recognition engine and a generation AI, can provide high-quality information according to the user's emotional state. This allows users to receive more satisfying responses, and realizes a sustainable information provision environment in which information providers can receive fair rewards.
[0832] The processing flow will be explained below.
[0833] Step 1:
[0834] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0835] Step 2:
[0836] The device analyzes the input from the user and sends the content, input speed, touch pattern, voice tone, etc. to the emotion recognition engine, which then analyzes the user's emotions based on this data.
[0837] Step 3:
[0838] The device receives the input question and emotion data from the emotion recognition engine. For example, the emotion data may indicate that the user is depressed.
[0839] Step 4:
[0840] The device converts the question data and emotion data into text format and prepares an HTTP POST request to send to the server. The destination URL is specified, and the question data and emotion data are included in the HTTP request body.
[0841] Step 5:
[0842] The device executes an HTTP POST request and sends the question data and emotion data to the server.
[0843] Step 6:
[0844] The server receives the HTTP request sent from the terminal and extracts the question data and emotion data from the request body.
[0845] Step 7:
[0846] The server uses generative AI (GPT model) to generate response data based on the question data and emotion data extracted by the server. For example, if the user is feeling down, the server generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0847] Step 8:
[0848] The server formats the generated response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[0849] Step 9:
[0850] The server sends the formatted response data to the terminal.
[0851] Step 10:
[0852] The terminal receives the HTTP response from the server and analyzes the response data.
[0853] Step 11:
[0854] The terminal analyzes the response data and displays it on the user interface, which is displayed according to the user's emotions.
[0855] Step 12:
[0856] The user can check the screen of their device and see the response displayed: "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0857] This series of processes allows optimal information to be provided according to the user's emotional state.
[0858] Example 2
[0859] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0860] Conventional vending machine-type information service systems have difficulty providing appropriate responses according to the user's individual emotional state. Because users need different information and responses depending on their emotional state at any given time, standard responses do not provide sufficient satisfaction. It is necessary to solve this problem and improve user satisfaction.
[0861] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0862] In this invention, the server includes means for generating artificial intelligence that generates response data based on received text data and emotional data, means for transmitting the response data to the terminal, and means for adjusting parameters for generating an optimal response based on the user's question and emotional state, thereby making it possible to provide response data that is adjusted according to the user's emotional state.
[0863] The "means for accepting information input from the user" refers to an interface that allows the user to input questions or information into the terminal. Specifically, it includes input devices such as a touch screen and a keyboard.
[0864] The "means for converting the information input into text format data" refers to a process and system for digitizing the input information and converting it into a format that can be processed as text data.
[0865] The "means for transmitting the text format data to the server" is a device that includes a network connection and a communication protocol for communicating text data from the terminal to the server.
[0866] A "server equipped with artificial intelligence that generates response data adjusted according to the user's emotional state based on received text data" is a server device that includes an AI engine for analyzing text data and emotional data and generating appropriate responses.
[0867] The "means for transmitting the response data to the terminal" is a device that includes a network connection and a communication protocol for communicating the generated response data from the server to the terminal.
[0868] The "means for displaying response data to the user on the terminal" refers to a system including a display device and display software for visually presenting the response results to the user on the terminal.
[0869] "Generative artificial intelligence that generates response data based on received text data and emotional data" refers to an artificial intelligence algorithm that generates optimal responses based on input text data and the results of emotional analysis.
[0870] "Means for adjusting parameters to generate optimal responses based on the user's question and emotional state" refers to processes and systems that modify internal settings to optimize the output of an AI model, taking into account input from the user and their emotional state.
[0871] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[0872] First, the user uses the device's touchscreen to input a question or piece of information. For example, they might input a question like, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and how they operate the device. The emotion recognition engine used uses a common machine learning algorithm, for example, and can analyze input speed, touch strength, voice tone, and other factors.
[0873] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent. Communication is performed using an internet connection and communication protocols (e.g., HTTP or HTTPS).
[0874] The server generates response data using generative AI (e.g., a GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling down, a response such as "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is generated. For example, OpenAI's GPT-3 model is used as the generative AI model.
[0875] The generated response data is converted into an appropriate format and sent to the terminal. The terminal displays the received response data on a user interface, providing it in a format that is easy for the user to understand. The display device of the terminal used can be a touch screen or a display.
[0876] For example, consider the following user actions and system behaviors:
[0877] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[0878] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[0879] 3. The server uses a generative AI to analyze the question and emotional data it receives and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0880] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[0881] 5. The terminal displays the received response to the user.
[0882] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[0883] An example of a prompt might be:
[0884] "Please tell me where I should go on my next holiday. I'm feeling low."
[0885] This invention enables the provision of high-quality information tailored to the user's emotional state, thereby increasing user satisfaction. By combining generative AI with an emotion recognition engine, this system can provide more personalized responses.
[0886] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0887] Step 1:
[0888] The user uses the vending machine terminal's touchscreen to input a question or information, such as "What are some recommended travel destinations for my next holiday?" The input is temporarily stored in the terminal's internal memory.
[0889] Input: User question: "What are some recommended travel destinations for my next holiday?"
[0890] Output: Questions as text data
[0891] Specific operation: The user enters a question by touching the touchscreen and typing. The entered information is internally converted to text format.
[0892] Step 2:
[0893] The device uses a built-in emotion recognition engine to analyze the user's emotions based on their input and operation methods, such as typing speed, touch strength, and tone, to determine whether the user is depressed.
[0894] Input: Questions as text data, and user interaction data such as typing speed and touch strength
[0895] Output: Emotion data (e.g., "depressed")
[0896] Specific operation: The emotion recognition engine analyzes the user's input characteristics and estimates their emotional state. As a result of the analysis, emotional data such as "depressed" is generated.
[0897] Step 3:
[0898] The device sends text data and emotion data to the server using communication protocols such as HTTP and HTTPS.
[0899] Input: Questions as text data, emotion data
[0900] Output: Data sent to the server
[0901] Specific operation: The device uses a network connection to send the analyzed data to the server, which includes text data and emotion data.
[0902] Step 4:
[0903] The server generates response data using a generative AI model (e.g., a GPT model) based on the received text data and emotion data. The generated response is adjusted according to the user's emotional state.
[0904] Input: Received text data, emotion data
[0905] Output: Generated response data
[0906] Specific operation: The server's AI model generates the optimal response for the user based on the text data "What is the recommended travel destination for your next holiday?" and the emotion data "I feel depressed." Example: "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[0907] Step 5:
[0908] The generated response data is sent from the server to the terminal, again using a communication protocol such as HTTP or HTTPS.
[0909] Input: Generated response data
[0910] Output: Data sent to the terminal
[0911] Specific operation: The server formats the response data appropriately and sends it to the terminal. The sent data includes the response in text format.
[0912] Step 6:
[0913] The device displays the received response data on the user interface. For example, the text "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is displayed on the touch screen.
[0914] Input: Received response data
[0915] Output: Information displayed to the user
[0916] Specific operation: The terminal displays the received response data on the screen and provides information to the user. The user can check the displayed information and use it as reference.
[0917] Through these steps, users can receive information tailored to their emotional state. This system combines a generative AI model and an emotion recognition engine, aiming to increase user satisfaction.
[0918] (Application example 2)
[0919] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0920] Autonomous vehicles are required to recognize passenger emotions in real time and provide optimal services based on those emotions. This will improve passenger comfort and safety and enable more effective service provision. To solve this problem, the present invention provides a system that combines emotion recognition functionality with a generative AI model.
[0921] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text-format data, means for transmitting the text-format data to the server, means for receiving response data from the server, means for displaying the response data to the user, means for recognizing passenger emotions, means for transmitting the emotion information to a generative AI model, and means for adjusting the response data based on the emotion information and providing a service. This makes it possible to provide optimal services according to passenger emotions, thereby improving comfort and safety.
[0922] The "means for accepting information input from the user" is an interface for the user to input information, and may use a touch screen or a voice input device.
[0923] The "means for analyzing the information input and converting it into text format data" refers to a process for analyzing the information input by the user and converting it into digital text data.
[0924] The "means for transmitting the text format data to the server" refers to a communication means for transmitting the converted text data to the server via a network.
[0925] The "means for receiving response data from the server" refers to a communication means for receiving response data sent from the server.
[0926] The "means for displaying the response data to the user" refers to a display device for visually presenting the received response data to the user.
[0927] "Means for recognizing passenger emotions" refers to an emotion recognition engine that analyzes emotions from passengers' facial expressions, tone of voice, etc.
[0928] "Means for transmitting the emotion information to the generative AI model" refers to a communication means for transmitting the recognized emotion information to the generative AI model as input.
[0929] The "means for adjusting response data based on the emotion information and providing a service" refers to a means for providing response data generated based on emotion information in a form suitable for passengers.
[0930] The present invention relates to a system in which a user inputs information through a system for an autonomous vehicle and an optimal response is provided by a generative AI model. The system is implemented in the following configuration.
[0931] First, the user inputs information using the in-car touchscreen or voice input device, and an emotion recognition engine is activated to analyze the passenger's emotions from facial expressions, tone of voice, etc., and collect emotional data.
[0932] Next, the emotion data analyzed by the emotion recognition engine and the information entered by the user are converted into text using natural language processing technology. The converted text data and emotion data are then sent to a server via a network.
[0933] The server uses a generative AI model (e.g., a GPT model) to analyze the received text and emotion data and generate an appropriate response. This response is tailored based on the passenger's emotional state. For example, if the passenger is feeling depressed, the server might generate a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[0934] The generated response data is transmitted in real time to a terminal inside the autonomous vehicle, which then displays the received response data on a user interface in an easy-to-understand format for passengers.
[0935] Specifically, the device is equipped with a camera and microphone, which use OpenCV and the SpeechRecognition library to perform emotion analysis, and the server side uses the Transformers pipeline (the Hugging Face GPT model) as a generative AI model to generate responses.
[0936] Examples of specific prompts for generative AI models include:
[0937] "User sentiment: Depressed. User asks: 'What's the best place to go for my next vacation?'. Answer:"
[0938] Based on this prompt, the generative AI model generates a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[0939] This system allows passengers to receive the most appropriate service according to their mood, improving comfort and safety.
[0940] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0941] Step 1:
[0942] Users input information using the car's touchscreen or voice input device, such as a question like, "What are some recommended travel destinations for my next vacation?" The input information is treated as text data.
[0943] Step 2:
[0944] The device's emotion recognition engine uses the camera and microphone to analyze emotion data from the user's facial expressions and tone of voice. This process utilizes OpenCV and the SpeechRecognition library. The input is the user's voice and video data, and the output is emotion data such as "depression" or "joy."
[0945] Step 3:
[0946] The device collects the user's input information (text data) and emotion data and sends them to the server. The input is text data and emotion data, and the output is data sent to the server.
[0947] Step 4:
[0948] The server generates response data using a generative AI model (GPT model) based on the received text data and emotion data. First, it analyzes the text data, and then adjusts the response taking into account the emotion data. The input is text data and emotion data, and the output is the optimal response data to the user's question.
[0949] Step 5:
[0950] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the transmission of data to the terminal.
[0951] Step 6:
[0952] The terminal displays the received response data on the user interface. Specifically, it uses a touch screen to display a response such as "We recommend a place surrounded by beautiful nature for you to refresh yourself." The input is the response data sent from the server, and the output is visual information provided to the user.
[0953] Through these processing steps, optimal information is provided based on the user's emotions, improving passenger comfort and safety.
[0954] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0955] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0956] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0957] [Fourth embodiment]
[0958] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0959] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0960] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0961] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0962] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0963] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0964] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0965] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0966] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0967] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0968] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0969] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0970] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0971] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. Specifically, it provides a mechanism in which the user inputs a question, the server generates a response to the question, and displays it to the user through the terminal.
[0972] First, a user inputs a question using a user interface such as a touch screen on a vending machine terminal, which can be specific, such as "What are some recommended travel destinations for my next holiday?"
[0973] The terminal analyzes the input question and converts it into text data, which is then sent to a server via a communication method such as the Internet.
[0974] The server generates response data using generative AI (GPT model) based on the received text data. The generated response data, such as "Kyoto is a recommended travel destination," is then sent back to the device via communication means.
[0975] The terminal displays the response data received from the server to the user, allowing the user to easily obtain the necessary information from the vending machine.
[0976] This system is expected to have the effect of establishing a culture among users that recognizes the value of information and pays for it. It will also contribute to the creation of a sustainable information provision environment in which information providers can receive fair compensation.
[0977] For example, consider the following user actions and system behaviors:
[0978] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[0979] 2. The device converts the question into text data and sends it to the server.
[0980] 3. The server analyzes the received question using the GPT model and generates a response such as "Kyoto is a recommended travel destination."
[0981] 4. The generated response is sent to the terminal, which displays the response to the user.
[0982] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[0983] The unique feature of this system is that the service is provided through a physical vending machine, without the need for an internet connection. This has the advantage that users can easily obtain the information they need wherever they are, and it is thought that it can also be used as part of information literacy education, which is particularly useful for young people.
[0984] The processing flow will be explained below.
[0985] Step 1:
[0986] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[0987] Step 2:
[0988] The terminal parses the input from the user and obtains the question data as a string, which is then converted into a text format.
[0989] Step 3:
[0990] The device prepares an HTTP POST request to send text-formatted question data to the server, specifying the destination URL and including the question data in the body of the HTTP request.
[0991] Step 4:
[0992] The device executes an HTTP POST request and sends the query data to the server.
[0993] Step 5:
[0994] The server receives the HTTP request sent from the terminal and extracts the question data from the request body.
[0995] Step 6:
[0996] The server passes the extracted question data to the generation AI (GPT model) and generates a response. At this time, the API key is used to send a request to the API endpoint of the GPT model.
[0997] Step 7:
[0998] The server receives the response data returned from the GPT model. For example, it receives the text "Kyoto is a recommended travel destination."
[0999] Step 8:
[1000] The server formats the response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[1001] Step 9:
[1002] The server sends the formatted response data to the terminal.
[1003] Step 10:
[1004] The terminal receives the HTTP response from the server and analyzes the response data.
[1005] Step 11:
[1006] The terminal displays the analyzed response data on the user interface.
[1007] Step 12:
[1008] The user can check the screen of the device and see the displayed response, "Kyoto is a recommended travel destination."
[1009] Example 1
[1010] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1011] In modern society, much information is digitized, and users want to obtain the information they need quickly and easily. However, conventional information acquisition methods have problems such as the need for an Internet connection and uncertainty about the accuracy of the information. The purpose of this invention is to solve these problems and build a system that quickly provides highly reliable information.
[1012] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1013] In this invention, the server includes means for generating response data using a generative AI model based on received text data, means for transmitting the text data to the server using a communication means, and means for displaying the response data on a user interface, thereby making it possible to provide advanced information in real time by utilizing the generative AI model.
[1014] "Means for accepting information input" refers to an interface for a user to input information and a mechanism for receiving that information.
[1015] "Means for converting into text format data" refers to a process for recognizing information entered by a user as a digital string and converting it into text data.
[1016] "Means for transmitting to a server using a communication means" refers to infrastructure and software for transmitting text-format data to a server via a communication network such as the Internet.
[1017] The "means for receiving response data from the server" refers to a communication protocol and hardware for receiving response data sent from the server.
[1018] "Means for displaying on a user interface" refers to a display and display software for displaying received response data in a form that can be viewed by a user.
[1019] "Generative artificial intelligence model" refers to the machine learning model and associated algorithms used to generate appropriate responses based on received text data.
[1020] "Means for adjusting the algorithm" refers to the process of dynamically changing and tuning input parameters and processing methods so that the generative artificial intelligence model can output the optimal response.
[1021] This invention relates to a system that allows users to purchase generative AI-based information services by operating a terminal on a physical vending machine. The system uses a terminal, a server, and a generative AI model to provide users with efficient and accurate information.
[1022] First, the user uses the touchscreen of the vending machine terminal to input a question, such as the prompt "Where would you recommend for my next holiday?" The terminal is equipped with a high-precision touchscreen and software to process the input.
[1023] The terminal converts the information entered by the user into text data using character recognition software to convert the input string into digital data. The converted text data is then sent to a server via a communication method (e.g., an internet connection). The communication protocol used is HTTP or HTTPS.
[1024] The server analyzes the received text data and generates response data using a generative AI model (for example, OpenAI's GPT-4). The server has high-performance computing resources and can run Node.js or Python-based application servers. As a concrete example, the response generated is "Kyoto is a recommended travel destination."
[1025] The generated response data is then sent back to the terminal via the communication means. The terminal analyzes the received response data and displays it on the user interface. The display typically uses an LCD panel or an organic EL panel. The user can check this and obtain the necessary information.
[1026] This system allows users to obtain fast and accurate information through a physical vending machine. Its unique feature is that it does not require an internet connection, making it possible to conveniently obtain information anytime, anywhere. Furthermore, by utilizing generative AI models, advanced information is provided in real time, improving user satisfaction.
[1027] The specific operation flow is as follows:
[1028] 1. The user types "What's your recommended travel destination for my next holiday?" into the vending machine terminal's touchscreen.
[1029] 2. The device converts the question into text data and sends it to the server.
[1030] 3. The server analyzes the received question using a generative AI model and generates a response such as, "Kyoto is a recommended travel destination."
[1031] 4. The generated response is sent to the terminal, which displays the response to the user.
[1032] 5. The user looks at the device screen and receives the information, "Kyoto is a recommended travel destination."
[1033] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1034] Step 1:
[1035] The user operates the touch screen of the vending machine terminal and inputs a question.
[1036] Input: A question that the user types into the touchscreen (e.g., "What are your recommended travel destinations for my next holiday?").
[1037] Output: Text data recognized by the device.
[1038] How it works: The user types "What are some recommended travel destinations for my next holiday?" into the touchscreen, and the input information is captured by the touchscreen's sensors.
[1039] Step 2:
[1040] The terminal converts the question entered by the user into text data.
[1041] Input: Touchscreen input data.
[1042] Output: Data in text format (digital string).
[1043] Specific operation: The device's internal software analyzes the input data obtained from the sensor and converts it into text data in UTF-8 format.
[1044] Step 3:
[1045] The terminal transmits the converted text data to the server using a communication means.
[1046] Input: Data in text format.
[1047] Output: The HTTP POST request sent to the server.
[1048] Specific operation: An HTTP POST request is generated from the terminal to the server and sent to the server via the Internet.
[1049] Step 4:
[1050] The server analyzes the received text data and generates response data using a generative artificial intelligence model.
[1051] Input: Text data to be included in the HTTP POST request.
[1052] Output: The generated response data.
[1053] How it works: The server uses a Node.js or Python-based application server to send text data to the GPT-4 API and generate an appropriate response.
[1054] Step 5:
[1055] The server transmits the generated response data to the terminal using the communication means.
[1056] Input: The generated response data.
[1057] Output: The response data as an HTTP response.
[1058] Specific operation: The server generates response data in JSON format as an HTTP response and sends it to the terminal.
[1059] Step 6:
[1060] The terminal analyzes the response data received from the server and displays it on the user interface.
[1061] Input: The HTTP response data sent by the server.
[1062] Output: The text that appears in the user interface.
[1063] Specific operation: The device's software analyzes the HTTP response and displays "Kyoto is a recommended travel destination" on the touchscreen.
[1064] Step 7:
[1065] The user checks and acquires the information displayed on the screen of the terminal.
[1066] Input: The response data displayed in the user interface.
[1067] Output: The information obtained by the user.
[1068] Specific operation: The user visually recognizes the information displayed on the touch screen, "Kyoto is a recommended travel destination," and obtains the necessary information.
[1069] (Application example 1)
[1070] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1071] In conventional information provision systems using vending machines, users had to use a physical keyboard or touch screen to input questions, which made the input process cumbersome. In addition, there were limitations to the ways in which users could visually obtain information, so a more intuitive and faster way to obtain information was needed.
[1072] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1073] In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text data, means for transmitting the text data to the server, means for receiving response data from the server, means for displaying the response data to the user, speech recognition means for accepting voice input and converting it into text data, and means for visually displaying the response data, thereby enabling a user to ask questions using voice input and obtain responses in an intuitive manner.
[1074] "Means for accepting information input" refers to a device or method for accepting questions or instructions from a user.
[1075] "Means for analyzing and converting into text format data" refers to a device or function that analyzes input information and converts it into text format data that can be processed by a computer.
[1076] The "means for transmitting to a server" refers to a communication device or system for transmitting the converted text format data to a remote server.
[1077] The "means for receiving response data" refers to a device or method for receiving response data sent from a server.
[1078] The term "means for displaying response data to a user" refers to a device or system that visually displays the received response data to a user.
[1079] "Speech recognition means" refers to a device or technology that converts a user's voice input into digital data and then into an analyzable text format.
[1080] "Visual display means" refers to a device or system that displays data on a screen or display in a format that is easy for a user to understand.
[1081] This invention relates to a system that uses a generative AI model to provide information based on user input. The system for implementing the invention mainly includes the following components:
[1082] 1. Information input method
[1083] Users input questions by voice using devices such as smart glasses or smartphones. Using speech recognition technology (e.g., the speech_recognition library), the voice can be converted into text data.
[1084] 2. Analysis and text conversion methods
[1085] The voice input data is automatically converted into text format. For example, Google's speech recognition API is used to convert voice data into text data.
[1086] 3. Means of communication
[1087] The converted text data is sent to a remote server via the Internet using an HTTP request library (e.g., the requests library).
[1088] 4. Response Data Generation Method
[1089] On the server side, appropriate response data is generated using a generative AI model (for example, a GPT model) based on the received text data, and this response data is again sent to the device via the network.
[1090] 5. Response data display method
[1091] The device visually displays the received response data to the user: in the case of smart glasses, this is done by using a display that shows the information in the user's field of view, while in the case of smartphones, this is done by displaying the information on a screen.
[1092] Specific examples
[1093] For example, if a user wears smart glasses and speaks a question such as "What is a recommended travel destination for my next holiday?", the question is converted into text data by a speech recognition means and sent to a server. The server uses the GPT model to generate response data such as "Kyoto is a recommended travel destination" and returns it to the device. In this way, the user can visually obtain the information "Kyoto is a recommended travel destination" on the display of the smart glasses.
[1094] Prompt Sentence Examples
[1095] The user enters the following prompt sentence as an example of a question:
[1096] "What's your recommended destination for your next holiday?"
[1097] The GPT model, upon seeing this prompt, generates the optimal response based on the specified conditions and presents it to the user.
[1098] By combining the above components, the present invention allows users to obtain information intuitively and quickly, contributing to more efficient information provision and improved user experience.
[1099] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1100] Step 1:
[1101] The user uses the smart glasses to input a question by voice.
[1102] (Input) The user's spoken question.
[1103] (Operation) A voice recognition means captures a voice query.
[1104] (Output) The captured audio data.
[1105] Step 2:
[1106] The terminal converts the voice data into text data.
[1107] (Input) The captured audio data.
[1108] (Operation) Analyze the voice data using a voice recognition method (for example, Google's voice recognition API) and convert it into text data.
[1109] (Output) Question data in text format.
[1110] Step 3:
[1111] The terminal transmits question data in text format to the server.
[1112] (Input) Question data in text format.
[1113] (Operation) Use an HTTP request library (for example, the requests library) to send text-formatted question data to the server.
[1114] (Output) The question data sent to the server in text format.
[1115] Step 4:
[1116] The server generates response data using a generative AI model based on the received text data.
[1117] (Input) The textual question data sent to the server.
[1118] (Operation) The server analyzes the text data using a generative AI model (GPT model) and generates optimal response data.
[1119] (Output) The generated response data (e.g., "Kyoto is a recommended travel destination").
[1120] Step 5:
[1121] The server sends the generated response data to the terminal.
[1122] (Input) The generated response data.
[1123] (Operation) The server sends response data to the terminal.
[1124] (Output) Response data sent to the terminal.
[1125] Step 6:
[1126] The response data received by the terminal is visually displayed to the user.
[1127] (Input) Response data sent to the terminal.
[1128] (Operation) The terminal display is used to display the response data within the user's field of view.
[1129] (Output) Response data that the user visually obtains (e.g., "Kyoto is a recommended travel destination").
[1130] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1131] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[1132] First, the user inputs a question or information using the device's touchscreen. For example, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and operation.
[1133] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent.
[1134] The server generates response data using generative AI (GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling depressed, the generated response might be something like, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[1135] The generated response data is converted into an appropriate format and sent to the terminal, which then displays the received response data on a user interface in a format that is easy for the user to understand.
[1136] For example, consider the following user actions and system behaviors:
[1137] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[1138] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[1139] 3. The server uses a GPT model to analyze the received question and sentiment data and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[1140] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[1141] 5. The terminal displays the received response to the user.
[1142] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[1143] In this way, the system of the present invention, by combining an emotion recognition engine and a generation AI, can provide high-quality information according to the user's emotional state. This allows users to receive more satisfying responses, and realizes a sustainable information provision environment in which information providers can receive fair rewards.
[1144] The processing flow will be explained below.
[1145] Step 1:
[1146] The user operates the touchscreen of the vending machine terminal and inputs a question, for example, "What are some recommended travel destinations for my next holiday?"
[1147] Step 2:
[1148] The device analyzes the input from the user and sends the content, input speed, touch pattern, voice tone, etc. to the emotion recognition engine, which then analyzes the user's emotions based on this data.
[1149] Step 3:
[1150] The device receives the input question and emotion data from the emotion recognition engine. For example, the emotion data may indicate that the user is depressed.
[1151] Step 4:
[1152] The device converts the question data and emotion data into text format and prepares an HTTP POST request to send to the server. The destination URL is specified, and the question data and emotion data are included in the HTTP request body.
[1153] Step 5:
[1154] The device executes an HTTP POST request and sends the question data and emotion data to the server.
[1155] Step 6:
[1156] The server receives the HTTP request sent from the terminal and extracts the question data and emotion data from the request body.
[1157] Step 7:
[1158] The server uses generative AI (GPT model) to generate response data based on the question data and emotion data extracted by the server. For example, if the user is feeling down, the server generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[1159] Step 8:
[1160] The server formats the generated response data in an appropriate format (e.g., JSON) and prepares it as an HTTP response.
[1161] Step 9:
[1162] The server sends the formatted response data to the terminal.
[1163] Step 10:
[1164] The terminal receives the HTTP response from the server and analyzes the response data.
[1165] Step 11:
[1166] The terminal analyzes the response data and displays it on the user interface, which is displayed according to the user's emotions.
[1167] Step 12:
[1168] The user can check the screen of their device and see the response displayed: "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[1169] This series of processes allows optimal information to be provided according to the user's emotional state.
[1170] Example 2
[1171] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1172] Conventional vending machine-type information service systems have difficulty providing appropriate responses according to the user's individual emotional state. Because users need different information and responses depending on their emotional state at any given time, standard responses do not provide sufficient satisfaction. It is necessary to solve this problem and improve user satisfaction.
[1173] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1174] In this invention, the server includes means for generating artificial intelligence that generates response data based on received text data and emotional data, means for transmitting the response data to the terminal, and means for adjusting parameters for generating an optimal response based on the user's question and emotional state, thereby making it possible to provide response data that is adjusted according to the user's emotional state.
[1175] The "means for accepting information input from the user" refers to an interface that allows the user to input questions or information into the terminal. Specifically, it includes input devices such as a touch screen and a keyboard.
[1176] The "means for converting the information input into text format data" refers to a process and system for digitizing the input information and converting it into a format that can be processed as text data.
[1177] The "means for transmitting the text format data to the server" is a device that includes a network connection and a communication protocol for communicating text data from the terminal to the server.
[1178] A "server equipped with artificial intelligence that generates response data adjusted according to the user's emotional state based on received text data" is a server device that includes an AI engine for analyzing text data and emotional data and generating appropriate responses.
[1179] The "means for transmitting the response data to the terminal" is a device that includes a network connection and a communication protocol for communicating the generated response data from the server to the terminal.
[1180] The "means for displaying response data to the user on the terminal" refers to a system including a display device and display software for visually presenting the response results to the user on the terminal.
[1181] "Generative artificial intelligence that generates response data based on received text data and emotional data" refers to an artificial intelligence algorithm that generates optimal responses based on input text data and the results of emotional analysis.
[1182] "Means for adjusting parameters to generate optimal responses based on the user's question and emotional state" refers to processes and systems that modify internal settings to optimize the output of an AI model, taking into account input from the user and their emotional state.
[1183] This invention relates to a system that allows users to operate a vending machine terminal and provides generative AI-based information services in combination with an emotion recognition engine. The system aims to input information from the user, convert it into text data, send it to a server, generate response data, display it to the user, and further recognize the user's emotions and generate an optimal response that incorporates that information.
[1184] First, the user uses the device's touchscreen to input a question or piece of information. For example, they might input a question like, "What's the best place to travel to on my next holiday?" The device is equipped with an emotion recognition engine that analyzes the user's emotions based on their input and how they operate the device. The emotion recognition engine used uses a common machine learning algorithm, for example, and can analyze input speed, touch strength, voice tone, and other factors.
[1185] Next, this input information is converted into text data and sent to the server along with emotional data. For example, if the user is feeling depressed, emotional data indicating that state is also sent. Communication is performed using an internet connection and communication protocols (e.g., HTTP or HTTPS).
[1186] The server generates response data using generative AI (e.g., a GPT model) based on the received text data. The generated response is adjusted taking into account emotional data, providing the optimal response according to the user's emotions. For example, if the user is feeling down, a response such as "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is generated. For example, OpenAI's GPT-3 model is used as the generative AI model.
[1187] The generated response data is converted into an appropriate format and sent to the terminal. The terminal displays the received response data on a user interface, providing it in a format that is easy for the user to understand. The display device of the terminal used can be a touch screen or a display.
[1188] For example, consider the following user actions and system behaviors:
[1189] 1. A user types "What's the best place to go for my next holiday?" into the touchscreen of a vending machine terminal. The emotion recognition engine detects a depressed emotion from the user's tone and typing speed.
[1190] 2. The device converts the question into text data and sends it to the server along with emotional data (e.g., "I feel depressed").
[1191] 3. The server uses a generative AI to analyze the question and emotional data it receives and generates a response such as, "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[1192] 4. The generated response is adjusted taking into account the emotional data and sent to the device.
[1193] 5. The terminal displays the received response to the user.
[1194] 6. The user checks the screen of their device and receives the information, "We recommend Kyoto, surrounded by beautiful nature, to refresh yourself," and is convinced.
[1195] An example of a prompt might be:
[1196] "Please tell me where I should go on my next holiday. I'm feeling low."
[1197] This invention enables the provision of high-quality information tailored to the user's emotional state, thereby increasing user satisfaction. By combining generative AI with an emotion recognition engine, this system can provide more personalized responses.
[1198] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1199] Step 1:
[1200] The user uses the vending machine terminal's touchscreen to input a question or information, such as "What are some recommended travel destinations for my next holiday?" The input is temporarily stored in the terminal's internal memory.
[1201] Input: User question: "What are some recommended travel destinations for my next holiday?"
[1202] Output: Questions as text data
[1203] Specific operation: The user enters a question by touching the touchscreen and typing. The entered information is internally converted to text format.
[1204] Step 2:
[1205] The device uses a built-in emotion recognition engine to analyze the user's emotions based on their input and operation methods, such as typing speed, touch strength, and tone, to determine whether the user is depressed.
[1206] Input: Questions as text data, and user interaction data such as typing speed and touch strength
[1207] Output: Emotion data (e.g., "depressed")
[1208] Specific operation: The emotion recognition engine analyzes the user's input characteristics and estimates their emotional state. As a result of the analysis, emotional data such as "depressed" is generated.
[1209] Step 3:
[1210] The device sends text data and emotion data to the server using communication protocols such as HTTP and HTTPS.
[1211] Input: Questions as text data, emotion data
[1212] Output: Data sent to the server
[1213] Specific operation: The device uses a network connection to send the analyzed data to the server, which includes text data and emotion data.
[1214] Step 4:
[1215] The server generates response data using a generative AI model (e.g., a GPT model) based on the received text data and emotion data. The generated response is adjusted according to the user's emotional state.
[1216] Input: Received text data, emotion data
[1217] Output: Generated response data
[1218] Specific operation: The server's AI model generates the optimal response for the user based on the text data "What is the recommended travel destination for your next holiday?" and the emotion data "I feel depressed." Example: "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself."
[1219] Step 5:
[1220] The generated response data is sent from the server to the terminal, again using a communication protocol such as HTTP or HTTPS.
[1221] Input: Generated response data
[1222] Output: Data sent to the terminal
[1223] Specific operation: The server formats the response data appropriately and sends it to the terminal. The sent data includes the response in text format.
[1224] Step 6:
[1225] The device displays the received response data on the user interface. For example, the text "I recommend Kyoto, surrounded by beautiful nature, to refresh yourself" is displayed on the touch screen.
[1226] Input: Received response data
[1227] Output: Information displayed to the user
[1228] Specific operation: The terminal displays the received response data on the screen and provides information to the user. The user can check the displayed information and use it as reference.
[1229] Through these steps, users can receive information tailored to their emotional state. This system combines a generative AI model and an emotion recognition engine, aiming to increase user satisfaction.
[1230] (Application example 2)
[1231] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1232] Autonomous vehicles are required to recognize passenger emotions in real time and provide optimal services based on those emotions. This will improve passenger comfort and safety and enable more effective service provision. To solve this problem, the present invention provides a system that combines emotion recognition functionality with a generative AI model.
[1233] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for accepting information input from a user, means for analyzing the information input and converting it into text-format data, means for transmitting the text-format data to the server, means for receiving response data from the server, means for displaying the response data to the user, means for recognizing passenger emotions, means for transmitting the emotion information to a generative AI model, and means for adjusting the response data based on the emotion information and providing a service. This makes it possible to provide optimal services according to passenger emotions, thereby improving comfort and safety.
[1234] The "means for accepting information input from the user" is an interface for the user to input information, and may use a touch screen or a voice input device.
[1235] The "means for analyzing the information input and converting it into text format data" refers to a process for analyzing the information input by the user and converting it into digital text data.
[1236] The "means for transmitting the text format data to the server" refers to a communication means for transmitting the converted text data to the server via a network.
[1237] The "means for receiving response data from the server" refers to a communication means for receiving response data sent from the server.
[1238] The "means for displaying the response data to the user" refers to a display device for visually presenting the received response data to the user.
[1239] "Means for recognizing passenger emotions" refers to an emotion recognition engine that analyzes emotions from passengers' facial expressions, tone of voice, etc.
[1240] "Means for transmitting the emotion information to the generative AI model" refers to a communication means for transmitting the recognized emotion information to the generative AI model as input.
[1241] The "means for adjusting response data based on the emotion information and providing a service" refers to a means for providing response data generated based on emotion information in a form suitable for passengers.
[1242] The present invention relates to a system in which a user inputs information through a system for an autonomous vehicle and an optimal response is provided by a generative AI model. The system is implemented in the following configuration.
[1243] First, the user inputs information using the in-car touchscreen or voice input device, and an emotion recognition engine is activated to analyze the passenger's emotions from facial expressions, tone of voice, etc., and collect emotional data.
[1244] Next, the emotion data analyzed by the emotion recognition engine and the information entered by the user are converted into text using natural language processing technology. The converted text data and emotion data are then sent to a server via a network.
[1245] The server uses a generative AI model (e.g., a GPT model) to analyze the received text and emotion data and generate an appropriate response. This response is tailored based on the passenger's emotional state. For example, if the passenger is feeling depressed, the server might generate a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[1246] The generated response data is transmitted in real time to a terminal inside the autonomous vehicle, which then displays the received response data on a user interface in an easy-to-understand format for passengers.
[1247] Specifically, the device is equipped with a camera and microphone, which use OpenCV and the SpeechRecognition library to perform emotion analysis, and the server side uses the Transformers pipeline (the Hugging Face GPT model) as a generative AI model to generate responses.
[1248] Examples of specific prompts for generative AI models include:
[1249] "User sentiment: Depressed. User asks: 'What's the best place to go for my next vacation?'. Answer:"
[1250] Based on this prompt, the generative AI model generates a response such as, "I recommend a place surrounded by beautiful nature to refresh yourself."
[1251] This system allows passengers to receive the most appropriate service according to their mood, improving comfort and safety.
[1252] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1253] Step 1:
[1254] Users input information using the car's touchscreen or voice input device, such as a question like, "What are some recommended travel destinations for my next vacation?" The input information is treated as text data.
[1255] Step 2:
[1256] The device's emotion recognition engine uses the camera and microphone to analyze emotion data from the user's facial expressions and tone of voice. This process utilizes OpenCV and the SpeechRecognition library. The input is the user's voice and video data, and the output is emotion data such as "depression" or "joy."
[1257] Step 3:
[1258] The device collects the user's input information (text data) and emotion data and sends them to the server. The input is text data and emotion data, and the output is data sent to the server.
[1259] Step 4:
[1260] The server generates response data using a generative AI model (GPT model) based on the received text data and emotion data. First, it analyzes the text data, and then adjusts the response taking into account the emotion data. The input is text data and emotion data, and the output is the optimal response data to the user's question.
[1261] Step 5:
[1262] The server sends the generated response data to the terminal. The input is the generated response data, and the output is the transmission of data to the terminal.
[1263] Step 6:
[1264] The terminal displays the received response data on the user interface. Specifically, it uses a touch screen to display a response such as "We recommend a place surrounded by beautiful nature for you to refresh yourself." The input is the response data sent from the server, and the output is visual information provided to the user.
[1265] Through these processing steps, optimal information is provided based on the user's emotions, improving passenger comfort and safety.
[1266] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1267] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1268] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1269] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1270] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1271] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1272] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1273] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1274] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1275] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1276] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1277] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1278] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1279] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1280] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1281] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1282] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1283] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1284] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1285] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1286] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1287] The following is further disclosed regarding the above embodiment.
[1288] (Claim 1)
[1289] means for accepting information input from a user;
[1290] means for analyzing the information input and converting it into text format data;
[1291] means for transmitting the text format data to a server;
[1292] means for receiving response data from the server;
[1293] The system includes means for displaying said response data to a user.
[1294] (Claim 2)
[1295] The system according to claim 1, characterized in that the server generates response data using generative artificial intelligence (GPT model) based on the received text data and transmits the response data to the terminal.
[1296] (Claim 3)
[1297] The system of claim 1, further comprising means for adjusting parameters for generating an optimal response based on a user's question when using the generative artificial intelligence (GPT model).
[1298] "Example 1"
[1299] (Claim 1)
[1300] means for accepting information input from a user;
[1301] means for analyzing the information input and converting it into text format data;
[1302] means for transmitting the text format data to a server using a communication means;
[1303] means for receiving response data from the server;
[1304] The system includes means for displaying the response data in a user interface.
[1305] (Claim 2)
[1306] 2. The system according to claim 1, wherein the server generates response data using a generative artificial intelligence model based on the received text data and transmits the response data to the terminal.
[1307] (Claim 3)
[1308] 10. The system of claim 1, further comprising means for adjusting an algorithm for generating an optimal response based on a user's question when using the generative artificial intelligence model.
[1309] "Application Example 1"
[1310] (Claim 1)
[1311] means for accepting information input from a user;
[1312] means for analyzing the information input and converting it into text format data;
[1313] means for transmitting the text format data to a server;
[1314] means for receiving response data from the server;
[1315] means for displaying the response data to a user;
[1316] Including a voice recognition means for accepting voice input and converting it into text data;
[1317] and means for visually displaying said response data.
[1318] (Claim 2)
[1319] The system according to claim 1, characterized in that the server generates response data using generative artificial intelligence (GPT model) based on the received text data and transmits the response data to the terminal.
[1320] (Claim 3)
[1321] The system of claim 1, further comprising means for adjusting parameters for generating an optimal response based on a user's question when using the generative artificial intelligence (GPT model).
[1322] "Example 2: Combining Emotion Engines"
[1323] (Claim 1)
[1324] means for accepting information input from a user;
[1325] means for converting the information input into text format data;
[1326] means for transmitting the text format data to a server;
[1327] a server having a generation artificial intelligence that generates response data adjusted according to the user's emotional state based on the received text data;
[1328] means for transmitting the response data to a terminal;
[1329] The system includes means for displaying response data to a user at said terminal.
[1330] (Claim 2)
[1331] 2. The system according to claim 1, wherein the server generates response data using a generative artificial intelligence based on the received text data and emotion data, and transmits the response data to the terminal.
[1332] (Claim 3)
[1333] 2. The system of claim 1, further comprising means for adjusting parameters for generating an optimal response based on a user's question and emotional state when using the generative artificial intelligence.
[1334] "Application example 2 when combining emotion engines"
[1335] (Claim 1)
[1336] means for accepting information input from a user;
[1337] means for analyzing the information input and converting it into text format data;
[1338] means for transmitting the text format data to a server;
[1339] means for receiving response data from the server;
[1340] means for displaying the response data to a user;
[1341] a means for recognizing passenger emotions;
[1342] means for transmitting the emotion information to a generative AI model;
[1343] The system includes means for adjusting response data based on the emotion information and providing a service.
[1344] (Claim 2)
[1345] The system according to claim 1, characterized in that the server generates response data using a generative artificial intelligence (generative AI model) based on the received text data and emotional information, and transmits the response data to the terminal.
[1346] (Claim 3)
[1347] The system of claim 1, further comprising means for adjusting parameters for generating an optimal response based on a user's question and emotional information when using the generative artificial intelligence (generative AI model). [Explanation of symbols]
[1348] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for accepting information input from a user; means for analyzing the information input and converting it into text format data; means for transmitting the text format data to a server; means for receiving response data from the server; The system includes means for displaying said response data to a user.
2. 2. The system according to claim 1, wherein the server generates response data using artificial intelligence based on the received text data and transmits the response data to the terminal.
3. 2. The system of claim 1, further comprising means for adjusting parameters for generating an optimal response based on a user's question when using said generative artificial intelligence.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A