System
The system addresses the limitations of large-scale language models by collecting and structuring user and external data for real-time, personalized, and reliable responses without the need for retraining.
Patent Information
- Application Number
- JP2024133575
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Current large-scale language models face challenges in quickly updating to provide the latest and personalized information, leading to potential inaccuracies and reliability issues in user responses.
A system that collects user conversation logs and reliable external information, stores them in a structured database, and uses a large-scale language model to generate responses, allowing for real-time updates without retraining.
Enables the provision of accurate, personalized, and up-to-date information by efficiently integrating user logs and external data, ensuring timely and reliable responses.
Smart Images

Figure 2026030591000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current large-scale language models have problems such as difficulty in quickly updating the latest information and providing personalized responses. Furthermore, the generated information may lack reliability, which increases the risk of users receiving incorrect information. There is a need to provide a system that can solve these problems and ensure users always receive the latest, reliable, and customized information. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for collecting user conversation logs, a means for collecting reliable external information, a means for storing these conversation logs and external information in a structured database, a means for generating responses using a large-scale language model based on the stored data, and a means for transmitting the generated responses to the user's terminal. This system can respond to individual user needs and always provide the latest and most reliable information. Furthermore, the use of a structured database allows updates to information to be reflected without the need for re-learning, so the latest information can be provided to the user quickly and efficiently.
[0006] "User conversation log" refers to the history of interactions with a user, such as text, voice, and images entered by the user, and is data converted into a format that can be processed by the system.
[0007] "Reliable external information" refers to accurate and up-to-date data provided by third parties, such as weather information, news, and restaurant reservation status.
[0008] A "structured database" is a database in which the user's conversation logs and reliable external information are systematically organized and stored in a format that is easy to search and use.
[0009] A "large-scale language model" refers to a machine learning model that can learn from massive amounts of text data and understand and generate natural language.
[0010] "Generating a response" refers to the process of generating and outputting appropriate information or answers based on user input and stored data.
[0011] "Terminal" refers to a device used by a user, such as a computer, smartphone, or tablet, which provides an interface between the system and the user.
[0012] "Data-based" refers to processing and responding according to the contents of a stored structured database.
[0013] "Without the need for retraining" refers to a state in which a trained model can directly use new information to generate responses without retraining, even when new information is added. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention relates to a system that collects user conversation logs and reliable external information, stores them in a structured database, generates responses using a large-scale language model (LLM) based on the stored data, and transmits the generated responses to the user's device. By using this system, it is possible to respond to individual needs and always provide the latest and most reliable information.
[0036] System configuration
[0037] 1. Collecting user conversation logs
[0038] Users interact with the system on a daily basis via their smartphones or computers. For example, a user might input a question such as, "Which restaurants are open this afternoon?" This interaction is digitally recorded on the user's device and periodically sent to the server.
[0039] 2. Gathering reliable external information
[0040] The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[0041] 3. Data structuring and storage
[0042] The server analyzes the user's conversation log and external information and stores them in a structured database. The conversation log is categorized by user, and reliable external information is organized by category. For example, a user's conversation log includes "question content," "date and time," and "frequency," while external information includes "weather information" and "restaurant availability."
[0043] 4. Generating the Response
[0044] The server retrieves the necessary information from a structured database and provides it to a large-scale language model. LLM generates the optimal response based on this information. For example, if a user asks about restaurant availability, LLM generates the response "Italian, French, and Japanese restaurants are open in the afternoon."
[0045] 5. Sending the Response
[0046] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[0047] Specific examples
[0048] 1. User Questions
[0049] User: "Which restaurants are open this afternoon?"
[0050] 2. Sending conversation logs
[0051] The terminal sends this question to the server.
[0052] 3. Search for related information
[0053] The server searches the database for the user's past conversation log and the latest restaurant vacancy information.
[0054] 4. Generating the Response
[0055] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[0056] 5. Sending and Displaying Responses
[0057] The server sends the generated response to the terminal, which displays the response to the user.
[0058] In this way, the system can provide customized responses in real time, and the use of a structured database allows it to quickly update without the need for retraining.
[0059] The processing flow will be explained below.
[0060] Step 1:
[0061] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[0062] Step 2:
[0063] The device records the conversation in real time, saves it as text data, and then periodically sends the conversation log to a server.
[0064] Step 3:
[0065] The server analyzes the received conversation logs and stores them in a database for each user. The analysis process includes extracting important information such as question content, time, and frequency.
[0066] Step 4:
[0067] The server periodically retrieves the latest data from external trusted sources (e.g., weather information APIs or restaurant reservation sites), which is also preprocessed and stored in a structured database.
[0068] Step 5:
[0069] When the device receives a new question from the user, it sends the content to the server, where it is processed in real time.
[0070] Step 6:
[0071] Based on the received question, the server searches the database for relevant conversation logs and external information, for example, past conversation history and the latest restaurant information.
[0072] Step 7:
[0073] The server then passes the search results to a large-scale language model (LLM) to generate the best possible response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[0074] Step 8:
[0075] The server sends the generated response to the terminal.
[0076] Step 9:
[0077] The terminal displays the received response to the user, who then decides on the next action to take.
[0078] Step 10:
[0079] The server periodically updates the external information and stores it in the database again, so that the latest information is always available within the system.
[0080] In this way, the system provides fast and accurate responses to user questions, and by utilizing a structured database and large-scale language models, it is possible to reflect the latest information in real time without the need for retraining.
[0081] Example 1
[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0083] Conventional systems have had difficulty efficiently collecting and organizing user conversation logs and combining them with reliable external information to provide optimal responses in real time. Furthermore, because the external information is not updated frequently enough, there is no guarantee that the information provided is up-to-date and accurate. Furthermore, generating effective prompts has been a challenge when utilizing large-scale language models (LLMs).
[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0085] In this invention, the server includes means for digitally collecting user conversation logs and transmitting them from the terminal to the server, means for periodically collecting reliable external information from APIs and information sites, and means for analyzing and classifying the conversation logs and external information and storing them in a structured database. This enables the generation of accurate responses that meet individual needs based on the latest, most reliable information in real time.
[0086] "User conversation log" refers to the content of communication such as questions and requests made by the user to the system.
[0087] "Digital format" refers to a format in which information, such as sound or text, can be recorded and processed electronically.
[0088] "Terminal" refers to a device, such as a smartphone or computer, that a user uses to access the system.
[0089] "Server" refers to a computer system for receiving and processing data sent from a user's terminal and generating a response.
[0090] "Reliable external information" refers to information obtained from reliable sources such as weather information APIs and restaurant reservation sites.
[0091] "API" stands for Application Programming Interface and refers to a protocol for exchanging information between different software programs.
[0092] An "information site" refers to a web page or online service that provides specific information.
[0093] "Analysis" refers to the process of breaking down data into more detailed pieces to make it easier to understand.
[0094] "Classification" refers to the process of grouping collected data based on specific criteria.
[0095] A "structured database" refers to a database in which data is systematically organized so that it can be efficiently stored and searched.
[0096] A "large-scale language model (LLM)" refers to a natural language processing model trained on large amounts of text data.
[0097] A "prompt sentence" refers to a guided text that is input into a large-scale language model to generate an appropriate response.
[0098] "Response" refers to the answer or information generated or provided by the system in response to a user's question.
[0099] The present invention is a system that collects user conversation logs and reliable external information, stores them in a structured database, and processes them. This system generates responses based on the collected data using a large-scale language model (LLM), and transmits the generated responses to the user's terminal for display. A specific embodiment of the system is described below.
[0100] Hardware and Software Configuration
[0101] 1. Hardware
[0102] User terminal: A device through which a user accesses the system, such as a smartphone or computer.
[0103] Server: A computer system, such as a cloud server, that receives data sent from a user's device, processes it, and generates a response.
[0104] 2. Software
[0105] Conversation log collection software: Software for collecting user conversations in digital form.
[0106] External data collection software: Software for collecting reliable external information from APIs and information sites.
[0107] Structured database software: Software for analyzing, classifying, and storing collected data.
[0108] Large-scale language model (LLM) software: Software that uses natural language processing techniques to generate responses based on collected data.
[0109] System configuration and operation
[0110] The system operates in the following steps.
[0111] 1. Collecting user conversation logs
[0112] User Asks a Question: A user uses a smartphone or computer to send a question or request to the system. Example: "Which restaurants are open this afternoon?"
[0113] Device processing: The device digitally records the user's conversations and periodically transmits them to a server, in particular speech recognition software that may convert speech to text.
[0114] 2. Gathering external information
[0115] Server information collection: The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. Example: Obtaining "Tokyo weather information" or "restaurant seat availability information."
[0116] 3. Data structuring and storage
[0117] Server data analysis: The server analyzes the received conversation logs and external information, and classifies and stores them in a structured database. Specifically, tokenization and tagging are performed.
[0118] Conversation log classification: User conversation logs are classified by user ID, and external information is organized by category.
[0119] 4. Generating the Response
[0120] Server data retrieval: The server retrieves the information required by the user from a structured database and provides it to the LLM.
[0121] LLM response generation: LLM generates the best response based on the data provided. Example: "Italian, French, and Japanese restaurants are open in the afternoon."
[0122] 5. Sending the Response
[0123] Transmission from server to terminal: The generated response is transmitted to the user's terminal via the server.
[0124] Display on the terminal: The terminal displays the received response to the user, who can then ask further questions based on the displayed information.
[0125] Specific examples
[0126] Below are some examples of specific prompt sentences.
[0127] 1. User Questions
[0128] User: "Which restaurants are open this afternoon?"
[0129] 2. Sending the device
[0130] The terminal sends this question in text format to the server.
[0131] 3. Searching for a server
[0132] The server retrieves the conversation log and the latest restaurant information from a database.
[0133] 4. Generating the Response
[0134] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[0135] 5. Sending and Displaying Responses
[0136] The server sends the generated response to the terminal, which displays the response to the user.
[0137] The system's features include the ability to respond quickly and accurately to user needs, collecting and providing the latest information in real time, and utilizing a structured database to generate responses with high efficiency without the need for re-learning.
[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0139] Step 1:
[0140] A user enters a question or request into the system.
[0141] Input: User question (e.g., "Which restaurants are open this afternoon?")
[0142] Specific operation: The user accesses the system using a smartphone or computer and enters a question.
[0143] Step 2:
[0144] The device collects the user's conversation log in digital format and sends it to the server.
[0145] Input: User conversation log
[0146] Output: Digital conversation log
[0147] Specific operation: The device converts the conversation log into text format, adds metadata such as the user ID, question content, date and time, and sends it to the server periodically.
[0148] Step 3:
[0149] The server periodically collects reliable external information from APIs and information sites.
[0150] Input: URL of external information API endpoint or information site
[0151] Output: External information data (e.g., weather information, restaurant seat availability information)
[0152] Specific operation: The server periodically retrieves the latest data from external sources such as weather information APIs and restaurant reservation sites.
[0153] Step 4:
[0154] The server analyzes the conversation log and external information and stores it in a structured database.
[0155] Input: Digital conversation logs and external information data
[0156] Output: Parsed and classified data
[0157] How it works: The server uses natural language processing technology to tokenize the conversation logs and organize external information into categories. The analyzed and categorized data is then stored in a structured database.
[0158] Step 5:
[0159] The server retrieves the necessary information from a structured database and provides it to the LLM.
[0160] Input: User conversation logs and external information
[0161] Output: Data provided to LLM
[0162] Specific operation: The server searches the user's past conversation history and the latest external information from a structured database and provides them to the LLM.
[0163] Step 6:
[0164] LLM generates the best response based on the data provided.
[0165] Input: User conversation logs and external information
[0166] Output: Generated response text (e.g., "Italian, French, and Japanese restaurants are open in the afternoon.")
[0167] How it works: LLM uses natural language processing algorithms to generate optimal responses based on tokenized data.
[0168] Step 7:
[0169] The server sends the generated response to the user's terminal.
[0170] Input: Generated response text
[0171] Output: The response sent to the user's device
[0172] Specific operation: The server sends the generated response to the user's device, transferring data in real time using a communication protocol.
[0173] Step 8:
[0174] The terminal displays the received response to the user.
[0175] Input: Response text sent by the server
[0176] Output: The response message to be displayed
[0177] Specific operation: The terminal displays the received response on the user interface, and the user can ask more detailed questions based on this information.
[0178] Through these steps, the system is able to provide accurate and appropriate responses to user questions in real time.
[0179] (Application example 1)
[0180] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0181] Modern food delivery services are required to respond quickly and accurately to diverse user needs. However, conventional systems have difficulty providing personalized suggestions based on user preferences or real-time restaurant information. This has led to issues such as users being unable to quickly make choices that suit their preferences and current circumstances, resulting in reduced convenience.
[0182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0183] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response based on the stored data using a large-scale language model, means for transmitting the generated response to the user's terminal, and means for providing customized recommendations based on the user's past ordering history and preferences. This allows users to receive real-time suggestions that match their preferences, significantly improving the convenience of food delivery.
[0184] A "conversation log" is text data generated through a user's interaction with the system, and is information that includes the user's questions and responses.
[0185] "External information" refers to the latest data the system collects from trusted external sources, such as restaurant menus, seat availability, and weather information.
[0186] A "structured database" is a database for organizing and storing conversation logs and external information, in which each piece of information is organized by classification and category.
[0187] A "large-scale language model" is an advanced machine learning model for natural language processing that generates optimal responses based on data provided by a structured database.
[0188] A "terminal" is a digital device used by a user, such as a smartphone or computer, on which the generated response is displayed.
[0189] "Customized Recommendations" are suggestions that are specifically tailored based on a user's past ordering history and preferences, and are information designed to address a user's individual needs.
[0190] A system embodying the present invention includes the following functions.
[0191] The server first collects the user's conversation log. The user interacts with the system via a smartphone or computer, and the conversation is recorded digitally. For example, the user may enter a question such as, "Which restaurants are available now?" This question is collected and sent to the server.
[0192] Next, the server collects reliable external information, such as restaurant menus, seat availability, and weather information, which it periodically obtains via online APIs. For example, the server obtains current seat availability information from a restaurant reservation site.
[0193] The server then stores the collected conversation logs and external information in a structured database. Conversation logs are organized by user, and external information is categorized. Specifically, a user's conversation log includes "question content," "date and time," and "frequency," while reliable external information includes "restaurant seat availability information" and "menu information."
[0194] The server then uses the stored data to generate a response using a large-scale language model (LLM). The server searches a structured database for the necessary information and sends a prompt to the LLM based on this information. For example, if a user asks about restaurant availability, the LLM can generate a response such as, "As of 10:00 AM, an Italian restaurant and a Chinese restaurant are available."
[0195] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[0196] Additionally, the server provides customized recommendations based on the user's past ordering history and preferences, allowing the user to quickly find options that suit their preferences.
[0197] For example, the following prompts might be used for a generative AI model:
[0198] text
[0199] User asks: "What Italian restaurant would you recommend?"
[0200] Past conversation: [{ "timestamp": "2023-10-08T12:00:00", "conversation": "I like pizza"}, {...}]
[0201] External Data: [{ "name": "Italian Restaurant A", "rating": 4.5, "availability": "open"}, {...}]
[0202] Based on this prompt, the LLM processes past conversation data and external information to generate a response such as, "Based on your preferences, we recommend Italian restaurant A." In this way, the present invention can provide optimal information tailored to the user's individual needs.
[0203] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0204] Step 1:
[0205] The user enters a question and the terminal receives the question.
[0206] Specifically, when a user uses a smartphone or computer to enter a question such as "Which restaurants are open now?", the text data is sent to the device and saved.
[0207] Input: User question
[0208] Output: Question data in text format
[0209] Step 2:
[0210] The terminal transmits the saved question data to the server.
[0211] This data is sent to the server along with the user ID, and the server receives it and records it as a conversation log.
[0212] Input: Text question data, user ID
[0213] Output: Conversation log sent to the server
[0214] Step 3:
[0215] The server collects trusted external information.
[0216] Specifically, the server accesses a pre-configured API endpoint (e.g., the API of a restaurant reservation system) and obtains the latest data such as restaurant seat availability and menu information.
[0217] Input: API endpoint
[0218] Output: External information data (e.g., restaurant seat availability, menu information)
[0219] Step 4:
[0220] The server stores the collected conversation logs and external information in a structured database.
[0221] Conversation logs are categorized by user, and external information is organized by category. For example, restaurant vacancy information is organized into categories such as "restaurant" and "vacancy information."
[0222] Input: Conversation log, external information
[0223] Output: Structured database entries
[0224] Step 5:
[0225] The server searches for the necessary information from a structured database and uses it to generate and send prompts to a large-scale language model (LLM).
[0226] The prompt includes the user's question, past conversation logs, and external information. The prompt is created and sent to the LLM to generate the optimal response.
[0227] Input: Structured database information, user questions
[0228] Output: Generated prompt, response from LLM
[0229] Step 6:
[0230] The server sends the generated response to the user's terminal.
[0231] The terminal displays the received response to the user, thereby providing the user with an answer to their question.
[0232] Input: The response generated by the LLM
[0233] Output: The response displayed on the user's terminal
[0234] Step 7:
[0235] The server generates and provides customized recommendations to the user based on the user's past ordering history and preferences.
[0236] By analyzing past conversation logs and order data, the system suggests restaurants and menus that suit the user's preferences.
[0237] Input: User's past order history, preference data
[0238] Output: Customized recommendations
[0239] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0240] This invention relates to a system that collects user conversation logs and reliable external information, generates responses based on these using a large-scale language model (LLM), and combines this with an emotion engine that recognizes the user's emotions. This system is capable of providing customized responses in real time according to the user's individual needs and emotional state. Furthermore, by utilizing a structured database, it is possible to quickly reflect the latest information without the need for re-learning.
[0241] System configuration
[0242] 1. Collecting user conversation logs
[0243] Users interact with the system using a smartphone or computer. For example, they might input a question like, "Which restaurants are open this afternoon?" The content of this interaction is recorded in real time on the device and periodically sent to the server.
[0244] 2. Gathering reliable external information
[0245] The server periodically obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[0246] 3. Operation of the Emotion Engine
[0247] When the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[0248] 4. Data structuring and storage
[0249] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[0250] 5. Generating the Response
[0251] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[0252] 6. Sending the Response
[0253] The generated response is sent via the server to the terminal and displayed to the user, who then decides what to do next.
[0254] Specific examples
[0255] 1. User Questions
[0256] User: "Which restaurants are open this afternoon?"
[0257] 2. Sending conversation logs
[0258] The terminal sends this question to the server.
[0259] 3. Sentiment Analysis
[0260] The server uses an emotion engine to identify the emotion "interested" from the user's question and stores it in a database.
[0261] 4. Searching for related information
[0262] The server searches a database for user conversation logs, emotional data, and the latest restaurant seat availability information.
[0263] 5. Generating the Response
[0264] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[0265] 6. Sending and Displaying Responses
[0266] The server sends the generated response to the terminal, which displays the response to the user.
[0267] In this way, the system can provide customized responses in real time according to the user's emotional state. The introduction of an emotion engine further improves the user experience. It can provide information quickly and efficiently because it can reflect the latest information in real time without the need for retraining.
[0268] The processing flow will be explained below.
[0269] Step 1:
[0270] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[0271] Step 2:
[0272] The device records the conversation in real time, saves it as text data, and then sends the conversation log to a server.
[0273] Step 3:
[0274] The server passes the received conversation log to the emotion engine, which analyzes the user's emotions. The emotion engine identifies emotions such as "interest," "anxiety," and "joy" from the text and voice data.
[0275] Step 4:
[0276] The server stores the analyzed emotional information along with the conversation log in a database, allowing the content of conversations and emotional states to be organized for each user.
[0277] Step 5:
[0278] The server periodically retrieves the latest data from external trusted sources, such as weather information APIs and restaurant reservation sites, to get the latest weather information and restaurant availability information.
[0279] Step 6:
[0280] The server preprocesses the external information it acquires and stores it in a structured database by category, for example, by organizing it into a "weather information table" or a "restaurant information table."
[0281] Step 7:
[0282] The terminal receives a new question from the user, for example, "Which restaurants are open in the afternoon?"
[0283] Step 8:
[0284] The terminal sends the received question to the server.
[0285] Step 9:
[0286] The server analyzes the question and searches a database for relevant conversation logs, emotion data, and the latest external information based on the analysis.
[0287] Step 10:
[0288] The server passes the search results to a large-scale language model (LLM) to generate the best response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[0289] Step 11:
[0290] Based on the responses generated by the LLM, the server tailors a customized response that takes into account the user's emotional state: if the user is determined to be "interested," it uses more positive language.
[0291] Step 12:
[0292] The server sends the generated customized response to the terminal.
[0293] Step 13:
[0294] The terminal displays the received response to the user, who then decides what to do next based on the response. For example, if the user sees a response saying "Italian, French, and Japanese restaurants are open in the afternoon," they can choose from those restaurants.
[0295] Step 14:
[0296] The server periodically updates the external information and adds the new information to the structured database, ensuring that the entire system always provides up-to-date information.
[0297] Example 2
[0298] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0299] Conventional conversation systems have difficulty generating responses that take the user's emotions into account, resulting in an unsatisfactory user experience. Furthermore, because the collection and updating of conversation logs and external information is not automated, it is difficult to quickly reflect the latest information. As a result, users are unable to receive appropriate information, which can lead to dissatisfaction.
[0300] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0301] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response using a large-scale language model based on the stored data, means for analyzing the user's emotions, means for adjusting the response based on the results of the emotion analysis, and means for transmitting the generated response to the user's terminal. This makes it possible to understand the user's emotions and provide customized responses in real time according to individual needs. Furthermore, the latest information can be quickly reflected without the need for re-learning, thereby improving user satisfaction.
[0302] A "conversation log" is data that records the content of a conversation, such as text or voice data, that a user inputs to the system.
[0303] "External information" is information obtained from trusted external data sources, such as weather information or restaurant availability information.
[0304] A "structured database" is a database in which conversation logs, external information, etc. are organized and stored in a specific structure.
[0305] A "large-scale language model" is an advanced algorithm or model that is trained based on large amounts of data in natural language processing and is capable of understanding and generating language.
[0306] An "emotion engine" is software or algorithm that analyzes emotions from a user's text or voice data and processes the results to identify them.
[0307] A "terminal" is a device that a user uses to interact with the system, including, for example, a smartphone or computer.
[0308] A "response" is information or a message that is generated based on the user's conversation log, external information, and emotional information, and is provided to the user.
[0309] "Collection methods" are mechanisms and processes for acquiring and storing various data (conversation logs, external information, etc.).
[0310] This system generates responses based on a large-scale language model (LLM), which collects user conversation logs and reliable external information. By combining this with an emotion engine that recognizes the user's emotions, it can provide customized responses in real time that correspond to the user's individual needs and emotional state.
[0311] 1. Collecting user conversation logs
[0312] A user interacts with the system using a smartphone or computer. For example, the user types a question such as, "Which restaurants are open this afternoon?" The device records this question in real time and sends it to the server at regular intervals.
[0313] 2. Gathering reliable external information
[0314] The server periodically retrieves the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server retrieves "Tokyo weather information" from a weather API and collects "seat availability information" from a restaurant reservation site.
[0315] 3. Operation of the Emotion Engine
[0316] Once the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[0317] 4. Data structuring and storage
[0318] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[0319] 5. Generating the Response
[0320] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[0321] 6. Sending and Displaying Responses
[0322] The generated response is sent to the terminal via the server and displayed to the user, who then decides on the next course of action based on the displayed information.
[0323] Specific examples
[0324] User Questions
[0325] User: "Which restaurants are open this afternoon?"
[0326] Sending transcripts
[0327] The terminal sends a query to the server.
[0328] Sentiment analysis
[0329] The server uses an emotion engine to identify the emotion "interested" and stores it in a database.
[0330] Searching for information
[0331] The server retrieves the user's conversation log, emotion data, and the latest restaurant seat availability information.
[0332] Generating a response
[0333] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[0334] Sending and Displaying Responses
[0335] The server sends the generated response to the terminal, which displays the response to the user.
[0336] The system can accurately analyze user sentiment and provide customized responses based on individual needs in real time, and the use of a structured database allows it to quickly reflect the latest information without the need for retraining.
[0337] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0338] Step 1:
[0339] The user inputs the conversation into the system. For example, the user uses a smartphone or computer to type, "Which restaurants are open this afternoon?" This input is recorded in real time on the device and saved as a conversation log. The input data is in text format, and the specific operation involves the user typing a message into the target device using a keyboard or voice recognition.
[0340] Step 2:
[0341] The device sends the user's conversation log to the server. The device then processes the collected conversation log to send it to the server at regular intervals. The input is the user's text data, and the output is the conversation log sent to the server. Specifically, an API is called to send the conversation log to the server using an HTTPS request.
[0342] Step 3:
[0343] The server collects external information. The server obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. The input is the API request parameters, and the output is the obtained weather information and seat availability information. Specifically, the server periodically sends requests to the external API, receives responses, and analyzes them.
[0344] Step 4:
[0345] The server analyzes the conversation log using an emotion engine. The emotion engine is used on the received conversation log to identify the user's emotions. The input is the text data of the conversation log, and the output is the emotion analysis results. Specifically, it runs an emotion analysis algorithm (for example, a natural language processing model) to identify the user's emotions (such as "joy," "anger," or "sadness").
[0346] Step 5:
[0347] The server stores the conversation log, external information, and emotional information in a structured database. The server analyzes the data and stores it in an organized form in the structured database. The input is the conversation log, external information, and emotional information, and the output is structured data stored in the database. Specifically, it uses SQL queries to organize the conversation log by user, and organizes the external information by category and inserts it into the database.
[0348] Step 6:
[0349] The server searches for the required information and provides it to a large-scale language model (LLM). The server searches for the required information from a database and provides it as input to the LLM. The input is a database query related to the user's question, and the output is the input data to the LLM. Specifically, the information searched for by the query is formatted into an appropriate format such as JSON and provided to the LLM.
[0350] Step 7:
[0351] The server uses the LLM to generate the optimal response. The LLM generates a response to the user based on the provided data. The input is the searched information and the user's emotional data, and the output is the generated response. Specifically, it sets a prompt in the LLM and generates a response in text format.
[0352] Step 8:
[0353] The server sends the generated response to the terminal. The server sends the generated response to the user's terminal, and the terminal displays this response to the user. The input is the generated response text, and the output is a message to the user displayed on the terminal. In concrete terms, the response is sent to the terminal using an HTTPS request, and the system on the terminal displays the message to the user.
[0354] (Application example 2)
[0355] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0356] Conventional user response systems generate responses primarily based on conversation logs and external information, but because they do not consider the user's emotional state, they often fail to respond appropriately. Furthermore, while there is a demand for improved stress management and work efficiency for operators, particularly in the industrial field, current systems have difficulty meeting these demands. Therefore, there is a growing need for a system that can recognize the user's emotional state in real time and generate appropriate responses based on that information.
[0357] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0358] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, emotion recognition means for analyzing the user's emotion based on the conversation log and the external information, means for generating a response using a large-scale language model based on the stored data, means for transmitting the generated response to the user's terminal, and means for customizing the response based on the analyzed emotion information, thereby making it possible to provide an optimal response in real time according to the user's emotional state.
[0359] A "user conversation log" is a record of the text and voice data entered by a user while interacting with the system.
[0360] "Reliable external information" refers to the latest data obtained from reliable external sources, such as weather information APIs and restaurant reservation sites.
[0361] A "structured database" is a database that organizes collected data such as conversation logs and external information, making it possible to search and use it efficiently.
[0362] A "large-scale language model (LLM)" is a machine learning model trained to perform natural language processing based on massive amounts of text data.
[0363] An "emotion recognition means" is an algorithm or system for identifying emotions such as joy, anger, sadness, etc. from text or voice data entered by a user.
[0364] The "means for generating a response" is a process that uses a large-scale language model to generate an optimal response based on the user's conversation log, external information, and emotional data.
[0365] The "means for customizing responses" is a function that adjusts and provides responses generated by a large-scale language model according to the user's emotional state.
[0366] "User terminal" refers to a device used by a user to interact with the system, including a smartphone, computer, robot, etc.
[0367] The system for implementing this invention collects a user's conversation log and incorporates reliable external information to generate customized responses based on the user's emotional state. Specifically, the system performs the following processes.
[0368] Program Overview
[0369] The system is configured using the following hardware and software:
[0370] Hardware:
[0371] Terminals or robots in the factory
[0372] Internet-connected server
[0373] software:
[0374] Large-scale Language Model (LLM): GPT-2 model using the Hugging Face Transformers library
[0375] Emotion Recognition Engine: TextBlob Library
[0376] External information acquisition API: requests library
[0377] What the program does
[0378] 1. Obtaining user conversation logs
[0379] Terminals and robots within the factory acquire text and voice data entered by users (operators) in real time.
[0380] 2. Sentiment analysis
[0381] The server uses TextBlob to extract emotions from user input and store the emotion information in a database. For example, it detects stress from an input such as "This machine seems to be malfunctioning since last night."
[0382] 3. Obtaining external information
[0383] The server uses an external API to obtain the latest status and operation status of factory machines, which is also stored in a database and updated regularly.
[0384] 4. Response Generation
[0385] The server uses a large-scale language model (GPT-2) to generate responses based on the user's conversation log, emotional data, and external information. Specifically, the following prompt sentences are used to provide data to the model:
[0386] Examples:
[0387] Worker: "This machine seems to be acting up since last night."
[0388] Emotion: -0.2 (worried)
[0389] Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheated"}
[0390] assistant:
[0391] Based on the prompt sentence above, the GPT-2 model generates a response.
[0392] 5. Sending the Response
[0393] The generated response is sent via the server to the device or robot being used by the user and displayed to the user. For example, a response such as "Motor overheating detected. We recommend checking the cooling fan" may be generated.
[0394] In this way, by providing real-time responses according to the user's emotional state, it is possible to manage the operator's stress and improve work efficiency. The introduction of an emotion engine also further improves the user experience. Since the latest information can be reflected in real time without the need for re-learning, it is possible to provide information quickly and efficiently.
[0395] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0396] Step 1:
[0397] Users use terminals or robots in the factory to input text or voice data into the system. For example, a user might say, "This machine seems to have been acting up since last night." This becomes the user's conversation log.
[0398] Input: User text or voice data
[0399] Output: Conversation log
[0400] Step 2:
[0401] The device sends the acquired conversation log in real time to the server, which receives the conversation log and stores it in a database.
[0402] Input: conversation log
[0403] Output: Conversation logs stored in a database
[0404] Step 3:
[0405] The server uses the TextBlob library to analyze emotions from the received conversation log. For example, stress can be detected from the text, "This machine seems to be malfunctioning since last night." The analyzed emotional information is stored in a database.
[0406] Input: conversation log
[0407] Data Processing: Sentiment Analysis using TextBlob
[0408] Output: Emotional information (e.g., stress)
[0409] Step 4:
[0410] The server uses an external API (e.g., machine status API) to obtain the latest machine operating status and warning information. The obtained information is stored in a database and updated regularly.
[0411] Input: Machine data request from external API
[0412] Data processing: Data acquisition from API
[0413] Output: External information stored in a database
[0414] Step 5:
[0415] The server provides prompts to the generative AI model (GPT-2) based on the user's conversation log, emotional data, and external information. For example, it generates the following prompt sentence:
[0416] "User Input: "This machine has been acting up since last night." Emotion: Stress Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheating" Assistant: "
[0417] Input: User conversation logs, emotion data, external information
[0418] Data processing: Prompt sentence generation
[0419] Output: Prompt sentence to the generative AI model
[0420] Step 6:
[0421] The generative AI model (GPT-2) generates the optimal response based on the provided prompt, for example, "Motor overheating detected. We recommend checking the cooling fan."
[0422] Input: prompt statement
[0423] Data Computing: Response Generation Using Generative AI Models
[0424] Output: The generated response
[0425] Step 7:
[0426] The server sends the generated response to the terminal, which displays the response to the user, who can review the response and decide what to do next.
[0427] Input: Generated response
[0428] Output: The response displayed on the user's terminal
[0429] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0430] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0431] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0432] [Second embodiment]
[0433] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0434] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0435] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0436] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0437] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0438] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0439] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0440] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0441] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0442] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0443] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0444] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0445] This invention relates to a system that collects user conversation logs and reliable external information, stores them in a structured database, generates responses using a large-scale language model (LLM) based on the stored data, and transmits the generated responses to the user's device. By using this system, it is possible to respond to individual needs and always provide the latest and most reliable information.
[0446] System configuration
[0447] 1. Collecting user conversation logs
[0448] Users interact with the system on a daily basis via their smartphones or computers. For example, a user might input a question such as, "Which restaurants are open this afternoon?" This interaction is digitally recorded on the user's device and periodically sent to the server.
[0449] 2. Gathering reliable external information
[0450] The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[0451] 3. Data structuring and storage
[0452] The server analyzes the user's conversation log and external information and stores them in a structured database. The conversation log is categorized by user, and reliable external information is organized by category. For example, a user's conversation log includes "question content," "date and time," and "frequency," while external information includes "weather information" and "restaurant availability."
[0453] 4. Generating the Response
[0454] The server retrieves the necessary information from a structured database and provides it to a large-scale language model. LLM generates the optimal response based on this information. For example, if a user asks about restaurant availability, LLM generates the response "Italian, French, and Japanese restaurants are open in the afternoon."
[0455] 5. Sending the Response
[0456] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[0457] Specific examples
[0458] 1. User Questions
[0459] User: "Which restaurants are open this afternoon?"
[0460] 2. Sending conversation logs
[0461] The terminal sends this question to the server.
[0462] 3. Search for related information
[0463] The server searches the database for the user's past conversation log and the latest restaurant vacancy information.
[0464] 4. Generating the Response
[0465] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[0466] 5. Sending and Displaying Responses
[0467] The server sends the generated response to the terminal, which displays the response to the user.
[0468] In this way, the system can provide customized responses in real time, and the use of a structured database allows it to quickly update without the need for retraining.
[0469] The processing flow will be explained below.
[0470] Step 1:
[0471] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[0472] Step 2:
[0473] The device records the conversation in real time, saves it as text data, and then periodically sends the conversation log to a server.
[0474] Step 3:
[0475] The server analyzes the received conversation logs and stores them in a database for each user. The analysis process includes extracting important information such as question content, time, and frequency.
[0476] Step 4:
[0477] The server periodically retrieves the latest data from external trusted sources (e.g., weather information APIs or restaurant reservation sites), which is also preprocessed and stored in a structured database.
[0478] Step 5:
[0479] When the device receives a new question from the user, it sends the content to the server, where it is processed in real time.
[0480] Step 6:
[0481] Based on the received question, the server searches the database for relevant conversation logs and external information, for example, past conversation history and the latest restaurant information.
[0482] Step 7:
[0483] The server then passes the search results to a large-scale language model (LLM) to generate the best possible response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[0484] Step 8:
[0485] The server sends the generated response to the terminal.
[0486] Step 9:
[0487] The terminal displays the received response to the user, who then decides on the next action to take.
[0488] Step 10:
[0489] The server periodically updates the external information and stores it in the database again, so that the latest information is always available within the system.
[0490] In this way, the system provides fast and accurate responses to user questions, and by utilizing a structured database and large-scale language models, it is possible to reflect the latest information in real time without the need for retraining.
[0491] Example 1
[0492] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0493] Conventional systems have had difficulty efficiently collecting and organizing user conversation logs and combining them with reliable external information to provide optimal responses in real time. Furthermore, because the external information is not updated frequently enough, there is no guarantee that the information provided is up-to-date and accurate. Furthermore, generating effective prompts has been a challenge when utilizing large-scale language models (LLMs).
[0494] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0495] In this invention, the server includes means for digitally collecting user conversation logs and transmitting them from the terminal to the server, means for periodically collecting reliable external information from APIs and information sites, and means for analyzing and classifying the conversation logs and external information and storing them in a structured database. This enables the generation of accurate responses that meet individual needs based on the latest, most reliable information in real time.
[0496] "User conversation log" refers to the content of communication such as questions and requests made by the user to the system.
[0497] "Digital format" refers to a format in which information, such as sound or text, can be recorded and processed electronically.
[0498] "Terminal" refers to a device, such as a smartphone or computer, that a user uses to access the system.
[0499] "Server" refers to a computer system for receiving and processing data sent from a user's terminal and generating a response.
[0500] "Reliable external information" refers to information obtained from reliable sources such as weather information APIs and restaurant reservation sites.
[0501] "API" stands for Application Programming Interface and refers to a protocol for exchanging information between different software programs.
[0502] An "information site" refers to a web page or online service that provides specific information.
[0503] "Analysis" refers to the process of breaking down data into more detailed pieces to make it easier to understand.
[0504] "Classification" refers to the process of grouping collected data based on specific criteria.
[0505] A "structured database" refers to a database in which data is systematically organized so that it can be efficiently stored and searched.
[0506] A "large-scale language model (LLM)" refers to a natural language processing model trained on large amounts of text data.
[0507] A "prompt sentence" refers to a guided text that is input into a large-scale language model to generate an appropriate response.
[0508] "Response" refers to the answer or information generated or provided by the system in response to a user's question.
[0509] The present invention is a system that collects user conversation logs and reliable external information, stores them in a structured database, and processes them. This system generates responses based on the collected data using a large-scale language model (LLM), and transmits the generated responses to the user's terminal for display. A specific embodiment of the system is described below.
[0510] Hardware and Software Configuration
[0511] 1. Hardware
[0512] User terminal: A device through which a user accesses the system, such as a smartphone or computer.
[0513] Server: A computer system, such as a cloud server, that receives data sent from a user's device, processes it, and generates a response.
[0514] 2. Software
[0515] Conversation log collection software: Software for collecting user conversations in digital form.
[0516] External data collection software: Software for collecting reliable external information from APIs and information sites.
[0517] Structured database software: Software for analyzing, classifying, and storing collected data.
[0518] Large-scale language model (LLM) software: Software that uses natural language processing techniques to generate responses based on collected data.
[0519] System configuration and operation
[0520] The system operates in the following steps.
[0521] 1. Collecting user conversation logs
[0522] User Asks a Question: A user uses a smartphone or computer to send a question or request to the system. Example: "Which restaurants are open this afternoon?"
[0523] Device processing: The device digitally records the user's conversations and periodically transmits them to a server, in particular speech recognition software that may convert speech to text.
[0524] 2. Gathering external information
[0525] Server information collection: The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. Example: Obtaining "Tokyo weather information" or "restaurant seat availability information."
[0526] 3. Data structuring and storage
[0527] Server data analysis: The server analyzes the received conversation logs and external information, and classifies and stores them in a structured database. Specifically, tokenization and tagging are performed.
[0528] Conversation log classification: User conversation logs are classified by user ID, and external information is organized by category.
[0529] 4. Generating the Response
[0530] Server data retrieval: The server retrieves the information required by the user from a structured database and provides it to the LLM.
[0531] LLM response generation: LLM generates the best response based on the data provided. Example: "Italian, French, and Japanese restaurants are open in the afternoon."
[0532] 5. Sending the Response
[0533] Transmission from server to terminal: The generated response is transmitted to the user's terminal via the server.
[0534] Display on the terminal: The terminal displays the received response to the user, who can then ask further questions based on the displayed information.
[0535] Specific examples
[0536] Below are some examples of specific prompt sentences.
[0537] 1. User Questions
[0538] User: "Which restaurants are open this afternoon?"
[0539] 2. Sending the device
[0540] The terminal sends this question in text format to the server.
[0541] 3. Searching for a server
[0542] The server retrieves the conversation log and the latest restaurant information from a database.
[0543] 4. Generating the Response
[0544] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[0545] 5. Sending and Displaying Responses
[0546] The server sends the generated response to the terminal, which displays the response to the user.
[0547] The system's features include the ability to respond quickly and accurately to user needs, collecting and providing the latest information in real time, and utilizing a structured database to generate responses with high efficiency without the need for re-learning.
[0548] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0549] Step 1:
[0550] A user enters a question or request into the system.
[0551] Input: User question (e.g., "Which restaurants are open this afternoon?")
[0552] Specific operation: The user accesses the system using a smartphone or computer and enters a question.
[0553] Step 2:
[0554] The device collects the user's conversation log in digital format and sends it to the server.
[0555] Input: User conversation log
[0556] Output: Digital conversation log
[0557] Specific operation: The device converts the conversation log into text format, adds metadata such as the user ID, question content, date and time, and sends it to the server periodically.
[0558] Step 3:
[0559] The server periodically collects reliable external information from APIs and information sites.
[0560] Input: URL of external information API endpoint or information site
[0561] Output: External information data (e.g., weather information, restaurant seat availability information)
[0562] Specific operation: The server periodically retrieves the latest data from external sources such as weather information APIs and restaurant reservation sites.
[0563] Step 4:
[0564] The server analyzes the conversation log and external information and stores it in a structured database.
[0565] Input: Digital conversation logs and external information data
[0566] Output: Parsed and classified data
[0567] How it works: The server uses natural language processing technology to tokenize the conversation logs and organize external information into categories. The analyzed and categorized data is then stored in a structured database.
[0568] Step 5:
[0569] The server retrieves the necessary information from a structured database and provides it to the LLM.
[0570] Input: User conversation logs and external information
[0571] Output: Data provided to LLM
[0572] Specific operation: The server searches the user's past conversation history and the latest external information from a structured database and provides them to the LLM.
[0573] Step 6:
[0574] LLM generates the best response based on the data provided.
[0575] Input: User conversation logs and external information
[0576] Output: Generated response text (e.g., "Italian, French, and Japanese restaurants are open in the afternoon.")
[0577] How it works: LLM uses natural language processing algorithms to generate optimal responses based on tokenized data.
[0578] Step 7:
[0579] The server sends the generated response to the user's terminal.
[0580] Input: Generated response text
[0581] Output: The response sent to the user's device
[0582] Specific operation: The server sends the generated response to the user's device, transferring data in real time using a communication protocol.
[0583] Step 8:
[0584] The terminal displays the received response to the user.
[0585] Input: Response text sent by the server
[0586] Output: The response message to be displayed
[0587] Specific operation: The terminal displays the received response on the user interface, and the user can ask more detailed questions based on this information.
[0588] Through these steps, the system is able to provide accurate and appropriate responses to user questions in real time.
[0589] (Application example 1)
[0590] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0591] Modern food delivery services are required to respond quickly and accurately to diverse user needs. However, conventional systems have difficulty providing personalized suggestions based on user preferences or real-time restaurant information. This has led to issues such as users being unable to quickly make choices that suit their preferences and current circumstances, resulting in reduced convenience.
[0592] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0593] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response based on the stored data using a large-scale language model, means for transmitting the generated response to the user's terminal, and means for providing customized recommendations based on the user's past ordering history and preferences. This allows users to receive real-time suggestions that match their preferences, significantly improving the convenience of food delivery.
[0594] A "conversation log" is text data generated through a user's interaction with the system, and is information that includes the user's questions and responses.
[0595] "External information" refers to the latest data the system collects from trusted external sources, such as restaurant menus, seat availability, and weather information.
[0596] A "structured database" is a database for organizing and storing conversation logs and external information, in which each piece of information is organized by classification and category.
[0597] A "large-scale language model" is an advanced machine learning model for natural language processing that generates optimal responses based on data provided by a structured database.
[0598] A "terminal" is a digital device used by a user, such as a smartphone or computer, on which the generated response is displayed.
[0599] "Customized Recommendations" are suggestions that are specifically tailored based on a user's past ordering history and preferences, and are information designed to address a user's individual needs.
[0600] A system embodying the present invention includes the following functions.
[0601] The server first collects the user's conversation log. The user interacts with the system via a smartphone or computer, and the conversation is recorded digitally. For example, the user may enter a question such as, "Which restaurants are available now?" This question is collected and sent to the server.
[0602] Next, the server collects reliable external information, such as restaurant menus, seat availability, and weather information, which it periodically obtains via online APIs. For example, the server obtains current seat availability information from a restaurant reservation site.
[0603] The server then stores the collected conversation logs and external information in a structured database. Conversation logs are organized by user, and external information is categorized. Specifically, a user's conversation log includes "question content," "date and time," and "frequency," while reliable external information includes "restaurant seat availability information" and "menu information."
[0604] The server then uses the stored data to generate a response using a large-scale language model (LLM). The server searches a structured database for the necessary information and sends a prompt to the LLM based on this information. For example, if a user asks about restaurant availability, the LLM can generate a response such as, "As of 10:00 AM, an Italian restaurant and a Chinese restaurant are available."
[0605] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[0606] Additionally, the server provides customized recommendations based on the user's past ordering history and preferences, allowing the user to quickly find options that suit their preferences.
[0607] For example, the following prompts might be used for a generative AI model:
[0608] text
[0609] User asks: "What Italian restaurant would you recommend?"
[0610] Past conversation: [{ "timestamp": "2023-10-08T12:00:00", "conversation": "I like pizza"}, {...}]
[0611] External Data: [{ "name": "Italian Restaurant A", "rating": 4.5, "availability": "open"}, {...}]
[0612] Based on this prompt, the LLM processes past conversation data and external information to generate a response such as, "Based on your preferences, we recommend Italian restaurant A." In this way, the present invention can provide optimal information tailored to the user's individual needs.
[0613] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0614] Step 1:
[0615] The user enters a question and the terminal receives the question.
[0616] Specifically, when a user uses a smartphone or computer to enter a question such as "Which restaurants are open now?", the text data is sent to the device and saved.
[0617] Input: User question
[0618] Output: Question data in text format
[0619] Step 2:
[0620] The terminal transmits the saved question data to the server.
[0621] This data is sent to the server along with the user ID, and the server receives it and records it as a conversation log.
[0622] Input: Text question data, user ID
[0623] Output: Conversation log sent to the server
[0624] Step 3:
[0625] The server collects trusted external information.
[0626] Specifically, the server accesses a pre-configured API endpoint (e.g., the API of a restaurant reservation system) and obtains the latest data such as restaurant seat availability and menu information.
[0627] Input: API endpoint
[0628] Output: External information data (e.g., restaurant seat availability, menu information)
[0629] Step 4:
[0630] The server stores the collected conversation logs and external information in a structured database.
[0631] Conversation logs are categorized by user, and external information is organized by category. For example, restaurant vacancy information is organized into categories such as "restaurant" and "vacancy information."
[0632] Input: Conversation log, external information
[0633] Output: Structured database entries
[0634] Step 5:
[0635] The server searches for the necessary information from a structured database and uses it to generate and send prompts to a large-scale language model (LLM).
[0636] The prompt includes the user's question, past conversation logs, and external information. The prompt is created and sent to the LLM to generate the optimal response.
[0637] Input: Structured database information, user questions
[0638] Output: Generated prompt, response from LLM
[0639] Step 6:
[0640] The server sends the generated response to the user's terminal.
[0641] The terminal displays the received response to the user, thereby providing the user with an answer to their question.
[0642] Input: The response generated by the LLM
[0643] Output: The response displayed on the user's terminal
[0644] Step 7:
[0645] The server generates and provides customized recommendations to the user based on the user's past ordering history and preferences.
[0646] By analyzing past conversation logs and order data, the system suggests restaurants and menus that suit the user's preferences.
[0647] Input: User's past order history, preference data
[0648] Output: Customized recommendations
[0649] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0650] This invention relates to a system that collects user conversation logs and reliable external information, generates responses based on these using a large-scale language model (LLM), and combines this with an emotion engine that recognizes the user's emotions. This system is capable of providing customized responses in real time according to the user's individual needs and emotional state. Furthermore, by utilizing a structured database, it is possible to quickly reflect the latest information without the need for re-learning.
[0651] System configuration
[0652] 1. Collecting user conversation logs
[0653] Users interact with the system using a smartphone or computer. For example, they might input a question like, "Which restaurants are open this afternoon?" The content of this interaction is recorded in real time on the device and periodically sent to the server.
[0654] 2. Gathering reliable external information
[0655] The server periodically obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[0656] 3. Operation of the Emotion Engine
[0657] When the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[0658] 4. Data structuring and storage
[0659] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[0660] 5. Generating the Response
[0661] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[0662] 6. Sending the Response
[0663] The generated response is sent via the server to the terminal and displayed to the user, who then decides what to do next.
[0664] Specific examples
[0665] 1. User Questions
[0666] User: "Which restaurants are open this afternoon?"
[0667] 2. Sending conversation logs
[0668] The terminal sends this question to the server.
[0669] 3. Sentiment Analysis
[0670] The server uses an emotion engine to identify the emotion "interested" from the user's question and stores it in a database.
[0671] 4. Searching for related information
[0672] The server searches a database for user conversation logs, emotional data, and the latest restaurant seat availability information.
[0673] 5. Generating the Response
[0674] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[0675] 6. Sending and Displaying Responses
[0676] The server sends the generated response to the terminal, which displays the response to the user.
[0677] In this way, the system can provide customized responses in real time according to the user's emotional state. The introduction of an emotion engine further improves the user experience. It can provide information quickly and efficiently because it can reflect the latest information in real time without the need for retraining.
[0678] The processing flow will be explained below.
[0679] Step 1:
[0680] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[0681] Step 2:
[0682] The device records the conversation in real time, saves it as text data, and then sends the conversation log to a server.
[0683] Step 3:
[0684] The server passes the received conversation log to the emotion engine, which analyzes the user's emotions. The emotion engine identifies emotions such as "interest," "anxiety," and "joy" from the text and voice data.
[0685] Step 4:
[0686] The server stores the analyzed emotional information along with the conversation log in a database, allowing the content of conversations and emotional states to be organized for each user.
[0687] Step 5:
[0688] The server periodically retrieves the latest data from external trusted sources, such as weather information APIs and restaurant reservation sites, to get the latest weather information and restaurant availability information.
[0689] Step 6:
[0690] The server preprocesses the external information it acquires and stores it in a structured database by category, for example, by organizing it into a "weather information table" or a "restaurant information table."
[0691] Step 7:
[0692] The terminal receives a new question from the user, for example, "Which restaurants are open in the afternoon?"
[0693] Step 8:
[0694] The terminal sends the received question to the server.
[0695] Step 9:
[0696] The server analyzes the question and searches a database for relevant conversation logs, emotion data, and the latest external information based on the analysis.
[0697] Step 10:
[0698] The server passes the search results to a large-scale language model (LLM) to generate the best response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[0699] Step 11:
[0700] Based on the responses generated by the LLM, the server tailors a customized response that takes into account the user's emotional state: if the user is determined to be "interested," it uses more positive language.
[0701] Step 12:
[0702] The server sends the generated customized response to the terminal.
[0703] Step 13:
[0704] The terminal displays the received response to the user, who then decides what to do next based on the response. For example, if the user sees a response saying "Italian, French, and Japanese restaurants are open in the afternoon," they can choose from those restaurants.
[0705] Step 14:
[0706] The server periodically updates the external information and adds the new information to the structured database, ensuring that the entire system always provides up-to-date information.
[0707] Example 2
[0708] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0709] Conventional conversation systems have difficulty generating responses that take the user's emotions into account, resulting in an unsatisfactory user experience. Furthermore, because the collection and updating of conversation logs and external information is not automated, it is difficult to quickly reflect the latest information. As a result, users are unable to receive appropriate information, which can lead to dissatisfaction.
[0710] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0711] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response using a large-scale language model based on the stored data, means for analyzing the user's emotions, means for adjusting the response based on the results of the emotion analysis, and means for transmitting the generated response to the user's terminal. This makes it possible to understand the user's emotions and provide customized responses in real time according to individual needs. Furthermore, the latest information can be quickly reflected without the need for re-learning, thereby improving user satisfaction.
[0712] A "conversation log" is data that records the content of a conversation, such as text or voice data, that a user inputs to the system.
[0713] "External information" is information obtained from trusted external data sources, such as weather information or restaurant availability information.
[0714] A "structured database" is a database in which conversation logs, external information, etc. are organized and stored in a specific structure.
[0715] A "large-scale language model" is an advanced algorithm or model that is trained based on large amounts of data in natural language processing and is capable of understanding and generating language.
[0716] An "emotion engine" is software or algorithm that analyzes emotions from a user's text or voice data and processes the results to identify them.
[0717] A "terminal" is a device that a user uses to interact with the system, including, for example, a smartphone or computer.
[0718] A "response" is information or a message that is generated based on the user's conversation log, external information, and emotional information, and is provided to the user.
[0719] "Collection methods" are mechanisms and processes for acquiring and storing various data (conversation logs, external information, etc.).
[0720] This system generates responses based on a large-scale language model (LLM), which collects user conversation logs and reliable external information. By combining this with an emotion engine that recognizes the user's emotions, it can provide customized responses in real time that correspond to the user's individual needs and emotional state.
[0721] 1. Collecting user conversation logs
[0722] A user interacts with the system using a smartphone or computer. For example, the user types a question such as, "Which restaurants are open this afternoon?" The device records this question in real time and sends it to the server at regular intervals.
[0723] 2. Gathering reliable external information
[0724] The server periodically retrieves the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server retrieves "Tokyo weather information" from a weather API and collects "seat availability information" from a restaurant reservation site.
[0725] 3. Operation of the Emotion Engine
[0726] Once the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[0727] 4. Data structuring and storage
[0728] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[0729] 5. Generating the Response
[0730] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[0731] 6. Sending and Displaying Responses
[0732] The generated response is sent to the terminal via the server and displayed to the user, who then decides on the next course of action based on the displayed information.
[0733] Specific examples
[0734] User Questions
[0735] User: "Which restaurants are open this afternoon?"
[0736] Sending transcripts
[0737] The terminal sends a query to the server.
[0738] Sentiment analysis
[0739] The server uses an emotion engine to identify the emotion "interested" and stores it in a database.
[0740] Searching for information
[0741] The server retrieves the user's conversation log, emotion data, and the latest restaurant seat availability information.
[0742] Generating a response
[0743] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[0744] Sending and Displaying Responses
[0745] The server sends the generated response to the terminal, which displays the response to the user.
[0746] The system can accurately analyze user sentiment and provide customized responses based on individual needs in real time, and the use of a structured database allows it to quickly reflect the latest information without the need for retraining.
[0747] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0748] Step 1:
[0749] The user inputs the conversation into the system. For example, the user uses a smartphone or computer to type, "Which restaurants are open this afternoon?" This input is recorded in real time on the device and saved as a conversation log. The input data is in text format, and the specific operation involves the user typing a message into the target device using a keyboard or voice recognition.
[0750] Step 2:
[0751] The device sends the user's conversation log to the server. The device then processes the collected conversation log to send it to the server at regular intervals. The input is the user's text data, and the output is the conversation log sent to the server. Specifically, an API is called to send the conversation log to the server using an HTTPS request.
[0752] Step 3:
[0753] The server collects external information. The server obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. The input is the API request parameters, and the output is the obtained weather information and seat availability information. Specifically, the server periodically sends requests to the external API, receives responses, and analyzes them.
[0754] Step 4:
[0755] The server analyzes the conversation log using an emotion engine. The emotion engine is used on the received conversation log to identify the user's emotions. The input is the text data of the conversation log, and the output is the emotion analysis results. Specifically, it runs an emotion analysis algorithm (for example, a natural language processing model) to identify the user's emotions (such as "joy," "anger," or "sadness").
[0756] Step 5:
[0757] The server stores the conversation log, external information, and emotional information in a structured database. The server analyzes the data and stores it in an organized form in the structured database. The input is the conversation log, external information, and emotional information, and the output is structured data stored in the database. Specifically, it uses SQL queries to organize the conversation log by user, and organizes the external information by category and inserts it into the database.
[0758] Step 6:
[0759] The server searches for the required information and provides it to a large-scale language model (LLM). The server searches for the required information from a database and provides it as input to the LLM. The input is a database query related to the user's question, and the output is the input data to the LLM. Specifically, the information searched for by the query is formatted into an appropriate format such as JSON and provided to the LLM.
[0760] Step 7:
[0761] The server uses the LLM to generate the optimal response. The LLM generates a response to the user based on the provided data. The input is the searched information and the user's emotional data, and the output is the generated response. Specifically, it sets a prompt in the LLM and generates a response in text format.
[0762] Step 8:
[0763] The server sends the generated response to the terminal. The server sends the generated response to the user's terminal, and the terminal displays this response to the user. The input is the generated response text, and the output is a message to the user displayed on the terminal. In concrete terms, the response is sent to the terminal using an HTTPS request, and the system on the terminal displays the message to the user.
[0764] (Application example 2)
[0765] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0766] Conventional user response systems generate responses primarily based on conversation logs and external information, but because they do not consider the user's emotional state, they often fail to respond appropriately. Furthermore, while there is a demand for improved stress management and work efficiency for operators, particularly in the industrial field, current systems have difficulty meeting these demands. Therefore, there is a growing need for a system that can recognize the user's emotional state in real time and generate appropriate responses based on that information.
[0767] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0768] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, emotion recognition means for analyzing the user's emotion based on the conversation log and the external information, means for generating a response using a large-scale language model based on the stored data, means for transmitting the generated response to the user's terminal, and means for customizing the response based on the analyzed emotion information, thereby making it possible to provide an optimal response in real time according to the user's emotional state.
[0769] A "user conversation log" is a record of the text and voice data entered by a user while interacting with the system.
[0770] "Reliable external information" refers to the latest data obtained from reliable external sources, such as weather information APIs and restaurant reservation sites.
[0771] A "structured database" is a database that organizes collected data such as conversation logs and external information, making it possible to search and use it efficiently.
[0772] A "large-scale language model (LLM)" is a machine learning model trained to perform natural language processing based on massive amounts of text data.
[0773] An "emotion recognition means" is an algorithm or system for identifying emotions such as joy, anger, sadness, etc. from text or voice data entered by a user.
[0774] The "means for generating a response" is a process that uses a large-scale language model to generate an optimal response based on the user's conversation log, external information, and emotional data.
[0775] The "means for customizing responses" is a function that adjusts and provides responses generated by a large-scale language model according to the user's emotional state.
[0776] "User terminal" refers to a device used by a user to interact with the system, including a smartphone, computer, robot, etc.
[0777] The system for implementing this invention collects a user's conversation log and incorporates reliable external information to generate customized responses based on the user's emotional state. Specifically, the system performs the following processes.
[0778] Program Overview
[0779] The system is configured using the following hardware and software:
[0780] Hardware:
[0781] Terminals or robots in the factory
[0782] Internet-connected server
[0783] software:
[0784] Large-scale Language Model (LLM): GPT-2 model using the Hugging Face Transformers library
[0785] Emotion Recognition Engine: TextBlob Library
[0786] External information acquisition API: requests library
[0787] What the program does
[0788] 1. Obtaining user conversation logs
[0789] Terminals and robots within the factory acquire text and voice data entered by users (operators) in real time.
[0790] 2. Sentiment analysis
[0791] The server uses TextBlob to extract emotions from user input and store the emotion information in a database. For example, it detects stress from an input such as "This machine seems to be malfunctioning since last night."
[0792] 3. Obtaining external information
[0793] The server uses an external API to obtain the latest status and operation status of factory machines, which is also stored in a database and updated regularly.
[0794] 4. Response Generation
[0795] The server uses a large-scale language model (GPT-2) to generate responses based on the user's conversation log, emotional data, and external information. Specifically, the following prompt sentences are used to provide data to the model:
[0796] Examples:
[0797] Worker: "This machine seems to be acting up since last night."
[0798] Emotion: -0.2 (worried)
[0799] Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheated"}
[0800] assistant:
[0801] Based on the prompt sentence above, the GPT-2 model generates a response.
[0802] 5. Sending the Response
[0803] The generated response is sent via the server to the device or robot being used by the user and displayed to the user. For example, a response such as "Motor overheating detected. We recommend checking the cooling fan" may be generated.
[0804] In this way, by providing real-time responses according to the user's emotional state, it is possible to manage the operator's stress and improve work efficiency. The introduction of an emotion engine also further improves the user experience. Since the latest information can be reflected in real time without the need for re-learning, it is possible to provide information quickly and efficiently.
[0805] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0806] Step 1:
[0807] Users use terminals or robots in the factory to input text or voice data into the system. For example, a user might say, "This machine seems to have been acting up since last night." This becomes the user's conversation log.
[0808] Input: User text or voice data
[0809] Output: Conversation log
[0810] Step 2:
[0811] The device sends the acquired conversation log in real time to the server, which receives the conversation log and stores it in a database.
[0812] Input: conversation log
[0813] Output: Conversation logs stored in a database
[0814] Step 3:
[0815] The server uses the TextBlob library to analyze emotions from the received conversation log. For example, stress can be detected from the text, "This machine seems to be malfunctioning since last night." The analyzed emotional information is stored in a database.
[0816] Input: conversation log
[0817] Data Processing: Sentiment Analysis using TextBlob
[0818] Output: Emotional information (e.g., stress)
[0819] Step 4:
[0820] The server uses an external API (e.g., machine status API) to obtain the latest machine operating status and warning information. The obtained information is stored in a database and updated regularly.
[0821] Input: Machine data request from external API
[0822] Data processing: Data acquisition from API
[0823] Output: External information stored in a database
[0824] Step 5:
[0825] The server provides prompts to the generative AI model (GPT-2) based on the user's conversation log, emotional data, and external information. For example, it generates the following prompt sentence:
[0826] "User Input: "This machine has been acting up since last night." Emotion: Stress Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheating" Assistant: "
[0827] Input: User conversation logs, emotion data, external information
[0828] Data processing: Prompt sentence generation
[0829] Output: Prompt sentence to the generative AI model
[0830] Step 6:
[0831] The generative AI model (GPT-2) generates the optimal response based on the provided prompt, for example, "Motor overheating detected. We recommend checking the cooling fan."
[0832] Input: prompt statement
[0833] Data Computing: Response Generation Using Generative AI Models
[0834] Output: The generated response
[0835] Step 7:
[0836] The server sends the generated response to the terminal, which displays the response to the user, who can review the response and decide what to do next.
[0837] Input: Generated response
[0838] Output: The response displayed on the user's terminal
[0839] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0840] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0841] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0842] [Third embodiment]
[0843] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0844] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0845] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0846] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0847] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0848] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0849] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0850] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0851] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0852] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0853] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0854] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0855] This invention relates to a system that collects user conversation logs and reliable external information, stores them in a structured database, generates responses using a large-scale language model (LLM) based on the stored data, and transmits the generated responses to the user's device. By using this system, it is possible to respond to individual needs and always provide the latest and most reliable information.
[0856] System configuration
[0857] 1. Collecting user conversation logs
[0858] Users interact with the system on a daily basis via their smartphones or computers. For example, a user might input a question such as, "Which restaurants are open this afternoon?" This interaction is digitally recorded on the user's device and periodically sent to the server.
[0859] 2. Gathering reliable external information
[0860] The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[0861] 3. Data structuring and storage
[0862] The server analyzes the user's conversation log and external information and stores them in a structured database. The conversation log is categorized by user, and reliable external information is organized by category. For example, a user's conversation log includes "question content," "date and time," and "frequency," while external information includes "weather information" and "restaurant availability."
[0863] 4. Generating the Response
[0864] The server retrieves the necessary information from a structured database and provides it to a large-scale language model. LLM generates the optimal response based on this information. For example, if a user asks about restaurant availability, LLM generates the response "Italian, French, and Japanese restaurants are open in the afternoon."
[0865] 5. Sending the Response
[0866] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[0867] Specific examples
[0868] 1. User Questions
[0869] User: "Which restaurants are open this afternoon?"
[0870] 2. Sending conversation logs
[0871] The terminal sends this question to the server.
[0872] 3. Search for related information
[0873] The server searches the database for the user's past conversation log and the latest restaurant vacancy information.
[0874] 4. Generating the Response
[0875] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[0876] 5. Sending and Displaying Responses
[0877] The server sends the generated response to the terminal, which displays the response to the user.
[0878] In this way, the system can provide customized responses in real time, and the use of a structured database allows it to quickly update without the need for retraining.
[0879] The processing flow will be explained below.
[0880] Step 1:
[0881] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[0882] Step 2:
[0883] The device records the conversation in real time, saves it as text data, and then periodically sends the conversation log to a server.
[0884] Step 3:
[0885] The server analyzes the received conversation logs and stores them in a database for each user. The analysis process includes extracting important information such as question content, time, and frequency.
[0886] Step 4:
[0887] The server periodically retrieves the latest data from external trusted sources (e.g., weather information APIs or restaurant reservation sites), which is also preprocessed and stored in a structured database.
[0888] Step 5:
[0889] When the device receives a new question from the user, it sends the content to the server, where it is processed in real time.
[0890] Step 6:
[0891] Based on the received question, the server searches the database for relevant conversation logs and external information, for example, past conversation history and the latest restaurant information.
[0892] Step 7:
[0893] The server then passes the search results to a large-scale language model (LLM) to generate the best possible response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[0894] Step 8:
[0895] The server sends the generated response to the terminal.
[0896] Step 9:
[0897] The terminal displays the received response to the user, who then decides on the next action to take.
[0898] Step 10:
[0899] The server periodically updates the external information and stores it in the database again, so that the latest information is always available within the system.
[0900] In this way, the system provides fast and accurate responses to user questions, and by utilizing a structured database and large-scale language models, it is possible to reflect the latest information in real time without the need for retraining.
[0901] Example 1
[0902] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0903] Conventional systems have had difficulty efficiently collecting and organizing user conversation logs and combining them with reliable external information to provide optimal responses in real time. Furthermore, because the external information is not updated frequently enough, there is no guarantee that the information provided is up-to-date and accurate. Furthermore, generating effective prompts has been a challenge when utilizing large-scale language models (LLMs).
[0904] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0905] In this invention, the server includes means for digitally collecting user conversation logs and transmitting them from the terminal to the server, means for periodically collecting reliable external information from APIs and information sites, and means for analyzing and classifying the conversation logs and external information and storing them in a structured database. This enables the generation of accurate responses that meet individual needs based on the latest, most reliable information in real time.
[0906] "User conversation log" refers to the content of communication such as questions and requests made by the user to the system.
[0907] "Digital format" refers to a format in which information, such as sound or text, can be recorded and processed electronically.
[0908] "Terminal" refers to a device, such as a smartphone or computer, that a user uses to access the system.
[0909] "Server" refers to a computer system for receiving and processing data sent from a user's terminal and generating a response.
[0910] "Reliable external information" refers to information obtained from reliable sources such as weather information APIs and restaurant reservation sites.
[0911] "API" stands for Application Programming Interface and refers to a protocol for exchanging information between different software programs.
[0912] An "information site" refers to a web page or online service that provides specific information.
[0913] "Analysis" refers to the process of breaking down data into more detailed pieces to make it easier to understand.
[0914] "Classification" refers to the process of grouping collected data based on specific criteria.
[0915] A "structured database" refers to a database in which data is systematically organized so that it can be efficiently stored and searched.
[0916] A "large-scale language model (LLM)" refers to a natural language processing model trained on large amounts of text data.
[0917] A "prompt sentence" refers to a guided text that is input into a large-scale language model to generate an appropriate response.
[0918] "Response" refers to the answer or information generated or provided by the system in response to a user's question.
[0919] The present invention is a system that collects user conversation logs and reliable external information, stores them in a structured database, and processes them. This system generates responses based on the collected data using a large-scale language model (LLM), and transmits the generated responses to the user's terminal for display. A specific embodiment of the system is described below.
[0920] Hardware and Software Configuration
[0921] 1. Hardware
[0922] User terminal: A device through which a user accesses the system, such as a smartphone or computer.
[0923] Server: A computer system, such as a cloud server, that receives data sent from a user's device, processes it, and generates a response.
[0924] 2. Software
[0925] Conversation log collection software: Software for collecting user conversations in digital form.
[0926] External data collection software: Software for collecting reliable external information from APIs and information sites.
[0927] Structured database software: Software for analyzing, classifying, and storing collected data.
[0928] Large-scale language model (LLM) software: Software that uses natural language processing techniques to generate responses based on collected data.
[0929] System configuration and operation
[0930] The system operates in the following steps.
[0931] 1. Collecting user conversation logs
[0932] User Asks a Question: A user uses a smartphone or computer to send a question or request to the system. Example: "Which restaurants are open this afternoon?"
[0933] Device processing: The device digitally records the user's conversations and periodically transmits them to a server, in particular speech recognition software that may convert speech to text.
[0934] 2. Gathering external information
[0935] Server information collection: The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. Example: Obtaining "Tokyo weather information" or "restaurant seat availability information."
[0936] 3. Data structuring and storage
[0937] Server data analysis: The server analyzes the received conversation logs and external information, and classifies and stores them in a structured database. Specifically, tokenization and tagging are performed.
[0938] Conversation log classification: User conversation logs are classified by user ID, and external information is organized by category.
[0939] 4. Generating the Response
[0940] Server data retrieval: The server retrieves the information required by the user from a structured database and provides it to the LLM.
[0941] LLM response generation: LLM generates the best response based on the data provided. Example: "Italian, French, and Japanese restaurants are open in the afternoon."
[0942] 5. Sending the Response
[0943] Transmission from server to terminal: The generated response is transmitted to the user's terminal via the server.
[0944] Display on the terminal: The terminal displays the received response to the user, who can then ask further questions based on the displayed information.
[0945] Specific examples
[0946] Below are some examples of specific prompt sentences.
[0947] 1. User Questions
[0948] User: "Which restaurants are open this afternoon?"
[0949] 2. Sending the device
[0950] The terminal sends this question in text format to the server.
[0951] 3. Searching for a server
[0952] The server retrieves the conversation log and the latest restaurant information from a database.
[0953] 4. Generating the Response
[0954] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[0955] 5. Sending and Displaying Responses
[0956] The server sends the generated response to the terminal, which displays the response to the user.
[0957] The system's features include the ability to respond quickly and accurately to user needs, collecting and providing the latest information in real time, and utilizing a structured database to generate responses with high efficiency without the need for re-learning.
[0958] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0959] Step 1:
[0960] A user enters a question or request into the system.
[0961] Input: User question (e.g., "Which restaurants are open this afternoon?")
[0962] Specific operation: The user accesses the system using a smartphone or computer and enters a question.
[0963] Step 2:
[0964] The device collects the user's conversation log in digital format and sends it to the server.
[0965] Input: User conversation log
[0966] Output: Digital conversation log
[0967] Specific operation: The device converts the conversation log into text format, adds metadata such as the user ID, question content, date and time, and sends it to the server periodically.
[0968] Step 3:
[0969] The server periodically collects reliable external information from APIs and information sites.
[0970] Input: URL of external information API endpoint or information site
[0971] Output: External information data (e.g., weather information, restaurant seat availability information)
[0972] Specific operation: The server periodically retrieves the latest data from external sources such as weather information APIs and restaurant reservation sites.
[0973] Step 4:
[0974] The server analyzes the conversation log and external information and stores it in a structured database.
[0975] Input: Digital conversation logs and external information data
[0976] Output: Parsed and classified data
[0977] How it works: The server uses natural language processing technology to tokenize the conversation logs and organize external information into categories. The analyzed and categorized data is then stored in a structured database.
[0978] Step 5:
[0979] The server retrieves the necessary information from a structured database and provides it to the LLM.
[0980] Input: User conversation logs and external information
[0981] Output: Data provided to LLM
[0982] Specific operation: The server searches the user's past conversation history and the latest external information from a structured database and provides them to the LLM.
[0983] Step 6:
[0984] LLM generates the best response based on the data provided.
[0985] Input: User conversation logs and external information
[0986] Output: Generated response text (e.g., "Italian, French, and Japanese restaurants are open in the afternoon.")
[0987] How it works: LLM uses natural language processing algorithms to generate optimal responses based on tokenized data.
[0988] Step 7:
[0989] The server sends the generated response to the user's terminal.
[0990] Input: Generated response text
[0991] Output: The response sent to the user's device
[0992] Specific operation: The server sends the generated response to the user's device, transferring data in real time using a communication protocol.
[0993] Step 8:
[0994] The terminal displays the received response to the user.
[0995] Input: Response text sent by the server
[0996] Output: The response message to be displayed
[0997] Specific operation: The terminal displays the received response on the user interface, and the user can ask more detailed questions based on this information.
[0998] Through these steps, the system is able to provide accurate and appropriate responses to user questions in real time.
[0999] (Application example 1)
[1000] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1001] Modern food delivery services are required to respond quickly and accurately to diverse user needs. However, conventional systems have difficulty providing personalized suggestions based on user preferences or real-time restaurant information. This has led to issues such as users being unable to quickly make choices that suit their preferences and current circumstances, resulting in reduced convenience.
[1002] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1003] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response based on the stored data using a large-scale language model, means for transmitting the generated response to the user's terminal, and means for providing customized recommendations based on the user's past ordering history and preferences. This allows users to receive real-time suggestions that match their preferences, significantly improving the convenience of food delivery.
[1004] A "conversation log" is text data generated through a user's interaction with the system, and is information that includes the user's questions and responses.
[1005] "External information" refers to the latest data the system collects from trusted external sources, such as restaurant menus, seat availability, and weather information.
[1006] A "structured database" is a database for organizing and storing conversation logs and external information, in which each piece of information is organized by classification and category.
[1007] A "large-scale language model" is an advanced machine learning model for natural language processing that generates optimal responses based on data provided by a structured database.
[1008] A "terminal" is a digital device used by a user, such as a smartphone or computer, on which the generated response is displayed.
[1009] "Customized Recommendations" are suggestions that are specifically tailored based on a user's past ordering history and preferences, and are information designed to address a user's individual needs.
[1010] A system embodying the present invention includes the following functions.
[1011] The server first collects the user's conversation log. The user interacts with the system via a smartphone or computer, and the conversation is recorded digitally. For example, the user may enter a question such as, "Which restaurants are available now?" This question is collected and sent to the server.
[1012] Next, the server collects reliable external information, such as restaurant menus, seat availability, and weather information, which it periodically obtains via online APIs. For example, the server obtains current seat availability information from a restaurant reservation site.
[1013] The server then stores the collected conversation logs and external information in a structured database. Conversation logs are organized by user, and external information is categorized. Specifically, a user's conversation log includes "question content," "date and time," and "frequency," while reliable external information includes "restaurant seat availability information" and "menu information."
[1014] The server then uses the stored data to generate a response using a large-scale language model (LLM). The server searches a structured database for the necessary information and sends a prompt to the LLM based on this information. For example, if a user asks about restaurant availability, the LLM can generate a response such as, "As of 10:00 AM, an Italian restaurant and a Chinese restaurant are available."
[1015] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[1016] Additionally, the server provides customized recommendations based on the user's past ordering history and preferences, allowing the user to quickly find options that suit their preferences.
[1017] For example, the following prompts might be used for a generative AI model:
[1018] text
[1019] User asks: "What Italian restaurant would you recommend?"
[1020] Past conversation: [{ "timestamp": "2023-10-08T12:00:00", "conversation": "I like pizza"}, {...}]
[1021] External Data: [{ "name": "Italian Restaurant A", "rating": 4.5, "availability": "open"}, {...}]
[1022] Based on this prompt, the LLM processes past conversation data and external information to generate a response such as, "Based on your preferences, we recommend Italian restaurant A." In this way, the present invention can provide optimal information tailored to the user's individual needs.
[1023] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1024] Step 1:
[1025] The user enters a question and the terminal receives the question.
[1026] Specifically, when a user uses a smartphone or computer to enter a question such as "Which restaurants are open now?", the text data is sent to the device and saved.
[1027] Input: User question
[1028] Output: Question data in text format
[1029] Step 2:
[1030] The terminal transmits the saved question data to the server.
[1031] This data is sent to the server along with the user ID, and the server receives it and records it as a conversation log.
[1032] Input: Text question data, user ID
[1033] Output: Conversation log sent to the server
[1034] Step 3:
[1035] The server collects trusted external information.
[1036] Specifically, the server accesses a pre-configured API endpoint (e.g., the API of a restaurant reservation system) and obtains the latest data such as restaurant seat availability and menu information.
[1037] Input: API endpoint
[1038] Output: External information data (e.g., restaurant seat availability, menu information)
[1039] Step 4:
[1040] The server stores the collected conversation logs and external information in a structured database.
[1041] Conversation logs are categorized by user, and external information is organized by category. For example, restaurant vacancy information is organized into categories such as "restaurant" and "vacancy information."
[1042] Input: Conversation log, external information
[1043] Output: Structured database entries
[1044] Step 5:
[1045] The server searches for the necessary information from a structured database and uses it to generate and send prompts to a large-scale language model (LLM).
[1046] The prompt includes the user's question, past conversation logs, and external information. The prompt is created and sent to the LLM to generate the optimal response.
[1047] Input: Structured database information, user questions
[1048] Output: Generated prompt, response from LLM
[1049] Step 6:
[1050] The server sends the generated response to the user's terminal.
[1051] The terminal displays the received response to the user, thereby providing the user with an answer to their question.
[1052] Input: The response generated by the LLM
[1053] Output: The response displayed on the user's terminal
[1054] Step 7:
[1055] The server generates and provides customized recommendations to the user based on the user's past ordering history and preferences.
[1056] By analyzing past conversation logs and order data, the system suggests restaurants and menus that suit the user's preferences.
[1057] Input: User's past order history, preference data
[1058] Output: Customized recommendations
[1059] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1060] This invention relates to a system that collects user conversation logs and reliable external information, generates responses based on these using a large-scale language model (LLM), and combines this with an emotion engine that recognizes the user's emotions. This system is capable of providing customized responses in real time according to the user's individual needs and emotional state. Furthermore, by utilizing a structured database, it is possible to quickly reflect the latest information without the need for re-learning.
[1061] System configuration
[1062] 1. Collecting user conversation logs
[1063] Users interact with the system using a smartphone or computer. For example, they might input a question like, "Which restaurants are open this afternoon?" The content of this interaction is recorded in real time on the device and periodically sent to the server.
[1064] 2. Gathering reliable external information
[1065] The server periodically obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[1066] 3. Operation of the Emotion Engine
[1067] When the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[1068] 4. Data structuring and storage
[1069] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[1070] 5. Generating the Response
[1071] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[1072] 6. Sending the Response
[1073] The generated response is sent via the server to the terminal and displayed to the user, who then decides what to do next.
[1074] Specific examples
[1075] 1. User Questions
[1076] User: "Which restaurants are open this afternoon?"
[1077] 2. Sending conversation logs
[1078] The terminal sends this question to the server.
[1079] 3. Sentiment Analysis
[1080] The server uses an emotion engine to identify the emotion "interested" from the user's question and stores it in a database.
[1081] 4. Searching for related information
[1082] The server searches a database for user conversation logs, emotional data, and the latest restaurant seat availability information.
[1083] 5. Generating the Response
[1084] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[1085] 6. Sending and Displaying Responses
[1086] The server sends the generated response to the terminal, which displays the response to the user.
[1087] In this way, the system can provide customized responses in real time according to the user's emotional state. The introduction of an emotion engine further improves the user experience. It can provide information quickly and efficiently because it can reflect the latest information in real time without the need for retraining.
[1088] The processing flow will be explained below.
[1089] Step 1:
[1090] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[1091] Step 2:
[1092] The device records the conversation in real time, saves it as text data, and then sends the conversation log to a server.
[1093] Step 3:
[1094] The server passes the received conversation log to the emotion engine, which analyzes the user's emotions. The emotion engine identifies emotions such as "interest," "anxiety," and "joy" from the text and voice data.
[1095] Step 4:
[1096] The server stores the analyzed emotional information along with the conversation log in a database, allowing the content of conversations and emotional states to be organized for each user.
[1097] Step 5:
[1098] The server periodically retrieves the latest data from external trusted sources, such as weather information APIs and restaurant reservation sites, to get the latest weather information and restaurant availability information.
[1099] Step 6:
[1100] The server preprocesses the external information it acquires and stores it in a structured database by category, for example, by organizing it into a "weather information table" or a "restaurant information table."
[1101] Step 7:
[1102] The terminal receives a new question from the user, for example, "Which restaurants are open in the afternoon?"
[1103] Step 8:
[1104] The terminal sends the received question to the server.
[1105] Step 9:
[1106] The server analyzes the question and searches a database for relevant conversation logs, emotion data, and the latest external information based on the analysis.
[1107] Step 10:
[1108] The server passes the search results to a large-scale language model (LLM) to generate the best response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[1109] Step 11:
[1110] Based on the responses generated by the LLM, the server tailors a customized response that takes into account the user's emotional state: if the user is determined to be "interested," it uses more positive language.
[1111] Step 12:
[1112] The server sends the generated customized response to the terminal.
[1113] Step 13:
[1114] The terminal displays the received response to the user, who then decides what to do next based on the response. For example, if the user sees a response saying "Italian, French, and Japanese restaurants are open in the afternoon," they can choose from those restaurants.
[1115] Step 14:
[1116] The server periodically updates the external information and adds the new information to the structured database, ensuring that the entire system always provides up-to-date information.
[1117] Example 2
[1118] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1119] Conventional conversation systems have difficulty generating responses that take the user's emotions into account, resulting in an unsatisfactory user experience. Furthermore, because the collection and updating of conversation logs and external information is not automated, it is difficult to quickly reflect the latest information. As a result, users are unable to receive appropriate information, which can lead to dissatisfaction.
[1120] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1121] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response using a large-scale language model based on the stored data, means for analyzing the user's emotions, means for adjusting the response based on the results of the emotion analysis, and means for transmitting the generated response to the user's terminal. This makes it possible to understand the user's emotions and provide customized responses in real time according to individual needs. Furthermore, the latest information can be quickly reflected without the need for re-learning, thereby improving user satisfaction.
[1122] A "conversation log" is data that records the content of a conversation, such as text or voice data, that a user inputs to the system.
[1123] "External information" is information obtained from trusted external data sources, such as weather information or restaurant availability information.
[1124] A "structured database" is a database in which conversation logs, external information, etc. are organized and stored in a specific structure.
[1125] A "large-scale language model" is an advanced algorithm or model that is trained based on large amounts of data in natural language processing and is capable of understanding and generating language.
[1126] An "emotion engine" is software or algorithm that analyzes emotions from a user's text or voice data and processes the results to identify them.
[1127] A "terminal" is a device that a user uses to interact with the system, including, for example, a smartphone or computer.
[1128] A "response" is information or a message that is generated based on the user's conversation log, external information, and emotional information, and is provided to the user.
[1129] "Collection methods" are mechanisms and processes for acquiring and storing various data (conversation logs, external information, etc.).
[1130] This system generates responses based on a large-scale language model (LLM), which collects user conversation logs and reliable external information. By combining this with an emotion engine that recognizes the user's emotions, it can provide customized responses in real time that correspond to the user's individual needs and emotional state.
[1131] 1. Collecting user conversation logs
[1132] A user interacts with the system using a smartphone or computer. For example, the user types a question such as, "Which restaurants are open this afternoon?" The device records this question in real time and sends it to the server at regular intervals.
[1133] 2. Gathering reliable external information
[1134] The server periodically retrieves the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server retrieves "Tokyo weather information" from a weather API and collects "seat availability information" from a restaurant reservation site.
[1135] 3. Operation of the Emotion Engine
[1136] Once the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[1137] 4. Data structuring and storage
[1138] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[1139] 5. Generating the Response
[1140] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[1141] 6. Sending and Displaying Responses
[1142] The generated response is sent to the terminal via the server and displayed to the user, who then decides on the next course of action based on the displayed information.
[1143] Specific examples
[1144] User Questions
[1145] User: "Which restaurants are open this afternoon?"
[1146] Sending transcripts
[1147] The terminal sends a query to the server.
[1148] Sentiment analysis
[1149] The server uses an emotion engine to identify the emotion "interested" and stores it in a database.
[1150] Searching for information
[1151] The server retrieves the user's conversation log, emotion data, and the latest restaurant seat availability information.
[1152] Generating a response
[1153] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[1154] Sending and Displaying Responses
[1155] The server sends the generated response to the terminal, which displays the response to the user.
[1156] The system can accurately analyze user sentiment and provide customized responses based on individual needs in real time, and the use of a structured database allows it to quickly reflect the latest information without the need for retraining.
[1157] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1158] Step 1:
[1159] The user inputs the conversation into the system. For example, the user uses a smartphone or computer to type, "Which restaurants are open this afternoon?" This input is recorded in real time on the device and saved as a conversation log. The input data is in text format, and the specific operation involves the user typing a message into the target device using a keyboard or voice recognition.
[1160] Step 2:
[1161] The device sends the user's conversation log to the server. The device then processes the collected conversation log to send it to the server at regular intervals. The input is the user's text data, and the output is the conversation log sent to the server. Specifically, an API is called to send the conversation log to the server using an HTTPS request.
[1162] Step 3:
[1163] The server collects external information. The server obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. The input is the API request parameters, and the output is the obtained weather information and seat availability information. Specifically, the server periodically sends requests to the external API, receives responses, and analyzes them.
[1164] Step 4:
[1165] The server analyzes the conversation log using an emotion engine. The emotion engine is used on the received conversation log to identify the user's emotions. The input is the text data of the conversation log, and the output is the emotion analysis results. Specifically, it runs an emotion analysis algorithm (for example, a natural language processing model) to identify the user's emotions (such as "joy," "anger," or "sadness").
[1166] Step 5:
[1167] The server stores the conversation log, external information, and emotional information in a structured database. The server analyzes the data and stores it in an organized form in the structured database. The input is the conversation log, external information, and emotional information, and the output is structured data stored in the database. Specifically, it uses SQL queries to organize the conversation log by user, and organizes the external information by category and inserts it into the database.
[1168] Step 6:
[1169] The server searches for the required information and provides it to a large-scale language model (LLM). The server searches for the required information from a database and provides it as input to the LLM. The input is a database query related to the user's question, and the output is the input data to the LLM. Specifically, the information searched for by the query is formatted into an appropriate format such as JSON and provided to the LLM.
[1170] Step 7:
[1171] The server uses the LLM to generate the optimal response. The LLM generates a response to the user based on the provided data. The input is the searched information and the user's emotional data, and the output is the generated response. Specifically, it sets a prompt in the LLM and generates a response in text format.
[1172] Step 8:
[1173] The server sends the generated response to the terminal. The server sends the generated response to the user's terminal, and the terminal displays this response to the user. The input is the generated response text, and the output is a message to the user displayed on the terminal. In concrete terms, the response is sent to the terminal using an HTTPS request, and the system on the terminal displays the message to the user.
[1174] (Application example 2)
[1175] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1176] Conventional user response systems generate responses primarily based on conversation logs and external information, but because they do not consider the user's emotional state, they often fail to respond appropriately. Furthermore, while there is a demand for improved stress management and work efficiency for operators, particularly in the industrial field, current systems have difficulty meeting these demands. Therefore, there is a growing need for a system that can recognize the user's emotional state in real time and generate appropriate responses based on that information.
[1177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1178] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, emotion recognition means for analyzing the user's emotion based on the conversation log and the external information, means for generating a response using a large-scale language model based on the stored data, means for transmitting the generated response to the user's terminal, and means for customizing the response based on the analyzed emotion information, thereby making it possible to provide an optimal response in real time according to the user's emotional state.
[1179] A "user conversation log" is a record of the text and voice data entered by a user while interacting with the system.
[1180] "Reliable external information" refers to the latest data obtained from reliable external sources, such as weather information APIs and restaurant reservation sites.
[1181] A "structured database" is a database that organizes collected data such as conversation logs and external information, making it possible to search and use it efficiently.
[1182] A "large-scale language model (LLM)" is a machine learning model trained to perform natural language processing based on massive amounts of text data.
[1183] An "emotion recognition means" is an algorithm or system for identifying emotions such as joy, anger, sadness, etc. from text or voice data entered by a user.
[1184] The "means for generating a response" is a process that uses a large-scale language model to generate an optimal response based on the user's conversation log, external information, and emotional data.
[1185] The "means for customizing responses" is a function that adjusts and provides responses generated by a large-scale language model according to the user's emotional state.
[1186] "User terminal" refers to a device used by a user to interact with the system, including a smartphone, computer, robot, etc.
[1187] The system for implementing this invention collects a user's conversation log and incorporates reliable external information to generate customized responses based on the user's emotional state. Specifically, the system performs the following processes.
[1188] Program Overview
[1189] The system is configured using the following hardware and software:
[1190] Hardware:
[1191] Terminals or robots in the factory
[1192] Internet-connected server
[1193] software:
[1194] Large-scale Language Model (LLM): GPT-2 model using the Hugging Face Transformers library
[1195] Emotion Recognition Engine: TextBlob Library
[1196] External information acquisition API: requests library
[1197] What the program does
[1198] 1. Obtaining user conversation logs
[1199] Terminals and robots within the factory acquire text and voice data entered by users (operators) in real time.
[1200] 2. Sentiment analysis
[1201] The server uses TextBlob to extract emotions from user input and store the emotion information in a database. For example, it detects stress from an input such as "This machine seems to be malfunctioning since last night."
[1202] 3. Obtaining external information
[1203] The server uses an external API to obtain the latest status and operation status of factory machines, which is also stored in a database and updated regularly.
[1204] 4. Response Generation
[1205] The server uses a large-scale language model (GPT-2) to generate responses based on the user's conversation log, emotional data, and external information. Specifically, the following prompt sentences are used to provide data to the model:
[1206] Examples:
[1207] Worker: "This machine seems to be acting up since last night."
[1208] Emotion: -0.2 (worried)
[1209] Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheated"}
[1210] assistant:
[1211] Based on the prompt sentence above, the GPT-2 model generates a response.
[1212] 5. Sending the Response
[1213] The generated response is sent via the server to the device or robot being used by the user and displayed to the user. For example, a response such as "Motor overheating detected. We recommend checking the cooling fan" may be generated.
[1214] In this way, by providing real-time responses according to the user's emotional state, it is possible to manage the operator's stress and improve work efficiency. The introduction of an emotion engine also further improves the user experience. Since the latest information can be reflected in real time without the need for re-learning, it is possible to provide information quickly and efficiently.
[1215] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1216] Step 1:
[1217] Users use terminals or robots in the factory to input text or voice data into the system. For example, a user might say, "This machine seems to have been acting up since last night." This becomes the user's conversation log.
[1218] Input: User text or voice data
[1219] Output: Conversation log
[1220] Step 2:
[1221] The device sends the acquired conversation log in real time to the server, which receives the conversation log and stores it in a database.
[1222] Input: conversation log
[1223] Output: Conversation logs stored in a database
[1224] Step 3:
[1225] The server uses the TextBlob library to analyze emotions from the received conversation log. For example, stress can be detected from the text, "This machine seems to be malfunctioning since last night." The analyzed emotional information is stored in a database.
[1226] Input: conversation log
[1227] Data Processing: Sentiment Analysis using TextBlob
[1228] Output: Emotional information (e.g., stress)
[1229] Step 4:
[1230] The server uses an external API (e.g., machine status API) to obtain the latest machine operating status and warning information. The obtained information is stored in a database and updated regularly.
[1231] Input: Machine data request from external API
[1232] Data processing: Data acquisition from API
[1233] Output: External information stored in a database
[1234] Step 5:
[1235] The server provides prompts to the generative AI model (GPT-2) based on the user's conversation log, emotional data, and external information. For example, it generates the following prompt sentence:
[1236] "User Input: "This machine has been acting up since last night." Emotion: Stress Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheating" Assistant: "
[1237] Input: User conversation logs, emotion data, external information
[1238] Data processing: Prompt sentence generation
[1239] Output: Prompt sentence to the generative AI model
[1240] Step 6:
[1241] The generative AI model (GPT-2) generates the optimal response based on the provided prompt, for example, "Motor overheating detected. We recommend checking the cooling fan."
[1242] Input: prompt statement
[1243] Data Computing: Response Generation Using Generative AI Models
[1244] Output: The generated response
[1245] Step 7:
[1246] The server sends the generated response to the terminal, which displays the response to the user, who can review the response and decide what to do next.
[1247] Input: Generated response
[1248] Output: The response displayed on the user's terminal
[1249] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1250] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1251] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1252] [Fourth embodiment]
[1253] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1254] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1255] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1256] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1257] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1258] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1259] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1260] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1261] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1262] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1263] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1264] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1265] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1266] This invention relates to a system that collects user conversation logs and reliable external information, stores them in a structured database, generates responses using a large-scale language model (LLM) based on the stored data, and transmits the generated responses to the user's device. By using this system, it is possible to respond to individual needs and always provide the latest and most reliable information.
[1267] System configuration
[1268] 1. Collecting user conversation logs
[1269] Users interact with the system on a daily basis via their smartphones or computers. For example, a user might input a question such as, "Which restaurants are open this afternoon?" This interaction is digitally recorded on the user's device and periodically sent to the server.
[1270] 2. Gathering reliable external information
[1271] The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[1272] 3. Data structuring and storage
[1273] The server analyzes the user's conversation log and external information and stores them in a structured database. The conversation log is categorized by user, and reliable external information is organized by category. For example, a user's conversation log includes "question content," "date and time," and "frequency," while external information includes "weather information" and "restaurant availability."
[1274] 4. Generating the Response
[1275] The server retrieves the necessary information from a structured database and provides it to a large-scale language model. LLM generates the optimal response based on this information. For example, if a user asks about restaurant availability, LLM generates the response "Italian, French, and Japanese restaurants are open in the afternoon."
[1276] 5. Sending the Response
[1277] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[1278] Specific examples
[1279] 1. User Questions
[1280] User: "Which restaurants are open this afternoon?"
[1281] 2. Sending conversation logs
[1282] The terminal sends this question to the server.
[1283] 3. Search for related information
[1284] The server searches the database for the user's past conversation log and the latest restaurant vacancy information.
[1285] 4. Generating the Response
[1286] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[1287] 5. Sending and Displaying Responses
[1288] The server sends the generated response to the terminal, which displays the response to the user.
[1289] In this way, the system can provide customized responses in real time, and the use of a structured database allows it to quickly update without the need for retraining.
[1290] The processing flow will be explained below.
[1291] Step 1:
[1292] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[1293] Step 2:
[1294] The device records the conversation in real time, saves it as text data, and then periodically sends the conversation log to a server.
[1295] Step 3:
[1296] The server analyzes the received conversation logs and stores them in a database for each user. The analysis process includes extracting important information such as question content, time, and frequency.
[1297] Step 4:
[1298] The server periodically retrieves the latest data from external trusted sources (e.g., weather information APIs or restaurant reservation sites), which is also preprocessed and stored in a structured database.
[1299] Step 5:
[1300] When the device receives a new question from the user, it sends the content to the server, where it is processed in real time.
[1301] Step 6:
[1302] Based on the received question, the server searches the database for relevant conversation logs and external information, for example, past conversation history and the latest restaurant information.
[1303] Step 7:
[1304] The server then passes the search results to a large-scale language model (LLM) to generate the best possible response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[1305] Step 8:
[1306] The server sends the generated response to the terminal.
[1307] Step 9:
[1308] The terminal displays the received response to the user, who then decides on the next action to take.
[1309] Step 10:
[1310] The server periodically updates the external information and stores it in the database again, so that the latest information is always available within the system.
[1311] In this way, the system provides fast and accurate responses to user questions, and by utilizing a structured database and large-scale language models, it is possible to reflect the latest information in real time without the need for retraining.
[1312] Example 1
[1313] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1314] Conventional systems have had difficulty efficiently collecting and organizing user conversation logs and combining them with reliable external information to provide optimal responses in real time. Furthermore, because the external information is not updated frequently enough, there is no guarantee that the information provided is up-to-date and accurate. Furthermore, generating effective prompts has been a challenge when utilizing large-scale language models (LLMs).
[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1316] In this invention, the server includes means for digitally collecting user conversation logs and transmitting them from the terminal to the server, means for periodically collecting reliable external information from APIs and information sites, and means for analyzing and classifying the conversation logs and external information and storing them in a structured database. This enables the generation of accurate responses that meet individual needs based on the latest, most reliable information in real time.
[1317] "User conversation log" refers to the content of communication such as questions and requests made by the user to the system.
[1318] "Digital format" refers to a format in which information, such as sound or text, can be recorded and processed electronically.
[1319] "Terminal" refers to a device, such as a smartphone or computer, that a user uses to access the system.
[1320] "Server" refers to a computer system for receiving and processing data sent from a user's terminal and generating a response.
[1321] "Reliable external information" refers to information obtained from reliable sources such as weather information APIs and restaurant reservation sites.
[1322] "API" stands for Application Programming Interface and refers to a protocol for exchanging information between different software programs.
[1323] An "information site" refers to a web page or online service that provides specific information.
[1324] "Analysis" refers to the process of breaking down data into more detailed pieces to make it easier to understand.
[1325] "Classification" refers to the process of grouping collected data based on specific criteria.
[1326] A "structured database" refers to a database in which data is systematically organized so that it can be efficiently stored and searched.
[1327] A "large-scale language model (LLM)" refers to a natural language processing model trained on large amounts of text data.
[1328] A "prompt sentence" refers to a guided text that is input into a large-scale language model to generate an appropriate response.
[1329] "Response" refers to the answer or information generated or provided by the system in response to a user's question.
[1330] The present invention is a system that collects user conversation logs and reliable external information, stores them in a structured database, and processes them. This system generates responses based on the collected data using a large-scale language model (LLM), and transmits the generated responses to the user's terminal for display. A specific embodiment of the system is described below.
[1331] Hardware and Software Configuration
[1332] 1. Hardware
[1333] User terminal: A device through which a user accesses the system, such as a smartphone or computer.
[1334] Server: A computer system, such as a cloud server, that receives data sent from a user's device, processes it, and generates a response.
[1335] 2. Software
[1336] Conversation log collection software: Software for collecting user conversations in digital form.
[1337] External data collection software: Software for collecting reliable external information from APIs and information sites.
[1338] Structured database software: Software for analyzing, classifying, and storing collected data.
[1339] Large-scale language model (LLM) software: Software that uses natural language processing techniques to generate responses based on collected data.
[1340] System configuration and operation
[1341] The system operates in the following steps.
[1342] 1. Collecting user conversation logs
[1343] User Asks a Question: A user uses a smartphone or computer to send a question or request to the system. Example: "Which restaurants are open this afternoon?"
[1344] Device processing: The device digitally records the user's conversations and periodically transmits them to a server, in particular speech recognition software that may convert speech to text.
[1345] 2. Gathering external information
[1346] Server information collection: The server periodically collects the latest data from external information sources such as weather information APIs and restaurant reservation sites. Example: Obtaining "Tokyo weather information" or "restaurant seat availability information."
[1347] 3. Data structuring and storage
[1348] Server data analysis: The server analyzes the received conversation logs and external information, and classifies and stores them in a structured database. Specifically, tokenization and tagging are performed.
[1349] Conversation log classification: User conversation logs are classified by user ID, and external information is organized by category.
[1350] 4. Generating the Response
[1351] Server data retrieval: The server retrieves the information required by the user from a structured database and provides it to the LLM.
[1352] LLM response generation: LLM generates the best response based on the data provided. Example: "Italian, French, and Japanese restaurants are open in the afternoon."
[1353] 5. Sending the Response
[1354] Transmission from server to terminal: The generated response is transmitted to the user's terminal via the server.
[1355] Display on the terminal: The terminal displays the received response to the user, who can then ask further questions based on the displayed information.
[1356] Specific examples
[1357] Below are some examples of specific prompt sentences.
[1358] 1. User Questions
[1359] User: "Which restaurants are open this afternoon?"
[1360] 2. Sending the device
[1361] The terminal sends this question in text format to the server.
[1362] 3. Searching for a server
[1363] The server retrieves the conversation log and the latest restaurant information from a database.
[1364] 4. Generating the Response
[1365] The server provides the LLM with "afternoon restaurant availability information" and "past conversation history," which generates a response saying, "Italian, French, and Japanese restaurants are available in the afternoon."
[1366] 5. Sending and Displaying Responses
[1367] The server sends the generated response to the terminal, which displays the response to the user.
[1368] The system's features include the ability to respond quickly and accurately to user needs, collecting and providing the latest information in real time, and utilizing a structured database to generate responses with high efficiency without the need for re-learning.
[1369] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1370] Step 1:
[1371] A user enters a question or request into the system.
[1372] Input: User question (e.g., "Which restaurants are open this afternoon?")
[1373] Specific operation: The user accesses the system using a smartphone or computer and enters a question.
[1374] Step 2:
[1375] The device collects the user's conversation log in digital format and sends it to the server.
[1376] Input: User conversation log
[1377] Output: Digital conversation log
[1378] Specific operation: The device converts the conversation log into text format, adds metadata such as the user ID, question content, date and time, and sends it to the server periodically.
[1379] Step 3:
[1380] The server periodically collects reliable external information from APIs and information sites.
[1381] Input: URL of external information API endpoint or information site
[1382] Output: External information data (e.g., weather information, restaurant seat availability information)
[1383] Specific operation: The server periodically retrieves the latest data from external sources such as weather information APIs and restaurant reservation sites.
[1384] Step 4:
[1385] The server analyzes the conversation log and external information and stores it in a structured database.
[1386] Input: Digital conversation logs and external information data
[1387] Output: Parsed and classified data
[1388] How it works: The server uses natural language processing technology to tokenize the conversation logs and organize external information into categories. The analyzed and categorized data is then stored in a structured database.
[1389] Step 5:
[1390] The server retrieves the necessary information from a structured database and provides it to the LLM.
[1391] Input: User conversation logs and external information
[1392] Output: Data provided to LLM
[1393] Specific operation: The server searches the user's past conversation history and the latest external information from a structured database and provides them to the LLM.
[1394] Step 6:
[1395] LLM generates the best response based on the data provided.
[1396] Input: User conversation logs and external information
[1397] Output: Generated response text (e.g., "Italian, French, and Japanese restaurants are open in the afternoon.")
[1398] How it works: LLM uses natural language processing algorithms to generate optimal responses based on tokenized data.
[1399] Step 7:
[1400] The server sends the generated response to the user's terminal.
[1401] Input: Generated response text
[1402] Output: The response sent to the user's device
[1403] Specific operation: The server sends the generated response to the user's device, transferring data in real time using a communication protocol.
[1404] Step 8:
[1405] The terminal displays the received response to the user.
[1406] Input: Response text sent by the server
[1407] Output: The response message to be displayed
[1408] Specific operation: The terminal displays the received response on the user interface, and the user can ask more detailed questions based on this information.
[1409] Through these steps, the system is able to provide accurate and appropriate responses to user questions in real time.
[1410] (Application example 1)
[1411] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1412] Modern food delivery services are required to respond quickly and accurately to diverse user needs. However, conventional systems have difficulty providing personalized suggestions based on user preferences or real-time restaurant information. This has led to issues such as users being unable to quickly make choices that suit their preferences and current circumstances, resulting in reduced convenience.
[1413] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1414] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response based on the stored data using a large-scale language model, means for transmitting the generated response to the user's terminal, and means for providing customized recommendations based on the user's past ordering history and preferences. This allows users to receive real-time suggestions that match their preferences, significantly improving the convenience of food delivery.
[1415] A "conversation log" is text data generated through a user's interaction with the system, and is information that includes the user's questions and responses.
[1416] "External information" refers to the latest data the system collects from trusted external sources, such as restaurant menus, seat availability, and weather information.
[1417] A "structured database" is a database for organizing and storing conversation logs and external information, in which each piece of information is organized by classification and category.
[1418] A "large-scale language model" is an advanced machine learning model for natural language processing that generates optimal responses based on data provided by a structured database.
[1419] A "terminal" is a digital device used by a user, such as a smartphone or computer, on which the generated response is displayed.
[1420] "Customized Recommendations" are suggestions that are specifically tailored based on a user's past ordering history and preferences, and are information designed to address a user's individual needs.
[1421] A system embodying the present invention includes the following functions.
[1422] The server first collects the user's conversation log. The user interacts with the system via a smartphone or computer, and the conversation is recorded digitally. For example, the user may enter a question such as, "Which restaurants are available now?" This question is collected and sent to the server.
[1423] Next, the server collects reliable external information, such as restaurant menus, seat availability, and weather information, which it periodically obtains via online APIs. For example, the server obtains current seat availability information from a restaurant reservation site.
[1424] The server then stores the collected conversation logs and external information in a structured database. Conversation logs are organized by user, and external information is categorized. Specifically, a user's conversation log includes "question content," "date and time," and "frequency," while reliable external information includes "restaurant seat availability information" and "menu information."
[1425] The server then uses the stored data to generate a response using a large-scale language model (LLM). The server searches a structured database for the necessary information and sends a prompt to the LLM based on this information. For example, if a user asks about restaurant availability, the LLM can generate a response such as, "As of 10:00 AM, an Italian restaurant and a Chinese restaurant are available."
[1426] The generated response is sent to the user's terminal via the server and displayed to the user, who can then ask further detailed questions based on the response.
[1427] Additionally, the server provides customized recommendations based on the user's past ordering history and preferences, allowing the user to quickly find options that suit their preferences.
[1428] For example, the following prompts might be used for a generative AI model:
[1429] text
[1430] User asks: "What Italian restaurant would you recommend?"
[1431] Past conversation: [{ "timestamp": "2023-10-08T12:00:00", "conversation": "I like pizza"}, {...}]
[1432] External Data: [{ "name": "Italian Restaurant A", "rating": 4.5, "availability": "open"}, {...}]
[1433] Based on this prompt, the LLM processes past conversation data and external information to generate a response such as, "Based on your preferences, we recommend Italian restaurant A." In this way, the present invention can provide optimal information tailored to the user's individual needs.
[1434] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1435] Step 1:
[1436] The user enters a question and the terminal receives the question.
[1437] Specifically, when a user uses a smartphone or computer to enter a question such as "Which restaurants are open now?", the text data is sent to the device and saved.
[1438] Input: User question
[1439] Output: Question data in text format
[1440] Step 2:
[1441] The terminal transmits the saved question data to the server.
[1442] This data is sent to the server along with the user ID, and the server receives it and records it as a conversation log.
[1443] Input: Text question data, user ID
[1444] Output: Conversation log sent to the server
[1445] Step 3:
[1446] The server collects trusted external information.
[1447] Specifically, the server accesses a pre-configured API endpoint (e.g., the API of a restaurant reservation system) and obtains the latest data such as restaurant seat availability and menu information.
[1448] Input: API endpoint
[1449] Output: External information data (e.g., restaurant seat availability, menu information)
[1450] Step 4:
[1451] The server stores the collected conversation logs and external information in a structured database.
[1452] Conversation logs are categorized by user, and external information is organized by category. For example, restaurant vacancy information is organized into categories such as "restaurant" and "vacancy information."
[1453] Input: Conversation log, external information
[1454] Output: Structured database entries
[1455] Step 5:
[1456] The server searches for the necessary information from a structured database and uses it to generate and send prompts to a large-scale language model (LLM).
[1457] The prompt includes the user's question, past conversation logs, and external information. The prompt is created and sent to the LLM to generate the optimal response.
[1458] Input: Structured database information, user questions
[1459] Output: Generated prompt, response from LLM
[1460] Step 6:
[1461] The server sends the generated response to the user's terminal.
[1462] The terminal displays the received response to the user, thereby providing the user with an answer to their question.
[1463] Input: The response generated by the LLM
[1464] Output: The response displayed on the user's terminal
[1465] Step 7:
[1466] The server generates and provides customized recommendations to the user based on the user's past ordering history and preferences.
[1467] By analyzing past conversation logs and order data, the system suggests restaurants and menus that suit the user's preferences.
[1468] Input: User's past order history, preference data
[1469] Output: Customized recommendations
[1470] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1471] This invention relates to a system that collects user conversation logs and reliable external information, generates responses based on these using a large-scale language model (LLM), and combines this with an emotion engine that recognizes the user's emotions. This system is capable of providing customized responses in real time according to the user's individual needs and emotional state. Furthermore, by utilizing a structured database, it is possible to quickly reflect the latest information without the need for re-learning.
[1472] System configuration
[1473] 1. Collecting user conversation logs
[1474] Users interact with the system using a smartphone or computer. For example, they might input a question like, "Which restaurants are open this afternoon?" The content of this interaction is recorded in real time on the device and periodically sent to the server.
[1475] 2. Gathering reliable external information
[1476] The server periodically obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server obtains "Tokyo weather information" from a weather API and "seat availability information" from a restaurant reservation site.
[1477] 3. Operation of the Emotion Engine
[1478] When the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[1479] 4. Data structuring and storage
[1480] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[1481] 5. Generating the Response
[1482] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[1483] 6. Sending the Response
[1484] The generated response is sent via the server to the terminal and displayed to the user, who then decides what to do next.
[1485] Specific examples
[1486] 1. User Questions
[1487] User: "Which restaurants are open this afternoon?"
[1488] 2. Sending conversation logs
[1489] The terminal sends this question to the server.
[1490] 3. Sentiment Analysis
[1491] The server uses an emotion engine to identify the emotion "interested" from the user's question and stores it in a database.
[1492] 4. Searching for related information
[1493] The server searches a database for user conversation logs, emotional data, and the latest restaurant seat availability information.
[1494] 5. Generating the Response
[1495] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[1496] 6. Sending and Displaying Responses
[1497] The server sends the generated response to the terminal, which displays the response to the user.
[1498] In this way, the system can provide customized responses in real time according to the user's emotional state. The introduction of an emotion engine further improves the user experience. It can provide information quickly and efficiently because it can reflect the latest information in real time without the need for retraining.
[1499] The processing flow will be explained below.
[1500] Step 1:
[1501] A user interacts with the system using a terminal. For example, the user inputs a question such as, "Which restaurants are open in the afternoon?"
[1502] Step 2:
[1503] The device records the conversation in real time, saves it as text data, and then sends the conversation log to a server.
[1504] Step 3:
[1505] The server passes the received conversation log to the emotion engine, which analyzes the user's emotions. The emotion engine identifies emotions such as "interest," "anxiety," and "joy" from the text and voice data.
[1506] Step 4:
[1507] The server stores the analyzed emotional information along with the conversation log in a database, allowing the content of conversations and emotional states to be organized for each user.
[1508] Step 5:
[1509] The server periodically retrieves the latest data from external trusted sources, such as weather information APIs and restaurant reservation sites, to get the latest weather information and restaurant availability information.
[1510] Step 6:
[1511] The server preprocesses the external information it acquires and stores it in a structured database by category, for example, by organizing it into a "weather information table" or a "restaurant information table."
[1512] Step 7:
[1513] The terminal receives a new question from the user, for example, "Which restaurants are open in the afternoon?"
[1514] Step 8:
[1515] The terminal sends the received question to the server.
[1516] Step 9:
[1517] The server analyzes the question and searches a database for relevant conversation logs, emotion data, and the latest external information based on the analysis.
[1518] Step 10:
[1519] The server passes the search results to a large-scale language model (LLM) to generate the best response, such as "Italian, French, and Japanese restaurants are open in the afternoon."
[1520] Step 11:
[1521] Based on the responses generated by the LLM, the server tailors a customized response that takes into account the user's emotional state: if the user is determined to be "interested," it uses more positive language.
[1522] Step 12:
[1523] The server sends the generated customized response to the terminal.
[1524] Step 13:
[1525] The terminal displays the received response to the user, who then decides what to do next based on the response. For example, if the user sees a response saying "Italian, French, and Japanese restaurants are open in the afternoon," they can choose from those restaurants.
[1526] Step 14:
[1527] The server periodically updates the external information and adds the new information to the structured database, ensuring that the entire system always provides up-to-date information.
[1528] Example 2
[1529] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1530] Conventional conversation systems have difficulty generating responses that take the user's emotions into account, resulting in an unsatisfactory user experience. Furthermore, because the collection and updating of conversation logs and external information is not automated, it is difficult to quickly reflect the latest information. As a result, users are unable to receive appropriate information, which can lead to dissatisfaction.
[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1532] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, means for generating a response using a large-scale language model based on the stored data, means for analyzing the user's emotions, means for adjusting the response based on the results of the emotion analysis, and means for transmitting the generated response to the user's terminal. This makes it possible to understand the user's emotions and provide customized responses in real time according to individual needs. Furthermore, the latest information can be quickly reflected without the need for re-learning, thereby improving user satisfaction.
[1533] A "conversation log" is data that records the content of a conversation, such as text or voice data, that a user inputs to the system.
[1534] "External information" is information obtained from trusted external data sources, such as weather information or restaurant availability information.
[1535] A "structured database" is a database in which conversation logs, external information, etc. are organized and stored in a specific structure.
[1536] A "large-scale language model" is an advanced algorithm or model that is trained based on large amounts of data in natural language processing and is capable of understanding and generating language.
[1537] An "emotion engine" is software or algorithm that analyzes emotions from a user's text or voice data and processes the results to identify them.
[1538] A "terminal" is a device that a user uses to interact with the system, including, for example, a smartphone or computer.
[1539] A "response" is information or a message that is generated based on the user's conversation log, external information, and emotional information, and is provided to the user.
[1540] "Collection methods" are mechanisms and processes for acquiring and storing various data (conversation logs, external information, etc.).
[1541] This system generates responses based on a large-scale language model (LLM), which collects user conversation logs and reliable external information. By combining this with an emotion engine that recognizes the user's emotions, it can provide customized responses in real time that correspond to the user's individual needs and emotional state.
[1542] 1. Collecting user conversation logs
[1543] A user interacts with the system using a smartphone or computer. For example, the user types a question such as, "Which restaurants are open this afternoon?" The device records this question in real time and sends it to the server at regular intervals.
[1544] 2. Gathering reliable external information
[1545] The server periodically retrieves the latest data from external information sources such as weather information APIs and restaurant reservation sites. For example, the server retrieves "Tokyo weather information" from a weather API and collects "seat availability information" from a restaurant reservation site.
[1546] 3. Operation of the Emotion Engine
[1547] Once the conversation log is sent to the server, the server uses an emotion engine to analyze the user's emotions. For example, emotions such as "joy," "anger," and "sadness" can be identified from the user's text and voice data. This emotional information is also stored in a database.
[1548] 4. Data structuring and storage
[1549] The server analyzes the conversation logs, external information, and emotional information and stores them in a structured database. The conversation logs are organized by user, and reliable external information is organized by category. Emotional information is also stored in association with the corresponding conversation logs.
[1550] 5. Generating the Response
[1551] The server retrieves the necessary information from a structured database and provides it to a large-scale language model (LLM). The LLM generates the optimal response based on this information and the user's emotional state. For example, if a user asks, "Which restaurants are open in the afternoon?", the LLM will generate, "Italian, French, and Japanese restaurants are open in the afternoon." If the user expresses anxiety, the LLM will generate a more reassuring response.
[1552] 6. Sending and Displaying Responses
[1553] The generated response is sent to the terminal via the server and displayed to the user, who then decides on the next course of action based on the displayed information.
[1554] Specific examples
[1555] User Questions
[1556] User: "Which restaurants are open this afternoon?"
[1557] Sending transcripts
[1558] The terminal sends a query to the server.
[1559] Sentiment analysis
[1560] The server uses an emotion engine to identify the emotion "interested" and stores it in a database.
[1561] Searching for information
[1562] The server retrieves the user's conversation log, emotion data, and the latest restaurant seat availability information.
[1563] Generating a response
[1564] The server provides the LLM with "afternoon restaurant availability information," "past conversation history," and emotional data such as "user interest," and generates a response such as "Italian, French, and Japanese restaurants are available in the afternoon."
[1565] Sending and Displaying Responses
[1566] The server sends the generated response to the terminal, which displays the response to the user.
[1567] The system can accurately analyze user sentiment and provide customized responses based on individual needs in real time, and the use of a structured database allows it to quickly reflect the latest information without the need for retraining.
[1568] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1569] Step 1:
[1570] The user inputs the conversation into the system. For example, the user uses a smartphone or computer to type, "Which restaurants are open this afternoon?" This input is recorded in real time on the device and saved as a conversation log. The input data is in text format, and the specific operation involves the user typing a message into the target device using a keyboard or voice recognition.
[1571] Step 2:
[1572] The device sends the user's conversation log to the server. The device then processes the collected conversation log to send it to the server at regular intervals. The input is the user's text data, and the output is the conversation log sent to the server. Specifically, an API is called to send the conversation log to the server using an HTTPS request.
[1573] Step 3:
[1574] The server collects external information. The server obtains the latest data from external information sources such as weather information APIs and restaurant reservation sites. The input is the API request parameters, and the output is the obtained weather information and seat availability information. Specifically, the server periodically sends requests to the external API, receives responses, and analyzes them.
[1575] Step 4:
[1576] The server analyzes the conversation log using an emotion engine. The emotion engine is used on the received conversation log to identify the user's emotions. The input is the text data of the conversation log, and the output is the emotion analysis results. Specifically, it runs an emotion analysis algorithm (for example, a natural language processing model) to identify the user's emotions (such as "joy," "anger," or "sadness").
[1577] Step 5:
[1578] The server stores the conversation log, external information, and emotional information in a structured database. The server analyzes the data and stores it in an organized form in the structured database. The input is the conversation log, external information, and emotional information, and the output is structured data stored in the database. Specifically, it uses SQL queries to organize the conversation log by user, and organizes the external information by category and inserts it into the database.
[1579] Step 6:
[1580] The server searches for the required information and provides it to a large-scale language model (LLM). The server searches for the required information from a database and provides it as input to the LLM. The input is a database query related to the user's question, and the output is the input data to the LLM. Specifically, the information searched for by the query is formatted into an appropriate format such as JSON and provided to the LLM.
[1581] Step 7:
[1582] The server uses the LLM to generate the optimal response. The LLM generates a response to the user based on the provided data. The input is the searched information and the user's emotional data, and the output is the generated response. Specifically, it sets a prompt in the LLM and generates a response in text format.
[1583] Step 8:
[1584] The server sends the generated response to the terminal. The server sends the generated response to the user's terminal, and the terminal displays this response to the user. The input is the generated response text, and the output is a message to the user displayed on the terminal. In concrete terms, the response is sent to the terminal using an HTTPS request, and the system on the terminal displays the message to the user.
[1585] (Application example 2)
[1586] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1587] Conventional user response systems generate responses primarily based on conversation logs and external information, but because they do not consider the user's emotional state, they often fail to respond appropriately. Furthermore, while there is a demand for improved stress management and work efficiency for operators, particularly in the industrial field, current systems have difficulty meeting these demands. Therefore, there is a growing need for a system that can recognize the user's emotional state in real time and generate appropriate responses based on that information.
[1588] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1589] In this invention, the server includes means for collecting a user's conversation log, means for collecting reliable external information, means for storing the conversation log and the external information in a structured database, emotion recognition means for analyzing the user's emotion based on the conversation log and the external information, means for generating a response using a large-scale language model based on the stored data, means for transmitting the generated response to the user's terminal, and means for customizing the response based on the analyzed emotion information, thereby making it possible to provide an optimal response in real time according to the user's emotional state.
[1590] A "user conversation log" is a record of the text and voice data entered by a user while interacting with the system.
[1591] "Reliable external information" refers to the latest data obtained from reliable external sources, such as weather information APIs and restaurant reservation sites.
[1592] A "structured database" is a database that organizes collected data such as conversation logs and external information, making it possible to search and use it efficiently.
[1593] A "large-scale language model (LLM)" is a machine learning model trained to perform natural language processing based on massive amounts of text data.
[1594] An "emotion recognition means" is an algorithm or system for identifying emotions such as joy, anger, sadness, etc. from text or voice data entered by a user.
[1595] The "means for generating a response" is a process that uses a large-scale language model to generate an optimal response based on the user's conversation log, external information, and emotional data.
[1596] The "means for customizing responses" is a function that adjusts and provides responses generated by a large-scale language model according to the user's emotional state.
[1597] "User terminal" refers to a device used by a user to interact with the system, including a smartphone, computer, robot, etc.
[1598] The system for implementing this invention collects a user's conversation log and incorporates reliable external information to generate customized responses based on the user's emotional state. Specifically, the system performs the following processes.
[1599] Program Overview
[1600] The system is configured using the following hardware and software:
[1601] Hardware:
[1602] Terminals or robots in the factory
[1603] Internet-connected server
[1604] software:
[1605] Large-scale Language Model (LLM): GPT-2 model using the Hugging Face Transformers library
[1606] Emotion Recognition Engine: TextBlob Library
[1607] External information acquisition API: requests library
[1608] What the program does
[1609] 1. Obtaining user conversation logs
[1610] Terminals and robots within the factory acquire text and voice data entered by users (operators) in real time.
[1611] 2. Sentiment analysis
[1612] The server uses TextBlob to extract emotions from user input and store the emotion information in a database. For example, it detects stress from an input such as "This machine seems to be malfunctioning since last night."
[1613] 3. Obtaining external information
[1614] The server uses an external API to obtain the latest status and operation status of factory machines, which is also stored in a database and updated regularly.
[1615] 4. Response Generation
[1616] The server uses a large-scale language model (GPT-2) to generate responses based on the user's conversation log, emotional data, and external information. Specifically, the following prompt sentences are used to provide data to the model:
[1617] Examples:
[1618] Worker: "This machine seems to be acting up since last night."
[1619] Emotion: -0.2 (worried)
[1620] Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheated"}
[1621] assistant:
[1622] Based on the prompt sentence above, the GPT-2 model generates a response.
[1623] 5. Sending the Response
[1624] The generated response is sent via the server to the device or robot being used by the user and displayed to the user. For example, a response such as "Motor overheating detected. We recommend checking the cooling fan" may be generated.
[1625] In this way, by providing real-time responses according to the user's emotional state, it is possible to manage the operator's stress and improve work efficiency. The introduction of an emotion engine also further improves the user experience. Since the latest information can be reflected in real time without the need for re-learning, it is possible to provide information quickly and efficiently.
[1626] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1627] Step 1:
[1628] Users use terminals or robots in the factory to input text or voice data into the system. For example, a user might say, "This machine seems to have been acting up since last night." This becomes the user's conversation log.
[1629] Input: User text or voice data
[1630] Output: Conversation log
[1631] Step 2:
[1632] The device sends the acquired conversation log in real time to the server, which receives the conversation log and stores it in a database.
[1633] Input: conversation log
[1634] Output: Conversation logs stored in a database
[1635] Step 3:
[1636] The server uses the TextBlob library to analyze emotions from the received conversation log. For example, stress can be detected from the text, "This machine seems to be malfunctioning since last night." The analyzed emotional information is stored in a database.
[1637] Input: conversation log
[1638] Data Processing: Sentiment Analysis using TextBlob
[1639] Output: Emotional information (e.g., stress)
[1640] Step 4:
[1641] The server uses an external API (e.g., machine status API) to obtain the latest machine operating status and warning information. The obtained information is stored in a database and updated regularly.
[1642] Input: Machine data request from external API
[1643] Data processing: Data acquisition from API
[1644] Output: External information stored in a database
[1645] Step 5:
[1646] The server provides prompts to the generative AI model (GPT-2) based on the user's conversation log, emotional data, and external information. For example, it generates the following prompt sentence:
[1647] "User Input: "This machine has been acting up since last night." Emotion: Stress Machine Status: {"Machine ID": "ABC123", "Status": "Warning", "Details": "Motor overheating" Assistant: "
[1648] Input: User conversation logs, emotion data, external information
[1649] Data processing: Prompt sentence generation
[1650] Output: Prompt sentence to the generative AI model
[1651] Step 6:
[1652] The generative AI model (GPT-2) generates the optimal response based on the provided prompt, for example, "Motor overheating detected. We recommend checking the cooling fan."
[1653] Input: prompt statement
[1654] Data Computing: Response Generation Using Generative AI Models
[1655] Output: The generated response
[1656] Step 7:
[1657] The server sends the generated response to the terminal, which displays the response to the user, who can review the response and decide what to do next.
[1658] Input: Generated response
[1659] Output: The response displayed on the user's terminal
[1660] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1661] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1662] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1663] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1664] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1665] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1666] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1667] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1668] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1669] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1670] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1671] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1672] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1673] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1674] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1675] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1676] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1677] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1678] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1679] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1680] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1681] The following is further disclosed regarding the above embodiment.
[1682] (Claim 1)
[1683] A means for collecting user conversation logs;
[1684] A means of gathering reliable external information;
[1685] means for storing the conversation log and the external information in a structured database;
[1686] a means for generating responses using a large-scale language model based on the stored data;
[1687] means for transmitting the generated response to the user's terminal;
[1688] A system including:
[1689] (Claim 2)
[1690] 2. The system of claim 1, wherein the conversation logs are organized by user.
[1691] (Claim 3)
[1692] The system of claim 1 , wherein the external information is updated periodically.
[1693] (Claim 4)
[1694] 10. The system of claim 1, wherein the large-scale language model obtains information from a structured database to generate responses.
[1695] (Claim 5)
[1696] 10. The system of claim 1, wherein the generated response is customized based on user preferences and past conversation history.
[1697] "Example 1"
[1698] (Claim 1)
[1699] A means for collecting a user's conversation log in digital form and transmitting it from the terminal to a server;
[1700] A means of regularly collecting reliable external information from APIs and information sites,
[1701] means for analyzing and classifying the conversation log and the external information and storing them in a structured database;
[1702] A means for generating prompt sentences using a large-scale language model based on the stored data and generating optimal responses;
[1703] means for transmitting the generated response to a user's terminal for display;
[1704] A system including:
[1705] (Claim 2)
[1706] 2. The system of claim 1, wherein the conversation logs are categorized by user.
[1707] (Claim 3)
[1708] 2. The system according to claim 1, wherein the external information is updated periodically at a set timing.
[1709] "Application Example 1"
[1710] (Claim 1)
[1711] A means for collecting user conversation logs;
[1712] A means of gathering reliable external information;
[1713] means for storing the conversation log and the external information in a structured database;
[1714] a means for generating responses using a large-scale language model based on the stored data;
[1715] means for transmitting the generated response to the user's terminal;
[1716] means for providing customized recommendations based on the user's past ordering history and user preferences;
[1717] A system including:
[1718] (Claim 2)
[1719] 2. The system of claim 1, wherein the conversation logs are organized by user.
[1720] (Claim 3)
[1721] The system of claim 1 , wherein the external information is updated periodically.
[1722] "Example 2: Combining Emotion Engines"
[1723] (Claim 1)
[1724] A means for collecting user conversation logs;
[1725] A means of gathering reliable external information;
[1726] means for storing the conversation log and the external information in a structured database;
[1727] a means for generating responses using a large-scale language model based on the stored data;
[1728] means for analyzing user sentiment;
[1729] means for adjusting a response based on the results of said sentiment analysis;
[1730] means for transmitting the generated response to the user's terminal;
[1731] A system including:
[1732] (Claim 2)
[1733] 2. The system of claim 1, wherein the conversation logs are organized by user.
[1734] (Claim 3)
[1735] The system of claim 1 , wherein the external information is updated periodically.
[1736] "Application example 2 when combining emotion engines"
[1737] (Claim 1)
[1738] A means for collecting user conversation logs;
[1739] A means of gathering reliable external information;
[1740] means for storing the conversation log and the external information in a structured database;
[1741] a means for generating responses using a large-scale language model based on the stored data;
[1742] means for transmitting the generated response to the user's terminal;
[1743] emotion recognition means for analyzing the emotion of a user based on the conversation log and the external information;
[1744] a means of customizing responses based on the analyzed emotional information; and
[1745] A system including:
[1746] (Claim 2)
[1747] 2. The system of claim 1, wherein the conversation logs are organized by user.
[1748] (Claim 3)
[1749] The system of claim 1 , wherein the external information is updated periodically. [Explanation of symbols]
[1750] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for collecting user conversation logs; A means of gathering reliable external information; means for storing the conversation log and the external information in a structured database; a means for generating responses using a large-scale language model based on the stored data; means for transmitting the generated response to the user's terminal; A system including:
2. The system of claim 1 , wherein the conversation log is organized by user.
3. The system of claim 1 , wherein the external information is updated periodically.
4. The system of claim 1 , wherein the large-scale language model obtains information from a structured database to generate responses.
5. The system of claim 1 , wherein the generated response is customized based on user preferences and past conversation history.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A