System
The system addresses the inefficiencies of manual keyword searches by analyzing natural language questions, acquiring location, and summarizing social media posts to provide quick and accurate answers, enhancing user convenience and information retrieval.
Patent Information
- Application Number
- JP2024116387
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Current real-time search systems require users to manually search for information using keywords, which is time-consuming and inconvenient, especially when urgent information is needed, and they struggle to provide quick and accurate answers based on social networking service posts without complex operations.
A system that receives natural language questions, analyzes them using NLP, acquires user location, searches social networking services for relevant posts, summarizes the situation, and generates answers, incorporating generative AI for trend and keyword extraction.
Enables users to obtain accurate information instantly by simply entering a question in natural language, eliminating the need for complex searches and providing real-time, easy-to-understand answers.
Smart Images

Figure 2026014913000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Today, many users use social networking services to gather real-time information. However, current real-time search systems require users to manually search for relevant information by entering appropriate keywords. This process is time-consuming and extremely inconvenient, especially when information is needed urgently. Furthermore, to understand the detailed situation in each region, users must check each post one by one, making it difficult to obtain information quickly. To solve this problem, a system is needed that automatically retrieves and analyzes appropriate posts and explains the situation based on the results, simply by allowing users to enter a question in natural language. [Means for solving the problem]
[0005] The present invention provides a system including means for receiving a natural language question from a user, means for analyzing the question using natural language processing technology to identify the content of the question, means for acquiring the user's location information, means for searching for and acquiring related posts from a social networking service based on the acquired location information and the question analysis results, means for analyzing the acquired posts and summarizing a specific situation, means for generating an answer to the user's question based on the summarized situation, and means for providing the generated answer to the user. This system enables users to instantly obtain accurate information in response to natural language questions without going through a complex search process. Furthermore, the system also includes a function for extracting key keywords and trends from the analysis results and predicting weather changes and specific events, thereby enabling the provision of more accurate information. Furthermore, the system includes a function for retrieving posts in real time from the social networking service's API based on the user's question and location information, allowing the system to quickly obtain the latest information and provide appropriate answers.
[0006] "User" refers to any individual or group of people who wish to use this system to obtain information.
[0007] "Natural language" is a language used by humans on a daily basis, which expresses specific questions and requests as text.
[0008] A "question" is text entered in natural language by a user to obtain the information they want to know.
[0009] "Natural language processing technology" is a technology for analyzing human language and understanding its meaning, and is used to identify the content of a question.
[0010] "Location Information" is data that indicates a user's current geographic location.
[0011] "Social Networking Service" means a platform for sharing user-generated content online and interacting with other users.
[0012] A "post" is a message or content made public by a user on a social networking service.
[0013] "Searching" is the act or process of locating specific information on a social networking service.
[0014] "Harvesting" is the act or process of collecting the desired posts from a social networking service.
[0015] "Analysis" is the act or process of understanding the content of retrieved posts and extracting specific information.
[0016] "Summarization" is the act or process of concisely summarizing the analyzed content and extracting the main information.
[0017] "Answer" refers to information provided in response to a user's question.
[0018] The term "system" refers to an entire device or program for realizing all the functions encompassed by the present invention. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[0041] Overall system overview
[0042] The system mainly consists of the following components:
[0043] 1. The device that receives the user's questions
[0044] 2. Server that analyzes questions and searches for and analyzes information
[0045] 3. Device that generates and provides answers to questions
[0046] Overview of program processing
[0047] 1. Receive user questions
[0048] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[0049] 2. Question Analysis
[0050] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[0051] 3. Obtaining location information
[0052] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[0053] 4. Search and retrieve social media posts
[0054] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[0055] 5. Analysis and Summarization of Posts
[0056] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[0057] 6. Answer Generation
[0058] Based on the analysis results, the server generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to here as time passes."
[0059] 7. Submitting and Viewing Your Answers
[0060] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[0061] Specific examples
[0062] Here are some concrete examples:
[0063] A user types in a question: "The sky is getting dark. Is it going to rain?"
[0064] 1. User: Enters a question into the terminal.
[0065] 2. Terminal: Sends a query to the server.
[0066] 3. Server: Parses the question and identifies it as a weather question.
[0067] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0068] 5. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0069] 6. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0070] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0071] 8. Server: Sends the generated answer to the device.
[0072] 9. Terminal: Display the answer to the user.
[0073] In this way, the system of the present invention can provide quick and accurate information simply by asking a question in natural language, allowing users to obtain the information they need in a short time without having to perform complex search operations.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] User: Type a question into the device in natural language (e.g., "The sky is getting dark. Will it rain?").
[0077] Step 2:
[0078] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[0079] Step 3:
[0080] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. The obtained location information is sent to a server with the user's consent.
[0081] Step 4:
[0082] Device: Sends the user's question and location information to the server.
[0083] Step 5:
[0084] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[0085] Step 6:
[0086] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[0087] Step 7:
[0088] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[0089] Step 8:
[0090] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[0091] Step 9:
[0092] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[0093] Step 10:
[0094] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[0095] Step 11:
[0096] Server: Generates answers to user questions based on the analysis results. Specifically, the answer may be something like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0097] Step 12:
[0098] Server: Sends the generated answer to the user's device.
[0099] Step 13:
[0100] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[0101] Step 14:
[0102] Device: Show the user the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0103] In this way, through a series of processes, users can get quick and accurate answers to their questions.
[0104] Example 1
[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0106] Conventional information retrieval systems have difficulty providing appropriate answers to questions in real time, even when users input questions in natural language. Furthermore, they have limitations in analyzing the user's current location and providing real-time information using posts from social networking services. This has resulted in problems such as users being unable to quickly and accurately obtain the information they are looking for, and requiring complex search operations.
[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0108] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for searching for and acquiring related posts from social networking services based on the acquired location information and the question analysis results, means for analyzing the acquired posts and summarizing a specific situation using a generative AI model, means for generating an answer to the user's question based on the summarized situation, and means for providing the generated answer to the user. This enables a user to quickly acquire highly accurate information in real time by simply inputting a question in natural language.
[0109] The "means for receiving a natural language question from a user" refers to a device or function for receiving a natural language question entered by a user and passing the question on to subsequent processing.
[0110] "Means for analyzing questions using natural language processing technology and identifying the content of the question" refers to technology, devices, or functions for analyzing received natural language questions and understanding their content. Specifically, natural language processing technology is used to identify the topic and intent of the question.
[0111] "Means for acquiring user location information" refers to devices or functions for identifying and acquiring the user's current location. Specifically, location information is acquired using GPS, Wi-Fi data, etc.
[0112] "Means for searching and retrieving related posts from social networking services based on acquired location information and question analysis results" refers to technology, devices, or functions that search for and retrieve related posts from SNS based on the user's location information and question content.
[0113] "Means for analyzing acquired posts and summarizing specific situations using a generative AI model" refers to technology, devices, or functions that analyze acquired social media posts and summarize their content using a generative AI model.
[0114] The "means for generating an answer to a user's question based on the summarized situation" refers to a technology, device, or function for generating an appropriate answer to a user's question based on the summary result.
[0115] The "means for providing a generated answer to a user" is a device or function for transmitting and displaying a generated answer to a user.
[0116] MODE FOR CARRYING OUT THE INVENTION
[0117] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[0118] Overall system overview
[0119] The system consists of the following main components:
[0120] 1. The device that receives the user's questions
[0121] 2. Server that analyzes questions and searches for and analyzes information
[0122] 3. Device that generates and provides answers to questions
[0123] Process Overview
[0124] 1. Receive questions from users
[0125] The user inputs a question in natural language into the device. For example, the user inputs, "The sky is getting dark. Is it going to rain?" The device receives this question and sends it to the server.
[0126] 2. Question Analysis
[0127] The server receives the user's question and analyzes it using natural language processing technology, specifically using libraries such as TensorFlow and spaCy, to determine that the question is about the weather.
[0128] 3. Obtaining location information
[0129] The device acquires the user's current location using technologies such as GPS and Wi-Fi. After obtaining the user's consent, the device sends the acquired location information (e.g., Chuo Ward, Tokyo) to a server.
[0130] 4. Search and retrieve social media posts
[0131] The server uses the SNS API (e.g., Twitter's API) to search for and retrieve related posts based on the location information and the results of analyzing the question. Keywords such as "Western Tokyo, Rain" are typically used as the search query, and the retrieved data is handled in JSON format.
[0132] 5. Analysis and Summarization of Posts
[0133] The server analyzes the social media posts and summarizes them using a generative AI model (e.g., OpenAI's GPT-3). During this process, trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain") are identified.
[0134] 6. Answer Generation
[0135] The server generates a textual answer to the user's question based on the analysis results. For example, it might create a response like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0136] 7. Submitting and Viewing Your Answers
[0137] The server sends the generated answer to the terminal, which then displays it to the user, allowing the user to quickly obtain detailed and specific information in real time.
[0138] Adding specific examples
[0139] Here are some concrete examples:
[0140] Specific situations
[0141] User asks: "The sky is getting dark, is it going to rain?"
[0142] 1. User: Enters the question "The sky is getting dark. Is it going to rain?" into the device.
[0143] 2. Terminal: Sends the entered question to the server.
[0144] 3. Server: Analyzes the question using TensorFlow and spaCy and identifies it as a weather-related question.
[0145] 4. Device: Uses GPS or Wi-Fi to obtain current location information (e.g., Chuo-ku, Tokyo) and sends it to the server.
[0146] 5. Server: Using Twitter API etc., search for social media posts using the keyword "Tokyo Western Rain."
[0147] 6. Server: Analyze the acquired social media post data, summarize it using GPT-3, and identify trending keywords.
[0148] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0149] 8. Server: Sends the generated answer to the terminal.
[0150] 9. Terminal: Displays the received answer on the screen.
[0151] Example prompt sentences to use
[0152] An input prompt for a generative AI model takes the form:
[0153] User Question: "The sky is getting dark, is it going to rain?"
[0154] Location information: "Chuo-ku, Tokyo"
[0155] Social media post: "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer as time goes by."
[0156] Based on these prompts, the AI model generates appropriate answers to the user's questions.
[0157] This system allows users to quickly and accurately obtain the information they need simply by entering their questions in natural language, providing a high level of convenience by eliminating the complex search work previously required.
[0158] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0159] Step 1:
[0160] The user inputs a question in natural language into the terminal. The terminal receives the question (e.g., "The sky is getting dark. Will it rain?"). The input of this process is the natural language question input by the user, and the output is to store this question in internal memory.
[0161] Step 2:
[0162] The terminal sends the received question to the server. Specifically, it generates an HTTP POST request and sends a payload containing the question to the server. The input of this process is the question entered by the user and the server address information, and the output is the sending of an HTTP request containing the question.
[0163] Step 3:
[0164] The server receives questions sent by users. It extracts the received questions and analyzes the content of the questions using natural language processing technology. Specifically, it uses libraries such as TensorFlow and spaCy to analyze the meaning of the questions. The input to this process is the user's question text, and the output is the analysis result (e.g., identifying that the question is about the weather).
[0165] Step 4:
[0166] Based on the analysis results, the server identifies the question as weather-related. This identification process involves text classification to understand the intent of the question. The input to this process is the analysis results from natural language processing, and the output is information that classifies the question into a category related to "weather."
[0167] Step 5:
[0168] The device uses GPS and Wi-Fi data to obtain the user's location. After obtaining the user's consent, the device obtains the current location (e.g., latitude, longitude, and area name) and sends it to the server. The input is data from the device's location information acquisition function, and the output is the location information (e.g., Chuo Ward, Tokyo).
[0169] Step 6:
[0170] The server uses the SNS API to search for and retrieve related posts based on the acquired location information and question analysis results. For example, using the Twitter API, a search is performed for keywords such as "Tokyo Western Rain." The input for this process is the location information and the results of the question analysis, and the output is the SNS post data (in JSON format) as the search results.
[0171] Step 7:
[0172] The server analyzes the social media posts and uses a generative AI model (e.g., OpenAI's GPT-3) to summarize a specific situation. It extracts trends and key keywords and generates a summary. The input is the social media post data, and the output is text describing the summarized situation (e.g., "There are many posts reporting heavy rain in western Tokyo").
[0173] Step 8:
[0174] The server generates an answer to the user's question based on the summarized situation. This answer generation process also uses a generative AI model. The input is the analyzed and summarized data, and the output is a text answer to the user's question (e.g., "It appears that heavy rain is approaching us as time passes.").
[0175] Step 9:
[0176] The server sends the generated answer to the terminal. The answer is sent as an HTTP response, and the terminal receives it. The input is the generated answer text, and the output is the HTTP response from the server to the terminal.
[0177] Step 10:
[0178] The terminal displays the received answer to the user. The answer text is displayed in the application interface and the user confirms it. The input is the answer from the server, and the output is the answer text displayed on the user's screen.
[0179] (Application example 1)
[0180] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0181] Real-time road and weather information is an extremely important element in the operation of autonomous vehicles. However, there is a lack of means to quickly and accurately obtain relevant information from a variety of sources and provide it in a format that is easy for users to understand. In particular, there is a need for technology that makes this information available in real time via a voice interface. This will improve the safety and convenience of autonomous vehicles.
[0182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0183] In this invention, the server includes means for converting voice questions from users into text using voice recognition technology, means for acquiring user location information, means for searching and acquiring related posts from information provision services based on the acquired location information and question analysis results, means for analyzing the acquired posts and summarizing specific situations, and means for providing the generated answers to the user by voice using voice synthesis technology. This enables users to quickly and accurately acquire real-time road and weather information and receive answers by voice simply by asking questions by voice.
[0184] "Natural language" refers to a language used by humans on a daily basis, and is not limited to any particular computer language or format.
[0185] "Natural language processing technology" refers to all technologies that enable computers to understand, generate, and analyze human language, including text analysis and speech recognition.
[0186] "Location information" refers to the current geographic location of a user or device, and is data obtained using GPS, Wi-Fi, etc.
[0187] An "information provision service" is an online service that provides specific information to users, such as social networking services and news sites.
[0188] A "post" is any content such as text, images, or videos uploaded by a user to an information service.
[0189] "Speech recognition technology" refers to technology that converts a user's speech into text or other data formats.
[0190] "Speech synthesis technology" is a technology that generates natural speech from text and is used to output speech via a computer.
[0191] "Real-time" refers to responding immediately to user operations and inputs, and processing and providing data without delay.
[0192] This invention relates to a navigation assistant system that provides real-time road and weather information in autonomous vehicles. The system includes a series of processes that receive voice questions from users, convert them into text, and analyze them. It then collects and analyzes related information and provides answers to users via voice.
[0193] The system primarily includes the following hardware and software components:
[0194] 1. A microphone in the vehicle to receive the user's voice query
[0195] 2. Speech recognition technology that converts speech to text (e.g., Google Cloud Speech-to-Text API)
[0196] 3. GPS module to obtain the user's current location
[0197] 4. Internet connection and API access to retrieve relevant posts from information services (e.g., Twitter API)
[0198] 5. A server to analyze posts and summarize specific situations using a generative AI model (e.g., OpenAI's GPT-4 model)
[0199] 6. Speech synthesis technology that generates natural-sounding speech from text (e.g., Google Cloud Text-to-Speech API)
[0200] 7. On-board computer systems for autonomous vehicles
[0201] Explanation of system processing
[0202] Receiving and analyzing voice questions
[0203] The server receives voice questions from users through the vehicle's microphone. This voice is converted into text using the Google Cloud Speech-to-Text API. The converted text question is then analyzed using natural language processing technology. For example, if the question is "What's the weather like now?", it is identified as a weather-related question.
[0204] Obtaining location information
[0205] The device acquires the user's current location using the vehicle's GPS module, and sends the acquired location information along with the analyzed question to the server.
[0206] Collection and analysis of relevant information
[0207] The server uses information providers like the Twitter API to retrieve real-time posts related to the location and question, which are then analyzed using OpenAI's GPT-4 model and summarized based on key keywords and trends.
[0208] Generate and provide answers
[0209] The server generates an answer to the user's question based on the analysis results. This answer is converted into audio using the Google Cloud Text-to-Speech API and provided to the user through the vehicle's speakers. For example, the answer might be, "It's starting to rain near your current location. There is traffic congestion on nearby roads."
[0210] Examples of concrete examples and prompts
[0211] For example, if a user asks verbally, "What's the weather like today?", the server will input the following prompt into the generative AI model:
[0212] plaintext
[0213] Summarize the tweet below:
[0214] It's currently raining heavily in Tokyo. Drive carefully!
[0215] There's a traffic jam, but it looks like it's going to rain soon.
[0216] summary:
[0217] This allows users to quickly and accurately obtain real-time weather and traffic information and receive voice responses simply by asking voice questions.
[0218] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0219] Step 1:
[0220] Receive user voice questions
[0221] Users ask questions in natural language into a microphone in the vehicle, and the device receives the audio and stores it as an audio file.
[0222] Input: User's voice question
[0223] Output: Audio file
[0224] What it does: The microphone captures your voice and converts it into a digital audio file.
[0225] Step 2:
[0226] Convert speech to text
[0227] The device uses the Google Cloud Speech-to-Text API to convert the audio file saved in step 1 into text, which is then sent to the server as the user's question.
[0228] Input: Audio file
[0229] Output: Text data
[0230] Specific operation: The device sends the audio file to the Google Cloud Speech-to-Text API and receives natural language text data.
[0231] Step 3:
[0232] Get the user's location
[0233] The device uses the vehicle's GPS module to obtain the user's current location, which is then sent to the server along with the text data.
[0234] Input: Current geographic location (GPS signal)
[0235] Output: Location data (latitude and longitude)
[0236] Specific operation: The device obtains the latitude and longitude information of the current location from the GPS module and saves it as location data.
[0237] Step 4:
[0238] Get related posts from information services
[0239] Based on the user's question and location information, the server searches and retrieves relevant posts in real time from information providers such as the Twitter API.
[0240] Input: Text data, location data
[0241] Output: Related post data (tweets)
[0242] What it does: The server uses the Twitter API to retrieve posts based on a given location and keyword, for example, using a search query like "Tokyo weather."
[0243] Step 5:
[0244] Analyze and summarize the posts
[0245] The server uses OpenAI's GPT-4 model to analyze the acquired post data, extract and summarize key keywords and trends.
[0246] Input: Related post data
[0247] Output: Summary data (text)
[0248] Specific operation: The server analyzes the text data of the tweet, inputs a prompt sentence into the GPT-4 model, and obtains a summary result.
[0249] Example prompt:
[0250] plaintext
[0251] Summarize the tweet below:
[0252] It's currently raining heavily in Tokyo. Drive carefully!
[0253] There's a traffic jam, but it looks like it's going to rain soon.
[0254] summary:
[0255] Step 6:
[0256] Generate answers and convert them into audio
[0257] The server generates answers to the user's questions based on the summary data and converts them into audio using the Google Cloud Text-to-Speech API.
[0258] Input: Summary data
[0259] Output: Audio data (answer)
[0260] What it does: The server analyzes the summary data and generates an appropriate answer to the question in text form, then converts that text into audio using the Google Cloud Text-to-Speech API.
[0261] Step 7:
[0262] Providing answers to users
[0263] The device then provides the generated voice data to the user through the vehicle's speakers, allowing the user to receive voice responses to their questions in real time.
[0264] Input: Voice data (answer)
[0265] Output: Audio output to the user
[0266] Specific operation: The device routes the generated voice data to the car's speakers and provides it to the user audibly.
[0267] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0268] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[0269] Overall system overview
[0270] The system mainly consists of the following components:
[0271] 1. The device that receives the user's questions
[0272] 2. Server that analyzes questions and searches for and analyzes information
[0273] 3. Device that generates and provides answers to questions
[0274] 4. Emotion engine that recognizes user emotions
[0275] Overview of program processing
[0276] 1. Receive user questions
[0277] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[0278] 2. Question Analysis
[0279] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[0280] 3. Obtaining location information
[0281] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[0282] 4. Emotional Recognition
[0283] At the same time, the server uses an emotion engine to analyze the user's emotion from the question (e.g., "anxiety," "excitement," etc.). The emotion engine analyzes the text data of the question as input and identifies the emotion.
[0284] 5. Search and retrieve social media posts
[0285] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[0286] 6. Analysis and Summarization of Posts
[0287] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[0288] 7. Adjust your responses based on emotion
[0289] The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling anxious, it will prioritize providing more detailed and reassuring information.
[0290] 8. Answer Generation
[0291] The server takes into account the analysis results and emotional adjustments to generate a textual answer to the user's question. Specifically, it creates a response such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[0292] 9. Submitting and Viewing Your Answers
[0293] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[0294] Specific examples
[0295] Here are some concrete examples:
[0296] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[0297] 1. User: Enters a question into the terminal.
[0298] 2. Terminal: Sends a query to the server.
[0299] 3. Server: Parses the question and identifies it as a weather question.
[0300] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0301] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[0302] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0303] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0304] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[0305] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[0306] 10. Server: Sends the generated answer to the terminal.
[0307] 11. Terminal: Display the answer to the user.
[0308] In this way, the system of the present invention can provide prompt and appropriate information in response to a user's question, and by taking emotions into consideration, can realize a more user-friendly response.
[0309] The processing flow will be explained below.
[0310] Step 1:
[0311] User: Enters a question in natural language into the device (e.g., "The sky is getting dark. Is it going to rain? I'm worried.").
[0312] Step 2:
[0313] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[0314] Step 3:
[0315] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. With the user's consent, the obtained location information is prepared for transmission to the server.
[0316] Step 4:
[0317] Device: Sends the user's question and location information to the server.
[0318] Step 5:
[0319] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[0320] Step 6:
[0321] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[0322] Step 7:
[0323] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[0324] Step 8:
[0325] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[0326] Step 9:
[0327] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[0328] Step 10:
[0329] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[0330] Step 11:
[0331] Server: Uses an emotion engine to recognize emotions from the user's question. For example, identify that the question contains the emotion "anxiety."
[0332] Step 12:
[0333] Server: Adjusts the priority of search and analysis results based on the user's emotions recognized by the emotion engine. If the user is feeling anxious, it prioritizes providing more detailed and reassuring information.
[0334] Step 13:
[0335] Server: Generates a written answer to the user's question, taking into account the analysis results and emotional adjustments. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[0336] Step 14:
[0337] Server: Sends the generated answer to the user's device.
[0338] Step 15:
[0339] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[0340] Step 16:
[0341] Device: Display the following response to the user: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[0342] Through this process, users can receive prompt and appropriate answers to their questions, and by taking into account the user's emotions, the system can provide a more user-friendly response.
[0343] Example 2
[0344] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0345] In today's world, users want to obtain appropriate information in real time by asking questions in natural language. However, conventional information search systems are limited to providing information based on keywords, and it is difficult to generate answers that take into account the user's situation and emotions. This makes them insufficient to alleviate users' anxieties and questions, and there is a need for more accurate information provision systems.
[0346] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0347] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for recognizing the user's emotion from the question text, means for searching and acquiring related posts from an online platform based on the acquired location information and the question analysis result, means for analyzing the acquired posts and summarizing a specific situation, means for generating an answer to the user's question based on the summarized situation, means for adjusting the content of the answer based on the emotion recognition result, and means for providing the generated answer to the user. This makes it possible to provide more appropriate and specific information in real time in response to a natural language question entered by a user, taking into account the user's location information and emotion.
[0348] A "user" is a person using a terminal who enters a question in natural language to obtain information.
[0349] "Natural language" refers to language used in everyday life, sentences and words that do not require special interpretation by a computer program.
[0350] "Terminal" refers to a computing device, such as a smartphone, tablet, or PC, through which a user enters a question or receives a response.
[0351] A "server" is a computer device that receives and processes data sent from a terminal and provides appropriate information.
[0352] "Natural language processing technology" refers to all technologies that enable computers to understand, analyze, and generate natural language.
[0353] "Location Information" means data that indicates a user's geographic location and may include GPS data and Wi-Fi information.
[0354] An "emotion engine" refers to software or algorithms that analyze and identify user emotions based on text data.
[0355] "Online platform" refers to an online service that allows a large number of users to generate and share information and content, including social networking services.
[0356] A "generative AI model" is an artificial intelligence model that generates new text based on training data, such as GPT-4.
[0357] A "prompt" refers to text data that is input to a generative AI model to produce a specific output.
[0358] An "answer" is a sentence generated in response to a user's question based on the analysis results and acquired information.
[0359] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[0360] Hardware and Software
[0361] This system mainly uses the following hardware and software:
[0362] Device: A computing device used by a user, such as a smartphone, tablet, or PC.
[0363] Server: A computing device for receiving, processing, and providing data.
[0364] Natural language processing technology: Google Cloud Natural Language API, etc.
[0365] Sentiment engines: such as IBM Watson Tone Analyzer and Microsoft Azure Text Analytics.
[0366] Location information acquisition technology: Google Maps API, built-in GPS module, etc.
[0367] Social networking interfaces: Twitter API, etc.
[0368] Generative AI models: such as OpenAI's GPT-4.
[0369] Data processing and calculation
[0370] The main processing of the system involves the following data processing and data calculation.
[0371] 1. Natural language processing: The device sends the user's question to the server, which then uses natural language processing to analyze the question, for example, identifying that it is a question about the weather.
[0372] 2. Location information acquisition: The device acquires the user's current location information and sends it to the server. This location information is acquired using GPS or Wi-Fi.
[0373] 3. Emotion Recognition: The server uses an emotion engine to analyze and identify emotions from the user's question, thereby determining the emotion (e.g., "anxiety" or "excitement") the user is feeling when asking the question.
[0374] 4. Search and retrieve SNS posts: Based on the acquired location information and the question analysis results, the server uses the SNS interface to search and retrieve related posts. For example, a search can be performed using keywords such as "Tokyo West Rain."
[0375] 5. Post analysis: The server analyzes the social media posts and summarizes them using a generative AI model, identifying trends and key keywords from the content of the posts (e.g., "Western Tokyo" and "Heavy Rain").
[0376] 6. Emotion-based response adjustment: The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," the server will prioritize providing more detailed and reassuring information.
[0377] 7. Answer Generation: The server generates a textual answer to the user's question, taking into account the analysis results and sentiment-based adjustments. For example, it creates an answer such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time."
[0378] 8. Providing an answer: The server sends the generated answer to the terminal, and the terminal displays the answer to the user, allowing the user to quickly obtain specific and real-time information about the question.
[0379] Specific examples
[0380] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[0381] 1. User: Enters a question into the terminal.
[0382] 2. Terminal: Sends a query to the server.
[0383] 3. Server: Parses the question and identifies it as a weather question.
[0384] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0385] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[0386] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0387] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0388] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[0389] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[0390] 10. Server: Sends the generated answer to the terminal.
[0391] 11. Terminal: Display the answer to the user.
[0392] This allows the system of the present invention to provide prompt and appropriate information in response to user questions, and by taking emotions into consideration, it is possible to provide a more user-friendly response.
[0393] Prompt Sentence Examples
[0394] "Analyze questions entered in natural language, combine location information and emotion recognition to search for relevant information, and generate real-time answers. The user's question is 'The sky is getting dark. Will it rain? I'm worried.' The location information is 'Chuo Ward, Tokyo.' The emotion recognition result is 'Anxiety.'"
[0395] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0396] Step 1:
[0397] User: Enters a question into the terminal in natural language (e.g., "The sky is getting dark. Will it rain?"). The input data is a textual question.
[0398] Terminal: Receives the entered question, stores it as text data, and prepares it to be sent to the server. Specifically, when the user enters a question and presses the "Send" button, the terminal sends the question data to the server. The input is the user's text data, and the output is ready to be sent to the server.
[0399] Step 2:
[0400] Server: Analyzes the question data received from the user using natural language processing (NLP) technology. For example, it identifies that the question is about the weather. Specifically, it uses an NLP library (e.g., Google Cloud Natural Language API) to divide the question into tokens and analyzes the sentence structure, such as subject, predicate, and object. The input is the question text data, and the output is the analysis result (identification of the question content).
[0401] Step 3:
[0402] Device: Obtains the user's current location information. This is obtained using GPS or Wi-Fi. The obtained location information is sent to the server. Specifically, the device's location information acquisition function is called, and the current location is obtained as GPS coordinates or Wi-Fi information. The input is a request to obtain location information, and the output is the obtained location information (e.g., Chuo-ku, Tokyo).
[0403] Step 4:
[0404] Server: Analyzes text data (questions) and recognizes the user's emotions using an emotion engine. Specifically, it identifies emotions such as "anxiety" and "excitement." Specific operations involve inputting text data into an emotion engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Text Analytics) and obtaining an emotion score. The input is the question text data, and the output is the emotion recognition results.
[0405] Step 5:
[0406] Server: Based on the acquired location information and question analysis results, the server uses the online platform's API to search and retrieve related posts. For example, a search is performed using keywords such as "Tokyo West Rain." Specific operations include creating an API request and retrieving related posts. The input is location information and question analysis results, and the output is the retrieved post data.
[0407] Step 6:
[0408] Server: Analyzes the acquired social media post data using a generative AI model (e.g., OpenAI's GPT-4) and summarizes the situation. Identifies trends and key keywords from the content of the post. Specifically, it crawls the acquired post data, analyzes the text content, calculates the frequency distribution of keywords, and has the generative AI model summarize it. The input is the acquired post data, and the output is summarized situation information.
[0409] Step 7:
[0410] Server: Adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," it prioritizes providing more detailed and reassuring information. Specifically, it selects the most appropriate information from the analyzed information set based on the user's emotion score. The input is the emotion recognition result and summarized situation information, and the output is the adjusted answer content.
[0411] Step 8:
[0412] Server: Based on the above information, it generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time." Specific operations include inputting the necessary information as a prompt into the generative AI model and requesting it to generate a text. The generated text is reviewed and any necessary corrections are made. The input is the adjusted answer, and the output is the generated answer text.
[0413] Step 9:
[0414] Server: Sends the generated answer to the device.
[0415] Terminal: Displays the answer to the user. Specifically, the answer data received from the server is displayed on the user interface. The display format can be a pop-up, notification, chat format, etc. The input is the generated answer text, and the output is the answer displayed to the user.
[0416] In this way, the system of the present invention can quickly provide specific, real-time information in response to a user's question.
[0417] (Application example 2)
[0418] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0419] Conventional systems had the problem of making it difficult for users to obtain fast and accurate information in physical stores. In particular, when users had questions about products, there was a lack of a way to quickly and effectively provide detailed information, reviews, and inventory information related to the question in real time. Furthermore, responses did not take into consideration the user's feelings, which could lead to a decline in the quality of the user experience. This made it difficult to improve customer satisfaction.
[0420] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving a question in natural language from a user; means for analyzing the question using natural language processing technology and identifying the question content; emotion recognition means for analyzing the user's emotions; means for acquiring the user's location information; means for searching and acquiring related posts from social networking services based on the acquired location information and the question analysis results; means for analyzing the acquired posts and summarizing a specific situation; means for generating an answer to the user's question based on the summarized situation and the emotion analysis results; and means for providing the generated answer to the user. This allows users to easily obtain quick and accurate information related to products in physical stores and also enables responses that take user emotions into consideration. This improves customer satisfaction.
[0421] "Natural language questions" refer to questions asked by users using everyday words and expressions.
[0422] "Natural language processing technology" is a technology that allows computers to analyze and understand human language.
[0423] The "means for identifying the question content" is a mechanism for analyzing the received question and clarifying what the question is asking.
[0424] "Emotion recognition means" refers to a device or program that analyzes and identifies emotions and moods from text entered by a user.
[0425] "Means of obtaining location information" refers to a mechanism for obtaining the user's current location using technologies such as GPS and Wi-Fi.
[0426] A "social networking service" is an online communication platform that allows users to share information.
[0427] "Means for searching and retrieving related posts" refers to a system for searching and collecting posts that match specific keywords or conditions based on the acquired location information and analysis results.
[0428] "Means for analyzing posts and summarizing specific situations" refers to a system for analyzing collected post content, extracting important information from it, and summarizing it concisely.
[0429] The "means for generating an answer" is a mechanism for creating an appropriate response to a user's question in the form of text based on the analysis results and emotion recognition results.
[0430] The "means for providing an answer" is a device or program for presenting the generated answer to the user in a format that is easy to view.
[0431] The system that realizes this application example is a "smart shopping assistant" system that allows users to input product-related questions in natural language while shopping in a physical store and provides quick and appropriate answers.
[0432] System Program
[0433] This system is implemented using the following hardware and software:
[0434] Hardware
[0435] 1. User terminal: A device such as a smartphone, smart glasses, or head-mounted display.
[0436] 2. Server: A computer server with powerful computing power.
[0437] software
[0438] 1. Natural language processing libraries: Use NLTK (Natural Language Toolkit) and spaCy to analyze natural language questions from users.
[0439] 2. Emotion Recognition Engine: Analyzes user emotions using TextBlob and VADER (Valence Aware Dictionary and sEntiment Reasoner).
[0440] 3. SNS API: Use the API of social networking services (e.g. Twitter API) to search and retrieve related posts.
[0441] 4. Generative AI model: An artificial intelligence model that generates appropriate answers based on the information obtained.
[0442] What the program does
[0443] 1. Receiving Questions
[0444] A user types a question in natural language into a device such as a smartphone or smart glasses, for example, "Is the new model in stock?" This question is received by the device and sent to the server.
[0445] 2. Natural Language Processing
[0446] When the server receives a question, it uses a natural language processing library (e.g., spaCy) to analyze the question and extract information about specific categories and products.
[0447] 3. Emotion recognition
[0448] The server then uses an emotion recognition engine (e.g., TextBlob) to analyze the emotion from the user's question. For example, the emotion "anxiety" is identified from the question "I'm worried about stock availability."
[0449] 4. Obtaining location information
[0450] The user's device uses GPS and Wi-Fi to obtain its current location and sends it to the server, which can then obtain information about the store the user is in or nearby.
[0451] 5. Searching for Information
[0452] The server uses the acquired location information and the question analysis results to search and retrieve related posts using the SNS API, searching the social networking service for keywords such as "in stock" and "new model."
[0453] 6. Analysis and Summarization of Posts
[0454] The server analyzes the posts and uses a generative AI model to summarize important information, such as "latest model in stock" or "highly rated."
[0455] 7. Answer Generation
[0456] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, to a user who is feeling anxious, the server creates a response such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good."
[0457] 8. Providing answers
[0458] The generated answer is sent from the server to the device and displayed to the user, allowing the user to quickly obtain detailed information about the product in the physical store.
[0459] Specific examples
[0460] A user uses a smartphone in a physical store to ask, "Is the new model in stock?" At this time, the system analyzes the user's emotion as "anxiety" and collects the latest stock status and product review information from social media posts. Ultimately, it generates an answer such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good," and displays it on the user's device.
[0461] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0462] Step 1:
[0463] A user enters a question in natural language using a smartphone or smart glasses in a physical store. The entered question (e.g., "Is the new model in stock?") is received by the user device and sent to the server. In this step, the input is the user's question, and the output is the data that sends the question to the server.
[0464] Step 2:
[0465] The server analyzes the received question using a natural language processing library (e.g., spaCy). This analysis identifies the intent of the question and the target product category. Specifically, the server tokenizes the question, tags it with parts of speech, and performs semantic analysis. The input is the user's natural language question, and the output is the identified question (e.g., "new model" or "inventory").
[0466] Step 3:
[0467] At the same time, the server uses an emotion recognition engine (e.g., TextBlob) to analyze the user's emotion from the question. The emotion recognition engine receives text data as input, analyzes the emotion polarity, and identifies the emotion type (e.g., "anxiety," "relief," or "interest"). The input is the user's question, and the output is the analyzed emotion data (e.g., "anxiety").
[0468] Step 4:
[0469] The user device acquires its current location information using GPS or Wi-Fi and sends that location information to the server. The input is the location data of the user device, and the output is the location information data sent to the server (e.g., "Shinjuku-ku, Tokyo").
[0470] Step 5:
[0471] The server uses the API of the social networking service (SNS) to search and retrieve related posts based on the acquired location information and question analysis results. This search is performed by combining specific keywords and location information. The input is the question analysis results and location information, and the output is related post data (e.g., "SNS post: latest model in stock, many positive reviews").
[0472] Step 6:
[0473] The server analyzes the social media posts and uses a generative AI model to summarize important information. This summarization includes extracting trends and key keywords from the posts. The input is the social media post data, and the output is summarized information (e.g., "Likely Rated, In Stock").
[0474] Step 7:
[0475] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, for a user who is feeling anxious, the server might create a response such as, "This product is the latest model. It is currently in stock, so don't worry!" The input is summarized information and emotion data, and the output is the generated response text data.
[0476] Step 8:
[0477] The server sends the generated answer to the user's terminal, which then displays the answer. The user can view the answer and quickly obtain detailed information about the product. The input is the generated answer data, and the output is the display of the answer on the user's terminal.
[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0481] [Second embodiment]
[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0494] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[0495] Overall system overview
[0496] The system mainly consists of the following components:
[0497] 1. The device that receives the user's questions
[0498] 2. Server that analyzes questions and searches for and analyzes information
[0499] 3. Device that generates and provides answers to questions
[0500] Overview of program processing
[0501] 1. Receive user questions
[0502] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[0503] 2. Question Analysis
[0504] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[0505] 3. Obtaining location information
[0506] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[0507] 4. Search and retrieve social media posts
[0508] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[0509] 5. Analysis and Summarization of Posts
[0510] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[0511] 6. Answer Generation
[0512] Based on the analysis results, the server generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to here as time passes."
[0513] 7. Submitting and Viewing Your Answers
[0514] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[0515] Specific examples
[0516] Here are some concrete examples:
[0517] A user types in a question: "The sky is getting dark. Is it going to rain?"
[0518] 1. User: Enters a question into the terminal.
[0519] 2. Terminal: Sends a query to the server.
[0520] 3. Server: Parses the question and identifies it as a weather question.
[0521] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0522] 5. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0523] 6. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0524] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0525] 8. Server: Sends the generated answer to the device.
[0526] 9. Terminal: Display the answer to the user.
[0527] In this way, the system of the present invention can provide quick and accurate information simply by asking a question in natural language, allowing users to obtain the information they need in a short time without having to perform complex search operations.
[0528] The processing flow will be explained below.
[0529] Step 1:
[0530] User: Type a question into the device in natural language (e.g., "The sky is getting dark. Will it rain?").
[0531] Step 2:
[0532] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[0533] Step 3:
[0534] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. The obtained location information is sent to a server with the user's consent.
[0535] Step 4:
[0536] Device: Sends the user's question and location information to the server.
[0537] Step 5:
[0538] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[0539] Step 6:
[0540] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[0541] Step 7:
[0542] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[0543] Step 8:
[0544] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[0545] Step 9:
[0546] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[0547] Step 10:
[0548] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[0549] Step 11:
[0550] Server: Generates answers to user questions based on the analysis results. Specifically, the answer may be something like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0551] Step 12:
[0552] Server: Sends the generated answer to the user's device.
[0553] Step 13:
[0554] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[0555] Step 14:
[0556] Device: Show the user the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0557] In this way, through a series of processes, users can get quick and accurate answers to their questions.
[0558] Example 1
[0559] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0560] Conventional information retrieval systems have difficulty providing appropriate answers to questions in real time, even when users input questions in natural language. Furthermore, they have limitations in analyzing the user's current location and providing real-time information using posts from social networking services. This has resulted in problems such as users being unable to quickly and accurately obtain the information they are looking for, and requiring complex search operations.
[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0562] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for searching for and acquiring related posts from social networking services based on the acquired location information and the question analysis results, means for analyzing the acquired posts and summarizing a specific situation using a generative AI model, means for generating an answer to the user's question based on the summarized situation, and means for providing the generated answer to the user. This enables a user to quickly acquire highly accurate information in real time by simply inputting a question in natural language.
[0563] The "means for receiving a natural language question from a user" refers to a device or function for receiving a natural language question entered by a user and passing the question on to subsequent processing.
[0564] "Means for analyzing questions using natural language processing technology and identifying the content of the question" refers to technology, devices, or functions for analyzing received natural language questions and understanding their content. Specifically, natural language processing technology is used to identify the topic and intent of the question.
[0565] "Means for acquiring user location information" refers to devices or functions for identifying and acquiring the user's current location. Specifically, location information is acquired using GPS, Wi-Fi data, etc.
[0566] "Means for searching and retrieving related posts from social networking services based on acquired location information and question analysis results" refers to technology, devices, or functions that search for and retrieve related posts from SNS based on the user's location information and question content.
[0567] "Means for analyzing acquired posts and summarizing specific situations using a generative AI model" refers to technology, devices, or functions that analyze acquired social media posts and summarize their content using a generative AI model.
[0568] The "means for generating an answer to a user's question based on the summarized situation" refers to a technology, device, or function for generating an appropriate answer to a user's question based on the summary result.
[0569] The "means for providing a generated answer to a user" is a device or function for transmitting and displaying a generated answer to a user.
[0570] MODE FOR CARRYING OUT THE INVENTION
[0571] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[0572] Overall system overview
[0573] The system consists of the following main components:
[0574] 1. The device that receives the user's questions
[0575] 2. Server that analyzes questions and searches for and analyzes information
[0576] 3. Device that generates and provides answers to questions
[0577] Process Overview
[0578] 1. Receive questions from users
[0579] The user inputs a question in natural language into the device. For example, the user inputs, "The sky is getting dark. Is it going to rain?" The device receives this question and sends it to the server.
[0580] 2. Question Analysis
[0581] The server receives the user's question and analyzes it using natural language processing technology, specifically using libraries such as TensorFlow and spaCy, to determine that the question is about the weather.
[0582] 3. Obtaining location information
[0583] The device acquires the user's current location using technologies such as GPS and Wi-Fi. After obtaining the user's consent, the device sends the acquired location information (e.g., Chuo Ward, Tokyo) to a server.
[0584] 4. Search and retrieve social media posts
[0585] The server uses the SNS API (e.g., Twitter's API) to search for and retrieve related posts based on the location information and the results of analyzing the question. Keywords such as "Western Tokyo, Rain" are typically used as the search query, and the retrieved data is handled in JSON format.
[0586] 5. Analysis and Summarization of Posts
[0587] The server analyzes the social media posts and summarizes them using a generative AI model (e.g., OpenAI's GPT-3). During this process, trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain") are identified.
[0588] 6. Answer Generation
[0589] The server generates a textual answer to the user's question based on the analysis results. For example, it might create a response like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0590] 7. Submitting and Viewing Your Answers
[0591] The server sends the generated answer to the terminal, which then displays it to the user, allowing the user to quickly obtain detailed and specific information in real time.
[0592] Adding specific examples
[0593] Here are some concrete examples:
[0594] Specific situations
[0595] User asks: "The sky is getting dark, is it going to rain?"
[0596] 1. User: Enters the question "The sky is getting dark. Is it going to rain?" into the device.
[0597] 2. Terminal: Sends the entered question to the server.
[0598] 3. Server: Analyzes the question using TensorFlow and spaCy and identifies it as a weather-related question.
[0599] 4. Device: Uses GPS or Wi-Fi to obtain current location information (e.g., Chuo-ku, Tokyo) and sends it to the server.
[0600] 5. Server: Using Twitter API etc., search for social media posts using the keyword "Tokyo Western Rain."
[0601] 6. Server: Analyze the acquired social media post data, summarize it using GPT-3, and identify trending keywords.
[0602] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0603] 8. Server: Sends the generated answer to the terminal.
[0604] 9. Terminal: Displays the received answer on the screen.
[0605] Example prompt sentences to use
[0606] An input prompt for a generative AI model takes the form:
[0607] User Question: "The sky is getting dark, is it going to rain?"
[0608] Location information: "Chuo-ku, Tokyo"
[0609] Social media post: "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer as time goes by."
[0610] Based on these prompts, the AI model generates appropriate answers to the user's questions.
[0611] This system allows users to quickly and accurately obtain the information they need simply by entering their questions in natural language, providing a high level of convenience by eliminating the complex search work previously required.
[0612] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0613] Step 1:
[0614] The user inputs a question in natural language into the terminal. The terminal receives the question (e.g., "The sky is getting dark. Will it rain?"). The input of this process is the natural language question input by the user, and the output is to store this question in internal memory.
[0615] Step 2:
[0616] The terminal sends the received question to the server. Specifically, it generates an HTTP POST request and sends a payload containing the question to the server. The input of this process is the question entered by the user and the server address information, and the output is the sending of an HTTP request containing the question.
[0617] Step 3:
[0618] The server receives questions sent by users. It extracts the received questions and analyzes the content of the questions using natural language processing technology. Specifically, it uses libraries such as TensorFlow and spaCy to analyze the meaning of the questions. The input to this process is the user's question text, and the output is the analysis result (e.g., identifying that the question is about the weather).
[0619] Step 4:
[0620] Based on the analysis results, the server identifies the question as weather-related. This identification process involves text classification to understand the intent of the question. The input to this process is the analysis results from natural language processing, and the output is information that classifies the question into a category related to "weather."
[0621] Step 5:
[0622] The device uses GPS and Wi-Fi data to obtain the user's location. After obtaining the user's consent, the device obtains the current location (e.g., latitude, longitude, and area name) and sends it to the server. The input is data from the device's location information acquisition function, and the output is the location information (e.g., Chuo Ward, Tokyo).
[0623] Step 6:
[0624] The server uses the SNS API to search for and retrieve related posts based on the acquired location information and question analysis results. For example, using the Twitter API, a search is performed for keywords such as "Tokyo Western Rain." The input for this process is the location information and the results of the question analysis, and the output is the SNS post data (in JSON format) as the search results.
[0625] Step 7:
[0626] The server analyzes the social media posts and uses a generative AI model (e.g., OpenAI's GPT-3) to summarize a specific situation. It extracts trends and key keywords and generates a summary. The input is the social media post data, and the output is text describing the summarized situation (e.g., "There are many posts reporting heavy rain in western Tokyo").
[0627] Step 8:
[0628] The server generates an answer to the user's question based on the summarized situation. This answer generation process also uses a generative AI model. The input is the analyzed and summarized data, and the output is a text answer to the user's question (e.g., "It appears that heavy rain is approaching us as time passes.").
[0629] Step 9:
[0630] The server sends the generated answer to the terminal. The answer is sent as an HTTP response, and the terminal receives it. The input is the generated answer text, and the output is the HTTP response from the server to the terminal.
[0631] Step 10:
[0632] The terminal displays the received answer to the user. The answer text is displayed in the application interface and the user confirms it. The input is the answer from the server, and the output is the answer text displayed on the user's screen.
[0633] (Application example 1)
[0634] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0635] Real-time road and weather information is an extremely important element in the operation of autonomous vehicles. However, there is a lack of means to quickly and accurately obtain relevant information from a variety of sources and provide it in a format that is easy for users to understand. In particular, there is a need for technology that makes this information available in real time via a voice interface. This will improve the safety and convenience of autonomous vehicles.
[0636] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0637] In this invention, the server includes means for converting voice questions from users into text using voice recognition technology, means for acquiring user location information, means for searching and acquiring related posts from information provision services based on the acquired location information and question analysis results, means for analyzing the acquired posts and summarizing specific situations, and means for providing the generated answers to the user by voice using voice synthesis technology. This enables users to quickly and accurately acquire real-time road and weather information and receive answers by voice simply by asking questions by voice.
[0638] "Natural language" refers to a language used by humans on a daily basis, and is not limited to any particular computer language or format.
[0639] "Natural language processing technology" refers to all technologies that enable computers to understand, generate, and analyze human language, including text analysis and speech recognition.
[0640] "Location information" refers to the current geographic location of a user or device, and is data obtained using GPS, Wi-Fi, etc.
[0641] An "information provision service" is an online service that provides specific information to users, such as social networking services and news sites.
[0642] A "post" is any content such as text, images, or videos uploaded by a user to an information service.
[0643] "Speech recognition technology" refers to technology that converts a user's speech into text or other data formats.
[0644] "Speech synthesis technology" is a technology that generates natural speech from text and is used to output speech via a computer.
[0645] "Real-time" refers to responding immediately to user operations and inputs, and processing and providing data without delay.
[0646] This invention relates to a navigation assistant system that provides real-time road and weather information in autonomous vehicles. The system includes a series of processes that receive voice questions from users, convert them into text, and analyze them. It then collects and analyzes related information and provides answers to users via voice.
[0647] The system primarily includes the following hardware and software components:
[0648] 1. A microphone in the vehicle to receive the user's voice query
[0649] 2. Speech recognition technology that converts speech to text (e.g., Google Cloud Speech-to-Text API)
[0650] 3. GPS module to obtain the user's current location
[0651] 4. Internet connection and API access to retrieve relevant posts from information services (e.g., Twitter API)
[0652] 5. A server to analyze posts and summarize specific situations using a generative AI model (e.g., OpenAI's GPT-4 model)
[0653] 6. Speech synthesis technology that generates natural-sounding speech from text (e.g., Google Cloud Text-to-Speech API)
[0654] 7. On-board computer systems for autonomous vehicles
[0655] Explanation of system processing
[0656] Receiving and analyzing voice questions
[0657] The server receives voice questions from users through the vehicle's microphone. This voice is converted into text using the Google Cloud Speech-to-Text API. The converted text question is then analyzed using natural language processing technology. For example, if the question is "What's the weather like now?", it is identified as a weather-related question.
[0658] Obtaining location information
[0659] The device acquires the user's current location using the vehicle's GPS module, and sends the acquired location information along with the analyzed question to the server.
[0660] Collection and analysis of relevant information
[0661] The server uses information providers like the Twitter API to retrieve real-time posts related to the location and question, which are then analyzed using OpenAI's GPT-4 model and summarized based on key keywords and trends.
[0662] Generate and provide answers
[0663] The server generates an answer to the user's question based on the analysis results. This answer is converted into audio using the Google Cloud Text-to-Speech API and provided to the user through the vehicle's speakers. For example, the answer might be, "It's starting to rain near your current location. There is traffic congestion on nearby roads."
[0664] Examples of concrete examples and prompts
[0665] For example, if a user asks verbally, "What's the weather like today?", the server will input the following prompt into the generative AI model:
[0666] plaintext
[0667] Summarize the tweet below:
[0668] It's currently raining heavily in Tokyo. Drive carefully!
[0669] There's a traffic jam, but it looks like it's going to rain soon.
[0670] summary:
[0671] This allows users to quickly and accurately obtain real-time weather and traffic information and receive voice responses simply by asking voice questions.
[0672] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0673] Step 1:
[0674] Receive user voice questions
[0675] Users ask questions in natural language into a microphone in the vehicle, and the device receives the audio and stores it as an audio file.
[0676] Input: User's voice question
[0677] Output: Audio file
[0678] What it does: The microphone captures your voice and converts it into a digital audio file.
[0679] Step 2:
[0680] Convert speech to text
[0681] The device uses the Google Cloud Speech-to-Text API to convert the audio file saved in step 1 into text, which is then sent to the server as the user's question.
[0682] Input: Audio file
[0683] Output: Text data
[0684] Specific operation: The device sends the audio file to the Google Cloud Speech-to-Text API and receives natural language text data.
[0685] Step 3:
[0686] Get the user's location
[0687] The device uses the vehicle's GPS module to obtain the user's current location, which is then sent to the server along with the text data.
[0688] Input: Current geographic location (GPS signal)
[0689] Output: Location data (latitude and longitude)
[0690] Specific operation: The device obtains the latitude and longitude information of the current location from the GPS module and saves it as location data.
[0691] Step 4:
[0692] Get related posts from information services
[0693] Based on the user's question and location information, the server searches and retrieves relevant posts in real time from information providers such as the Twitter API.
[0694] Input: Text data, location data
[0695] Output: Related post data (tweets)
[0696] What it does: The server uses the Twitter API to retrieve posts based on a given location and keyword, for example, using a search query like "Tokyo weather."
[0697] Step 5:
[0698] Analyze and summarize the posts
[0699] The server uses OpenAI's GPT-4 model to analyze the acquired post data, extract and summarize key keywords and trends.
[0700] Input: Related post data
[0701] Output: Summary data (text)
[0702] Specific operation: The server analyzes the text data of the tweet, inputs a prompt sentence into the GPT-4 model, and obtains a summary result.
[0703] Example prompt:
[0704] plaintext
[0705] Summarize the tweet below:
[0706] It's currently raining heavily in Tokyo. Drive carefully!
[0707] There's a traffic jam, but it looks like it's going to rain soon.
[0708] summary:
[0709] Step 6:
[0710] Generate answers and convert them into audio
[0711] The server generates answers to the user's questions based on the summary data and converts them into audio using the Google Cloud Text-to-Speech API.
[0712] Input: Summary data
[0713] Output: Audio data (answer)
[0714] What it does: The server analyzes the summary data and generates an appropriate answer to the question in text form, then converts that text into audio using the Google Cloud Text-to-Speech API.
[0715] Step 7:
[0716] Providing answers to users
[0717] The device then provides the generated voice data to the user through the vehicle's speakers, allowing the user to receive voice responses to their questions in real time.
[0718] Input: Voice data (answer)
[0719] Output: Audio output to the user
[0720] Specific operation: The device routes the generated voice data to the car's speakers and provides it to the user audibly.
[0721] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0722] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[0723] Overall system overview
[0724] The system mainly consists of the following components:
[0725] 1. The device that receives the user's questions
[0726] 2. Server that analyzes questions and searches for and analyzes information
[0727] 3. Device that generates and provides answers to questions
[0728] 4. Emotion engine that recognizes user emotions
[0729] Overview of program processing
[0730] 1. Receive user questions
[0731] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[0732] 2. Question Analysis
[0733] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[0734] 3. Obtaining location information
[0735] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[0736] 4. Emotional Recognition
[0737] At the same time, the server uses an emotion engine to analyze the user's emotion from the question (e.g., "anxiety," "excitement," etc.). The emotion engine analyzes the text data of the question as input and identifies the emotion.
[0738] 5. Search and retrieve social media posts
[0739] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[0740] 6. Analysis and Summarization of Posts
[0741] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[0742] 7. Adjust your responses based on emotion
[0743] The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling anxious, it will prioritize providing more detailed and reassuring information.
[0744] 8. Answer Generation
[0745] The server takes into account the analysis results and emotional adjustments to generate a textual answer to the user's question. Specifically, it creates a response such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[0746] 9. Submitting and Viewing Your Answers
[0747] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[0748] Specific examples
[0749] Here are some concrete examples:
[0750] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[0751] 1. User: Enters a question into the terminal.
[0752] 2. Terminal: Sends a query to the server.
[0753] 3. Server: Parses the question and identifies it as a weather question.
[0754] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0755] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[0756] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0757] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0758] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[0759] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[0760] 10. Server: Sends the generated answer to the terminal.
[0761] 11. Terminal: Display the answer to the user.
[0762] In this way, the system of the present invention can provide prompt and appropriate information in response to a user's question, and by taking emotions into consideration, can realize a more user-friendly response.
[0763] The processing flow will be explained below.
[0764] Step 1:
[0765] User: Enters a question in natural language into the device (e.g., "The sky is getting dark. Is it going to rain? I'm worried.").
[0766] Step 2:
[0767] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[0768] Step 3:
[0769] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. With the user's consent, the obtained location information is prepared for transmission to the server.
[0770] Step 4:
[0771] Device: Sends the user's question and location information to the server.
[0772] Step 5:
[0773] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[0774] Step 6:
[0775] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[0776] Step 7:
[0777] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[0778] Step 8:
[0779] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[0780] Step 9:
[0781] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[0782] Step 10:
[0783] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[0784] Step 11:
[0785] Server: Uses an emotion engine to recognize emotions from the user's question. For example, identify that the question contains the emotion "anxiety."
[0786] Step 12:
[0787] Server: Adjusts the priority of search and analysis results based on the user's emotions recognized by the emotion engine. If the user is feeling anxious, it prioritizes providing more detailed and reassuring information.
[0788] Step 13:
[0789] Server: Generates a written answer to the user's question, taking into account the analysis results and emotional adjustments. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[0790] Step 14:
[0791] Server: Sends the generated answer to the user's device.
[0792] Step 15:
[0793] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[0794] Step 16:
[0795] Device: Display the following response to the user: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[0796] Through this process, users can receive prompt and appropriate answers to their questions, and by taking into account the user's emotions, the system can provide a more user-friendly response.
[0797] Example 2
[0798] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0799] In today's world, users want to obtain appropriate information in real time by asking questions in natural language. However, conventional information search systems are limited to providing information based on keywords, and it is difficult to generate answers that take into account the user's situation and emotions. This makes them insufficient to alleviate users' anxieties and questions, and there is a need for more accurate information provision systems.
[0800] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0801] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for recognizing the user's emotion from the question text, means for searching and acquiring related posts from an online platform based on the acquired location information and the question analysis result, means for analyzing the acquired posts and summarizing a specific situation, means for generating an answer to the user's question based on the summarized situation, means for adjusting the content of the answer based on the emotion recognition result, and means for providing the generated answer to the user. This makes it possible to provide more appropriate and specific information in real time in response to a natural language question entered by a user, taking into account the user's location information and emotion.
[0802] A "user" is a person using a terminal who enters a question in natural language to obtain information.
[0803] "Natural language" refers to language used in everyday life, sentences and words that do not require special interpretation by a computer program.
[0804] "Terminal" refers to a computing device, such as a smartphone, tablet, or PC, through which a user enters a question or receives a response.
[0805] A "server" is a computer device that receives and processes data sent from a terminal and provides appropriate information.
[0806] "Natural language processing technology" refers to all technologies that enable computers to understand, analyze, and generate natural language.
[0807] "Location Information" means data that indicates a user's geographic location and may include GPS data and Wi-Fi information.
[0808] An "emotion engine" refers to software or algorithms that analyze and identify user emotions based on text data.
[0809] "Online platform" refers to an online service that allows a large number of users to generate and share information and content, including social networking services.
[0810] A "generative AI model" is an artificial intelligence model that generates new text based on training data, such as GPT-4.
[0811] A "prompt" refers to text data that is input to a generative AI model to produce a specific output.
[0812] An "answer" is a sentence generated in response to a user's question based on the analysis results and acquired information.
[0813] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[0814] Hardware and Software
[0815] This system mainly uses the following hardware and software:
[0816] Device: A computing device used by a user, such as a smartphone, tablet, or PC.
[0817] Server: A computing device for receiving, processing, and providing data.
[0818] Natural language processing technology: Google Cloud Natural Language API, etc.
[0819] Sentiment engines: such as IBM Watson Tone Analyzer and Microsoft Azure Text Analytics.
[0820] Location information acquisition technology: Google Maps API, built-in GPS module, etc.
[0821] Social networking interfaces: Twitter API, etc.
[0822] Generative AI models: such as OpenAI's GPT-4.
[0823] Data processing and calculation
[0824] The main processing of the system involves the following data processing and data calculation.
[0825] 1. Natural language processing: The device sends the user's question to the server, which then uses natural language processing to analyze the question, for example, identifying that it is a question about the weather.
[0826] 2. Location information acquisition: The device acquires the user's current location information and sends it to the server. This location information is acquired using GPS or Wi-Fi.
[0827] 3. Emotion Recognition: The server uses an emotion engine to analyze and identify emotions from the user's question, thereby determining the emotion (e.g., "anxiety" or "excitement") the user is feeling when asking the question.
[0828] 4. Search and retrieve SNS posts: Based on the acquired location information and the question analysis results, the server uses the SNS interface to search and retrieve related posts. For example, a search can be performed using keywords such as "Tokyo West Rain."
[0829] 5. Post analysis: The server analyzes the social media posts and summarizes them using a generative AI model, identifying trends and key keywords from the content of the posts (e.g., "Western Tokyo" and "Heavy Rain").
[0830] 6. Emotion-based response adjustment: The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," the server will prioritize providing more detailed and reassuring information.
[0831] 7. Answer Generation: The server generates a textual answer to the user's question, taking into account the analysis results and sentiment-based adjustments. For example, it creates an answer such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time."
[0832] 8. Providing an answer: The server sends the generated answer to the terminal, and the terminal displays the answer to the user, allowing the user to quickly obtain specific and real-time information about the question.
[0833] Specific examples
[0834] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[0835] 1. User: Enters a question into the terminal.
[0836] 2. Terminal: Sends a query to the server.
[0837] 3. Server: Parses the question and identifies it as a weather question.
[0838] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0839] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[0840] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0841] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0842] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[0843] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[0844] 10. Server: Sends the generated answer to the terminal.
[0845] 11. Terminal: Display the answer to the user.
[0846] This allows the system of the present invention to provide prompt and appropriate information in response to user questions, and by taking emotions into consideration, it is possible to provide a more user-friendly response.
[0847] Prompt Sentence Examples
[0848] "Analyze questions entered in natural language, combine location information and emotion recognition to search for relevant information, and generate real-time answers. The user's question is 'The sky is getting dark. Will it rain? I'm worried.' The location information is 'Chuo Ward, Tokyo.' The emotion recognition result is 'Anxiety.'"
[0849] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0850] Step 1:
[0851] User: Enters a question into the terminal in natural language (e.g., "The sky is getting dark. Will it rain?"). The input data is a textual question.
[0852] Terminal: Receives the entered question, stores it as text data, and prepares it to be sent to the server. Specifically, when the user enters a question and presses the "Send" button, the terminal sends the question data to the server. The input is the user's text data, and the output is ready to be sent to the server.
[0853] Step 2:
[0854] Server: Analyzes the question data received from the user using natural language processing (NLP) technology. For example, it identifies that the question is about the weather. Specifically, it uses an NLP library (e.g., Google Cloud Natural Language API) to divide the question into tokens and analyzes the sentence structure, such as subject, predicate, and object. The input is the question text data, and the output is the analysis result (identification of the question content).
[0855] Step 3:
[0856] Device: Obtains the user's current location information. This is obtained using GPS or Wi-Fi. The obtained location information is sent to the server. Specifically, the device's location information acquisition function is called, and the current location is obtained as GPS coordinates or Wi-Fi information. The input is a request to obtain location information, and the output is the obtained location information (e.g., Chuo-ku, Tokyo).
[0857] Step 4:
[0858] Server: Analyzes text data (questions) and recognizes the user's emotions using an emotion engine. Specifically, it identifies emotions such as "anxiety" and "excitement." Specific operations involve inputting text data into an emotion engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Text Analytics) and obtaining an emotion score. The input is the question text data, and the output is the emotion recognition results.
[0859] Step 5:
[0860] Server: Based on the acquired location information and question analysis results, the server uses the online platform's API to search and retrieve related posts. For example, a search is performed using keywords such as "Tokyo West Rain." Specific operations include creating an API request and retrieving related posts. The input is location information and question analysis results, and the output is the retrieved post data.
[0861] Step 6:
[0862] Server: Analyzes the acquired social media post data using a generative AI model (e.g., OpenAI's GPT-4) and summarizes the situation. Identifies trends and key keywords from the content of the post. Specifically, it crawls the acquired post data, analyzes the text content, calculates the frequency distribution of keywords, and has the generative AI model summarize it. The input is the acquired post data, and the output is summarized situation information.
[0863] Step 7:
[0864] Server: Adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," it prioritizes providing more detailed and reassuring information. Specifically, it selects the most appropriate information from the analyzed information set based on the user's emotion score. The input is the emotion recognition result and summarized situation information, and the output is the adjusted answer content.
[0865] Step 8:
[0866] Server: Based on the above information, it generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time." Specific operations include inputting the necessary information as a prompt into the generative AI model and requesting it to generate a text. The generated text is reviewed and any necessary corrections are made. The input is the adjusted answer, and the output is the generated answer text.
[0867] Step 9:
[0868] Server: Sends the generated answer to the device.
[0869] Terminal: Displays the answer to the user. Specifically, the answer data received from the server is displayed on the user interface. The display format can be a pop-up, notification, chat format, etc. The input is the generated answer text, and the output is the answer displayed to the user.
[0870] In this way, the system of the present invention can quickly provide specific, real-time information in response to a user's question.
[0871] (Application example 2)
[0872] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0873] Conventional systems had the problem of making it difficult for users to obtain fast and accurate information in physical stores. In particular, when users had questions about products, there was a lack of a way to quickly and effectively provide detailed information, reviews, and inventory information related to the question in real time. Furthermore, responses did not take into consideration the user's feelings, which could lead to a decline in the quality of the user experience. This made it difficult to improve customer satisfaction.
[0874] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving a question in natural language from a user; means for analyzing the question using natural language processing technology and identifying the question content; emotion recognition means for analyzing the user's emotions; means for acquiring the user's location information; means for searching and acquiring related posts from social networking services based on the acquired location information and the question analysis results; means for analyzing the acquired posts and summarizing a specific situation; means for generating an answer to the user's question based on the summarized situation and the emotion analysis results; and means for providing the generated answer to the user. This allows users to easily obtain quick and accurate information related to products in physical stores and also enables responses that take user emotions into consideration. This improves customer satisfaction.
[0875] "Natural language questions" refer to questions asked by users using everyday words and expressions.
[0876] "Natural language processing technology" is a technology that allows computers to analyze and understand human language.
[0877] The "means for identifying the question content" is a mechanism for analyzing the received question and clarifying what the question is asking.
[0878] "Emotion recognition means" refers to a device or program that analyzes and identifies emotions and moods from text entered by a user.
[0879] "Means of obtaining location information" refers to a mechanism for obtaining the user's current location using technologies such as GPS and Wi-Fi.
[0880] A "social networking service" is an online communication platform that allows users to share information.
[0881] "Means for searching and retrieving related posts" refers to a system for searching and collecting posts that match specific keywords or conditions based on the acquired location information and analysis results.
[0882] "Means for analyzing posts and summarizing specific situations" refers to a system for analyzing collected post content, extracting important information from it, and summarizing it concisely.
[0883] The "means for generating an answer" is a mechanism for creating an appropriate response to a user's question in the form of text based on the analysis results and emotion recognition results.
[0884] The "means for providing an answer" is a device or program for presenting the generated answer to the user in a format that is easy to view.
[0885] The system that realizes this application example is a "smart shopping assistant" system that allows users to input product-related questions in natural language while shopping in a physical store and provides quick and appropriate answers.
[0886] System Program
[0887] This system is implemented using the following hardware and software:
[0888] Hardware
[0889] 1. User terminal: A device such as a smartphone, smart glasses, or head-mounted display.
[0890] 2. Server: A computer server with powerful computing power.
[0891] software
[0892] 1. Natural language processing libraries: Use NLTK (Natural Language Toolkit) and spaCy to analyze natural language questions from users.
[0893] 2. Emotion Recognition Engine: Analyzes user emotions using TextBlob and VADER (Valence Aware Dictionary and sEntiment Reasoner).
[0894] 3. SNS API: Use the API of social networking services (e.g. Twitter API) to search and retrieve related posts.
[0895] 4. Generative AI model: An artificial intelligence model that generates appropriate answers based on the information obtained.
[0896] What the program does
[0897] 1. Receiving Questions
[0898] A user types a question in natural language into a device such as a smartphone or smart glasses, for example, "Is the new model in stock?" This question is received by the device and sent to the server.
[0899] 2. Natural Language Processing
[0900] When the server receives a question, it uses a natural language processing library (e.g., spaCy) to analyze the question and extract information about specific categories and products.
[0901] 3. Emotion recognition
[0902] The server then uses an emotion recognition engine (e.g., TextBlob) to analyze the emotion from the user's question. For example, the emotion "anxiety" is identified from the question "I'm worried about stock availability."
[0903] 4. Obtaining location information
[0904] The user's device uses GPS and Wi-Fi to obtain its current location and sends it to the server, which can then obtain information about the store the user is in or nearby.
[0905] 5. Searching for Information
[0906] The server uses the acquired location information and the question analysis results to search and retrieve related posts using the SNS API, searching the social networking service for keywords such as "in stock" and "new model."
[0907] 6. Analysis and Summarization of Posts
[0908] The server analyzes the posts and uses a generative AI model to summarize important information, such as "latest model in stock" or "highly rated."
[0909] 7. Answer Generation
[0910] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, to a user who is feeling anxious, the server creates a response such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good."
[0911] 8. Providing answers
[0912] The generated answer is sent from the server to the device and displayed to the user, allowing the user to quickly obtain detailed information about the product in the physical store.
[0913] Specific examples
[0914] A user uses a smartphone in a physical store to ask, "Is the new model in stock?" At this time, the system analyzes the user's emotion as "anxiety" and collects the latest stock status and product review information from social media posts. Ultimately, it generates an answer such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good," and displays it on the user's device.
[0915] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0916] Step 1:
[0917] A user enters a question in natural language using a smartphone or smart glasses in a physical store. The entered question (e.g., "Is the new model in stock?") is received by the user device and sent to the server. In this step, the input is the user's question, and the output is the data that sends the question to the server.
[0918] Step 2:
[0919] The server analyzes the received question using a natural language processing library (e.g., spaCy). This analysis identifies the intent of the question and the target product category. Specifically, the server tokenizes the question, tags it with parts of speech, and performs semantic analysis. The input is the user's natural language question, and the output is the identified question (e.g., "new model" or "inventory").
[0920] Step 3:
[0921] At the same time, the server uses an emotion recognition engine (e.g., TextBlob) to analyze the user's emotion from the question. The emotion recognition engine receives text data as input, analyzes the emotion polarity, and identifies the emotion type (e.g., "anxiety," "relief," or "interest"). The input is the user's question, and the output is the analyzed emotion data (e.g., "anxiety").
[0922] Step 4:
[0923] The user device acquires its current location information using GPS or Wi-Fi and sends that location information to the server. The input is the location data of the user device, and the output is the location information data sent to the server (e.g., "Shinjuku-ku, Tokyo").
[0924] Step 5:
[0925] The server uses the API of the social networking service (SNS) to search and retrieve related posts based on the acquired location information and question analysis results. This search is performed by combining specific keywords and location information. The input is the question analysis results and location information, and the output is related post data (e.g., "SNS post: latest model in stock, many positive reviews").
[0926] Step 6:
[0927] The server analyzes the social media posts and uses a generative AI model to summarize important information. This summarization includes extracting trends and key keywords from the posts. The input is the social media post data, and the output is summarized information (e.g., "Likely Rated, In Stock").
[0928] Step 7:
[0929] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, for a user who is feeling anxious, the server might create a response such as, "This product is the latest model. It is currently in stock, so don't worry!" The input is summarized information and emotion data, and the output is the generated response text data.
[0930] Step 8:
[0931] The server sends the generated answer to the user's terminal, which then displays the answer. The user can view the answer and quickly obtain detailed information about the product. The input is the generated answer data, and the output is the display of the answer on the user's terminal.
[0932] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0933] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0934] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0935] [Third embodiment]
[0936] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0937] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0938] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0939] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0940] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0941] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0942] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0943] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0944] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0945] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0946] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0947] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0948] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[0949] Overall system overview
[0950] The system mainly consists of the following components:
[0951] 1. The device that receives the user's questions
[0952] 2. Server that analyzes questions and searches for and analyzes information
[0953] 3. Device that generates and provides answers to questions
[0954] Overview of program processing
[0955] 1. Receive user questions
[0956] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[0957] 2. Question Analysis
[0958] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[0959] 3. Obtaining location information
[0960] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[0961] 4. Search and retrieve social media posts
[0962] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[0963] 5. Analysis and Summarization of Posts
[0964] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[0965] 6. Answer Generation
[0966] Based on the analysis results, the server generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to here as time passes."
[0967] 7. Submitting and Viewing Your Answers
[0968] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[0969] Specific examples
[0970] Here are some concrete examples:
[0971] A user types in a question: "The sky is getting dark. Is it going to rain?"
[0972] 1. User: Enters a question into the terminal.
[0973] 2. Terminal: Sends a query to the server.
[0974] 3. Server: Parses the question and identifies it as a weather question.
[0975] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[0976] 5. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[0977] 6. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[0978] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[0979] 8. Server: Sends the generated answer to the device.
[0980] 9. Terminal: Display the answer to the user.
[0981] In this way, the system of the present invention can provide quick and accurate information simply by asking a question in natural language, allowing users to obtain the information they need in a short time without having to perform complex search operations.
[0982] The processing flow will be explained below.
[0983] Step 1:
[0984] User: Type a question into the device in natural language (e.g., "The sky is getting dark. Will it rain?").
[0985] Step 2:
[0986] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[0987] Step 3:
[0988] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. The obtained location information is sent to a server with the user's consent.
[0989] Step 4:
[0990] Device: Sends the user's question and location information to the server.
[0991] Step 5:
[0992] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[0993] Step 6:
[0994] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[0995] Step 7:
[0996] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[0997] Step 8:
[0998] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[0999] Step 9:
[1000] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[1001] Step 10:
[1002] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[1003] Step 11:
[1004] Server: Generates answers to user questions based on the analysis results. Specifically, the answer may be something like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1005] Step 12:
[1006] Server: Sends the generated answer to the user's device.
[1007] Step 13:
[1008] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[1009] Step 14:
[1010] Device: Show the user the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1011] In this way, through a series of processes, users can get quick and accurate answers to their questions.
[1012] Example 1
[1013] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1014] Conventional information retrieval systems have difficulty providing appropriate answers to questions in real time, even when users input questions in natural language. Furthermore, they have limitations in analyzing the user's current location and providing real-time information using posts from social networking services. This has resulted in problems such as users being unable to quickly and accurately obtain the information they are looking for, and requiring complex search operations.
[1015] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1016] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for searching for and acquiring related posts from social networking services based on the acquired location information and the question analysis results, means for analyzing the acquired posts and summarizing a specific situation using a generative AI model, means for generating an answer to the user's question based on the summarized situation, and means for providing the generated answer to the user. This enables a user to quickly acquire highly accurate information in real time by simply inputting a question in natural language.
[1017] The "means for receiving a natural language question from a user" refers to a device or function for receiving a natural language question entered by a user and passing the question on to subsequent processing.
[1018] "Means for analyzing questions using natural language processing technology and identifying the content of the question" refers to technology, devices, or functions for analyzing received natural language questions and understanding their content. Specifically, natural language processing technology is used to identify the topic and intent of the question.
[1019] "Means for acquiring user location information" refers to devices or functions for identifying and acquiring the user's current location. Specifically, location information is acquired using GPS, Wi-Fi data, etc.
[1020] "Means for searching and retrieving related posts from social networking services based on acquired location information and question analysis results" refers to technology, devices, or functions that search for and retrieve related posts from SNS based on the user's location information and question content.
[1021] "Means for analyzing acquired posts and summarizing specific situations using a generative AI model" refers to technology, devices, or functions that analyze acquired social media posts and summarize their content using a generative AI model.
[1022] The "means for generating an answer to a user's question based on the summarized situation" refers to a technology, device, or function for generating an appropriate answer to a user's question based on the summary result.
[1023] The "means for providing a generated answer to a user" is a device or function for transmitting and displaying a generated answer to a user.
[1024] MODE FOR CARRYING OUT THE INVENTION
[1025] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[1026] Overall system overview
[1027] The system consists of the following main components:
[1028] 1. The device that receives the user's questions
[1029] 2. Server that analyzes questions and searches for and analyzes information
[1030] 3. Device that generates and provides answers to questions
[1031] Process Overview
[1032] 1. Receive questions from users
[1033] The user inputs a question in natural language into the device. For example, the user inputs, "The sky is getting dark. Is it going to rain?" The device receives this question and sends it to the server.
[1034] 2. Question Analysis
[1035] The server receives the user's question and analyzes it using natural language processing technology, specifically using libraries such as TensorFlow and spaCy, to determine that the question is about the weather.
[1036] 3. Obtaining location information
[1037] The device acquires the user's current location using technologies such as GPS and Wi-Fi. After obtaining the user's consent, the device sends the acquired location information (e.g., Chuo Ward, Tokyo) to a server.
[1038] 4. Search and retrieve social media posts
[1039] The server uses the SNS API (e.g., Twitter's API) to search for and retrieve related posts based on the location information and the results of analyzing the question. Keywords such as "Western Tokyo, Rain" are typically used as the search query, and the retrieved data is handled in JSON format.
[1040] 5. Analysis and Summarization of Posts
[1041] The server analyzes the social media posts and summarizes them using a generative AI model (e.g., OpenAI's GPT-3). During this process, trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain") are identified.
[1042] 6. Answer Generation
[1043] The server generates a textual answer to the user's question based on the analysis results. For example, it might create a response like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1044] 7. Submitting and Viewing Your Answers
[1045] The server sends the generated answer to the terminal, which then displays it to the user, allowing the user to quickly obtain detailed and specific information in real time.
[1046] Adding specific examples
[1047] Here are some concrete examples:
[1048] Specific situations
[1049] User asks: "The sky is getting dark, is it going to rain?"
[1050] 1. User: Enters the question "The sky is getting dark. Is it going to rain?" into the device.
[1051] 2. Terminal: Sends the entered question to the server.
[1052] 3. Server: Analyzes the question using TensorFlow and spaCy and identifies it as a weather-related question.
[1053] 4. Device: Uses GPS or Wi-Fi to obtain current location information (e.g., Chuo-ku, Tokyo) and sends it to the server.
[1054] 5. Server: Using Twitter API etc., search for social media posts using the keyword "Tokyo Western Rain."
[1055] 6. Server: Analyze the acquired social media post data, summarize it using GPT-3, and identify trending keywords.
[1056] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1057] 8. Server: Sends the generated answer to the terminal.
[1058] 9. Terminal: Displays the received answer on the screen.
[1059] Example prompt sentences to use
[1060] An input prompt for a generative AI model takes the form:
[1061] User Question: "The sky is getting dark, is it going to rain?"
[1062] Location information: "Chuo-ku, Tokyo"
[1063] Social media post: "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer as time goes by."
[1064] Based on these prompts, the AI model generates appropriate answers to the user's questions.
[1065] This system allows users to quickly and accurately obtain the information they need simply by entering their questions in natural language, providing a high level of convenience by eliminating the complex search work previously required.
[1066] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1067] Step 1:
[1068] The user inputs a question in natural language into the terminal. The terminal receives the question (e.g., "The sky is getting dark. Will it rain?"). The input of this process is the natural language question input by the user, and the output is to store this question in internal memory.
[1069] Step 2:
[1070] The terminal sends the received question to the server. Specifically, it generates an HTTP POST request and sends a payload containing the question to the server. The input of this process is the question entered by the user and the server address information, and the output is the sending of an HTTP request containing the question.
[1071] Step 3:
[1072] The server receives questions sent by users. It extracts the received questions and analyzes the content of the questions using natural language processing technology. Specifically, it uses libraries such as TensorFlow and spaCy to analyze the meaning of the questions. The input to this process is the user's question text, and the output is the analysis result (e.g., identifying that the question is about the weather).
[1073] Step 4:
[1074] Based on the analysis results, the server identifies the question as weather-related. This identification process involves text classification to understand the intent of the question. The input to this process is the analysis results from natural language processing, and the output is information that classifies the question into a category related to "weather."
[1075] Step 5:
[1076] The device uses GPS and Wi-Fi data to obtain the user's location. After obtaining the user's consent, the device obtains the current location (e.g., latitude, longitude, and area name) and sends it to the server. The input is data from the device's location information acquisition function, and the output is the location information (e.g., Chuo Ward, Tokyo).
[1077] Step 6:
[1078] The server uses the SNS API to search for and retrieve related posts based on the acquired location information and question analysis results. For example, using the Twitter API, a search is performed for keywords such as "Tokyo Western Rain." The input for this process is the location information and the results of the question analysis, and the output is the SNS post data (in JSON format) as the search results.
[1079] Step 7:
[1080] The server analyzes the social media posts and uses a generative AI model (e.g., OpenAI's GPT-3) to summarize a specific situation. It extracts trends and key keywords and generates a summary. The input is the social media post data, and the output is text describing the summarized situation (e.g., "There are many posts reporting heavy rain in western Tokyo").
[1081] Step 8:
[1082] The server generates an answer to the user's question based on the summarized situation. This answer generation process also uses a generative AI model. The input is the analyzed and summarized data, and the output is a text answer to the user's question (e.g., "It appears that heavy rain is approaching us as time passes.").
[1083] Step 9:
[1084] The server sends the generated answer to the terminal. The answer is sent as an HTTP response, and the terminal receives it. The input is the generated answer text, and the output is the HTTP response from the server to the terminal.
[1085] Step 10:
[1086] The terminal displays the received answer to the user. The answer text is displayed in the application interface and the user confirms it. The input is the answer from the server, and the output is the answer text displayed on the user's screen.
[1087] (Application example 1)
[1088] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1089] Real-time road and weather information is an extremely important element in the operation of autonomous vehicles. However, there is a lack of means to quickly and accurately obtain relevant information from a variety of sources and provide it in a format that is easy for users to understand. In particular, there is a need for technology that makes this information available in real time via a voice interface. This will improve the safety and convenience of autonomous vehicles.
[1090] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1091] In this invention, the server includes means for converting voice questions from users into text using voice recognition technology, means for acquiring user location information, means for searching and acquiring related posts from information provision services based on the acquired location information and question analysis results, means for analyzing the acquired posts and summarizing specific situations, and means for providing the generated answers to the user by voice using voice synthesis technology. This enables users to quickly and accurately acquire real-time road and weather information and receive answers by voice simply by asking questions by voice.
[1092] "Natural language" refers to a language used by humans on a daily basis, and is not limited to any particular computer language or format.
[1093] "Natural language processing technology" refers to all technologies that enable computers to understand, generate, and analyze human language, including text analysis and speech recognition.
[1094] "Location information" refers to the current geographic location of a user or device, and is data obtained using GPS, Wi-Fi, etc.
[1095] An "information provision service" is an online service that provides specific information to users, such as social networking services and news sites.
[1096] A "post" is any content such as text, images, or videos uploaded by a user to an information service.
[1097] "Speech recognition technology" refers to technology that converts a user's speech into text or other data formats.
[1098] "Speech synthesis technology" is a technology that generates natural speech from text and is used to output speech via a computer.
[1099] "Real-time" refers to responding immediately to user operations and inputs, and processing and providing data without delay.
[1100] This invention relates to a navigation assistant system that provides real-time road and weather information in autonomous vehicles. The system includes a series of processes that receive voice questions from users, convert them into text, and analyze them. It then collects and analyzes related information and provides answers to users via voice.
[1101] The system primarily includes the following hardware and software components:
[1102] 1. A microphone in the vehicle to receive the user's voice query
[1103] 2. Speech recognition technology that converts speech to text (e.g., Google Cloud Speech-to-Text API)
[1104] 3. GPS module to obtain the user's current location
[1105] 4. Internet connection and API access to retrieve relevant posts from information services (e.g., Twitter API)
[1106] 5. A server to analyze posts and summarize specific situations using a generative AI model (e.g., OpenAI's GPT-4 model)
[1107] 6. Speech synthesis technology that generates natural-sounding speech from text (e.g., Google Cloud Text-to-Speech API)
[1108] 7. On-board computer systems for autonomous vehicles
[1109] Explanation of system processing
[1110] Receiving and analyzing voice questions
[1111] The server receives voice questions from users through the vehicle's microphone. This voice is converted into text using the Google Cloud Speech-to-Text API. The converted text question is then analyzed using natural language processing technology. For example, if the question is "What's the weather like now?", it is identified as a weather-related question.
[1112] Obtaining location information
[1113] The device acquires the user's current location using the vehicle's GPS module, and sends the acquired location information along with the analyzed question to the server.
[1114] Collection and analysis of relevant information
[1115] The server uses information providers like the Twitter API to retrieve real-time posts related to the location and question, which are then analyzed using OpenAI's GPT-4 model and summarized based on key keywords and trends.
[1116] Generate and provide answers
[1117] The server generates an answer to the user's question based on the analysis results. This answer is converted into audio using the Google Cloud Text-to-Speech API and provided to the user through the vehicle's speakers. For example, the answer might be, "It's starting to rain near your current location. There is traffic congestion on nearby roads."
[1118] Examples of concrete examples and prompts
[1119] For example, if a user asks verbally, "What's the weather like today?", the server will input the following prompt into the generative AI model:
[1120] plaintext
[1121] Summarize the tweet below:
[1122] It's currently raining heavily in Tokyo. Drive carefully!
[1123] There's a traffic jam, but it looks like it's going to rain soon.
[1124] summary:
[1125] This allows users to quickly and accurately obtain real-time weather and traffic information and receive voice responses simply by asking voice questions.
[1126] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1127] Step 1:
[1128] Receive user voice questions
[1129] Users ask questions in natural language into a microphone in the vehicle, and the device receives the audio and stores it as an audio file.
[1130] Input: User's voice question
[1131] Output: Audio file
[1132] What it does: The microphone captures your voice and converts it into a digital audio file.
[1133] Step 2:
[1134] Convert speech to text
[1135] The device uses the Google Cloud Speech-to-Text API to convert the audio file saved in step 1 into text, which is then sent to the server as the user's question.
[1136] Input: Audio file
[1137] Output: Text data
[1138] Specific operation: The device sends the audio file to the Google Cloud Speech-to-Text API and receives natural language text data.
[1139] Step 3:
[1140] Get the user's location
[1141] The device uses the vehicle's GPS module to obtain the user's current location, which is then sent to the server along with the text data.
[1142] Input: Current geographic location (GPS signal)
[1143] Output: Location data (latitude and longitude)
[1144] Specific operation: The device obtains the latitude and longitude information of the current location from the GPS module and saves it as location data.
[1145] Step 4:
[1146] Get related posts from information services
[1147] Based on the user's question and location information, the server searches and retrieves relevant posts in real time from information providers such as the Twitter API.
[1148] Input: Text data, location data
[1149] Output: Related post data (tweets)
[1150] What it does: The server uses the Twitter API to retrieve posts based on a given location and keyword, for example, using a search query like "Tokyo weather."
[1151] Step 5:
[1152] Analyze and summarize the posts
[1153] The server uses OpenAI's GPT-4 model to analyze the acquired post data, extract and summarize key keywords and trends.
[1154] Input: Related post data
[1155] Output: Summary data (text)
[1156] Specific operation: The server analyzes the text data of the tweet, inputs a prompt sentence into the GPT-4 model, and obtains a summary result.
[1157] Example prompt:
[1158] plaintext
[1159] Summarize the tweet below:
[1160] It's currently raining heavily in Tokyo. Drive carefully!
[1161] There's a traffic jam, but it looks like it's going to rain soon.
[1162] summary:
[1163] Step 6:
[1164] Generate answers and convert them into audio
[1165] The server generates answers to the user's questions based on the summary data and converts them into audio using the Google Cloud Text-to-Speech API.
[1166] Input: Summary data
[1167] Output: Audio data (answer)
[1168] What it does: The server analyzes the summary data and generates an appropriate answer to the question in text form, then converts that text into audio using the Google Cloud Text-to-Speech API.
[1169] Step 7:
[1170] Providing answers to users
[1171] The device then provides the generated voice data to the user through the vehicle's speakers, allowing the user to receive voice responses to their questions in real time.
[1172] Input: Voice data (answer)
[1173] Output: Audio output to the user
[1174] Specific operation: The device routes the generated voice data to the car's speakers and provides it to the user audibly.
[1175] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1176] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[1177] Overall system overview
[1178] The system mainly consists of the following components:
[1179] 1. The device that receives the user's questions
[1180] 2. Server that analyzes questions and searches for and analyzes information
[1181] 3. Device that generates and provides answers to questions
[1182] 4. Emotion engine that recognizes user emotions
[1183] Overview of program processing
[1184] 1. Receive user questions
[1185] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[1186] 2. Question Analysis
[1187] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[1188] 3. Obtaining location information
[1189] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[1190] 4. Emotional Recognition
[1191] At the same time, the server uses an emotion engine to analyze the user's emotion from the question (e.g., "anxiety," "excitement," etc.). The emotion engine analyzes the text data of the question as input and identifies the emotion.
[1192] 5. Search and retrieve social media posts
[1193] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[1194] 6. Analysis and Summarization of Posts
[1195] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[1196] 7. Adjust your responses based on emotion
[1197] The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling anxious, it will prioritize providing more detailed and reassuring information.
[1198] 8. Answer Generation
[1199] The server takes into account the analysis results and emotional adjustments to generate a textual answer to the user's question. Specifically, it creates a response such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[1200] 9. Submitting and Viewing Your Answers
[1201] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[1202] Specific examples
[1203] Here are some concrete examples:
[1204] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[1205] 1. User: Enters a question into the terminal.
[1206] 2. Terminal: Sends a query to the server.
[1207] 3. Server: Parses the question and identifies it as a weather question.
[1208] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[1209] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[1210] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[1211] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[1212] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[1213] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[1214] 10. Server: Sends the generated answer to the terminal.
[1215] 11. Terminal: Display the answer to the user.
[1216] In this way, the system of the present invention can provide prompt and appropriate information in response to a user's question, and by taking emotions into consideration, can realize a more user-friendly response.
[1217] The processing flow will be explained below.
[1218] Step 1:
[1219] User: Enters a question in natural language into the device (e.g., "The sky is getting dark. Is it going to rain? I'm worried.").
[1220] Step 2:
[1221] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[1222] Step 3:
[1223] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. With the user's consent, the obtained location information is prepared for transmission to the server.
[1224] Step 4:
[1225] Device: Sends the user's question and location information to the server.
[1226] Step 5:
[1227] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[1228] Step 6:
[1229] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[1230] Step 7:
[1231] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[1232] Step 8:
[1233] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[1234] Step 9:
[1235] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[1236] Step 10:
[1237] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[1238] Step 11:
[1239] Server: Uses an emotion engine to recognize emotions from the user's question. For example, identify that the question contains the emotion "anxiety."
[1240] Step 12:
[1241] Server: Adjusts the priority of search and analysis results based on the user's emotions recognized by the emotion engine. If the user is feeling anxious, it prioritizes providing more detailed and reassuring information.
[1242] Step 13:
[1243] Server: Generates a written answer to the user's question, taking into account the analysis results and emotional adjustments. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[1244] Step 14:
[1245] Server: Sends the generated answer to the user's device.
[1246] Step 15:
[1247] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[1248] Step 16:
[1249] Device: Display the following response to the user: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[1250] Through this process, users can receive prompt and appropriate answers to their questions, and by taking into account the user's emotions, the system can provide a more user-friendly response.
[1251] Example 2
[1252] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1253] In today's world, users want to obtain appropriate information in real time by asking questions in natural language. However, conventional information search systems are limited to providing information based on keywords, and it is difficult to generate answers that take into account the user's situation and emotions. This makes them insufficient to alleviate users' anxieties and questions, and there is a need for more accurate information provision systems.
[1254] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1255] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for recognizing the user's emotion from the question text, means for searching and acquiring related posts from an online platform based on the acquired location information and the question analysis result, means for analyzing the acquired posts and summarizing a specific situation, means for generating an answer to the user's question based on the summarized situation, means for adjusting the content of the answer based on the emotion recognition result, and means for providing the generated answer to the user. This makes it possible to provide more appropriate and specific information in real time in response to a natural language question entered by a user, taking into account the user's location information and emotion.
[1256] A "user" is a person using a terminal who enters a question in natural language to obtain information.
[1257] "Natural language" refers to language used in everyday life, sentences and words that do not require special interpretation by a computer program.
[1258] "Terminal" refers to a computing device, such as a smartphone, tablet, or PC, through which a user enters a question or receives a response.
[1259] A "server" is a computer device that receives and processes data sent from a terminal and provides appropriate information.
[1260] "Natural language processing technology" refers to all technologies that enable computers to understand, analyze, and generate natural language.
[1261] "Location Information" means data that indicates a user's geographic location and may include GPS data and Wi-Fi information.
[1262] An "emotion engine" refers to software or algorithms that analyze and identify user emotions based on text data.
[1263] "Online platform" refers to an online service that allows a large number of users to generate and share information and content, including social networking services.
[1264] A "generative AI model" is an artificial intelligence model that generates new text based on training data, such as GPT-4.
[1265] A "prompt" refers to text data that is input to a generative AI model to produce a specific output.
[1266] An "answer" is a sentence generated in response to a user's question based on the analysis results and acquired information.
[1267] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[1268] Hardware and Software
[1269] This system mainly uses the following hardware and software:
[1270] Device: A computing device used by a user, such as a smartphone, tablet, or PC.
[1271] Server: A computing device for receiving, processing, and providing data.
[1272] Natural language processing technology: Google Cloud Natural Language API, etc.
[1273] Sentiment engines: such as IBM Watson Tone Analyzer and Microsoft Azure Text Analytics.
[1274] Location information acquisition technology: Google Maps API, built-in GPS module, etc.
[1275] Social networking interfaces: Twitter API, etc.
[1276] Generative AI models: such as OpenAI's GPT-4.
[1277] Data processing and calculation
[1278] The main processing of the system involves the following data processing and data calculation.
[1279] 1. Natural language processing: The device sends the user's question to the server, which then uses natural language processing to analyze the question, for example, identifying that it is a question about the weather.
[1280] 2. Location information acquisition: The device acquires the user's current location information and sends it to the server. This location information is acquired using GPS or Wi-Fi.
[1281] 3. Emotion Recognition: The server uses an emotion engine to analyze and identify emotions from the user's question, thereby determining the emotion (e.g., "anxiety" or "excitement") the user is feeling when asking the question.
[1282] 4. Search and retrieve SNS posts: Based on the acquired location information and the question analysis results, the server uses the SNS interface to search and retrieve related posts. For example, a search can be performed using keywords such as "Tokyo West Rain."
[1283] 5. Post analysis: The server analyzes the social media posts and summarizes them using a generative AI model, identifying trends and key keywords from the content of the posts (e.g., "Western Tokyo" and "Heavy Rain").
[1284] 6. Emotion-based response adjustment: The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," the server will prioritize providing more detailed and reassuring information.
[1285] 7. Answer Generation: The server generates a textual answer to the user's question, taking into account the analysis results and sentiment-based adjustments. For example, it creates an answer such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time."
[1286] 8. Providing an answer: The server sends the generated answer to the terminal, and the terminal displays the answer to the user, allowing the user to quickly obtain specific and real-time information about the question.
[1287] Specific examples
[1288] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[1289] 1. User: Enters a question into the terminal.
[1290] 2. Terminal: Sends a query to the server.
[1291] 3. Server: Parses the question and identifies it as a weather question.
[1292] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[1293] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[1294] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[1295] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[1296] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[1297] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[1298] 10. Server: Sends the generated answer to the terminal.
[1299] 11. Terminal: Display the answer to the user.
[1300] This allows the system of the present invention to provide prompt and appropriate information in response to user questions, and by taking emotions into consideration, it is possible to provide a more user-friendly response.
[1301] Prompt Sentence Examples
[1302] "Analyze questions entered in natural language, combine location information and emotion recognition to search for relevant information, and generate real-time answers. The user's question is 'The sky is getting dark. Will it rain? I'm worried.' The location information is 'Chuo Ward, Tokyo.' The emotion recognition result is 'Anxiety.'"
[1303] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1304] Step 1:
[1305] User: Enters a question into the terminal in natural language (e.g., "The sky is getting dark. Will it rain?"). The input data is a textual question.
[1306] Terminal: Receives the entered question, stores it as text data, and prepares it to be sent to the server. Specifically, when the user enters a question and presses the "Send" button, the terminal sends the question data to the server. The input is the user's text data, and the output is ready to be sent to the server.
[1307] Step 2:
[1308] Server: Analyzes the question data received from the user using natural language processing (NLP) technology. For example, it identifies that the question is about the weather. Specifically, it uses an NLP library (e.g., Google Cloud Natural Language API) to divide the question into tokens and analyzes the sentence structure, such as subject, predicate, and object. The input is the question text data, and the output is the analysis result (identification of the question content).
[1309] Step 3:
[1310] Device: Obtains the user's current location information. This is obtained using GPS or Wi-Fi. The obtained location information is sent to the server. Specifically, the device's location information acquisition function is called, and the current location is obtained as GPS coordinates or Wi-Fi information. The input is a request to obtain location information, and the output is the obtained location information (e.g., Chuo-ku, Tokyo).
[1311] Step 4:
[1312] Server: Analyzes text data (questions) and recognizes the user's emotions using an emotion engine. Specifically, it identifies emotions such as "anxiety" and "excitement." Specific operations involve inputting text data into an emotion engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Text Analytics) and obtaining an emotion score. The input is the question text data, and the output is the emotion recognition results.
[1313] Step 5:
[1314] Server: Based on the acquired location information and question analysis results, the server uses the online platform's API to search and retrieve related posts. For example, a search is performed using keywords such as "Tokyo West Rain." Specific operations include creating an API request and retrieving related posts. The input is location information and question analysis results, and the output is the retrieved post data.
[1315] Step 6:
[1316] Server: Analyzes the acquired social media post data using a generative AI model (e.g., OpenAI's GPT-4) and summarizes the situation. Identifies trends and key keywords from the content of the post. Specifically, it crawls the acquired post data, analyzes the text content, calculates the frequency distribution of keywords, and has the generative AI model summarize it. The input is the acquired post data, and the output is summarized situation information.
[1317] Step 7:
[1318] Server: Adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," it prioritizes providing more detailed and reassuring information. Specifically, it selects the most appropriate information from the analyzed information set based on the user's emotion score. The input is the emotion recognition result and summarized situation information, and the output is the adjusted answer content.
[1319] Step 8:
[1320] Server: Based on the above information, it generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time." Specific operations include inputting the necessary information as a prompt into the generative AI model and requesting it to generate a text. The generated text is reviewed and any necessary corrections are made. The input is the adjusted answer, and the output is the generated answer text.
[1321] Step 9:
[1322] Server: Sends the generated answer to the device.
[1323] Terminal: Displays the answer to the user. Specifically, the answer data received from the server is displayed on the user interface. The display format can be a pop-up, notification, chat format, etc. The input is the generated answer text, and the output is the answer displayed to the user.
[1324] In this way, the system of the present invention can quickly provide specific, real-time information in response to a user's question.
[1325] (Application example 2)
[1326] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1327] Conventional systems had the problem of making it difficult for users to obtain fast and accurate information in physical stores. In particular, when users had questions about products, there was a lack of a way to quickly and effectively provide detailed information, reviews, and inventory information related to the question in real time. Furthermore, responses did not take into consideration the user's feelings, which could lead to a decline in the quality of the user experience. This made it difficult to improve customer satisfaction.
[1328] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving a question in natural language from a user; means for analyzing the question using natural language processing technology and identifying the question content; emotion recognition means for analyzing the user's emotions; means for acquiring the user's location information; means for searching and acquiring related posts from social networking services based on the acquired location information and the question analysis results; means for analyzing the acquired posts and summarizing a specific situation; means for generating an answer to the user's question based on the summarized situation and the emotion analysis results; and means for providing the generated answer to the user. This allows users to easily obtain quick and accurate information related to products in physical stores and also enables responses that take user emotions into consideration. This improves customer satisfaction.
[1329] "Natural language questions" refer to questions asked by users using everyday words and expressions.
[1330] "Natural language processing technology" is a technology that allows computers to analyze and understand human language.
[1331] The "means for identifying the question content" is a mechanism for analyzing the received question and clarifying what the question is asking.
[1332] "Emotion recognition means" refers to a device or program that analyzes and identifies emotions and moods from text entered by a user.
[1333] "Means of obtaining location information" refers to a mechanism for obtaining the user's current location using technologies such as GPS and Wi-Fi.
[1334] A "social networking service" is an online communication platform that allows users to share information.
[1335] "Means for searching and retrieving related posts" refers to a system for searching and collecting posts that match specific keywords or conditions based on the acquired location information and analysis results.
[1336] "Means for analyzing posts and summarizing specific situations" refers to a system for analyzing collected post content, extracting important information from it, and summarizing it concisely.
[1337] The "means for generating an answer" is a mechanism for creating an appropriate response to a user's question in the form of text based on the analysis results and emotion recognition results.
[1338] The "means for providing an answer" is a device or program for presenting the generated answer to the user in a format that is easy to view.
[1339] The system that realizes this application example is a "smart shopping assistant" system that allows users to input product-related questions in natural language while shopping in a physical store and provides quick and appropriate answers.
[1340] System Program
[1341] This system is implemented using the following hardware and software:
[1342] Hardware
[1343] 1. User terminal: A device such as a smartphone, smart glasses, or head-mounted display.
[1344] 2. Server: A computer server with powerful computing power.
[1345] software
[1346] 1. Natural language processing libraries: Use NLTK (Natural Language Toolkit) and spaCy to analyze natural language questions from users.
[1347] 2. Emotion Recognition Engine: Analyzes user emotions using TextBlob and VADER (Valence Aware Dictionary and sEntiment Reasoner).
[1348] 3. SNS API: Use the API of social networking services (e.g. Twitter API) to search and retrieve related posts.
[1349] 4. Generative AI model: An artificial intelligence model that generates appropriate answers based on the information obtained.
[1350] What the program does
[1351] 1. Receiving Questions
[1352] A user types a question in natural language into a device such as a smartphone or smart glasses, for example, "Is the new model in stock?" This question is received by the device and sent to the server.
[1353] 2. Natural Language Processing
[1354] When the server receives a question, it uses a natural language processing library (e.g., spaCy) to analyze the question and extract information about specific categories and products.
[1355] 3. Emotion recognition
[1356] The server then uses an emotion recognition engine (e.g., TextBlob) to analyze the emotion from the user's question. For example, the emotion "anxiety" is identified from the question "I'm worried about stock availability."
[1357] 4. Obtaining location information
[1358] The user's device uses GPS and Wi-Fi to obtain its current location and sends it to the server, which can then obtain information about the store the user is in or nearby.
[1359] 5. Searching for Information
[1360] The server uses the acquired location information and the question analysis results to search and retrieve related posts using the SNS API, searching the social networking service for keywords such as "in stock" and "new model."
[1361] 6. Analysis and Summarization of Posts
[1362] The server analyzes the posts and uses a generative AI model to summarize important information, such as "latest model in stock" or "highly rated."
[1363] 7. Answer Generation
[1364] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, to a user who is feeling anxious, the server creates a response such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good."
[1365] 8. Providing answers
[1366] The generated answer is sent from the server to the device and displayed to the user, allowing the user to quickly obtain detailed information about the product in the physical store.
[1367] Specific examples
[1368] A user uses a smartphone in a physical store to ask, "Is the new model in stock?" At this time, the system analyzes the user's emotion as "anxiety" and collects the latest stock status and product review information from social media posts. Ultimately, it generates an answer such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good," and displays it on the user's device.
[1369] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1370] Step 1:
[1371] A user enters a question in natural language using a smartphone or smart glasses in a physical store. The entered question (e.g., "Is the new model in stock?") is received by the user device and sent to the server. In this step, the input is the user's question, and the output is the data that sends the question to the server.
[1372] Step 2:
[1373] The server analyzes the received question using a natural language processing library (e.g., spaCy). This analysis identifies the intent of the question and the target product category. Specifically, the server tokenizes the question, tags it with parts of speech, and performs semantic analysis. The input is the user's natural language question, and the output is the identified question (e.g., "new model" or "inventory").
[1374] Step 3:
[1375] At the same time, the server uses an emotion recognition engine (e.g., TextBlob) to analyze the user's emotion from the question. The emotion recognition engine receives text data as input, analyzes the emotion polarity, and identifies the emotion type (e.g., "anxiety," "relief," or "interest"). The input is the user's question, and the output is the analyzed emotion data (e.g., "anxiety").
[1376] Step 4:
[1377] The user device acquires its current location information using GPS or Wi-Fi and sends that location information to the server. The input is the location data of the user device, and the output is the location information data sent to the server (e.g., "Shinjuku-ku, Tokyo").
[1378] Step 5:
[1379] The server uses the API of the social networking service (SNS) to search and retrieve related posts based on the acquired location information and question analysis results. This search is performed by combining specific keywords and location information. The input is the question analysis results and location information, and the output is related post data (e.g., "SNS post: latest model in stock, many positive reviews").
[1380] Step 6:
[1381] The server analyzes the social media posts and uses a generative AI model to summarize important information. This summarization includes extracting trends and key keywords from the posts. The input is the social media post data, and the output is summarized information (e.g., "Likely Rated, In Stock").
[1382] Step 7:
[1383] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, for a user who is feeling anxious, the server might create a response such as, "This product is the latest model. It is currently in stock, so don't worry!" The input is summarized information and emotion data, and the output is the generated response text data.
[1384] Step 8:
[1385] The server sends the generated answer to the user's terminal, which then displays the answer. The user can view the answer and quickly obtain detailed information about the product. The input is the generated answer data, and the output is the display of the answer on the user's terminal.
[1386] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1387] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1388] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1389] [Fourth embodiment]
[1390] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1391] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1392] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1393] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1394] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1395] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1396] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1397] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1398] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1399] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1400] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1401] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1402] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1403] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[1404] Overall system overview
[1405] The system mainly consists of the following components:
[1406] 1. The device that receives the user's questions
[1407] 2. Server that analyzes questions and searches for and analyzes information
[1408] 3. Device that generates and provides answers to questions
[1409] Overview of program processing
[1410] 1. Receive user questions
[1411] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[1412] 2. Question Analysis
[1413] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[1414] 3. Obtaining location information
[1415] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[1416] 4. Search and retrieve social media posts
[1417] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[1418] 5. Analysis and Summarization of Posts
[1419] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[1420] 6. Answer Generation
[1421] Based on the analysis results, the server generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to here as time passes."
[1422] 7. Submitting and Viewing Your Answers
[1423] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[1424] Specific examples
[1425] Here are some concrete examples:
[1426] A user types in a question: "The sky is getting dark. Is it going to rain?"
[1427] 1. User: Enters a question into the terminal.
[1428] 2. Terminal: Sends a query to the server.
[1429] 3. Server: Parses the question and identifies it as a weather question.
[1430] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[1431] 5. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[1432] 6. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[1433] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1434] 8. Server: Sends the generated answer to the device.
[1435] 9. Terminal: Display the answer to the user.
[1436] In this way, the system of the present invention can provide quick and accurate information simply by asking a question in natural language, allowing users to obtain the information they need in a short time without having to perform complex search operations.
[1437] The processing flow will be explained below.
[1438] Step 1:
[1439] User: Type a question into the device in natural language (e.g., "The sky is getting dark. Will it rain?").
[1440] Step 2:
[1441] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[1442] Step 3:
[1443] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. The obtained location information is sent to a server with the user's consent.
[1444] Step 4:
[1445] Device: Sends the user's question and location information to the server.
[1446] Step 5:
[1447] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[1448] Step 6:
[1449] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[1450] Step 7:
[1451] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[1452] Step 8:
[1453] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[1454] Step 9:
[1455] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[1456] Step 10:
[1457] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[1458] Step 11:
[1459] Server: Generates answers to user questions based on the analysis results. Specifically, the answer may be something like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1460] Step 12:
[1461] Server: Sends the generated answer to the user's device.
[1462] Step 13:
[1463] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[1464] Step 14:
[1465] Device: Show the user the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1466] In this way, through a series of processes, users can get quick and accurate answers to their questions.
[1467] Example 1
[1468] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1469] Conventional information retrieval systems have difficulty providing appropriate answers to questions in real time, even when users input questions in natural language. Furthermore, they have limitations in analyzing the user's current location and providing real-time information using posts from social networking services. This has resulted in problems such as users being unable to quickly and accurately obtain the information they are looking for, and requiring complex search operations.
[1470] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1471] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for searching for and acquiring related posts from social networking services based on the acquired location information and the question analysis results, means for analyzing the acquired posts and summarizing a specific situation using a generative AI model, means for generating an answer to the user's question based on the summarized situation, and means for providing the generated answer to the user. This enables a user to quickly acquire highly accurate information in real time by simply inputting a question in natural language.
[1472] The "means for receiving a natural language question from a user" refers to a device or function for receiving a natural language question entered by a user and passing the question on to subsequent processing.
[1473] "Means for analyzing questions using natural language processing technology and identifying the content of the question" refers to technology, devices, or functions for analyzing received natural language questions and understanding their content. Specifically, natural language processing technology is used to identify the topic and intent of the question.
[1474] "Means for acquiring user location information" refers to devices or functions for identifying and acquiring the user's current location. Specifically, location information is acquired using GPS, Wi-Fi data, etc.
[1475] "Means for searching and retrieving related posts from social networking services based on acquired location information and question analysis results" refers to technology, devices, or functions that search for and retrieve related posts from SNS based on the user's location information and question content.
[1476] "Means for analyzing acquired posts and summarizing specific situations using a generative AI model" refers to technology, devices, or functions that analyze acquired social media posts and summarize their content using a generative AI model.
[1477] The "means for generating an answer to a user's question based on the summarized situation" refers to a technology, device, or function for generating an appropriate answer to a user's question based on the summary result.
[1478] The "means for providing a generated answer to a user" is a device or function for transmitting and displaying a generated answer to a user.
[1479] MODE FOR CARRYING OUT THE INVENTION
[1480] This system retrieves relevant information in real time and provides easy-to-understand answers simply by inputting a question in natural language. This system is realized through interactions between a server, a terminal, and the user.
[1481] Overall system overview
[1482] The system consists of the following main components:
[1483] 1. The device that receives the user's questions
[1484] 2. Server that analyzes questions and searches for and analyzes information
[1485] 3. Device that generates and provides answers to questions
[1486] Process Overview
[1487] 1. Receive questions from users
[1488] The user inputs a question in natural language into the device. For example, the user inputs, "The sky is getting dark. Is it going to rain?" The device receives this question and sends it to the server.
[1489] 2. Question Analysis
[1490] The server receives the user's question and analyzes it using natural language processing technology, specifically using libraries such as TensorFlow and spaCy, to determine that the question is about the weather.
[1491] 3. Obtaining location information
[1492] The device acquires the user's current location using technologies such as GPS and Wi-Fi. After obtaining the user's consent, the device sends the acquired location information (e.g., Chuo Ward, Tokyo) to a server.
[1493] 4. Search and retrieve social media posts
[1494] The server uses the SNS API (e.g., Twitter's API) to search for and retrieve related posts based on the location information and the results of analyzing the question. Keywords such as "Western Tokyo, Rain" are typically used as the search query, and the retrieved data is handled in JSON format.
[1495] 5. Analysis and Summarization of Posts
[1496] The server analyzes the social media posts and summarizes them using a generative AI model (e.g., OpenAI's GPT-3). During this process, trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain") are identified.
[1497] 6. Answer Generation
[1498] The server generates a textual answer to the user's question based on the analysis results. For example, it might create a response like, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1499] 7. Submitting and Viewing Your Answers
[1500] The server sends the generated answer to the terminal, which then displays it to the user, allowing the user to quickly obtain detailed and specific information in real time.
[1501] Adding specific examples
[1502] Here are some concrete examples:
[1503] Specific situations
[1504] User asks: "The sky is getting dark, is it going to rain?"
[1505] 1. User: Enters the question "The sky is getting dark. Is it going to rain?" into the device.
[1506] 2. Terminal: Sends the entered question to the server.
[1507] 3. Server: Analyzes the question using TensorFlow and spaCy and identifies it as a weather-related question.
[1508] 4. Device: Uses GPS or Wi-Fi to obtain current location information (e.g., Chuo-ku, Tokyo) and sends it to the server.
[1509] 5. Server: Using Twitter API etc., search for social media posts using the keyword "Tokyo Western Rain."
[1510] 6. Server: Analyze the acquired social media post data, summarize it using GPT-3, and identify trending keywords.
[1511] 7. Server: Generate the answer, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes."
[1512] 8. Server: Sends the generated answer to the terminal.
[1513] 9. Terminal: Displays the received answer on the screen.
[1514] Example prompt sentences to use
[1515] An input prompt for a generative AI model takes the form:
[1516] User Question: "The sky is getting dark, is it going to rain?"
[1517] Location information: "Chuo-ku, Tokyo"
[1518] Social media post: "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer as time goes by."
[1519] Based on these prompts, the AI model generates appropriate answers to the user's questions.
[1520] This system allows users to quickly and accurately obtain the information they need simply by entering their questions in natural language, providing a high level of convenience by eliminating the complex search work previously required.
[1521] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1522] Step 1:
[1523] The user inputs a question in natural language into the terminal. The terminal receives the question (e.g., "The sky is getting dark. Will it rain?"). The input of this process is the natural language question input by the user, and the output is to store this question in internal memory.
[1524] Step 2:
[1525] The terminal sends the received question to the server. Specifically, it generates an HTTP POST request and sends a payload containing the question to the server. The input of this process is the question entered by the user and the server address information, and the output is the sending of an HTTP request containing the question.
[1526] Step 3:
[1527] The server receives questions sent by users. It extracts the received questions and analyzes the content of the questions using natural language processing technology. Specifically, it uses libraries such as TensorFlow and spaCy to analyze the meaning of the questions. The input to this process is the user's question text, and the output is the analysis result (e.g., identifying that the question is about the weather).
[1528] Step 4:
[1529] Based on the analysis results, the server identifies the question as weather-related. This identification process involves text classification to understand the intent of the question. The input to this process is the analysis results from natural language processing, and the output is information that classifies the question into a category related to "weather."
[1530] Step 5:
[1531] The device uses GPS and Wi-Fi data to obtain the user's location. After obtaining the user's consent, the device obtains the current location (e.g., latitude, longitude, and area name) and sends it to the server. The input is data from the device's location information acquisition function, and the output is the location information (e.g., Chuo Ward, Tokyo).
[1532] Step 6:
[1533] The server uses the SNS API to search for and retrieve related posts based on the acquired location information and question analysis results. For example, using the Twitter API, a search is performed for keywords such as "Tokyo Western Rain." The input for this process is the location information and the results of the question analysis, and the output is the SNS post data (in JSON format) as the search results.
[1534] Step 7:
[1535] The server analyzes the social media posts and uses a generative AI model (e.g., OpenAI's GPT-3) to summarize a specific situation. It extracts trends and key keywords and generates a summary. The input is the social media post data, and the output is text describing the summarized situation (e.g., "There are many posts reporting heavy rain in western Tokyo").
[1536] Step 8:
[1537] The server generates an answer to the user's question based on the summarized situation. This answer generation process also uses a generative AI model. The input is the analyzed and summarized data, and the output is a text answer to the user's question (e.g., "It appears that heavy rain is approaching us as time passes.").
[1538] Step 9:
[1539] The server sends the generated answer to the terminal. The answer is sent as an HTTP response, and the terminal receives it. The input is the generated answer text, and the output is the HTTP response from the server to the terminal.
[1540] Step 10:
[1541] The terminal displays the received answer to the user. The answer text is displayed in the application interface and the user confirms it. The input is the answer from the server, and the output is the answer text displayed on the user's screen.
[1542] (Application example 1)
[1543] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1544] Real-time road and weather information is an extremely important element in the operation of autonomous vehicles. However, there is a lack of means to quickly and accurately obtain relevant information from a variety of sources and provide it in a format that is easy for users to understand. In particular, there is a need for technology that makes this information available in real time via a voice interface. This will improve the safety and convenience of autonomous vehicles.
[1545] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1546] In this invention, the server includes means for converting voice questions from users into text using voice recognition technology, means for acquiring user location information, means for searching and acquiring related posts from information provision services based on the acquired location information and question analysis results, means for analyzing the acquired posts and summarizing specific situations, and means for providing the generated answers to the user by voice using voice synthesis technology. This enables users to quickly and accurately acquire real-time road and weather information and receive answers by voice simply by asking questions by voice.
[1547] "Natural language" refers to a language used by humans on a daily basis, and is not limited to any particular computer language or format.
[1548] "Natural language processing technology" refers to all technologies that enable computers to understand, generate, and analyze human language, including text analysis and speech recognition.
[1549] "Location information" refers to the current geographic location of a user or device, and is data obtained using GPS, Wi-Fi, etc.
[1550] An "information provision service" is an online service that provides specific information to users, such as social networking services and news sites.
[1551] A "post" is any content such as text, images, or videos uploaded by a user to an information service.
[1552] "Speech recognition technology" refers to technology that converts a user's speech into text or other data formats.
[1553] "Speech synthesis technology" is a technology that generates natural speech from text and is used to output speech via a computer.
[1554] "Real-time" refers to responding immediately to user operations and inputs, and processing and providing data without delay.
[1555] This invention relates to a navigation assistant system that provides real-time road and weather information in autonomous vehicles. The system includes a series of processes that receive voice questions from users, convert them into text, and analyze them. It then collects and analyzes related information and provides answers to users via voice.
[1556] The system primarily includes the following hardware and software components:
[1557] 1. A microphone in the vehicle to receive the user's voice query
[1558] 2. Speech recognition technology that converts speech to text (e.g., Google Cloud Speech-to-Text API)
[1559] 3. GPS module to obtain the user's current location
[1560] 4. Internet connection and API access to retrieve relevant posts from information services (e.g., Twitter API)
[1561] 5. A server to analyze posts and summarize specific situations using a generative AI model (e.g., OpenAI's GPT-4 model)
[1562] 6. Speech synthesis technology that generates natural-sounding speech from text (e.g., Google Cloud Text-to-Speech API)
[1563] 7. On-board computer systems for autonomous vehicles
[1564] Explanation of system processing
[1565] Receiving and analyzing voice questions
[1566] The server receives voice questions from users through the vehicle's microphone. This voice is converted into text using the Google Cloud Speech-to-Text API. The converted text question is then analyzed using natural language processing technology. For example, if the question is "What's the weather like now?", it is identified as a weather-related question.
[1567] Obtaining location information
[1568] The device acquires the user's current location using the vehicle's GPS module, and sends the acquired location information along with the analyzed question to the server.
[1569] Collection and analysis of relevant information
[1570] The server uses information providers like the Twitter API to retrieve real-time posts related to the location and question, which are then analyzed using OpenAI's GPT-4 model and summarized based on key keywords and trends.
[1571] Generate and provide answers
[1572] The server generates an answer to the user's question based on the analysis results. This answer is converted into audio using the Google Cloud Text-to-Speech API and provided to the user through the vehicle's speakers. For example, the answer might be, "It's starting to rain near your current location. There is traffic congestion on nearby roads."
[1573] Examples of concrete examples and prompts
[1574] For example, if a user asks verbally, "What's the weather like today?", the server will input the following prompt into the generative AI model:
[1575] plaintext
[1576] Summarize the tweet below:
[1577] It's currently raining heavily in Tokyo. Drive carefully!
[1578] There's a traffic jam, but it looks like it's going to rain soon.
[1579] summary:
[1580] This allows users to quickly and accurately obtain real-time weather and traffic information and receive voice responses simply by asking voice questions.
[1581] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1582] Step 1:
[1583] Receive user voice questions
[1584] Users ask questions in natural language into a microphone in the vehicle, and the device receives the audio and stores it as an audio file.
[1585] Input: User's voice question
[1586] Output: Audio file
[1587] What it does: The microphone captures your voice and converts it into a digital audio file.
[1588] Step 2:
[1589] Convert speech to text
[1590] The device uses the Google Cloud Speech-to-Text API to convert the audio file saved in step 1 into text, which is then sent to the server as the user's question.
[1591] Input: Audio file
[1592] Output: Text data
[1593] Specific operation: The device sends the audio file to the Google Cloud Speech-to-Text API and receives natural language text data.
[1594] Step 3:
[1595] Get the user's location
[1596] The device uses the vehicle's GPS module to obtain the user's current location, which is then sent to the server along with the text data.
[1597] Input: Current geographic location (GPS signal)
[1598] Output: Location data (latitude and longitude)
[1599] Specific operation: The device obtains the latitude and longitude information of the current location from the GPS module and saves it as location data.
[1600] Step 4:
[1601] Get related posts from information services
[1602] Based on the user's question and location information, the server searches and retrieves relevant posts in real time from information providers such as the Twitter API.
[1603] Input: Text data, location data
[1604] Output: Related post data (tweets)
[1605] What it does: The server uses the Twitter API to retrieve posts based on a given location and keyword, for example, using a search query like "Tokyo weather."
[1606] Step 5:
[1607] Analyze and summarize the posts
[1608] The server uses OpenAI's GPT-4 model to analyze the acquired post data, extract and summarize key keywords and trends.
[1609] Input: Related post data
[1610] Output: Summary data (text)
[1611] Specific operation: The server analyzes the text data of the tweet, inputs a prompt sentence into the GPT-4 model, and obtains a summary result.
[1612] Example prompt:
[1613] plaintext
[1614] Summarize the tweet below:
[1615] It's currently raining heavily in Tokyo. Drive carefully!
[1616] There's a traffic jam, but it looks like it's going to rain soon.
[1617] summary:
[1618] Step 6:
[1619] Generate answers and convert them into audio
[1620] The server generates answers to the user's questions based on the summary data and converts them into audio using the Google Cloud Text-to-Speech API.
[1621] Input: Summary data
[1622] Output: Audio data (answer)
[1623] What it does: The server analyzes the summary data and generates an appropriate answer to the question in text form, then converts that text into audio using the Google Cloud Text-to-Speech API.
[1624] Step 7:
[1625] Providing answers to users
[1626] The device then provides the generated voice data to the user through the vehicle's speakers, allowing the user to receive voice responses to their questions in real time.
[1627] Input: Voice data (answer)
[1628] Output: Audio output to the user
[1629] Specific operation: The device routes the generated voice data to the car's speakers and provides it to the user audibly.
[1630] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1631] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[1632] Overall system overview
[1633] The system mainly consists of the following components:
[1634] 1. The device that receives the user's questions
[1635] 2. Server that analyzes questions and searches for and analyzes information
[1636] 3. Device that generates and provides answers to questions
[1637] 4. Emotion engine that recognizes user emotions
[1638] Overview of program processing
[1639] 1. Receive user questions
[1640] The user inputs a question in natural language into the device (e.g., "The sky is getting dark. Will it rain?"). The device receives this question and prepares it to be sent to the server.
[1641] 2. Question Analysis
[1642] The server receives the user's question and analyzes it using natural language processing technology. Specifically, it identifies that the question is about the weather.
[1643] 3. Obtaining location information
[1644] The device acquires the user's current location information (e.g., Chuo-ku, Tokyo) and sends it to the server. With the user's consent, the location information is accurately acquired using GPS or Wi-Fi.
[1645] 4. Emotional Recognition
[1646] At the same time, the server uses an emotion engine to analyze the user's emotion from the question (e.g., "anxiety," "excitement," etc.). The emotion engine analyzes the text data of the question as input and identifies the emotion.
[1647] 5. Search and retrieve social media posts
[1648] The server uses the acquired location information and the results of the question analysis to search and retrieve related posts using the API of the social networking service (SNS). Specifically, it searches for keywords such as "Tokyo West Rain."
[1649] 6. Analysis and Summarization of Posts
[1650] The server analyzes the social media posts it receives and uses generative AI to summarize the situation. Based on the content of the posts, it identifies trends and key keywords (e.g., "Western Tokyo" and "Heavy Rain").
[1651] 7. Adjust your responses based on emotion
[1652] The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling anxious, it will prioritize providing more detailed and reassuring information.
[1653] 8. Answer Generation
[1654] The server takes into account the analysis results and emotional adjustments to generate a textual answer to the user's question. Specifically, it creates a response such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[1655] 9. Submitting and Viewing Your Answers
[1656] The server then sends the generated answer to the device, which then displays the answer to the user, allowing the user to quickly obtain specific, real-time information about their question.
[1657] Specific examples
[1658] Here are some concrete examples:
[1659] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[1660] 1. User: Enters a question into the terminal.
[1661] 2. Terminal: Sends a query to the server.
[1662] 3. Server: Parses the question and identifies it as a weather question.
[1663] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[1664] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[1665] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[1666] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[1667] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[1668] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[1669] 10. Server: Sends the generated answer to the terminal.
[1670] 11. Terminal: Display the answer to the user.
[1671] In this way, the system of the present invention can provide prompt and appropriate information in response to a user's question, and by taking emotions into consideration, can realize a more user-friendly response.
[1672] The processing flow will be explained below.
[1673] Step 1:
[1674] User: Enters a question in natural language into the device (e.g., "The sky is getting dark. Is it going to rain? I'm worried.").
[1675] Step 2:
[1676] Terminal: Receives the question entered by the user. It stores the entered text internally and prepares it to be sent to the server.
[1677] Step 3:
[1678] Device: Obtains the device's location information. Location information is obtained via GPS or Wi-Fi. With the user's consent, the obtained location information is prepared for transmission to the server.
[1679] Step 4:
[1680] Device: Sends the user's question and location information to the server.
[1681] Step 5:
[1682] Server: Receives questions and location information from users. Questions are received as text data, and location information is received as coordinate data.
[1683] Step 6:
[1684] Server: Leverages a natural language processing (NLP) engine to analyze the incoming question, specifically identifying the subject of the question and determining that it is a weather-related question.
[1685] Step 7:
[1686] Server: Generates a search query based on location information and the results of question analysis. For example, it creates a query containing specific keywords such as "Tokyo, western region, rain."
[1687] Step 8:
[1688] Server: Send a search request to the social networking service API using the generated query. For example, send a request in the format "https: / / api.socialnetwork.com / v2 / search?query=Tokyo Western Rain".
[1689] Step 9:
[1690] Server: Receives search results from the social networking service. The results are returned as multiple posts.
[1691] Step 10:
[1692] Server: Analyzes the received post data using generation AI. Extracts key keywords and trends and summarizes the situation. For example, extracts information such as "Heavy rain will fall in western Tokyo."
[1693] Step 11:
[1694] Server: Uses an emotion engine to recognize emotions from the user's question. For example, identify that the question contains the emotion "anxiety."
[1695] Step 12:
[1696] Server: Adjusts the priority of search and analysis results based on the user's emotions recognized by the emotion engine. If the user is feeling anxious, it prioritizes providing more detailed and reassuring information.
[1697] Step 13:
[1698] Server: Generates a written answer to the user's question, taking into account the analysis results and emotional adjustments. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Please rest assured, there are no evacuation notices at this time."
[1699] Step 14:
[1700] Server: Sends the generated answer to the user's device.
[1701] Step 15:
[1702] Terminal: The answer received from the server is formatted for display. Specifically, it is displayed in a format that matches the UI so that it is easy for the user to see.
[1703] Step 16:
[1704] Device: Display the following response to the user: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[1705] Through this process, users can receive prompt and appropriate answers to their questions, and by taking into account the user's emotions, the system can provide a more user-friendly response.
[1706] Example 2
[1707] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1708] In today's world, users want to obtain appropriate information in real time by asking questions in natural language. However, conventional information search systems are limited to providing information based on keywords, and it is difficult to generate answers that take into account the user's situation and emotions. This makes them insufficient to alleviate users' anxieties and questions, and there is a need for more accurate information provision systems.
[1709] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1710] In this invention, the server includes means for receiving a question in natural language from a user, means for analyzing the question using natural language processing technology and identifying the content of the question, means for acquiring the user's location information, means for recognizing the user's emotion from the question text, means for searching and acquiring related posts from an online platform based on the acquired location information and the question analysis result, means for analyzing the acquired posts and summarizing a specific situation, means for generating an answer to the user's question based on the summarized situation, means for adjusting the content of the answer based on the emotion recognition result, and means for providing the generated answer to the user. This makes it possible to provide more appropriate and specific information in real time in response to a natural language question entered by a user, taking into account the user's location information and emotion.
[1711] A "user" is a person using a terminal who enters a question in natural language to obtain information.
[1712] "Natural language" refers to language used in everyday life, sentences and words that do not require special interpretation by a computer program.
[1713] "Terminal" refers to a computing device, such as a smartphone, tablet, or PC, through which a user enters a question or receives a response.
[1714] A "server" is a computer device that receives and processes data sent from a terminal and provides appropriate information.
[1715] "Natural language processing technology" refers to all technologies that enable computers to understand, analyze, and generate natural language.
[1716] "Location Information" means data that indicates a user's geographic location and may include GPS data and Wi-Fi information.
[1717] An "emotion engine" refers to software or algorithms that analyze and identify user emotions based on text data.
[1718] "Online platform" refers to an online service that allows a large number of users to generate and share information and content, including social networking services.
[1719] A "generative AI model" is an artificial intelligence model that generates new text based on training data, such as GPT-4.
[1720] A "prompt" refers to text data that is input to a generative AI model to produce a specific output.
[1721] An "answer" is a sentence generated in response to a user's question based on the analysis results and acquired information.
[1722] This system searches and analyzes related information in real time in response to questions entered by users in natural language, and provides easy-to-understand answers based on the results. Its unique feature is its incorporation of an emotion engine that recognizes the user's emotions, enabling it to provide more appropriate information. This system is realized through interactions between the server, terminals, and users.
[1723] Hardware and Software
[1724] This system mainly uses the following hardware and software:
[1725] Device: A computing device used by a user, such as a smartphone, tablet, or PC.
[1726] Server: A computing device for receiving, processing, and providing data.
[1727] Natural language processing technology: Google Cloud Natural Language API, etc.
[1728] Sentiment engines: such as IBM Watson Tone Analyzer and Microsoft Azure Text Analytics.
[1729] Location information acquisition technology: Google Maps API, built-in GPS module, etc.
[1730] Social networking interfaces: Twitter API, etc.
[1731] Generative AI models: such as OpenAI's GPT-4.
[1732] Data processing and calculation
[1733] The main processing of the system involves the following data processing and data calculation.
[1734] 1. Natural language processing: The device sends the user's question to the server, which then uses natural language processing to analyze the question, for example, identifying that it is a question about the weather.
[1735] 2. Location information acquisition: The device acquires the user's current location information and sends it to the server. This location information is acquired using GPS or Wi-Fi.
[1736] 3. Emotion Recognition: The server uses an emotion engine to analyze and identify emotions from the user's question, thereby determining the emotion (e.g., "anxiety" or "excitement") the user is feeling when asking the question.
[1737] 4. Search and retrieve SNS posts: Based on the acquired location information and the question analysis results, the server uses the SNS interface to search and retrieve related posts. For example, a search can be performed using keywords such as "Tokyo West Rain."
[1738] 5. Post analysis: The server analyzes the social media posts and summarizes them using a generative AI model, identifying trends and key keywords from the content of the posts (e.g., "Western Tokyo" and "Heavy Rain").
[1739] 6. Emotion-based response adjustment: The server adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," the server will prioritize providing more detailed and reassuring information.
[1740] 7. Answer Generation: The server generates a textual answer to the user's question, taking into account the analysis results and sentiment-based adjustments. For example, it creates an answer such as, "There are many posts reporting heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time."
[1741] 8. Providing an answer: The server sends the generated answer to the terminal, and the terminal displays the answer to the user, allowing the user to quickly obtain specific and real-time information about the question.
[1742] Specific examples
[1743] A user types in a question: "The sky is getting dark. Is it going to rain? I'm worried."
[1744] 1. User: Enters a question into the terminal.
[1745] 2. Terminal: Sends a query to the server.
[1746] 3. Server: Parses the question and identifies it as a weather question.
[1747] 4. Device: Obtain the user's location information and send it to the server (e.g., Chuo-ku, Tokyo).
[1748] 5. Server: Uses the emotion engine to identify that the user is feeling “anxiety.”
[1749] 6. Server: Uses the SNS API to search and retrieve related posts (e.g., "Tokyo, Western, Rain").
[1750] 7. Server: Analyzes the retrieved posts and identifies the trends "Western Tokyo" and "Heavy Rain."
[1751] 8. Server: Prioritize providing detailed and reassuring content to users who feel anxious.
[1752] 9. Server: Generate the following response: "There are many posts reporting heavy rain in western Tokyo. It appears to be getting closer as time passes. Rest assured, there are no evacuation notices at this time."
[1753] 10. Server: Sends the generated answer to the terminal.
[1754] 11. Terminal: Display the answer to the user.
[1755] This allows the system of the present invention to provide prompt and appropriate information in response to user questions, and by taking emotions into consideration, it is possible to provide a more user-friendly response.
[1756] Prompt Sentence Examples
[1757] "Analyze questions entered in natural language, combine location information and emotion recognition to search for relevant information, and generate real-time answers. The user's question is 'The sky is getting dark. Will it rain? I'm worried.' The location information is 'Chuo Ward, Tokyo.' The emotion recognition result is 'Anxiety.'"
[1758] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1759] Step 1:
[1760] User: Enters a question into the terminal in natural language (e.g., "The sky is getting dark. Will it rain?"). The input data is a textual question.
[1761] Terminal: Receives the entered question, stores it as text data, and prepares it to be sent to the server. Specifically, when the user enters a question and presses the "Send" button, the terminal sends the question data to the server. The input is the user's text data, and the output is ready to be sent to the server.
[1762] Step 2:
[1763] Server: Analyzes the question data received from the user using natural language processing (NLP) technology. For example, it identifies that the question is about the weather. Specifically, it uses an NLP library (e.g., Google Cloud Natural Language API) to divide the question into tokens and analyzes the sentence structure, such as subject, predicate, and object. The input is the question text data, and the output is the analysis result (identification of the question content).
[1764] Step 3:
[1765] Device: Obtains the user's current location information. This is obtained using GPS or Wi-Fi. The obtained location information is sent to the server. Specifically, the device's location information acquisition function is called, and the current location is obtained as GPS coordinates or Wi-Fi information. The input is a request to obtain location information, and the output is the obtained location information (e.g., Chuo-ku, Tokyo).
[1766] Step 4:
[1767] Server: Analyzes text data (questions) and recognizes the user's emotions using an emotion engine. Specifically, it identifies emotions such as "anxiety" and "excitement." Specific operations involve inputting text data into an emotion engine (e.g., IBM Watson Tone Analyzer or Microsoft Azure Text Analytics) and obtaining an emotion score. The input is the question text data, and the output is the emotion recognition results.
[1768] Step 5:
[1769] Server: Based on the acquired location information and question analysis results, the server uses the online platform's API to search and retrieve related posts. For example, a search is performed using keywords such as "Tokyo West Rain." Specific operations include creating an API request and retrieving related posts. The input is location information and question analysis results, and the output is the retrieved post data.
[1770] Step 6:
[1771] Server: Analyzes the acquired social media post data using a generative AI model (e.g., OpenAI's GPT-4) and summarizes the situation. Identifies trends and key keywords from the content of the post. Specifically, it crawls the acquired post data, analyzes the text content, calculates the frequency distribution of keywords, and has the generative AI model summarize it. The input is the acquired post data, and the output is summarized situation information.
[1772] Step 7:
[1773] Server: Adjusts the priority of analysis results based on the user's emotions identified by the emotion engine. For example, if the user is feeling "anxious," it prioritizes providing more detailed and reassuring information. Specifically, it selects the most appropriate information from the analyzed information set based on the user's emotion score. The input is the emotion recognition result and summarized situation information, and the output is the adjusted answer content.
[1774] Step 8:
[1775] Server: Based on the above information, it generates a textual answer to the user's question. Specifically, it creates an answer such as, "There are many posts about heavy rain in western Tokyo. It seems to be getting closer to us as time passes. Don't worry, there are no evacuation notices at this time." Specific operations include inputting the necessary information as a prompt into the generative AI model and requesting it to generate a text. The generated text is reviewed and any necessary corrections are made. The input is the adjusted answer, and the output is the generated answer text.
[1776] Step 9:
[1777] Server: Sends the generated answer to the device.
[1778] Terminal: Displays the answer to the user. Specifically, the answer data received from the server is displayed on the user interface. The display format can be a pop-up, notification, chat format, etc. The input is the generated answer text, and the output is the answer displayed to the user.
[1779] In this way, the system of the present invention can quickly provide specific, real-time information in response to a user's question.
[1780] (Application example 2)
[1781] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1782] Conventional systems had the problem of making it difficult for users to obtain fast and accurate information in physical stores. In particular, when users had questions about products, there was a lack of a way to quickly and effectively provide detailed information, reviews, and inventory information related to the question in real time. Furthermore, responses did not take into consideration the user's feelings, which could lead to a decline in the quality of the user experience. This made it difficult to improve customer satisfaction.
[1783] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for receiving a question in natural language from a user; means for analyzing the question using natural language processing technology and identifying the question content; emotion recognition means for analyzing the user's emotions; means for acquiring the user's location information; means for searching and acquiring related posts from social networking services based on the acquired location information and the question analysis results; means for analyzing the acquired posts and summarizing a specific situation; means for generating an answer to the user's question based on the summarized situation and the emotion analysis results; and means for providing the generated answer to the user. This allows users to easily obtain quick and accurate information related to products in physical stores and also enables responses that take user emotions into consideration. This improves customer satisfaction.
[1784] "Natural language questions" refer to questions asked by users using everyday words and expressions.
[1785] "Natural language processing technology" is a technology that allows computers to analyze and understand human language.
[1786] The "means for identifying the question content" is a mechanism for analyzing the received question and clarifying what the question is asking.
[1787] "Emotion recognition means" refers to a device or program that analyzes and identifies emotions and moods from text entered by a user.
[1788] "Means of obtaining location information" refers to a mechanism for obtaining the user's current location using technologies such as GPS and Wi-Fi.
[1789] A "social networking service" is an online communication platform that allows users to share information.
[1790] "Means for searching and retrieving related posts" refers to a system for searching and collecting posts that match specific keywords or conditions based on the acquired location information and analysis results.
[1791] "Means for analyzing posts and summarizing specific situations" refers to a system for analyzing collected post content, extracting important information from it, and summarizing it concisely.
[1792] The "means for generating an answer" is a mechanism for creating an appropriate response to a user's question in the form of text based on the analysis results and emotion recognition results.
[1793] The "means for providing an answer" is a device or program for presenting the generated answer to the user in a format that is easy to view.
[1794] The system that realizes this application example is a "smart shopping assistant" system that allows users to input product-related questions in natural language while shopping in a physical store and provides quick and appropriate answers.
[1795] System Program
[1796] This system is implemented using the following hardware and software:
[1797] Hardware
[1798] 1. User terminal: A device such as a smartphone, smart glasses, or head-mounted display.
[1799] 2. Server: A computer server with powerful computing power.
[1800] software
[1801] 1. Natural language processing libraries: Use NLTK (Natural Language Toolkit) and spaCy to analyze natural language questions from users.
[1802] 2. Emotion Recognition Engine: Analyzes user emotions using TextBlob and VADER (Valence Aware Dictionary and sEntiment Reasoner).
[1803] 3. SNS API: Use the API of social networking services (e.g. Twitter API) to search and retrieve related posts.
[1804] 4. Generative AI model: An artificial intelligence model that generates appropriate answers based on the information obtained.
[1805] What the program does
[1806] 1. Receiving Questions
[1807] A user types a question in natural language into a device such as a smartphone or smart glasses, for example, "Is the new model in stock?" This question is received by the device and sent to the server.
[1808] 2. Natural Language Processing
[1809] When the server receives a question, it uses a natural language processing library (e.g., spaCy) to analyze the question and extract information about specific categories and products.
[1810] 3. Emotion recognition
[1811] The server then uses an emotion recognition engine (e.g., TextBlob) to analyze the emotion from the user's question. For example, the emotion "anxiety" is identified from the question "I'm worried about stock availability."
[1812] 4. Obtaining location information
[1813] The user's device uses GPS and Wi-Fi to obtain its current location and sends it to the server, which can then obtain information about the store the user is in or nearby.
[1814] 5. Searching for Information
[1815] The server uses the acquired location information and the question analysis results to search and retrieve related posts using the SNS API, searching the social networking service for keywords such as "in stock" and "new model."
[1816] 6. Analysis and Summarization of Posts
[1817] The server analyzes the posts and uses a generative AI model to summarize important information, such as "latest model in stock" or "highly rated."
[1818] 7. Answer Generation
[1819] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, to a user who is feeling anxious, the server creates a response such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good."
[1820] 8. Providing answers
[1821] The generated answer is sent from the server to the device and displayed to the user, allowing the user to quickly obtain detailed information about the product in the physical store.
[1822] Specific examples
[1823] A user uses a smartphone in a physical store to ask, "Is the new model in stock?" At this time, the system analyzes the user's emotion as "anxiety" and collects the latest stock status and product review information from social media posts. Ultimately, it generates an answer such as, "This product is the latest model. It is currently in stock, so don't worry. The reviews are also very good," and displays it on the user's device.
[1824] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1825] Step 1:
[1826] A user enters a question in natural language using a smartphone or smart glasses in a physical store. The entered question (e.g., "Is the new model in stock?") is received by the user device and sent to the server. In this step, the input is the user's question, and the output is the data that sends the question to the server.
[1827] Step 2:
[1828] The server analyzes the received question using a natural language processing library (e.g., spaCy). This analysis identifies the intent of the question and the target product category. Specifically, the server tokenizes the question, tags it with parts of speech, and performs semantic analysis. The input is the user's natural language question, and the output is the identified question (e.g., "new model" or "inventory").
[1829] Step 3:
[1830] At the same time, the server uses an emotion recognition engine (e.g., TextBlob) to analyze the user's emotion from the question. The emotion recognition engine receives text data as input, analyzes the emotion polarity, and identifies the emotion type (e.g., "anxiety," "relief," or "interest"). The input is the user's question, and the output is the analyzed emotion data (e.g., "anxiety").
[1831] Step 4:
[1832] The user device acquires its current location information using GPS or Wi-Fi and sends that location information to the server. The input is the location data of the user device, and the output is the location information data sent to the server (e.g., "Shinjuku-ku, Tokyo").
[1833] Step 5:
[1834] The server uses the API of the social networking service (SNS) to search and retrieve related posts based on the acquired location information and question analysis results. This search is performed by combining specific keywords and location information. The input is the question analysis results and location information, and the output is related post data (e.g., "SNS post: latest model in stock, many positive reviews").
[1835] Step 6:
[1836] The server analyzes the social media posts and uses a generative AI model to summarize important information. This summarization includes extracting trends and key keywords from the posts. The input is the social media post data, and the output is summarized information (e.g., "Likely Rated, In Stock").
[1837] Step 7:
[1838] The server generates an appropriate answer to the user's question based on the analysis results and emotion recognition results. For example, for a user who is feeling anxious, the server might create a response such as, "This product is the latest model. It is currently in stock, so don't worry!" The input is summarized information and emotion data, and the output is the generated response text data.
[1839] Step 8:
[1840] The server sends the generated answer to the user's terminal, which then displays the answer. The user can view the answer and quickly obtain detailed information about the product. The input is the generated answer data, and the output is the display of the answer on the user's terminal.
[1841] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1842] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1843] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1844] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1845] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1846] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1847] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1848] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1849] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1850] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1851] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1852] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1853] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1854] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1855] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1856] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1857] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1858] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1859] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1860] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1861] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1862] The following is further disclosed regarding the above embodiment.
[1863] (Claim 1)
[1864] means for receiving a natural language question from a user;
[1865] a means for analyzing the question using natural language processing technology and identifying the question content;
[1866] A means for obtaining user location information;
[1867] A means for searching and acquiring related posts from social networking services based on the acquired location information and question analysis results;
[1868] A means of analyzing the retrieved posts and summarizing specific situations;
[1869] means for generating an answer to a user's question based on the summarized situation;
[1870] a means for providing the generated answer to the user;
[1871] A system including:
[1872] (Claim 2)
[1873] It also includes a means to extract key keywords and trends from the analysis results and predict weather changes and specific events.
[1874] 10. The system of claim 1.
[1875] (Claim 3)
[1876] and further including means for retrieving posts in real time from the API of a social networking service based on the user's natural language question and location information.
[1877] 10. The system of claim 1.
[1878] "Example 1"
[1879] (Claim 1)
[1880] means for receiving a natural language question from a user;
[1881] a means for analyzing the question using natural language processing technology and identifying the question content;
[1882] A means for obtaining user location information;
[1883] A means for searching and acquiring related posts from social networking services based on the acquired location information and question analysis results;
[1884] A means of analyzing the captured posts and summarizing specific situations using a generative AI model;
[1885] means for generating an answer to a user's question based on the summarized situation;
[1886] a means for providing the generated answer to the user;
[1887] A system including:
[1888] (Claim 2)
[1889] It also includes a means to extract key keywords and trends from the analysis results and predict weather changes and specific events.
[1890] 10. The system of claim 1.
[1891] (Claim 3)
[1892] and further including means for retrieving posts in real time from the API of a social networking service based on the user's natural language question and location information.
[1893] 10. The system of claim 1.
[1894] "Application Example 1"
[1895] (Claim 1)
[1896] means for receiving a natural language question from a user;
[1897] a means for analyzing the question using natural language processing technology and identifying the question content;
[1898] A means for obtaining user location information;
[1899] A means for searching and acquiring related posts from information providing services based on the acquired location information and question analysis results;
[1900] A means of analyzing the retrieved posts and summarizing specific situations;
[1901] means for generating an answer to a user's question based on the summarized situation;
[1902] a means for providing the generated answer to the user;
[1903] A means of converting voice questions from users into text using voice recognition technology;
[1904] a means for providing the generated answer to the user audibly using speech synthesis technology;
[1905] A system including:
[1906] (Claim 2)
[1907] It also includes a means to extract key keywords and trends from the analysis results and predict weather changes and specific events.
[1908] 10. The system of claim 1.
[1909] (Claim 3)
[1910] and further including means for retrieving posts in real time from an API of an information service based on the user's natural language question and location information.
[1911] 10. The system of claim 1.
[1912] "Example 2: Combining Emotion Engines"
[1913] (Claim 1)
[1914] means for receiving a natural language question from a user;
[1915] a means for analyzing the question using natural language processing technology and identifying the question content;
[1916] A means for obtaining user location information;
[1917] means for recognizing a user's emotion from the question text;
[1918] A means to search and retrieve related posts from online platforms based on the acquired location information and question analysis results;
[1919] A means of analyzing the retrieved posts and summarizing specific situations;
[1920] means for generating an answer to a user's question based on the summarized situation;
[1921] a means for adjusting the content of the response based on the emotion recognition result;
[1922] a means for providing the generated answer to the user;
[1923] A system including:
[1924] (Claim 2)
[1925] It also includes a means to extract key keywords and trends from the analysis results and predict weather changes and specific events.
[1926] 10. The system of claim 1.
[1927] (Claim 3)
[1928] and further comprising means for retrieving posts in real time from an interface of the online platform based on the user's natural language question and location information.
[1929] 10. The system of claim 1.
[1930] "Application example 2 when combining emotion engines"
[1931] (Claim 1)
[1932] means for receiving a natural language question from a user;
[1933] a means for analyzing the question using natural language processing technology and identifying the question content;
[1934] An emotion recognition means for analyzing the emotion of a user;
[1935] A means for obtaining user location information;
[1936] A means for searching and acquiring related posts from social networking services based on the acquired location information and question analysis results;
[1937] A means of analyzing the retrieved posts and summarizing specific situations;
[1938] means for generating an answer to a user's question based on the summarized situation and emotion analysis results;
[1939] a means for providing the generated answer to the user;
[1940] A system including:
[1941] (Claim 2)
[1942] Further, it includes a means for extracting key keywords and trends from the analysis results and providing information for specific situations.
[1943] 10. The system of claim 1.
[1944] (Claim 3)
[1945] and further including means for retrieving posts from the API of a social networking service in real time based on the user's natural language question, location information, and sentiment analysis results.
[1946] 10. The system of claim 1. [Explanation of symbols]
[1947] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving a natural language question from a user; a means for analyzing the question using natural language processing technology and identifying the question content; A means for obtaining user location information; A means for searching and retrieving related posts from social networking services based on the acquired location information and question analysis results; A means of analyzing the retrieved posts and summarizing specific situations; means for generating an answer to a user's question based on the summarized situation; a means for providing the generated answer to the user; A system including:
2. It also includes a means to extract key keywords and trends from the analysis results and predict weather changes and specific events. The system of claim 1 .
3. and further including means for retrieving posts in real time from the API of a social networking service based on the user's natural language question and location information. The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A