system
The system addresses the inefficiencies of conventional information retrieval by allowing a single question input to generate answers and present related links, improving user convenience and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Conventional information retrieval systems require multiple steps and are time-consuming, with users needing to separately find information and collect related links, leading to poor user convenience.
A system that includes input, analysis, generation, search, and output means to allow users to input a single question, analyze it using natural language processing, generate an answer with a generative AI, and simultaneously present related links, simplifying the search process and improving efficiency.
Enables users to quickly obtain necessary information and related links with a single question, significantly simplifying the search process and enhancing user convenience.
Smart Images

Figure 2026064583000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In a conventional information retrieval system, there is a problem that a user has to go through multiple steps to find the information they need, and the retrieval process is complicated and time-consuming. The object of the present invention is to significantly improve the efficiency of information retrieval by allowing a user to input only a single question, providing an appropriate answer using a generative AI, and simultaneously presenting related links.
Means for Solving the Problems
[0005] The present invention solves the above problems by providing a system that includes an input means for inputting a question, an analysis means for analyzing the input question, a generation means for generating an answer based on the analyzed question content, a search means for searching for links related to the generated answer, and an output means for outputting the answer and related links together. Specifically, the system analyzes the question using a natural language processing engine and generates an answer using an artificial intelligence model. Furthermore, by searching for related links from the internet and providing them to the user in an integrated manner, the system simplifies the user's search process and realizes efficient information provision.
[0006] "Input means for entering questions" refers to a device or software that provides an interface for users to enter questions in text format.
[0007] "Analysis means for analyzing input questions" refers to a device or software that analyzes user-input questions using natural language processing technology and extracts keywords and meanings.
[0008] "Generating means for generating answers based on analyzed question content" refers to a device or software that generates appropriate answers using an artificial intelligence model based on the analysis results.
[0009] "A search method for searching for links related to generated answers" refers to a device or software that searches the internet for information related to generated answers and collects it as links.
[0010] "Output means for outputting answers and related links together" refers to a device or software for providing the generated answers and related links to the user as a single integrated package. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, the terms used in the following description will be explained.
[0014] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0017] In the following embodiments, a labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system is implemented as follows:
[0033] User
[0034] The user first enters a question in natural language into the terminal's interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server.
[0035] terminal
[0036] The terminal receives the user's input and sends it to the server as a request. The request includes the user's question, and the server receives this request.
[0037] server
[0038] The server processes the following steps.
[0039] 1. Analysis of the question:
[0040] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0041] 2. Generating the answer:
[0042] The server then calls a generation mechanism based on the analysis results and uses an artificial intelligence model to generate an appropriate answer. For example, it might generate an answer such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0043] 3. Search for related links:
[0044] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0045] 4. Summary and output of results:
[0046] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0047] Specific example
[0048] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0049] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0050] 2. The device sends the question to the server.
[0051] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0052] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the artificial intelligence model.
[0053] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0054] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[0055] 7. The device displays this answer and link to the user.
[0056] This invention allows users to quickly obtain the necessary information with a single question, significantly simplifying the search process and improving efficiency.
[0057] The following describes the processing flow.
[0058] Step 1:
[0059] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[0060] Step 2:
[0061] The terminal retrieves the input. The terminal retrieves the user's input and constructs it as a request.
[0062] Step 3:
[0063] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request that includes the user's question.
[0064] Step 4:
[0065] The server receives the request. The server receives the request from the terminal and prepares to analyze the question.
[0066] Step 5:
[0067] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, the keywords "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0068] Step 6:
[0069] The server sends a question to a generative AI. Based on the analysis results, the server sends a question to the generative AI (e.g., GPT-4®) and has it generate an appropriate answer.
[0070] Step 7:
[0071] A generative AI generates an answer. The generative AI generates the answer, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0072] Step 8:
[0073] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[0074] Step 9:
[0075] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[0076] Step 10:
[0077] The server sends a response to the terminal. The response contains the generated answer and related links.
[0078] Step 11:
[0079] The terminal receives a response. The terminal receives a response from the server.
[0080] Step 12:
[0081] The device displays the results to the user. The device displays the received answer and link to the user, showing the answer and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0082] This allows users to quickly obtain answers to their questions and relevant links.
[0083] (Example 1)
[0084] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0085] Traditional information retrieval systems required users to perform numerous manual operations and multiple searches to efficiently obtain the information they needed, resulting in a cumbersome and time-consuming process. Furthermore, collecting links related to the generated answers required a separate effort, leading to poor user convenience.
[0086] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0087] In this invention, the server includes an input means, a terminal for receiving questions entered by the user through the input means, an analysis means for analyzing the questions transmitted from the terminal, a generation means for generating answers based on the analyzed question content, a search means for searching for links related to the answers generated by the generation means, and an output means for sending the generated answers and related links together to the terminal. As a result, the user can quickly obtain the necessary information by simply entering a single question, and related links are also provided at the same time, enabling efficient information retrieval.
[0088] "Input method" refers to the interface through which a user enters a question into the system.
[0089] A "terminal" refers to a device used by a user to input a question and send it to a server for processing.
[0090] "Analysis means" refers to a function that analyzes questions sent by users through their devices and extracts their content and key keywords.
[0091] "Generation means" refers to a function that generates appropriate answers based on keywords and content extracted by the analysis means.
[0092] "Search means" refers to a function for searching the internet for and collecting links related to the answers generated by the generation means.
[0093] "Output means" refers to a function that aggregates the generated responses and collected related links, sends them to the terminal, and displays them to the user.
[0094] A "natural language processing engine" refers to software or algorithms that analyze an input question and extract key keywords and content.
[0095] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate appropriate answers to user questions.
[0096] An "HTTP request" refers to a communication protocol used to send data to a web server.
[0097] "HTTP response" refers to the communication protocol used to send back data received from a web server.
[0098] A "prompt sentence" refers to a sentence used as an instruction to input into a generative AI model.
[0099] This invention is an information retrieval system that provides an appropriate answer using a generative AI model and simultaneously presents related links, based on the user's input of a single question. Specific embodiments of this system are described below.
[0100] Hardware and software configuration
[0101] This system consists of the following main components:
[0102] Device: A device used by the user to enter questions. This includes PCs, smartphones, tablets, etc.
[0103] Server: The central processing unit that analyzes questions, generates answers, searches for relevant links, and provides responses to the user. Software running on the server includes natural language processing engines, generative AI models, and search engine APIs.
[0104] Detailed processing of the system
[0105] The server performs the following steps:
[0106] 1. Receiving and sending questions:
[0107] The user enters a question into the terminal's interface, such as, "What are some recommended ramen restaurants in Tokyo?"
[0108] The device retrieves this question and sends it to the server as an HTTP POST request.
[0109] 2. Analysis of the question:
[0110] The server processes the received question using parsing tools. Specifically, it tokenizes the question using a natural language processing engine (e.g., Spacy, NLTK) and extracts key keywords and meanings.
[0111] For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" are extracted.
[0112] 3. Generating the answer:
[0113] The server invokes a generation mechanism based on the analysis results and generates an appropriate response using a generation AI model (e.g., GPT-4, GPT-3®).
[0114] For example, the system might generate responses such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0115] 4. Search for related links:
[0116] The server searches for relevant links based on the generated response. It collects the appropriate links from the internet using search methods (e.g., Google® Search API, Bing Search API).
[0117] For example, we collect links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0118] 5. Generating and sending results:
[0119] The server combines the generated answers and collected links to create the final response.
[0120] For example, a response in the format of "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." is sent to the terminal as an HTTP response.
[0121] 6. Displaying the results:
[0122] The terminal receives a response from the server and displays it in the user interface.
[0123] Through this, users can view the answers and related links.
[0124] Examples of specific actions
[0125] For example, if a user enters "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0126] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0127] 2. The device sends the question to the server.
[0128] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0129] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the generation AI model.
[0130] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0131] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[0132] 7. The device displays this answer and link to the user.
[0133] Example of a prompt
[0134] Here are some specific examples of prompt statements to input into a generative AI model:
[0135] "What are some recommended ramen restaurants in Tokyo?"
[0136] "What are some good tourist spots in Kyoto?"
[0137] This allows users to quickly obtain the information they need with a single question, enabling efficient information retrieval.
[0138] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0139] Step 1: User enters question
[0140] The user enters a question in natural language into the terminal interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" into the text field. The terminal retrieves the user's input (question text) so that it can be processed in the next step.
[0141] Input: User's question (natural language text)
[0142] Output: Retrieving questions via the terminal
[0143] Specific operation: The user enters a question into the terminal's text field. The terminal stores this input in memory and prepares to send it in the next step.
[0144] Step 2: Sending a request via the terminal
[0145] The terminal sends the question received from the user to the server as an HTTP POST request. The request contains the question text.
[0146] Input: User's question (text stored on the device)
[0147] Output: HTTP POST request to the server
[0148] Specific operation: The terminal includes the retrieved question text in the payload of an HTTP POST request and sends it to a specific endpoint. The server receives this request.
[0149] Step 3: Server analyzes the question
[0150] The server uses a natural language processing engine (e.g., Spacy, NLTK) to analyze the received question text. This analysis extracts key keywords and meanings from the text.
[0151] Input: User's question (payload of HTTP POST request)
[0152] Output: Analyzed keywords and meanings
[0153] Specific operation: The server extracts the question text from the request payload and passes it to the natural language processing engine. The engine tokenizes the text and extracts keywords such as "Tokyo," "recommended," and "ramen restaurant."
[0154] Step 4: Server generates response
[0155] The server generates appropriate answers using a generative AI model (e.g., GPT-4, GPT-3) based on the analysis results. A prompt is input to the generative AI model, and the answer text is retrieved based on that input.
[0156] Input: Analyzed keywords and meanings
[0157] Output: Generated answer (text)
[0158] Specific operation: The server uses the analyzed keywords to create a prompt for the generative AI model and inputs a question such as "What are some recommended ramen restaurants in Tokyo?" into the model. The generative AI model responds by outputting text such as "Some recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0159] Step 5: Server searches for related links
[0160] The server searches the internet for relevant links based on the generated response. Links are collected using search methods (e.g., Google Search API, Bing Search API).
[0161] Input: Generated response (text)
[0162] Output: List of related links
[0163] Specific operation: The server extracts keywords from the generated response text and creates search queries such as "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". These queries are then fed into a search API to retrieve relevant links (e.g., official website, review site).
[0164] Step 6: Server summarizes and sends the results.
[0165] The server combines the generated answers and collected links to create a final response. This response is then sent to the terminal as an HTTP response.
[0166] Input: List of generated answers (text) and related links
[0167] Output: Final response (text and links)
[0168] Specific operation: The server combines the answer text and related links to prepare a response in the format, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." This is then sent to the terminal as an HTTP response.
[0169] Step 7: Display to the user via the device
[0170] The terminal displays the response received from the server in the user interface. Through this, the user can view the answer and related links.
[0171] Input: Final response from the server (text and links)
[0172] Output: Content displayed to the user
[0173] Specific operation: The terminal analyzes the received response and extracts the answer and link. These are then displayed using HTML or GUI components, making them viewable by the user.
[0174] As described above, users can quickly obtain the necessary information with a single question, and relevant links are provided simultaneously, enabling efficient information retrieval.
[0175] (Application Example 1)
[0176] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0177] Traditional information retrieval systems had the problem of requiring a lot of effort and time for users to find the right product even after entering a product name or specific characteristics. Furthermore, the limited functionality for displaying related links and product information in a single batch meant that user search efficiency was not improved.
[0178] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0179] In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question, a generation means for generating an answer based on the analyzed question content, a search means for searching for links related to the generated answer, an output means for outputting the answer and related links together, and a product link display means for displaying the corresponding product link along with the generated answer. This makes it possible to quickly display appropriate product suggestions and links to related products when a user inputs a question related to a product.
[0180] "Input means for entering a question" refers to a device or interface for a user to enter a question in natural language.
[0181] "Analysis means for analyzing input questions" refers to means that analyze input questions and use natural language processing techniques to extract key keywords and meanings.
[0182] "A means for generating answers based on analyzed question content" refers to a means that utilizes an artificial intelligence model to generate appropriate answers based on information extracted by the analysis means.
[0183] "A search method for finding links related to generated answers" refers to a method for searching the internet for information related to generated answers and collecting the relevant links.
[0184] "An output method for outputting answers and related links together" refers to a method for integrating the generated answers and collected related links and displaying them to the user.
[0185] "A means for displaying product links along with the generated response" refers to a means for additionally displaying links to products related to the generated response.
[0186] A "natural language processing engine" is a software engine that analyzes input natural language text and extracts its meaning and keywords.
[0187] An "artificial intelligence model" is a machine learning model that uses input data to perform reasoning like a human and generate appropriate answers.
[0188] This invention is an information retrieval system that can be applied to an e-commerce site application that provides appropriate answers and links to related products when a user enters a question related to a product.
[0189] System Configuration
[0190] 1. User Interface
[0191] The application includes an input method for users to enter questions in natural language using a smartphone application. For example, it would be a section where users can enter questions such as, "What smartphone do you recommend?"
[0192] 2. Analysis of the Question
[0193] The terminal retrieves the entered question and sends it to the server. The server uses a natural language processing engine (e.g., SpaCy) to analyze the entered question, extracting key keywords and meaning from it.
[0194] 3. Generating the answer
[0195] The server generates appropriate answers to questions using artificial intelligence models (e.g., OpenAI®'s GPT-4) based on the analyzed question content. This generation mechanism has the ability to perform real-time reasoning in response to user input and provide relevant information.
[0196] 4. Search for related links
[0197] The server includes a search mechanism for searching the internet for relevant product links based on the generated response. This search mechanism uses, for example, the BeautifulSoup or requests library to retrieve links to relevant products.
[0198] 5. Output of answers and links
[0199] The server has an output mechanism to combine the generated answer and related product links into a single response and send it to the terminal. The terminal displays this response to the user. This allows the user to view the answer to the question along with detailed links to related products all at once.
[0200] Hardware and software used
[0201] Hardware: Smartphone
[0202] software:
[0203] OpenAI API: GPT model for question analysis and answer generation
[0204] requests library: Used to retrieve links from websites.
[0205] BeautifulSoup: To parse links from retrieved HTML.
[0206] Natural language processing engines: SpaCy, etc.
[0207] Specific example
[0208] For example, if a user types the question "Can you recommend a smartphone?", the following steps will be taken:
[0209] 1. The user enters the question "What smartphone do you recommend?" into their device.
[0210] 2. The device sends the question to the server.
[0211] 3. The server analyzes the question and extracts the main keywords "recommendation" and "smartphone".
[0212] 4. Based on the analysis results, the server uses a generation method to generate the response "The latest iPhone (registered trademark) or Samsung Galaxy is recommended" from the artificial intelligence model.
[0213] 5. The server searches for links related to "iPhone" and "Samsung Galaxy".
[0214] 6. The server compiles the answers and links and sends them to the device in the format, "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details."
[0215] 7. The device displays this answer and link to the user.
[0216] Examples of prompts to input into a generative AI model:
[0217] Question: What smartphone do you recommend?
[0218] answer:
[0219] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0220] Step 1:
[0221] The user enters the question in natural language.
[0222] Input: The user enters the question "What smartphone do you recommend?" in a smartphone application.
[0223] Output: The entered question is sent to the terminal.
[0224] Specific action: The user enters a question into the app's interface and taps the submit button.
[0225] Step 2:
[0226] The terminal sends the entered question to the server.
[0227] Input: A natural language question entered by the user.
[0228] Output: The question content is sent to the server.
[0229] Specific operation: The terminal receives user input and forwards the question to the server as an HTTP request.
[0230] Step 3:
[0231] The server performs natural language processing on the question using parsing tools.
[0232] Input: A natural language question sent from the device.
[0233] Output: Analyzed keywords and their meanings.
[0234] Specific operation: The server uses a natural language processing engine (e.g., SpaCy) to analyze the question and extract keywords such as "recommended" and "smartphone".
[0235] Step 4:
[0236] The server generates an answer based on the analysis results.
[0237] Input: Keywords and meanings of the analyzed question.
[0238] Output: The generated answer.
[0239] Specific operation: The server inputs the analysis results into an artificial intelligence model (e.g., OpenAI's GPT-4) and generates responses such as, "We recommend the latest iPhone or Samsung Galaxy."
[0240] Step 5:
[0241] The server searches for relevant links based on the generated response.
[0242] Input: The content of the generated response.
[0243] Output: Related product links.
[0244] Specific operation: The server uses web scraping libraries (e.g., requests and BeautifulSoup) to collect product links related to "iPhone" and "Samsung Galaxy" from the internet.
[0245] Step 6:
[0246] The server outputs the answer and the link together.
[0247] Input: Generated responses and collected product links.
[0248] Output: The final response to be displayed to the user.
[0249] Specific operation: The server combines the answer text and each link into a single response and sends it back to the terminal as an HTTP response.
[0250] Step 7:
[0251] The device displays the response and link from the server.
[0252] Input: The final response sent from the server.
[0253] Output: Answers and links that users can view.
[0254] Specific action: The response received by the device is displayed on the user interface, and the message "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details." is presented.
[0255] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0256] This invention combines an information retrieval system that provides appropriate answers using generative AI and simultaneously presents related links based on a single question entered by the user, with an emotion engine that recognizes the user's emotions. This system is implemented in the following specific forms.
[0257] User
[0258] The user first inputs a question in natural language into the terminal's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server. It also recognizes emotions using the user's input and voice interface.
[0259] terminal
[0260] The terminal receives user input and sends it to the server as a request. The request includes the user's question and sentiment data, and the server receives this request.
[0261] server
[0262] The server processes the following steps.
[0263] 1. Analysis of the question:
[0264] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0265] 2. Recognition of emotions:
[0266] The server uses emotion recognition mechanisms to analyze the user's emotional data and recognize the user's emotional state. For example, emotions such as "joy," "sadness," and "surprise" can be identified.
[0267] 3. Generating the answer:
[0268] The server invokes a generation mechanism based on the analysis results and recognized emotions, and uses an artificial intelligence model to generate an appropriate response. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated with a tone and content that matches the user's emotions.
[0269] 4. Search for related links:
[0270] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0271] 5. Summary and output of results:
[0272] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0273] Specific example
[0274] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0275] 1. The user enters the question "What are some good tourist spots in Kyoto?" into the device. The user's input also detects the emotion "fun".
[0276] 2. The device sends the question and sentiment data to the server.
[0277] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0278] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[0279] 5. Based on the analysis results and sentiment data, the server uses a generation method to generate a sentiment-appropriate response from an artificial intelligence model, such as "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[0280] 6. The server searches for links related to the associated information "Kinkaku-ji", "Kiyomizu-dera", and "Arashiyama".
[0281] 7. The server summarizes the answers and links and sends them to the terminal in the form of "The recommended tourist attractions in Kyoto are Kinkaku-ji, Kiyomizu-dera, and Arashiyama. I think you will have a great time! Please refer to the following links for details."
[0282] 8. The terminal displays this answer and the links to the user.
[0283] According to the present invention, the user can not only quickly obtain the necessary information with a single question, but also get an answer that empathizes with the emotions, making the search process more user-friendly.
[0284] The following describes the processing flow.
[0285] Step 1:
[0286] The user enters a question. The user enters "What are the recommended ramen shops in Tokyo?" in the input field of the terminal.
[0287] Step 2:
[0288] The terminal obtains the input content. The terminal obtains the user's input content and, in addition, collects the emotion data obtained from the voice interface and constructs a request.
[0289] Step 3:
[0290] The terminal sends the request to the server. The terminal establishes a connection to the server and sends a request including the user's question and emotion data to the server.
[0291] Step 4:
[0292] The server receives the request. The server receives the request from the terminal and prepares to analyze the question content and sentiment data.
[0293] Step 5:
[0294] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0295] Step 6:
[0296] The server analyzes emotional data. Using emotion recognition tools, the server analyzes emotions from user input and voice data to recognize the user's emotional state. For example, emotions such as "excitement" and "anticipation" may be identified.
[0297] Step 7:
[0298] The server sends a question to a generative AI. Based on the analysis results and recognized emotions, the server sends the question to the generative AI (e.g., GPT-4) to generate an appropriate answer that is sensitive to those emotions.
[0299] Step 8:
[0300] A generative AI generates the answer. The generative AI generates an answer that matches the emotion, such as, "For ramen restaurants in Tokyo, I recommend Ichiran and Tsukemen Daioh. You're sure to have a wonderful time!"
[0301] Step 9:
[0302] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[0303] Step 10:
[0304] The server summarizes the results. The server combines the generated answer and the retrieved links to create the final response.
[0305] Step 11:
[0306] The server sends the response to the terminal. The response includes the generated answer and the related links.
[0307] Step 12:
[0308] The terminal receives the response. The terminal receives the response from the server.
[0309] Step 13:
[0310] The terminal displays the results to the user. The terminal displays the received answer and links to the user, and displays an answer and links such as "Popular ramen shops in Tokyo include Ichiran and Tsukemen Taikou. You're sure to have a great time! Please refer to the following links for details."
[0311] As a result, the user can not only quickly obtain the answer to the question and the related links, but also receive an answer that takes into account the emotions, making the search experience more user-friendly and increasing the satisfaction.
[0312] (Example 2)
[0313] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".
[0314] Traditional information retrieval systems required users to go through multiple steps to obtain appropriate answers after entering a question, resulting in a complex user experience. Furthermore, they lacked the ability to recognize and respond to user emotions, failing to adequately enhance user satisfaction. Therefore, there is a need for a system that provides appropriate answers and related information with a single question, and even provides answers that are sensitive to the user's emotions.
[0315] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0316] In this invention, the server includes an input means, a means for receiving questions and sentiment data transmitted from a terminal, an analysis means using a natural language processing engine for analyzing the input questions, an analysis means using an sentiment recognition engine for analyzing the sentiment data, a generation means using an artificial intelligence model for generating answers based on the analyzed question content and sentiment data, a search means using a search engine for searching the internet for links related to the generated answers, and an output means for sending the answers and related links together to the terminal. As a result, the user can obtain quick and appropriate answers and related information with a single question, and furthermore, answers that are sensitive to the user's emotions can be provided.
[0317] "Input method" refers to a device or software that provides an interface for users to input questions.
[0318] "Receiving means" refers to the function used to retrieve questions and sentiment data sent from a terminal on the server side.
[0319] A "natural language processing engine" refers to software or algorithms that analyze input natural language questions and extract their main keywords and meanings.
[0320] "Analysis means" refers to methods and techniques for analyzing received data and extracting necessary information.
[0321] An "emotion recognition engine" refers to software or algorithms that analyze and identify emotional states from user input data.
[0322] "Generation means" refers to functions and technologies that use artificial intelligence models to generate answers based on analyzed question content and sentiment data.
[0323] An "artificial intelligence model" refers to a system that has been trained using technologies such as machine learning and deep learning, and is capable of generating appropriate answers to questions.
[0324] A "search engine" refers to software or services used to search the internet for and collect links related to generated answers.
[0325] "Search methods" refer to methods and techniques for collecting relevant information.
[0326] "Output means" refers to methods and technologies for compiling generated answers and related links and providing them to the user.
[0327] A "terminal" refers to a device used by a user to input questions or view information received from a server.
[0328] A "question" refers to the content entered by the user in natural language regarding the information they want to know.
[0329] "Emotional data" refers to information indicating the emotional state, extracted from user input and voice.
[0330] This invention combines an information retrieval system that provides an appropriate answer using generative artificial intelligence (AI) and simultaneously presents related links based on the user's input of a single question, with an emotion engine that recognizes the user's emotions. A specific embodiment of this system is described below.
[0331] User actions
[0332] The user first inputs a question in natural language through the device's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" At this time, the device acquires the user's input and uses the voice interface to recognize the user's emotions. For example, the tone of voice and facial expressions while the user is inputting may be used to detect that the user is "happy."
[0333] Terminal operation
[0334] The terminal combines the user's input question and recognized sentiment data, and sends this to the server. Here, it plays the role of sending the question content and sentiment data as a single request to the server. Specific devices used include personal computers, smartphones, and tablets.
[0335] Server Processing
[0336] The server takes several steps to process an incoming request. First, it extracts key keywords and meanings by analyzing the question using a natural language processing engine (for example, the Google Cloud Natural Language API). For example, from the question "What are some recommended ramen restaurants in Tokyo?", the keywords extracted would be "Tokyo," "recommended," and "ramen restaurant."
[0337] Next, an emotion recognition engine (for example, IBM Watson® Tone Analyzer) is used to analyze the user's emotional data. For instance, emotions such as "happy" are identified from the user's voice tone and text.
[0338] Based on the analysis results, the server invokes a generation mechanism and uses an artificial intelligence model (e.g., OpenAI's GPT-3) to generate an appropriate response. This generated response is delivered in a content and tone that matches the user's emotions. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated in a way that resonates with the user's "happy" feelings.
[0339] Based on the generated response, the server uses search methods to find relevant links. Specifically, it uses a search engine (for example, the Google Search API) to collect links to relevant official websites and review sites. For example, it searches for links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0340] Finally, the server combines the generated answers and collected links to create a response to send to the terminal. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0341] Specific example
[0342] For example, consider a case where a user enters the question, "What are some good tourist spots in Kyoto?" The emotion "fun" is detected simultaneously with the user's input.
[0343] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0344] 2. The device sends the question and sentiment data to the server.
[0345] 3. The server analyzes the question and extracts the keywords "Kyoto," "tourist attractions," and "recommendations."
[0346] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[0347] 5. Based on the analysis results and sentiment data, the server uses an artificial intelligence model to generate the response, "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[0348] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0349] 7. The server compiles the answers and links and sends them to the device.
[0350] 8. The device displays this answer and link to the user.
[0351] Example of a prompt:
[0352] "What are some good tourist spots in Kyoto? I'm in a good mood."
[0353] This invention enables users to quickly obtain appropriate answers and related information with a single question, and further provides a search experience that is sensitive to the user's emotions.
[0354] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0355] Step 1:
[0356] The user enters a question into the terminal. Specifically, the user uses text input or voice input to enter a question into the interface, for example, "What are some recommended ramen restaurants in Tokyo?" The input is the text data of the question.
[0357] Step 2:
[0358] The device receives user input text and also acquires emotion data. In the case of voice input, it uses a speech recognition engine to convert it to text, and at the same time, an emotion recognition engine analyzes the emotion. For example, it might acquire the emotion "happy" as text data. The output consists of the converted text data and emotion data.
[0359] Step 3:
[0360] The device sends the question content and sentiment data to the server. The data sent includes the question text and request data containing sentiment data. Specifically, the request is sent to the server using the HTTP protocol.
[0361] Step 4:
[0362] The server receives the request data. The server analyzes the received question text and extracts key keywords. Specifically, it uses a natural language processing engine (e.g., Google Cloud Natural Language API) to extract keywords such as "Tokyo," "recommended," and "ramen restaurant" from the question text. The output is the extracted keywords.
[0363] Step 5:
[0364] The server analyzes emotional data to identify the user's emotions. Specifically, it uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the emotional data. The output is the user's emotional state. For example, it might retrieve data indicating "happy."
[0365] Step 6:
[0366] The server generates a response based on the analysis results (extracted keywords and sentiment data). An artificial intelligence model (e.g., OpenAI's GPT-3) is used as the generation method. The input consists of extracted keywords and sentiment data. For example, it might generate the response text: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The output is the generated response text.
[0367] Step 7:
[0368] The server searches for relevant links based on the generated response. It uses a search engine (e.g., Google Search API) to perform a web search using words included in the generated response. For example, it collects links for "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". The output is a list of relevant links.
[0369] Step 8:
[0370] The server combines the generated answers and collected links to create the final response. Specifically, it combines the answer text and the list of links into a single response data. For example, it might create a response in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." The output is the combined response data.
[0371] Step 9:
[0372] The server sends the final response data to the terminal. It sends the response data back to the terminal using a specific communication protocol. The output indicates that the transmission of the response data is complete.
[0373] Step 10:
[0374] The terminal displays response data obtained from the server to the user. Specifically, it either displays the response on the screen interface or reads it aloud. For example, it might display something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen! For more details, please refer to the link below." The output consists of the response displayed to the user and the link.
[0375] This specific processing flow allows users to quickly obtain appropriate answers and related information with a single question, and furthermore, to receive emotionally resonant responses.
[0376] (Application Example 2)
[0377] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0378] Traditional information retrieval systems provide answers to user questions, but they lack the ability to generate answers that take user emotions into consideration. Furthermore, their ability to provide related links to the generated answers is limited. As a result, users often experience low satisfaction when obtaining answers, and the search experience remains unimproved.
[0379] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question and the user's emotions, a generation means for generating an answer based on the analyzed question content and emotion data, a search means for searching for links related to the generated answer, and an output means for outputting the answer and related links together. This makes it possible to generate an answer that is sensitive to the user's emotions and to provide related information quickly.
[0380] A "question" is the content that users input in natural language to find out what they want to know.
[0381] "Input means" refers to the devices or software that users use to input questions.
[0382] "Analysis means" refers to a device or program for analyzing input questions and user sentiment data.
[0383] "Generation means" refers to a device or program that generates answers based on analyzed question content and sentiment data.
[0384] "Search means" refers to a device or program for searching the internet for links related to the generated answer.
[0385] "Output means" refers to a device or program for displaying answers and related links to the user.
[0386] "Emotional data" refers to emotional information extracted from user input.
[0387] An "artificial intelligence model" is a model trained based on machine learning or deep learning, and is used as a means of generation.
[0388] A "natural language processing engine" is a program that analyzes text input in natural language and extracts meaning and keywords.
[0389] "Related links" refer to URLs of web pages or information that are relevant to the generated response.
[0390] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system incorporates an emotion engine that recognizes the user's emotions and generates answers in a tone and content that corresponds to the user's emotions.
[0391] System Configuration
[0392] This system consists of the following main components:
[0393] 1. Input method: This refers to the means by which the user inputs questions, and includes devices such as smartphones and smart glasses.
[0394] 2. Analysis means: This refers to means for analyzing the input questions and sentiment data, and corresponds to natural language processing engines (e.g., Transformers, spaCy) or sentiment recognition libraries (e.g., emote).
[0395] 3. Generation means: This refers to means for generating answers based on the analyzed question content and sentiment data, and a generative AI model (e.g., GPT-3) falls under this category.
[0396] 4. Search methods: These are methods for searching the internet for links related to the generated answers, and web scraping libraries (BeautifulSoup, Scrapy) are examples of this.
[0397] 5. Output means: This refers to a means of displaying the answer and related links to the user, and this includes the device's display and notification function.
[0398] Processing flow
[0399] As a concrete example, the "emotion-recognition store navigator" application using smart glasses operates in the following steps.
[0400] 1. User input: The user uses smart glasses to input a question in natural language. For example, they might ask, "What product would suit my current mood?"
[0401] 2. Emotion Recognition: The device receives user input and analyzes emotions using an emotion recognition library (such as emote).
[0402] 3. Request to the server: Send the analyzed question and sentiment data to the server.
[0403] 4. Question Analysis: The server uses a natural language processing engine (such as Transformers or spaCy) to analyze the question and extract key keywords.
[0404] 5. Response generation: Use a generative AI model (such as GPT-3) to generate appropriate responses based on analysis results and sentiment data.
[0405] 6. Search for related links: Based on the generated answers, use a web scraping library (such as BeautifulSoup or Scrapy) to search the internet for related links.
[0406] 7. Displaying Results: The answers and related links are compiled and displayed on the smart glasses' screen.
[0407] Examples of specific cases and prompt statements
[0408] As a concrete example, if a user feels tired and asks, "What product would be perfect for how I'm feeling right now?", the following process would occur:
[0409] Input text: "Please recommend a product that perfectly matches my current mood."
[0410] Emotion recognition results: "Fatigue" and "Stress"
[0411] Server output: "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details."
[0412] Related links: "Aroma Candle Details Page", "Massage Chair Purchase Page"
[0413] Examples of prompt messages are as follows:
[0414] "What are some relaxing products suitable for relieving fatigue?"
[0415] "What are some good relaxation items for when you're tired?"
[0416] This system allows users to not only get quick answers to their questions, but also experience friendly and personalized responses tailored to their own emotions.
[0417] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0418] Step 1: User Input
[0419] The user uses smart glasses to input a question in natural language. The input might be something like, "Tell me what product would suit my current mood." The user asks the question via voice or text input. At this time, the smart glasses interface receives the question and retrieves the data.
[0420] Step 2: Recognizing Emotions
[0421] The device receives user input and analyzes emotions using an emotion recognition library (e.g., Emote). In this step, emotion data is extracted from the input text to recognize emotional states such as "fatigue" or "stress." The input data is in text format, and the output data consists of emotion labels and emotion scores.
[0422] Step 3: Request to the server
[0423] The terminal sends the analyzed question and sentiment data to the server. Specifically, it sends the question text and sentiment data, converted to JSON format, to the server via an HTTP POST request. The input data is in JSON format, and the output data is the response from the server.
[0424] Step 4: Analyzing the Question
[0425] The server uses a natural language processing engine (e.g., Transformers, spaCy) to analyze the received question and extract key keywords. For example, from the question "Tell me a product that suits my current mood," it extracts the keywords "mood" and "product." The input data is the question text in JSON format, and the output data is the analyzed keywords.
[0426] Step 5: Generating the answer
[0427] The server generates an appropriate answer using a generative AI model (e.g., GPT-3) based on the analysis results of the question and sentiment data. In this step, the generative AI model generates an answer based on the input data (analyzed keywords and sentiment data). For example, an answer such as "Relaxing aromatherapy candles and massage chairs are recommended" might be generated. The input data consists of keywords and sentiment data, and the output data is the generated answer text.
[0428] Step 6: Search for related links
[0429] The server uses a web scraping library (e.g., BeautifulSoup, Scrapy) to search the internet for relevant links based on the generated response. For example, it collects relevant links such as "details page for aromatherapy candles" and "purchase page for massage chairs." The input data is the generated response text, and the output data is a list of relevant links.
[0430] Step 7: Displaying the results
[0431] The server compiles the answers and related links and sends them to the device. The device displays the received answers and links to the user. At this time, the smart glasses display shows a message along with a link that reads, "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details." The input data is the response from the server, and the output data is information that the user can visually confirm.
[0432] As described above, a series of processes, from user questioning and sentiment recognition to natural language processing, the use of generative AI models, the search for related links, and output, are performed, making it possible to provide users with useful information that resonates with their emotions.
[0433] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0434] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0435] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0436] [Second Embodiment]
[0437] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0438] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0439] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0440] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0441] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0442] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0443] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0444] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0445] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0446] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0447] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0448] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0449] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system is implemented as follows:
[0450] User
[0451] The user first enters a question in natural language into the terminal's interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server.
[0452] terminal
[0453] The terminal receives the user's input and sends it to the server as a request. The request includes the user's question, and the server receives this request.
[0454] server
[0455] The server processes the following steps.
[0456] 1. Analysis of the question:
[0457] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0458] 2. Generating the answer:
[0459] The server then calls a generation mechanism based on the analysis results and uses an artificial intelligence model to generate an appropriate answer. For example, it might generate an answer such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0460] 3. Search for related links:
[0461] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0462] 4. Summary and output of results:
[0463] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0464] Specific example
[0465] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0466] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0467] 2. The device sends the question to the server.
[0468] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0469] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the artificial intelligence model.
[0470] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0471] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[0472] 7. The device displays this answer and link to the user.
[0473] This invention allows users to quickly obtain the necessary information with a single question, significantly simplifying the search process and improving efficiency.
[0474] The following describes the processing flow.
[0475] Step 1:
[0476] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[0477] Step 2:
[0478] The terminal retrieves the input. The terminal retrieves the user's input and constructs it as a request.
[0479] Step 3:
[0480] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request that includes the user's question.
[0481] Step 4:
[0482] The server receives the request. The server receives the request from the terminal and prepares to analyze the question.
[0483] Step 5:
[0484] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, the keywords "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0485] Step 6:
[0486] The server sends a question to a generative AI. Based on the analysis results, the server sends another question to the generative AI (e.g., GPT-4) to generate an appropriate answer.
[0487] Step 7:
[0488] A generative AI generates an answer. The generative AI generates the answer, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0489] Step 8:
[0490] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[0491] Step 9:
[0492] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[0493] Step 10:
[0494] The server sends a response to the terminal. The response contains the generated answer and related links.
[0495] Step 11:
[0496] The terminal receives a response. The terminal receives a response from the server.
[0497] Step 12:
[0498] The device displays the results to the user. The device displays the received answer and link to the user, showing the answer and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0499] This allows users to quickly obtain answers to their questions and relevant links.
[0500] (Example 1)
[0501] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0502] Traditional information retrieval systems required users to perform numerous manual operations and multiple searches to efficiently obtain the information they needed, resulting in a cumbersome and time-consuming process. Furthermore, collecting links related to the generated answers required a separate effort, leading to poor user convenience.
[0503] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0504] In this invention, the server includes an input means, a terminal for receiving questions entered by the user through the input means, an analysis means for analyzing the questions transmitted from the terminal, a generation means for generating answers based on the analyzed question content, a search means for searching for links related to the answers generated by the generation means, and an output means for sending the generated answers and related links together to the terminal. As a result, the user can quickly obtain the necessary information by simply entering a single question, and related links are also provided at the same time, enabling efficient information retrieval.
[0505] "Input method" refers to the interface through which a user enters a question into the system.
[0506] A "terminal" refers to a device used by a user to input a question and send it to a server for processing.
[0507] "Analysis means" refers to a function that analyzes questions sent by users through their devices and extracts their content and key keywords.
[0508] "Generation means" refers to a function that generates appropriate answers based on keywords and content extracted by the analysis means.
[0509] "Search means" refers to a function for searching the internet for and collecting links related to the answers generated by the generation means.
[0510] "Output means" refers to a function that aggregates the generated responses and collected related links, sends them to the terminal, and displays them to the user.
[0511] A "natural language processing engine" refers to software or algorithms that analyze an input question and extract key keywords and content.
[0512] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate appropriate answers to user questions.
[0513] An "HTTP request" refers to a communication protocol used to send data to a web server.
[0514] "HTTP response" refers to the communication protocol used to send back data received from a web server.
[0515] A "prompt sentence" refers to a sentence used as an instruction to input into a generative AI model.
[0516] This invention is an information retrieval system that provides an appropriate answer using a generative AI model and simultaneously presents related links, based on the user's input of a single question. Specific embodiments of this system are described below.
[0517] Hardware and software configuration
[0518] This system consists of the following main components:
[0519] Device: A device used by the user to enter questions. This includes PCs, smartphones, tablets, etc.
[0520] Server: The central processing unit that analyzes questions, generates answers, searches for relevant links, and provides responses to the user. Software running on the server includes natural language processing engines, generative AI models, and search engine APIs.
[0521] Detailed processing of the system
[0522] The server performs the following steps:
[0523] 1. Receiving and sending questions:
[0524] The user enters a question into the terminal's interface, such as, "What are some recommended ramen restaurants in Tokyo?"
[0525] The device retrieves this question and sends it to the server as an HTTP POST request.
[0526] 2. Analysis of the question:
[0527] The server processes the received question using parsing tools. Specifically, it tokenizes the question using a natural language processing engine (e.g., Spacy, NLTK) and extracts key keywords and meanings.
[0528] For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" are extracted.
[0529] 3. Generating the answer:
[0530] The server invokes a generation mechanism based on the analysis results and uses a generation AI model (e.g., GPT-4, GPT-3) to generate an appropriate response.
[0531] For example, the system might generate responses such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0532] 4. Search for related links:
[0533] The server searches for relevant links based on the generated response. It collects the appropriate links from the internet using search methods (e.g., Google Search API, Bing Search API).
[0534] For example, we collect links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0535] 5. Generating and sending results:
[0536] The server combines the generated answers and collected links to create the final response.
[0537] For example, a response in the format of "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." is sent to the terminal as an HTTP response.
[0538] 6. Displaying the results:
[0539] The terminal receives a response from the server and displays it in the user interface.
[0540] Through this, users can view the answers and related links.
[0541] Examples of specific actions
[0542] For example, if a user enters "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0543] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0544] 2. The device sends the question to the server.
[0545] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0546] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the generation AI model.
[0547] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0548] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[0549] 7. The device displays this answer and link to the user.
[0550] Example of a prompt
[0551] Here are some specific examples of prompt statements to input into a generative AI model:
[0552] "What are some recommended ramen restaurants in Tokyo?"
[0553] "What are some good tourist spots in Kyoto?"
[0554] This allows users to quickly obtain the information they need with a single question, enabling efficient information retrieval.
[0555] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0556] Step 1: User enters question
[0557] The user enters a question in natural language into the terminal interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" into the text field. The terminal retrieves the user's input (question text) so that it can be processed in the next step.
[0558] Input: User's question (natural language text)
[0559] Output: Retrieving questions via the terminal
[0560] Specific operation: The user enters a question into the terminal's text field. The terminal stores this input in memory and prepares to send it in the next step.
[0561] Step 2: Sending a request via the terminal
[0562] The terminal sends the question received from the user to the server as an HTTP POST request. The request contains the question text.
[0563] Input: User's question (text stored on the device)
[0564] Output: HTTP POST request to the server
[0565] Specific operation: The terminal includes the retrieved question text in the payload of an HTTP POST request and sends it to a specific endpoint. The server receives this request.
[0566] Step 3: Server analyzes the question
[0567] The server uses a natural language processing engine (e.g., Spacy, NLTK) to analyze the received question text. This analysis extracts key keywords and meanings from the text.
[0568] Input: User's question (payload of HTTP POST request)
[0569] Output: Analyzed keywords and meanings
[0570] Specific operation: The server extracts the question text from the request payload and passes it to the natural language processing engine. The engine tokenizes the text and extracts keywords such as "Tokyo," "recommended," and "ramen restaurant."
[0571] Step 4: Server generates response
[0572] The server generates appropriate answers using a generative AI model (e.g., GPT-4, GPT-3) based on the analysis results. A prompt is input to the generative AI model, and the answer text is retrieved based on that input.
[0573] Input: Analyzed keywords and meanings
[0574] Output: Generated answer (text)
[0575] Specific operation: The server uses the analyzed keywords to create a prompt for the generative AI model and inputs a question such as "What are some recommended ramen restaurants in Tokyo?" into the model. The generative AI model responds by outputting text such as "Some recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0576] Step 5: Server searches for related links
[0577] The server searches the internet for relevant links based on the generated response. Links are collected using search methods (e.g., Google Search API, Bing Search API).
[0578] Input: Generated response (text)
[0579] Output: List of related links
[0580] Specific operation: The server extracts keywords from the generated response text and creates search queries such as "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". These queries are then fed into a search API to retrieve relevant links (e.g., official website, review site).
[0581] Step 6: Server summarizes and sends the results.
[0582] The server combines the generated answers and collected links to create a final response. This response is then sent to the terminal as an HTTP response.
[0583] Input: List of generated answers (text) and related links
[0584] Output: Final response (text and links)
[0585] Specific operation: The server combines the answer text and related links to prepare a response in the format, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." This is then sent to the terminal as an HTTP response.
[0586] Step 7: Display to the user via the device
[0587] The terminal displays the response received from the server in the user interface. Through this, the user can view the answer and related links.
[0588] Input: Final response from the server (text and links)
[0589] Output: Content displayed to the user
[0590] Specific operation: The terminal analyzes the received response and extracts the answer and link. These are then displayed using HTML or GUI components, making them viewable by the user.
[0591] As described above, users can quickly obtain the necessary information with a single question, and relevant links are provided simultaneously, enabling efficient information retrieval.
[0592] (Application Example 1)
[0593] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0594] Traditional information retrieval systems had the problem of requiring a lot of effort and time for users to find the right product even after entering a product name or specific characteristics. Furthermore, the limited functionality for displaying related links and product information in a single batch meant that user search efficiency was not improved.
[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0596] In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question, a generation means for generating an answer based on the analyzed question content, a search means for searching for links related to the generated answer, an output means for outputting the answer and related links together, and a product link display means for displaying the corresponding product link along with the generated answer. This makes it possible to quickly display appropriate product suggestions and links to related products when a user inputs a question related to a product.
[0597] "Input means for entering a question" refers to a device or interface for a user to enter a question in natural language.
[0598] "Analysis means for analyzing input questions" refers to means that analyze input questions and use natural language processing techniques to extract key keywords and meanings.
[0599] "A means for generating answers based on analyzed question content" refers to a means that utilizes an artificial intelligence model to generate appropriate answers based on information extracted by the analysis means.
[0600] "A search method for finding links related to generated answers" refers to a method for searching the internet for information related to generated answers and collecting the relevant links.
[0601] "An output method for outputting answers and related links together" refers to a method for integrating the generated answers and collected related links and displaying them to the user.
[0602] "A means for displaying product links along with the generated response" refers to a means for additionally displaying links to products related to the generated response.
[0603] A "natural language processing engine" is a software engine that analyzes input natural language text and extracts its meaning and keywords.
[0604] An "artificial intelligence model" is a machine learning model that uses input data to perform reasoning like a human and generate appropriate answers.
[0605] This invention is an information retrieval system that can be applied to an e-commerce site application that provides appropriate answers and links to related products when a user enters a question related to a product.
[0606] System Configuration
[0607] 1. User Interface
[0608] The application includes an input method for users to enter questions in natural language using a smartphone application. For example, it would be a section where users can enter questions such as, "What smartphone do you recommend?"
[0609] 2. Analysis of the Question
[0610] The terminal retrieves the entered question and sends it to the server. The server uses a natural language processing engine (e.g., SpaCy) to analyze the entered question, extracting key keywords and meaning from it.
[0611] 3. Generating the answer
[0612] The server generates appropriate answers to questions using artificial intelligence models (e.g., OpenAI's GPT-4) based on the analyzed question content. This generation mechanism has the ability to perform real-time reasoning in response to user input and provide relevant information.
[0613] 4. Search for related links
[0614] The server includes a search mechanism for searching the internet for relevant product links based on the generated response. This search mechanism uses, for example, the BeautifulSoup or requests library to retrieve links to relevant products.
[0615] 5. Output of answers and links
[0616] The server has an output mechanism to combine the generated answer and related product links into a single response and send it to the terminal. The terminal displays this response to the user. This allows the user to view the answer to the question along with detailed links to related products all at once.
[0617] Hardware and software used
[0618] Hardware: Smartphone
[0619] software:
[0620] OpenAI API: GPT model for question analysis and answer generation
[0621] requests library: Used to retrieve links from websites.
[0622] BeautifulSoup: To parse links from retrieved HTML.
[0623] Natural language processing engines: SpaCy, etc.
[0624] Specific example
[0625] For example, if a user types the question "Can you recommend a smartphone?", the following steps will be taken:
[0626] 1. The user enters the question "What smartphone do you recommend?" into their device.
[0627] 2. The device sends the question to the server.
[0628] 3. The server analyzes the question and extracts the main keywords "recommendation" and "smartphone".
[0629] 4. Based on the analysis results, the server uses a generation method to generate the response "The latest iPhone or Samsung Galaxy is recommended" from the artificial intelligence model.
[0630] 5. The server searches for links related to "iPhone" and "Samsung Galaxy".
[0631] 6. The server compiles the answers and links and sends them to the device in the format, "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details."
[0632] 7. The device displays this answer and link to the user.
[0633] Examples of prompts to input into a generative AI model:
[0634] Question: What smartphone do you recommend?
[0635] answer:
[0636] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0637] Step 1:
[0638] The user enters the question in natural language.
[0639] Input: The user enters the question "What smartphone do you recommend?" in a smartphone application.
[0640] Output: The entered question is sent to the terminal.
[0641] Specific action: The user enters a question into the app's interface and taps the submit button.
[0642] Step 2:
[0643] The terminal sends the entered question to the server.
[0644] Input: A natural language question entered by the user.
[0645] Output: The question content is sent to the server.
[0646] Specific operation: The terminal receives user input and forwards the question to the server as an HTTP request.
[0647] Step 3:
[0648] The server performs natural language processing on the question using parsing tools.
[0649] Input: A natural language question sent from the device.
[0650] Output: Analyzed keywords and their meanings.
[0651] Specific operation: The server uses a natural language processing engine (e.g., SpaCy) to analyze the question and extract keywords such as "recommended" and "smartphone".
[0652] Step 4:
[0653] The server generates an answer based on the analysis results.
[0654] Input: Keywords and meanings of the analyzed question.
[0655] Output: The generated answer.
[0656] Specific operation: The server inputs the analysis results into an artificial intelligence model (e.g., OpenAI's GPT-4) and generates responses such as, "We recommend the latest iPhone or Samsung Galaxy."
[0657] Step 5:
[0658] The server searches for relevant links based on the generated response.
[0659] Input: The content of the generated response.
[0660] Output: Related product links.
[0661] Specific operation: The server uses web scraping libraries (e.g., requests and BeautifulSoup) to collect product links related to "iPhone" and "Samsung Galaxy" from the internet.
[0662] Step 6:
[0663] The server outputs the answer and the link together.
[0664] Input: Generated responses and collected product links.
[0665] Output: The final response to be displayed to the user.
[0666] Specific operation: The server combines the answer text and each link into a single response and sends it back to the terminal as an HTTP response.
[0667] Step 7:
[0668] The device displays the response and link from the server.
[0669] Input: The final response sent from the server.
[0670] Output: Answers and links that users can view.
[0671] Specific action: The response received by the device is displayed on the user interface, and the message "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details." is presented.
[0672] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0673] This invention combines an information retrieval system that provides appropriate answers using generative AI and simultaneously presents related links based on a single question entered by the user, with an emotion engine that recognizes the user's emotions. This system is implemented in the following specific forms.
[0674] User
[0675] The user first inputs a question in natural language into the terminal's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server. It also recognizes emotions using the user's input and voice interface.
[0676] terminal
[0677] The terminal receives user input and sends it to the server as a request. The request includes the user's question and sentiment data, and the server receives this request.
[0678] server
[0679] The server processes the following steps.
[0680] 1. Analysis of the question:
[0681] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0682] 2. Recognition of emotions:
[0683] The server uses emotion recognition mechanisms to analyze the user's emotional data and recognize the user's emotional state. For example, emotions such as "joy," "sadness," and "surprise" can be identified.
[0684] 3. Generating the answer:
[0685] The server invokes a generation mechanism based on the analysis results and recognized emotions, and uses an artificial intelligence model to generate an appropriate response. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated with a tone and content that matches the user's emotions.
[0686] 4. Search for related links:
[0687] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0688] 5. Summary and output of results:
[0689] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0690] Specific example
[0691] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0692] 1. The user enters the question "What are some good tourist spots in Kyoto?" into the device. The user's input also detects the emotion "fun".
[0693] 2. The device sends the question and sentiment data to the server.
[0694] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0695] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[0696] 5. Based on the analysis results and sentiment data, the server uses a generation method to generate a sentiment-appropriate response from an artificial intelligence model, such as "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[0697] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0698] 7. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. You're sure to have a wonderful time! Please see the link below for more details."
[0699] 8. The device displays this answer and link to the user.
[0700] This invention allows users to quickly obtain the necessary information with a single question, receive empathetic answers, and make the search process more user-friendly.
[0701] The following describes the processing flow.
[0702] Step 1:
[0703] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[0704] Step 2:
[0705] The terminal retrieves the input. The terminal retrieves the user's input, and in addition, collects sentiment data obtained from the voice interface to construct the request.
[0706] Step 3:
[0707] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request to the server containing the user's question and sentiment data.
[0708] Step 4:
[0709] The server receives the request. The server receives the request from the terminal and prepares to analyze the question content and sentiment data.
[0710] Step 5:
[0711] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0712] Step 6:
[0713] The server analyzes emotional data. Using emotion recognition tools, the server analyzes emotions from user input and voice data to recognize the user's emotional state. For example, emotions such as "excitement" and "anticipation" may be identified.
[0714] Step 7:
[0715] The server sends a question to a generative AI. Based on the analysis results and recognized emotions, the server sends the question to the generative AI (e.g., GPT-4) to generate an appropriate answer that is sensitive to those emotions.
[0716] Step 8:
[0717] A generative AI generates the answer. The generative AI generates an answer that matches the emotion, such as, "For ramen restaurants in Tokyo, I recommend Ichiran and Tsukemen Daioh. You're sure to have a wonderful time!"
[0718] Step 9:
[0719] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[0720] Step 10:
[0721] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[0722] Step 11:
[0723] The server sends a response to the terminal. The response contains the generated answer and related links.
[0724] Step 12:
[0725] The terminal receives a response. The terminal receives a response from the server.
[0726] Step 13:
[0727] The device displays the results to the user. The device displays the received response and link to the user, showing the response and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. You're sure to have a wonderful time! Please see the link below for details."
[0728] This allows users to quickly obtain answers and relevant links to their questions, and also provides emotionally resonant responses, making the search experience more user-friendly and satisfying.
[0729] (Example 2)
[0730] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0731] Traditional information retrieval systems required users to go through multiple steps to obtain appropriate answers after entering a question, resulting in a complex user experience. Furthermore, they lacked the ability to recognize and respond to user emotions, failing to adequately enhance user satisfaction. Therefore, there is a need for a system that provides appropriate answers and related information with a single question, and even provides answers that are sensitive to the user's emotions.
[0732] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0733] In this invention, the server includes an input means, a means for receiving questions and sentiment data transmitted from a terminal, an analysis means using a natural language processing engine for analyzing the input questions, an analysis means using an sentiment recognition engine for analyzing the sentiment data, a generation means using an artificial intelligence model for generating answers based on the analyzed question content and sentiment data, a search means using a search engine for searching the internet for links related to the generated answers, and an output means for sending the answers and related links together to the terminal. As a result, the user can obtain quick and appropriate answers and related information with a single question, and furthermore, answers that are sensitive to the user's emotions can be provided.
[0734] "Input method" refers to a device or software that provides an interface for users to input questions.
[0735] "Receiving means" refers to the function used to retrieve questions and sentiment data sent from a terminal on the server side.
[0736] A "natural language processing engine" refers to software or algorithms that analyze input natural language questions and extract their main keywords and meanings.
[0737] "Analysis means" refers to methods and techniques for analyzing received data and extracting necessary information.
[0738] An "emotion recognition engine" refers to software or algorithms that analyze and identify emotional states from user input data.
[0739] "Generation means" refers to functions and technologies that use artificial intelligence models to generate answers based on analyzed question content and sentiment data.
[0740] An "artificial intelligence model" refers to a system that has been trained using technologies such as machine learning and deep learning, and is capable of generating appropriate answers to questions.
[0741] A "search engine" refers to software or services used to search the internet for and collect links related to generated answers.
[0742] "Search methods" refer to methods and techniques for collecting relevant information.
[0743] "Output means" refers to methods and technologies for compiling generated answers and related links and providing them to the user.
[0744] A "terminal" refers to a device used by a user to input questions or view information received from a server.
[0745] A "question" refers to the content entered by the user in natural language regarding the information they want to know.
[0746] "Emotional data" refers to information indicating the emotional state, extracted from user input and voice.
[0747] This invention combines an information retrieval system that provides an appropriate answer using generative artificial intelligence (AI) and simultaneously presents related links based on the user's input of a single question, with an emotion engine that recognizes the user's emotions. A specific embodiment of this system is described below.
[0748] User actions
[0749] The user first inputs a question in natural language through the device's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" At this time, the device acquires the user's input and uses the voice interface to recognize the user's emotions. For example, the tone of voice and facial expressions while the user is inputting may be used to detect that the user is "happy."
[0750] Terminal operation
[0751] The terminal combines the user's input question and recognized sentiment data, and sends this to the server. Here, it plays the role of sending the question content and sentiment data as a single request to the server. Specific devices used include personal computers, smartphones, and tablets.
[0752] Server Processing
[0753] The server takes several steps to process an incoming request. First, it extracts key keywords and meanings by analyzing the question using a natural language processing engine (for example, the Google Cloud Natural Language API). For example, from the question "What are some recommended ramen restaurants in Tokyo?", the keywords extracted would be "Tokyo," "recommended," and "ramen restaurant."
[0754] Next, an emotion recognition engine (for example, IBM Watson Tone Analyzer) is used to analyze the user's emotional data. For instance, emotions such as "happy" are identified from the user's voice tone and text.
[0755] Based on the analysis results, the server invokes a generation mechanism and uses an artificial intelligence model (e.g., OpenAI's GPT-3) to generate an appropriate response. This generated response is delivered in a content and tone that matches the user's emotions. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated in a way that resonates with the user's "happy" feelings.
[0756] Based on the generated response, the server uses search methods to find relevant links. Specifically, it uses a search engine (for example, the Google Search API) to collect links to relevant official websites and review sites. For example, it searches for links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0757] Finally, the server combines the generated answers and collected links to create a response to send to the terminal. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0758] Specific example
[0759] For example, consider a case where a user enters the question, "What are some good tourist spots in Kyoto?" The emotion "fun" is detected simultaneously with the user's input.
[0760] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0761] 2. The device sends the question and sentiment data to the server.
[0762] 3. The server analyzes the question and extracts the keywords "Kyoto," "tourist attractions," and "recommendations."
[0763] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[0764] 5. Based on the analysis results and sentiment data, the server uses an artificial intelligence model to generate the response, "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[0765] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0766] 7. The server compiles the answers and links and sends them to the device.
[0767] 8. The device displays this answer and link to the user.
[0768] Example of a prompt:
[0769] "What are some good tourist spots in Kyoto? I'm in a good mood."
[0770] This invention enables users to quickly obtain appropriate answers and related information with a single question, and further provides a search experience that is sensitive to the user's emotions.
[0771] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0772] Step 1:
[0773] The user enters a question into the terminal. Specifically, the user uses text input or voice input to enter a question into the interface, for example, "What are some recommended ramen restaurants in Tokyo?" The input is the text data of the question.
[0774] Step 2:
[0775] The device receives user input text and also acquires emotion data. In the case of voice input, it uses a speech recognition engine to convert it to text, and at the same time, an emotion recognition engine analyzes the emotion. For example, it might acquire the emotion "happy" as text data. The output consists of the converted text data and emotion data.
[0776] Step 3:
[0777] The device sends the question content and sentiment data to the server. The data sent includes the question text and request data containing sentiment data. Specifically, the request is sent to the server using the HTTP protocol.
[0778] Step 4:
[0779] The server receives the request data. The server analyzes the received question text and extracts key keywords. Specifically, it uses a natural language processing engine (e.g., Google Cloud Natural Language API) to extract keywords such as "Tokyo," "recommended," and "ramen restaurant" from the question text. The output is the extracted keywords.
[0780] Step 5:
[0781] The server analyzes emotional data to identify the user's emotions. Specifically, it uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the emotional data. The output is the user's emotional state. For example, it might retrieve data indicating "happy."
[0782] Step 6:
[0783] The server generates a response based on the analysis results (extracted keywords and sentiment data). An artificial intelligence model (e.g., OpenAI's GPT-3) is used as the generation method. The input consists of extracted keywords and sentiment data. For example, it might generate the response text: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The output is the generated response text.
[0784] Step 7:
[0785] The server searches for relevant links based on the generated response. It uses a search engine (e.g., Google Search API) to perform a web search using words included in the generated response. For example, it collects links for "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". The output is a list of relevant links.
[0786] Step 8:
[0787] The server combines the generated answers and collected links to create the final response. Specifically, it combines the answer text and the list of links into a single response data. For example, it might create a response in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." The output is the combined response data.
[0788] Step 9:
[0789] The server sends the final response data to the terminal. It sends the response data back to the terminal using a specific communication protocol. The output indicates that the transmission of the response data is complete.
[0790] Step 10:
[0791] The terminal displays response data obtained from the server to the user. Specifically, it either displays the response on the screen interface or reads it aloud. For example, it might display something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen! For more details, please refer to the link below." The output consists of the response displayed to the user and the link.
[0792] This specific processing flow allows users to quickly obtain appropriate answers and related information with a single question, and furthermore, to receive emotionally resonant responses.
[0793] (Application Example 2)
[0794] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0795] Traditional information retrieval systems provide answers to user questions, but they lack the ability to generate answers that take user emotions into consideration. Furthermore, their ability to provide related links to the generated answers is limited. As a result, users often experience low satisfaction when obtaining answers, and the search experience remains unimproved.
[0796] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question and the user's emotions, a generation means for generating an answer based on the analyzed question content and emotion data, a search means for searching for links related to the generated answer, and an output means for outputting the answer and related links together. This makes it possible to generate an answer that is sensitive to the user's emotions and to provide related information quickly.
[0797] A "question" is the content that users input in natural language to find out what they want to know.
[0798] "Input means" refers to the devices or software that users use to input questions.
[0799] "Analysis means" refers to a device or program for analyzing input questions and user sentiment data.
[0800] "Generation means" refers to a device or program that generates answers based on analyzed question content and sentiment data.
[0801] "Search means" refers to a device or program for searching the internet for links related to the generated answer.
[0802] "Output means" refers to a device or program for displaying answers and related links to the user.
[0803] "Emotional data" refers to emotional information extracted from user input.
[0804] An "artificial intelligence model" is a model trained based on machine learning or deep learning, and is used as a means of generation.
[0805] A "natural language processing engine" is a program that analyzes text input in natural language and extracts meaning and keywords.
[0806] "Related links" refer to URLs of web pages or information that are relevant to the generated response.
[0807] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system incorporates an emotion engine that recognizes the user's emotions and generates answers in a tone and content that corresponds to the user's emotions.
[0808] System Configuration
[0809] This system consists of the following main components:
[0810] 1. Input method: This refers to the means by which the user inputs questions, and includes devices such as smartphones and smart glasses.
[0811] 2. Analysis means: This refers to means for analyzing the input questions and sentiment data, and corresponds to natural language processing engines (e.g., Transformers, spaCy) or sentiment recognition libraries (e.g., emote).
[0812] 3. Generation means: This refers to means for generating answers based on the analyzed question content and sentiment data, and a generative AI model (e.g., GPT-3) falls under this category.
[0813] 4. Search methods: These are methods for searching the internet for links related to the generated answers, and web scraping libraries (BeautifulSoup, Scrapy) are examples of this.
[0814] 5. Output means: This refers to a means of displaying the answer and related links to the user, and this includes the device's display and notification function.
[0815] Processing flow
[0816] As a concrete example, the "emotion-recognition store navigator" application using smart glasses operates in the following steps.
[0817] 1. User input: The user uses smart glasses to input a question in natural language. For example, they might ask, "What product would suit my current mood?"
[0818] 2. Emotion Recognition: The device receives user input and analyzes emotions using an emotion recognition library (such as emote).
[0819] 3. Request to the server: Send the analyzed question and sentiment data to the server.
[0820] 4. Question Analysis: The server uses a natural language processing engine (such as Transformers or spaCy) to analyze the question and extract key keywords.
[0821] 5. Response generation: Use a generative AI model (such as GPT-3) to generate appropriate responses based on analysis results and sentiment data.
[0822] 6. Search for related links: Based on the generated answers, use a web scraping library (such as BeautifulSoup or Scrapy) to search the internet for related links.
[0823] 7. Displaying Results: The answers and related links are compiled and displayed on the smart glasses' screen.
[0824] Examples of specific cases and prompt statements
[0825] As a concrete example, if a user feels tired and asks, "What product would be perfect for how I'm feeling right now?", the following process would occur:
[0826] Input text: "Please recommend a product that perfectly matches my current mood."
[0827] Emotion recognition results: "Fatigue" and "Stress"
[0828] Server output: "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details."
[0829] Related links: "Aroma Candle Details Page", "Massage Chair Purchase Page"
[0830] Examples of prompt messages are as follows:
[0831] "What are some relaxing products suitable for relieving fatigue?"
[0832] "What are some good relaxation items for when you're tired?"
[0833] This system allows users to not only get quick answers to their questions, but also experience friendly and personalized responses tailored to their own emotions.
[0834] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0835] Step 1: User Input
[0836] The user uses smart glasses to input a question in natural language. The input might be something like, "Tell me what product would suit my current mood." The user asks the question via voice or text input. At this time, the smart glasses interface receives the question and retrieves the data.
[0837] Step 2: Recognizing Emotions
[0838] The device receives user input and analyzes emotions using an emotion recognition library (e.g., Emote). In this step, emotion data is extracted from the input text to recognize emotional states such as "fatigue" or "stress." The input data is in text format, and the output data consists of emotion labels and emotion scores.
[0839] Step 3: Request to the server
[0840] The terminal sends the analyzed question and sentiment data to the server. Specifically, it sends the question text and sentiment data, converted to JSON format, to the server via an HTTP POST request. The input data is in JSON format, and the output data is the response from the server.
[0841] Step 4: Analyzing the Question
[0842] The server uses a natural language processing engine (e.g., Transformers, spaCy) to analyze the received question and extract key keywords. For example, from the question "Tell me a product that suits my current mood," it extracts the keywords "mood" and "product." The input data is the question text in JSON format, and the output data is the analyzed keywords.
[0843] Step 5: Generating the answer
[0844] The server generates an appropriate answer using a generative AI model (e.g., GPT-3) based on the analysis results of the question and sentiment data. In this step, the generative AI model generates an answer based on the input data (analyzed keywords and sentiment data). For example, an answer such as "Relaxing aromatherapy candles and massage chairs are recommended" might be generated. The input data consists of keywords and sentiment data, and the output data is the generated answer text.
[0845] Step 6: Search for related links
[0846] The server uses a web scraping library (e.g., BeautifulSoup, Scrapy) to search the internet for relevant links based on the generated response. For example, it collects relevant links such as "details page for aromatherapy candles" and "purchase page for massage chairs." The input data is the generated response text, and the output data is a list of relevant links.
[0847] Step 7: Displaying the results
[0848] The server compiles the answers and related links and sends them to the device. The device displays the received answers and links to the user. At this time, the smart glasses display shows a message along with a link that reads, "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details." The input data is the response from the server, and the output data is information that the user can visually confirm.
[0849] As described above, a series of processes, from user questioning and sentiment recognition to natural language processing, the use of generative AI models, the search for related links, and output, are performed, making it possible to provide users with useful information that resonates with their emotions.
[0850] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0851] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0852] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0853] [Third Embodiment]
[0854] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0855] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0856] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0857] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0858] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0859] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0860] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0861] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0862] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0863] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0864] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0865] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0866] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system is implemented as follows:
[0867] User
[0868] The user first enters a question in natural language into the terminal's interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server.
[0869] terminal
[0870] The terminal receives the user's input and sends it to the server as a request. The request includes the user's question, and the server receives this request.
[0871] server
[0872] The server processes the following steps.
[0873] 1. Analysis of the question:
[0874] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0875] 2. Generating the answer:
[0876] The server then calls a generation mechanism based on the analysis results and uses an artificial intelligence model to generate an appropriate answer. For example, it might generate an answer such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0877] 3. Search for related links:
[0878] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0879] 4. Summary and output of results:
[0880] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0881] Specific example
[0882] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0883] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0884] 2. The device sends the question to the server.
[0885] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0886] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the artificial intelligence model.
[0887] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0888] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[0889] 7. The device displays this answer and link to the user.
[0890] This invention allows users to quickly obtain the necessary information with a single question, significantly simplifying the search process and improving efficiency.
[0891] The following describes the processing flow.
[0892] Step 1:
[0893] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[0894] Step 2:
[0895] The terminal retrieves the input. The terminal retrieves the user's input and constructs it as a request.
[0896] Step 3:
[0897] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request that includes the user's question.
[0898] Step 4:
[0899] The server receives the request. The server receives the request from the terminal and prepares to analyze the question.
[0900] Step 5:
[0901] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, the keywords "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[0902] Step 6:
[0903] The server sends a question to a generative AI. Based on the analysis results, the server sends another question to the generative AI (e.g., GPT-4) to generate an appropriate answer.
[0904] Step 7:
[0905] A generative AI generates an answer. The generative AI generates the answer, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0906] Step 8:
[0907] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[0908] Step 9:
[0909] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[0910] Step 10:
[0911] The server sends a response to the terminal. The response contains the generated answer and related links.
[0912] Step 11:
[0913] The terminal receives a response. The terminal receives a response from the server.
[0914] Step 12:
[0915] The device displays the results to the user. The device displays the received answer and link to the user, showing the answer and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[0916] This allows users to quickly obtain answers to their questions and relevant links.
[0917] (Example 1)
[0918] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0919] Traditional information retrieval systems required users to perform numerous manual operations and multiple searches to efficiently obtain the information they needed, resulting in a cumbersome and time-consuming process. Furthermore, collecting links related to the generated answers required a separate effort, leading to poor user convenience.
[0920] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0921] In this invention, the server includes an input means, a terminal for receiving questions entered by the user through the input means, an analysis means for analyzing the questions transmitted from the terminal, a generation means for generating answers based on the analyzed question content, a search means for searching for links related to the answers generated by the generation means, and an output means for sending the generated answers and related links together to the terminal. As a result, the user can quickly obtain the necessary information by simply entering a single question, and related links are also provided at the same time, enabling efficient information retrieval.
[0922] "Input method" refers to the interface through which a user enters a question into the system.
[0923] A "terminal" refers to a device used by a user to input a question and send it to a server for processing.
[0924] "Analysis means" refers to a function that analyzes questions sent by users through their devices and extracts their content and key keywords.
[0925] "Generation means" refers to a function that generates appropriate answers based on keywords and content extracted by the analysis means.
[0926] "Search means" refers to a function for searching the internet for and collecting links related to the answers generated by the generation means.
[0927] "Output means" refers to a function that aggregates the generated responses and collected related links, sends them to the terminal, and displays them to the user.
[0928] A "natural language processing engine" refers to software or algorithms that analyze an input question and extract key keywords and content.
[0929] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate appropriate answers to user questions.
[0930] An "HTTP request" refers to a communication protocol used to send data to a web server.
[0931] "HTTP response" refers to the communication protocol used to send back data received from a web server.
[0932] A "prompt sentence" refers to a sentence used as an instruction to input into a generative AI model.
[0933] This invention is an information retrieval system that provides an appropriate answer using a generative AI model and simultaneously presents related links, based on the user's input of a single question. Specific embodiments of this system are described below.
[0934] Hardware and software configuration
[0935] This system consists of the following main components:
[0936] Device: A device used by the user to enter questions. This includes PCs, smartphones, tablets, etc.
[0937] Server: The central processing unit that analyzes questions, generates answers, searches for relevant links, and provides responses to the user. Software running on the server includes natural language processing engines, generative AI models, and search engine APIs.
[0938] Detailed processing of the system
[0939] The server performs the following steps:
[0940] 1. Receiving and sending questions:
[0941] The user enters a question into the terminal's interface, such as, "What are some recommended ramen restaurants in Tokyo?"
[0942] The device retrieves this question and sends it to the server as an HTTP POST request.
[0943] 2. Analysis of the question:
[0944] The server processes the received question using parsing tools. Specifically, it tokenizes the question using a natural language processing engine (e.g., Spacy, NLTK) and extracts key keywords and meanings.
[0945] For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" are extracted.
[0946] 3. Generating the answer:
[0947] The server invokes a generation mechanism based on the analysis results and uses a generation AI model (e.g., GPT-4, GPT-3) to generate an appropriate response.
[0948] For example, the system might generate responses such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0949] 4. Search for related links:
[0950] The server searches for relevant links based on the generated response. It collects the appropriate links from the internet using search methods (e.g., Google Search API, Bing Search API).
[0951] For example, we collect links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[0952] 5. Generating and sending results:
[0953] The server combines the generated answers and collected links to create the final response.
[0954] For example, a response in the format of "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." is sent to the terminal as an HTTP response.
[0955] 6. Displaying the results:
[0956] The terminal receives a response from the server and displays it in the user interface.
[0957] Through this, users can view the answers and related links.
[0958] Examples of specific actions
[0959] For example, if a user enters "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[0960] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[0961] 2. The device sends the question to the server.
[0962] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[0963] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the generation AI model.
[0964] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[0965] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[0966] 7. The device displays this answer and link to the user.
[0967] Example of a prompt
[0968] Here are some specific examples of prompt statements to input into a generative AI model:
[0969] "What are some recommended ramen restaurants in Tokyo?"
[0970] "What are some good tourist spots in Kyoto?"
[0971] This allows users to quickly obtain the information they need with a single question, enabling efficient information retrieval.
[0972] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0973] Step 1: User enters question
[0974] The user enters a question in natural language into the terminal interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" into the text field. The terminal retrieves the user's input (question text) so that it can be processed in the next step.
[0975] Input: User's question (natural language text)
[0976] Output: Retrieving questions via the terminal
[0977] Specific operation: The user enters a question into the terminal's text field. The terminal stores this input in memory and prepares to send it in the next step.
[0978] Step 2: Sending a request via the terminal
[0979] The terminal sends the question received from the user to the server as an HTTP POST request. The request contains the question text.
[0980] Input: User's question (text stored on the device)
[0981] Output: HTTP POST request to the server
[0982] Specific operation: The terminal includes the retrieved question text in the payload of an HTTP POST request and sends it to a specific endpoint. The server receives this request.
[0983] Step 3: Server analyzes the question
[0984] The server uses a natural language processing engine (e.g., Spacy, NLTK) to analyze the received question text. This analysis extracts key keywords and meanings from the text.
[0985] Input: User's question (payload of HTTP POST request)
[0986] Output: Analyzed keywords and meanings
[0987] Specific operation: The server extracts the question text from the request payload and passes it to the natural language processing engine. The engine tokenizes the text and extracts keywords such as "Tokyo," "recommended," and "ramen restaurant."
[0988] Step 4: Server generates response
[0989] The server generates appropriate answers using a generative AI model (e.g., GPT-4, GPT-3) based on the analysis results. A prompt is input to the generative AI model, and the answer text is retrieved based on that input.
[0990] Input: Analyzed keywords and meanings
[0991] Output: Generated answer (text)
[0992] Specific operation: The server uses the analyzed keywords to create a prompt for the generative AI model and inputs a question such as "What are some recommended ramen restaurants in Tokyo?" into the model. The generative AI model responds by outputting text such as "Some recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[0993] Step 5: Server searches for related links
[0994] The server searches the internet for relevant links based on the generated response. Links are collected using search methods (e.g., Google Search API, Bing Search API).
[0995] Input: Generated response (text)
[0996] Output: List of related links
[0997] Specific operation: The server extracts keywords from the generated response text and creates search queries such as "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". These queries are then fed into a search API to retrieve relevant links (e.g., official website, review site).
[0998] Step 6: Server summarizes and sends the results.
[0999] The server combines the generated answers and collected links to create a final response. This response is then sent to the terminal as an HTTP response.
[1000] Input: List of generated answers (text) and related links
[1001] Output: Final response (text and links)
[1002] Specific operation: The server combines the answer text and related links to prepare a response in the format, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." This is then sent to the terminal as an HTTP response.
[1003] Step 7: Display to the user via the device
[1004] The terminal displays the response received from the server in the user interface. Through this, the user can view the answer and related links.
[1005] Input: Final response from the server (text and links)
[1006] Output: Content displayed to the user
[1007] Specific operation: The terminal analyzes the received response and extracts the answer and link. These are then displayed using HTML or GUI components, making them viewable by the user.
[1008] As described above, users can quickly obtain the necessary information with a single question, and relevant links are provided simultaneously, enabling efficient information retrieval.
[1009] (Application Example 1)
[1010] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1011] Traditional information retrieval systems had the problem of requiring a lot of effort and time for users to find the right product even after entering a product name or specific characteristics. Furthermore, the limited functionality for displaying related links and product information in a single batch meant that user search efficiency was not improved.
[1012] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1013] In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question, a generation means for generating an answer based on the analyzed question content, a search means for searching for links related to the generated answer, an output means for outputting the answer and related links together, and a product link display means for displaying the corresponding product link along with the generated answer. This makes it possible to quickly display appropriate product suggestions and links to related products when a user inputs a question related to a product.
[1014] "Input means for entering a question" refers to a device or interface for a user to enter a question in natural language.
[1015] "Analysis means for analyzing input questions" refers to means that analyze input questions and use natural language processing techniques to extract key keywords and meanings.
[1016] "A means for generating answers based on analyzed question content" refers to a means that utilizes an artificial intelligence model to generate appropriate answers based on information extracted by the analysis means.
[1017] "A search method for finding links related to generated answers" refers to a method for searching the internet for information related to generated answers and collecting the relevant links.
[1018] "An output method for outputting answers and related links together" refers to a method for integrating the generated answers and collected related links and displaying them to the user.
[1019] "A means for displaying product links along with the generated response" refers to a means for additionally displaying links to products related to the generated response.
[1020] A "natural language processing engine" is a software engine that analyzes input natural language text and extracts its meaning and keywords.
[1021] An "artificial intelligence model" is a machine learning model that uses input data to perform reasoning like a human and generate appropriate answers.
[1022] This invention is an information retrieval system that can be applied to an e-commerce site application that provides appropriate answers and links to related products when a user enters a question related to a product.
[1023] System Configuration
[1024] 1. User Interface
[1025] The application includes an input method for users to enter questions in natural language using a smartphone application. For example, it would be a section where users can enter questions such as, "What smartphone do you recommend?"
[1026] 2. Analysis of the Question
[1027] The terminal retrieves the entered question and sends it to the server. The server uses a natural language processing engine (e.g., SpaCy) to analyze the entered question, extracting key keywords and meaning from it.
[1028] 3. Generating the answer
[1029] The server generates appropriate answers to questions using artificial intelligence models (e.g., OpenAI's GPT-4) based on the analyzed question content. This generation mechanism has the ability to perform real-time reasoning in response to user input and provide relevant information.
[1030] 4. Search for related links
[1031] The server includes a search mechanism for searching the internet for relevant product links based on the generated response. This search mechanism uses, for example, the BeautifulSoup or requests library to retrieve links to relevant products.
[1032] 5. Output of answers and links
[1033] The server has an output mechanism to combine the generated answer and related product links into a single response and send it to the terminal. The terminal displays this response to the user. This allows the user to view the answer to the question along with detailed links to related products all at once.
[1034] Hardware and software used
[1035] Hardware: Smartphone
[1036] software:
[1037] OpenAI API: GPT model for question analysis and answer generation
[1038] requests library: Used to retrieve links from websites.
[1039] BeautifulSoup: To parse links from retrieved HTML.
[1040] Natural language processing engines: SpaCy, etc.
[1041] Specific example
[1042] For example, if a user types the question "Can you recommend a smartphone?", the following steps will be taken:
[1043] 1. The user enters the question "What smartphone do you recommend?" into their device.
[1044] 2. The device sends the question to the server.
[1045] 3. The server analyzes the question and extracts the main keywords "recommendation" and "smartphone".
[1046] 4. Based on the analysis results, the server uses a generation method to generate the response "The latest iPhone or Samsung Galaxy is recommended" from the artificial intelligence model.
[1047] 5. The server searches for links related to "iPhone" and "Samsung Galaxy".
[1048] 6. The server compiles the answers and links and sends them to the device in the format, "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details."
[1049] 7. The device displays this answer and link to the user.
[1050] Examples of prompts to input into a generative AI model:
[1051] Question: What smartphone do you recommend?
[1052] answer:
[1053] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1054] Step 1:
[1055] The user enters the question in natural language.
[1056] Input: The user enters the question "What smartphone do you recommend?" in a smartphone application.
[1057] Output: The entered question is sent to the terminal.
[1058] Specific action: The user enters a question into the app's interface and taps the submit button.
[1059] Step 2:
[1060] The terminal sends the entered question to the server.
[1061] Input: A natural language question entered by the user.
[1062] Output: The question content is sent to the server.
[1063] Specific operation: The terminal receives user input and forwards the question to the server as an HTTP request.
[1064] Step 3:
[1065] The server performs natural language processing on the question using parsing tools.
[1066] Input: A natural language question sent from the device.
[1067] Output: Analyzed keywords and their meanings.
[1068] Specific operation: The server uses a natural language processing engine (e.g., SpaCy) to analyze the question and extract keywords such as "recommended" and "smartphone".
[1069] Step 4:
[1070] The server generates an answer based on the analysis results.
[1071] Input: Keywords and meanings of the analyzed question.
[1072] Output: The generated answer.
[1073] Specific operation: The server inputs the analysis results into an artificial intelligence model (e.g., OpenAI's GPT-4) and generates responses such as, "We recommend the latest iPhone or Samsung Galaxy."
[1074] Step 5:
[1075] The server searches for relevant links based on the generated response.
[1076] Input: The content of the generated response.
[1077] Output: Related product links.
[1078] Specific operation: The server uses web scraping libraries (e.g., requests and BeautifulSoup) to collect product links related to "iPhone" and "Samsung Galaxy" from the internet.
[1079] Step 6:
[1080] The server outputs the answer and the link together.
[1081] Input: Generated responses and collected product links.
[1082] Output: The final response to be displayed to the user.
[1083] Specific operation: The server combines the answer text and each link into a single response and sends it back to the terminal as an HTTP response.
[1084] Step 7:
[1085] The device displays the response and link from the server.
[1086] Input: The final response sent from the server.
[1087] Output: Answers and links that users can view.
[1088] Specific action: The response received by the device is displayed on the user interface, and the message "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details." is presented.
[1089] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1090] This invention combines an information retrieval system that provides appropriate answers using generative AI and simultaneously presents related links based on a single question entered by the user, with an emotion engine that recognizes the user's emotions. This system is implemented in the following specific forms.
[1091] User
[1092] The user first inputs a question in natural language into the terminal's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server. It also recognizes emotions using the user's input and voice interface.
[1093] terminal
[1094] The terminal receives user input and sends it to the server as a request. The request includes the user's question and sentiment data, and the server receives this request.
[1095] server
[1096] The server processes the following steps.
[1097] 1. Analysis of the question:
[1098] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[1099] 2. Recognition of emotions:
[1100] The server uses emotion recognition mechanisms to analyze the user's emotional data and recognize the user's emotional state. For example, emotions such as "joy," "sadness," and "surprise" can be identified.
[1101] 3. Generating the answer:
[1102] The server invokes a generation mechanism based on the analysis results and recognized emotions, and uses an artificial intelligence model to generate an appropriate response. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated with a tone and content that matches the user's emotions.
[1103] 4. Search for related links:
[1104] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[1105] 5. Summary and output of results:
[1106] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[1107] Specific example
[1108] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[1109] 1. The user enters the question "What are some good tourist spots in Kyoto?" into the device. The user's input also detects the emotion "fun".
[1110] 2. The device sends the question and sentiment data to the server.
[1111] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[1112] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[1113] 5. Based on the analysis results and sentiment data, the server uses a generation method to generate a sentiment-appropriate response from an artificial intelligence model, such as "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[1114] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[1115] 7. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. You're sure to have a wonderful time! Please see the link below for more details."
[1116] 8. The device displays this answer and link to the user.
[1117] This invention allows users to quickly obtain the necessary information with a single question, receive empathetic answers, and make the search process more user-friendly.
[1118] The following describes the processing flow.
[1119] Step 1:
[1120] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[1121] Step 2:
[1122] The terminal retrieves the input. The terminal retrieves the user's input, and in addition, collects sentiment data obtained from the voice interface to construct the request.
[1123] Step 3:
[1124] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request to the server containing the user's question and sentiment data.
[1125] Step 4:
[1126] The server receives the request. The server receives the request from the terminal and prepares to analyze the question content and sentiment data.
[1127] Step 5:
[1128] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[1129] Step 6:
[1130] The server analyzes emotional data. Using emotion recognition tools, the server analyzes emotions from user input and voice data to recognize the user's emotional state. For example, emotions such as "excitement" and "anticipation" may be identified.
[1131] Step 7:
[1132] The server sends a question to a generative AI. Based on the analysis results and recognized emotions, the server sends the question to the generative AI (e.g., GPT-4) to generate an appropriate answer that is sensitive to those emotions.
[1133] Step 8:
[1134] A generative AI generates the answer. The generative AI generates an answer that matches the emotion, such as, "For ramen restaurants in Tokyo, I recommend Ichiran and Tsukemen Daioh. You're sure to have a wonderful time!"
[1135] Step 9:
[1136] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[1137] Step 10:
[1138] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[1139] Step 11:
[1140] The server sends a response to the terminal. The response contains the generated answer and related links.
[1141] Step 12:
[1142] The terminal receives a response. The terminal receives a response from the server.
[1143] Step 13:
[1144] The device displays the results to the user. The device displays the received response and link to the user, showing the response and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. You're sure to have a wonderful time! Please see the link below for details."
[1145] This allows users to quickly obtain answers and relevant links to their questions, and also provides emotionally resonant responses, making the search experience more user-friendly and satisfying.
[1146] (Example 2)
[1147] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1148] Traditional information retrieval systems required users to go through multiple steps to obtain appropriate answers after entering a question, resulting in a complex user experience. Furthermore, they lacked the ability to recognize and respond to user emotions, failing to adequately enhance user satisfaction. Therefore, there is a need for a system that provides appropriate answers and related information with a single question, and even provides answers that are sensitive to the user's emotions.
[1149] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1150] In this invention, the server includes an input means, a means for receiving questions and sentiment data transmitted from a terminal, an analysis means using a natural language processing engine for analyzing the input questions, an analysis means using an sentiment recognition engine for analyzing the sentiment data, a generation means using an artificial intelligence model for generating answers based on the analyzed question content and sentiment data, a search means using a search engine for searching the internet for links related to the generated answers, and an output means for sending the answers and related links together to the terminal. As a result, the user can obtain quick and appropriate answers and related information with a single question, and furthermore, answers that are sensitive to the user's emotions can be provided.
[1151] "Input method" refers to a device or software that provides an interface for users to input questions.
[1152] "Receiving means" refers to the function used to retrieve questions and sentiment data sent from a terminal on the server side.
[1153] A "natural language processing engine" refers to software or algorithms that analyze input natural language questions and extract their main keywords and meanings.
[1154] "Analysis means" refers to methods and techniques for analyzing received data and extracting necessary information.
[1155] An "emotion recognition engine" refers to software or algorithms that analyze and identify emotional states from user input data.
[1156] "Generation means" refers to functions and technologies that use artificial intelligence models to generate answers based on analyzed question content and sentiment data.
[1157] An "artificial intelligence model" refers to a system that has been trained using technologies such as machine learning and deep learning, and is capable of generating appropriate answers to questions.
[1158] A "search engine" refers to software or services used to search the internet for and collect links related to generated answers.
[1159] "Search methods" refer to methods and techniques for collecting relevant information.
[1160] "Output means" refers to methods and technologies for compiling generated answers and related links and providing them to the user.
[1161] A "terminal" refers to a device used by a user to input questions or view information received from a server.
[1162] A "question" refers to the content entered by the user in natural language regarding the information they want to know.
[1163] "Emotional data" refers to information indicating the emotional state, extracted from user input and voice.
[1164] This invention combines an information retrieval system that provides an appropriate answer using generative artificial intelligence (AI) and simultaneously presents related links based on the user's input of a single question, with an emotion engine that recognizes the user's emotions. A specific embodiment of this system is described below.
[1165] User actions
[1166] The user first inputs a question in natural language through the device's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" At this time, the device acquires the user's input and uses the voice interface to recognize the user's emotions. For example, the tone of voice and facial expressions while the user is inputting may be used to detect that the user is "happy."
[1167] Terminal operation
[1168] The terminal combines the user's input question and recognized sentiment data, and sends this to the server. Here, it plays the role of sending the question content and sentiment data as a single request to the server. Specific devices used include personal computers, smartphones, and tablets.
[1169] Server Processing
[1170] The server takes several steps to process an incoming request. First, it extracts key keywords and meanings by analyzing the question using a natural language processing engine (for example, the Google Cloud Natural Language API). For example, from the question "What are some recommended ramen restaurants in Tokyo?", the keywords extracted would be "Tokyo," "recommended," and "ramen restaurant."
[1171] Next, an emotion recognition engine (for example, IBM Watson Tone Analyzer) is used to analyze the user's emotional data. For instance, emotions such as "happy" are identified from the user's voice tone and text.
[1172] Based on the analysis results, the server invokes a generation mechanism and uses an artificial intelligence model (e.g., OpenAI's GPT-3) to generate an appropriate response. This generated response is delivered in a content and tone that matches the user's emotions. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated in a way that resonates with the user's "happy" feelings.
[1173] Based on the generated response, the server uses search methods to find relevant links. Specifically, it uses a search engine (for example, the Google Search API) to collect links to relevant official websites and review sites. For example, it searches for links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[1174] Finally, the server combines the generated answers and collected links to create a response to send to the terminal. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[1175] Specific example
[1176] For example, consider a case where a user enters the question, "What are some good tourist spots in Kyoto?" The emotion "fun" is detected simultaneously with the user's input.
[1177] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[1178] 2. The device sends the question and sentiment data to the server.
[1179] 3. The server analyzes the question and extracts the keywords "Kyoto," "tourist attractions," and "recommendations."
[1180] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[1181] 5. Based on the analysis results and sentiment data, the server uses an artificial intelligence model to generate the response, "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[1182] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[1183] 7. The server compiles the answers and links and sends them to the device.
[1184] 8. The device displays this answer and link to the user.
[1185] Example of a prompt:
[1186] "What are some good tourist spots in Kyoto? I'm in a good mood."
[1187] This invention enables users to quickly obtain appropriate answers and related information with a single question, and further provides a search experience that is sensitive to the user's emotions.
[1188] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1189] Step 1:
[1190] The user enters a question into the terminal. Specifically, the user uses text input or voice input to enter a question into the interface, for example, "What are some recommended ramen restaurants in Tokyo?" The input is the text data of the question.
[1191] Step 2:
[1192] The device receives user input text and also acquires emotion data. In the case of voice input, it uses a speech recognition engine to convert it to text, and at the same time, an emotion recognition engine analyzes the emotion. For example, it might acquire the emotion "happy" as text data. The output consists of the converted text data and emotion data.
[1193] Step 3:
[1194] The device sends the question content and sentiment data to the server. The data sent includes the question text and request data containing sentiment data. Specifically, the request is sent to the server using the HTTP protocol.
[1195] Step 4:
[1196] The server receives the request data. The server analyzes the received question text and extracts key keywords. Specifically, it uses a natural language processing engine (e.g., Google Cloud Natural Language API) to extract keywords such as "Tokyo," "recommended," and "ramen restaurant" from the question text. The output is the extracted keywords.
[1197] Step 5:
[1198] The server analyzes emotional data to identify the user's emotions. Specifically, it uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the emotional data. The output is the user's emotional state. For example, it might retrieve data indicating "happy."
[1199] Step 6:
[1200] The server generates a response based on the analysis results (extracted keywords and sentiment data). An artificial intelligence model (e.g., OpenAI's GPT-3) is used as the generation method. The input consists of extracted keywords and sentiment data. For example, it might generate the response text: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The output is the generated response text.
[1201] Step 7:
[1202] The server searches for relevant links based on the generated response. It uses a search engine (e.g., Google Search API) to perform a web search using words included in the generated response. For example, it collects links for "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". The output is a list of relevant links.
[1203] Step 8:
[1204] The server combines the generated answers and collected links to create the final response. Specifically, it combines the answer text and the list of links into a single response data. For example, it might create a response in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." The output is the combined response data.
[1205] Step 9:
[1206] The server sends the final response data to the terminal. It sends the response data back to the terminal using a specific communication protocol. The output indicates that the transmission of the response data is complete.
[1207] Step 10:
[1208] The terminal displays response data obtained from the server to the user. Specifically, it either displays the response on the screen interface or reads it aloud. For example, it might display something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen! For more details, please refer to the link below." The output consists of the response displayed to the user and the link.
[1209] This specific processing flow allows users to quickly obtain appropriate answers and related information with a single question, and furthermore, to receive emotionally resonant responses.
[1210] (Application Example 2)
[1211] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1212] Traditional information retrieval systems provide answers to user questions, but they lack the ability to generate answers that take user emotions into consideration. Furthermore, their ability to provide related links to the generated answers is limited. As a result, users often experience low satisfaction when obtaining answers, and the search experience remains unimproved.
[1213] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question and the user's emotions, a generation means for generating an answer based on the analyzed question content and emotion data, a search means for searching for links related to the generated answer, and an output means for outputting the answer and related links together. This makes it possible to generate an answer that is sensitive to the user's emotions and to provide related information quickly.
[1214] A "question" is the content that users input in natural language to find out what they want to know.
[1215] "Input means" refers to the devices or software that users use to input questions.
[1216] "Analysis means" refers to a device or program for analyzing input questions and user sentiment data.
[1217] "Generation means" refers to a device or program that generates answers based on analyzed question content and sentiment data.
[1218] "Search means" refers to a device or program for searching the internet for links related to the generated answer.
[1219] "Output means" refers to a device or program for displaying answers and related links to the user.
[1220] "Emotional data" refers to emotional information extracted from user input.
[1221] An "artificial intelligence model" is a model trained based on machine learning or deep learning, and is used as a means of generation.
[1222] A "natural language processing engine" is a program that analyzes text input in natural language and extracts meaning and keywords.
[1223] "Related links" refer to URLs of web pages or information that are relevant to the generated response.
[1224] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system incorporates an emotion engine that recognizes the user's emotions and generates answers in a tone and content that corresponds to the user's emotions.
[1225] System Configuration
[1226] This system consists of the following main components:
[1227] 1. Input method: This refers to the means by which the user inputs questions, and includes devices such as smartphones and smart glasses.
[1228] 2. Analysis means: This refers to means for analyzing the input questions and sentiment data, and corresponds to natural language processing engines (e.g., Transformers, spaCy) or sentiment recognition libraries (e.g., emote).
[1229] 3. Generation means: This refers to means for generating answers based on the analyzed question content and sentiment data, and a generative AI model (e.g., GPT-3) falls under this category.
[1230] 4. Search methods: These are methods for searching the internet for links related to the generated answers, and web scraping libraries (BeautifulSoup, Scrapy) are examples of this.
[1231] 5. Output means: This refers to a means of displaying the answer and related links to the user, and this includes the device's display and notification function.
[1232] Processing flow
[1233] As a concrete example, the "emotion-recognition store navigator" application using smart glasses operates in the following steps.
[1234] 1. User input: The user uses smart glasses to input a question in natural language. For example, they might ask, "What product would suit my current mood?"
[1235] 2. Emotion Recognition: The device receives user input and analyzes emotions using an emotion recognition library (such as emote).
[1236] 3. Request to the server: Send the analyzed question and sentiment data to the server.
[1237] 4. Question Analysis: The server uses a natural language processing engine (such as Transformers or spaCy) to analyze the question and extract key keywords.
[1238] 5. Response generation: Use a generative AI model (such as GPT-3) to generate appropriate responses based on analysis results and sentiment data.
[1239] 6. Search for related links: Based on the generated answers, use a web scraping library (such as BeautifulSoup or Scrapy) to search the internet for related links.
[1240] 7. Displaying Results: The answers and related links are compiled and displayed on the smart glasses' screen.
[1241] Examples of specific cases and prompt statements
[1242] As a concrete example, if a user feels tired and asks, "What product would be perfect for how I'm feeling right now?", the following process would occur:
[1243] Input text: "Please recommend a product that perfectly matches my current mood."
[1244] Emotion recognition results: "Fatigue" and "Stress"
[1245] Server output: "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details."
[1246] Related links: "Aroma Candle Details Page", "Massage Chair Purchase Page"
[1247] Examples of prompt messages are as follows:
[1248] "What are some relaxing products suitable for relieving fatigue?"
[1249] "What are some good relaxation items for when you're tired?"
[1250] This system allows users to not only get quick answers to their questions, but also experience friendly and personalized responses tailored to their own emotions.
[1251] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1252] Step 1: User Input
[1253] The user uses smart glasses to input a question in natural language. The input might be something like, "Tell me what product would suit my current mood." The user asks the question via voice or text input. At this time, the smart glasses interface receives the question and retrieves the data.
[1254] Step 2: Recognizing Emotions
[1255] The device receives user input and analyzes emotions using an emotion recognition library (e.g., Emote). In this step, emotion data is extracted from the input text to recognize emotional states such as "fatigue" or "stress." The input data is in text format, and the output data consists of emotion labels and emotion scores.
[1256] Step 3: Request to the server
[1257] The terminal sends the analyzed question and sentiment data to the server. Specifically, it sends the question text and sentiment data, converted to JSON format, to the server via an HTTP POST request. The input data is in JSON format, and the output data is the response from the server.
[1258] Step 4: Analyzing the Question
[1259] The server uses a natural language processing engine (e.g., Transformers, spaCy) to analyze the received question and extract key keywords. For example, from the question "Tell me a product that suits my current mood," it extracts the keywords "mood" and "product." The input data is the question text in JSON format, and the output data is the analyzed keywords.
[1260] Step 5: Generating the answer
[1261] The server generates an appropriate answer using a generative AI model (e.g., GPT-3) based on the analysis results of the question and sentiment data. In this step, the generative AI model generates an answer based on the input data (analyzed keywords and sentiment data). For example, an answer such as "Relaxing aromatherapy candles and massage chairs are recommended" might be generated. The input data consists of keywords and sentiment data, and the output data is the generated answer text.
[1262] Step 6: Search for related links
[1263] The server uses a web scraping library (e.g., BeautifulSoup, Scrapy) to search the internet for relevant links based on the generated response. For example, it collects relevant links such as "details page for aromatherapy candles" and "purchase page for massage chairs." The input data is the generated response text, and the output data is a list of relevant links.
[1264] Step 7: Displaying the results
[1265] The server compiles the answers and related links and sends them to the device. The device displays the received answers and links to the user. At this time, the smart glasses display shows a message along with a link that reads, "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details." The input data is the response from the server, and the output data is information that the user can visually confirm.
[1266] As described above, a series of processes, from user questioning and sentiment recognition to natural language processing, the use of generative AI models, the search for related links, and output, are performed, making it possible to provide users with useful information that resonates with their emotions.
[1267] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1268] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1269] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1270] [Fourth Embodiment]
[1271] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1272] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1273] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1274] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1275] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1276] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1277] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1278] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1279] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1280] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1281] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1282] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1283] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1284] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system is implemented as follows:
[1285] User
[1286] The user first enters a question in natural language into the terminal's interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server.
[1287] terminal
[1288] The terminal receives the user's input and sends it to the server as a request. The request includes the user's question, and the server receives this request.
[1289] server
[1290] The server processes the following steps.
[1291] 1. Analysis of the question:
[1292] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[1293] 2. Generating the answer:
[1294] The server then calls a generation mechanism based on the analysis results and uses an artificial intelligence model to generate an appropriate answer. For example, it might generate an answer such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[1295] 3. Search for related links:
[1296] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[1297] 4. Summary and output of results:
[1298] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[1299] Specific example
[1300] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[1301] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[1302] 2. The device sends the question to the server.
[1303] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[1304] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the artificial intelligence model.
[1305] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[1306] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[1307] 7. The device displays this answer and link to the user.
[1308] This invention allows users to quickly obtain the necessary information with a single question, significantly simplifying the search process and improving efficiency.
[1309] The following describes the processing flow.
[1310] Step 1:
[1311] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[1312] Step 2:
[1313] The terminal retrieves the input. The terminal retrieves the user's input and constructs it as a request.
[1314] Step 3:
[1315] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request that includes the user's question.
[1316] Step 4:
[1317] The server receives the request. The server receives the request from the terminal and prepares to analyze the question.
[1318] Step 5:
[1319] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, the keywords "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[1320] Step 6:
[1321] The server sends a question to a generative AI. Based on the analysis results, the server sends another question to the generative AI (e.g., GPT-4) to generate an appropriate answer.
[1322] Step 7:
[1323] A generative AI generates an answer. The generative AI generates the answer, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[1324] Step 8:
[1325] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[1326] Step 9:
[1327] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[1328] Step 10:
[1329] The server sends a response to the terminal. The response contains the generated answer and related links.
[1330] Step 11:
[1331] The terminal receives a response. The terminal receives a response from the server.
[1332] Step 12:
[1333] The device displays the results to the user. The device displays the received answer and link to the user, showing the answer and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[1334] This allows users to quickly obtain answers to their questions and relevant links.
[1335] (Example 1)
[1336] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1337] Traditional information retrieval systems required users to perform numerous manual operations and multiple searches to efficiently obtain the information they needed, resulting in a cumbersome and time-consuming process. Furthermore, collecting links related to the generated answers required a separate effort, leading to poor user convenience.
[1338] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1339] In this invention, the server includes an input means, a terminal for receiving questions entered by the user through the input means, an analysis means for analyzing the questions transmitted from the terminal, a generation means for generating answers based on the analyzed question content, a search means for searching for links related to the answers generated by the generation means, and an output means for sending the generated answers and related links together to the terminal. As a result, the user can quickly obtain the necessary information by simply entering a single question, and related links are also provided at the same time, enabling efficient information retrieval.
[1340] "Input method" refers to the interface through which a user enters a question into the system.
[1341] A "terminal" refers to a device used by a user to input a question and send it to a server for processing.
[1342] "Analysis means" refers to a function that analyzes questions sent by users through their devices and extracts their content and key keywords.
[1343] "Generation means" refers to a function that generates appropriate answers based on keywords and content extracted by the analysis means.
[1344] "Search means" refers to a function for searching the internet for and collecting links related to the answers generated by the generation means.
[1345] "Output means" refers to a function that aggregates the generated responses and collected related links, sends them to the terminal, and displays them to the user.
[1346] A "natural language processing engine" refers to software or algorithms that analyze an input question and extract key keywords and content.
[1347] A "generative AI model" refers to a machine learning model that uses artificial intelligence technology to generate appropriate answers to user questions.
[1348] An "HTTP request" refers to a communication protocol used to send data to a web server.
[1349] "HTTP response" refers to the communication protocol used to send back data received from a web server.
[1350] A "prompt sentence" refers to a sentence used as an instruction to input into a generative AI model.
[1351] This invention is an information retrieval system that provides an appropriate answer using a generative AI model and simultaneously presents related links, based on the user's input of a single question. Specific embodiments of this system are described below.
[1352] Hardware and software configuration
[1353] This system consists of the following main components:
[1354] Device: A device used by the user to enter questions. This includes PCs, smartphones, tablets, etc.
[1355] Server: The central processing unit that analyzes questions, generates answers, searches for relevant links, and provides responses to the user. Software running on the server includes natural language processing engines, generative AI models, and search engine APIs.
[1356] Detailed processing of the system
[1357] The server performs the following steps:
[1358] 1. Receiving and sending questions:
[1359] The user enters a question into the terminal's interface, such as, "What are some recommended ramen restaurants in Tokyo?"
[1360] The device retrieves this question and sends it to the server as an HTTP POST request.
[1361] 2. Analysis of the question:
[1362] The server processes the received question using parsing tools. Specifically, it tokenizes the question using a natural language processing engine (e.g., Spacy, NLTK) and extracts key keywords and meanings.
[1363] For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" are extracted.
[1364] 3. Generating the answer:
[1365] The server invokes a generation mechanism based on the analysis results and uses a generation AI model (e.g., GPT-4, GPT-3) to generate an appropriate response.
[1366] For example, the system might generate responses such as, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[1367] 4. Search for related links:
[1368] The server searches for relevant links based on the generated response. It collects the appropriate links from the internet using search methods (e.g., Google Search API, Bing Search API).
[1369] For example, we collect links to official websites and review sites related to "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[1370] 5. Generating and sending results:
[1371] The server combines the generated answers and collected links to create the final response.
[1372] For example, a response in the format of "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." is sent to the terminal as an HTTP response.
[1373] 6. Displaying the results:
[1374] The terminal receives a response from the server and displays it in the user interface.
[1375] Through this, users can view the answers and related links.
[1376] Examples of specific actions
[1377] For example, if a user enters "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[1378] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[1379] 2. The device sends the question to the server.
[1380] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[1381] 4. Based on the analysis results, the server uses a generation method to generate the answer "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama" from the generation AI model.
[1382] 5. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[1383] 6. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist attractions in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. Please see the link below for details."
[1384] 7. The device displays this answer and link to the user.
[1385] Example of a prompt
[1386] Here are some specific examples of prompt statements to input into a generative AI model:
[1387] "What are some recommended ramen restaurants in Tokyo?"
[1388] "What are some good tourist spots in Kyoto?"
[1389] This allows users to quickly obtain the information they need with a single question, enabling efficient information retrieval.
[1390] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1391] Step 1: User enters question
[1392] The user enters a question in natural language into the terminal interface. For example, they might enter a question like, "What are some recommended ramen restaurants in Tokyo?" into the text field. The terminal retrieves the user's input (question text) so that it can be processed in the next step.
[1393] Input: User's question (natural language text)
[1394] Output: Retrieving questions via the terminal
[1395] Specific operation: The user enters a question into the terminal's text field. The terminal stores this input in memory and prepares to send it in the next step.
[1396] Step 2: Sending a request via the terminal
[1397] The terminal sends the question received from the user to the server as an HTTP POST request. The request contains the question text.
[1398] Input: User's question (text stored on the device)
[1399] Output: HTTP POST request to the server
[1400] Specific operation: The terminal includes the retrieved question text in the payload of an HTTP POST request and sends it to a specific endpoint. The server receives this request.
[1401] Step 3: Server analyzes the question
[1402] The server uses a natural language processing engine (e.g., Spacy, NLTK) to analyze the received question text. This analysis extracts key keywords and meanings from the text.
[1403] Input: User's question (payload of HTTP POST request)
[1404] Output: Analyzed keywords and meanings
[1405] Specific operation: The server extracts the question text from the request payload and passes it to the natural language processing engine. The engine tokenizes the text and extracts keywords such as "Tokyo," "recommended," and "ramen restaurant."
[1406] Step 4: Server generates response
[1407] The server generates appropriate answers using a generative AI model (e.g., GPT-4, GPT-3) based on the analysis results. A prompt is input to the generative AI model, and the answer text is retrieved based on that input.
[1408] Input: Analyzed keywords and meanings
[1409] Output: Generated answer (text)
[1410] Specific operation: The server uses the analyzed keywords to create a prompt for the generative AI model and inputs a question such as "What are some recommended ramen restaurants in Tokyo?" into the model. The generative AI model responds by outputting text such as "Some recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh."
[1411] Step 5: Server searches for related links
[1412] The server searches the internet for relevant links based on the generated response. Links are collected using search methods (e.g., Google Search API, Bing Search API).
[1413] Input: Generated response (text)
[1414] Output: List of related links
[1415] Specific operation: The server extracts keywords from the generated response text and creates search queries such as "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". These queries are then fed into a search API to retrieve relevant links (e.g., official website, review site).
[1416] Step 6: Server summarizes and sends the results.
[1417] The server combines the generated answers and collected links to create a final response. This response is then sent to the terminal as an HTTP response.
[1418] Input: List of generated answers (text) and related links
[1419] Output: Final response (text and links)
[1420] Specific operation: The server combines the answer text and related links to prepare a response in the format, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." This is then sent to the terminal as an HTTP response.
[1421] Step 7: Display to the user via the device
[1422] The terminal displays the response received from the server in the user interface. Through this, the user can view the answer and related links.
[1423] Input: Final response from the server (text and links)
[1424] Output: Content displayed to the user
[1425] Specific operation: The terminal analyzes the received response and extracts the answer and link. These are then displayed using HTML or GUI components, making them viewable by the user.
[1426] As described above, users can quickly obtain the necessary information with a single question, and relevant links are provided simultaneously, enabling efficient information retrieval.
[1427] (Application Example 1)
[1428] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1429] Traditional information retrieval systems had the problem of requiring a lot of effort and time for users to find the right product even after entering a product name or specific characteristics. Furthermore, the limited functionality for displaying related links and product information in a single batch meant that user search efficiency was not improved.
[1430] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1431] In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question, a generation means for generating an answer based on the analyzed question content, a search means for searching for links related to the generated answer, an output means for outputting the answer and related links together, and a product link display means for displaying the corresponding product link along with the generated answer. This makes it possible to quickly display appropriate product suggestions and links to related products when a user inputs a question related to a product.
[1432] "Input means for entering a question" refers to a device or interface for a user to enter a question in natural language.
[1433] "Analysis means for analyzing input questions" refers to means that analyze input questions and use natural language processing techniques to extract key keywords and meanings.
[1434] "A means for generating answers based on analyzed question content" refers to a means that utilizes an artificial intelligence model to generate appropriate answers based on information extracted by the analysis means.
[1435] "A search method for finding links related to generated answers" refers to a method for searching the internet for information related to generated answers and collecting the relevant links.
[1436] "An output method for outputting answers and related links together" refers to a method for integrating the generated answers and collected related links and displaying them to the user.
[1437] "A means for displaying product links along with the generated response" refers to a means for additionally displaying links to products related to the generated response.
[1438] A "natural language processing engine" is a software engine that analyzes input natural language text and extracts its meaning and keywords.
[1439] An "artificial intelligence model" is a machine learning model that uses input data to perform reasoning like a human and generate appropriate answers.
[1440] This invention is an information retrieval system that can be applied to an e-commerce site application that provides appropriate answers and links to related products when a user enters a question related to a product.
[1441] System Configuration
[1442] 1. User Interface
[1443] The application includes an input method for users to enter questions in natural language using a smartphone application. For example, it would be a section where users can enter questions such as, "What smartphone do you recommend?"
[1444] 2. Analysis of the Question
[1445] The terminal retrieves the entered question and sends it to the server. The server uses a natural language processing engine (e.g., SpaCy) to analyze the entered question, extracting key keywords and meaning from it.
[1446] 3. Generating the answer
[1447] The server generates appropriate answers to questions using artificial intelligence models (e.g., OpenAI's GPT-4) based on the analyzed question content. This generation mechanism has the ability to perform real-time reasoning in response to user input and provide relevant information.
[1448] 4. Search for related links
[1449] The server includes a search mechanism for searching the internet for relevant product links based on the generated response. This search mechanism uses, for example, the BeautifulSoup or requests library to retrieve links to relevant products.
[1450] 5. Output of answers and links
[1451] The server has an output mechanism to combine the generated answer and related product links into a single response and send it to the terminal. The terminal displays this response to the user. This allows the user to view the answer to the question along with detailed links to related products all at once.
[1452] Hardware and software used
[1453] Hardware: Smartphone
[1454] software:
[1455] OpenAI API: GPT model for question analysis and answer generation
[1456] requests library: Used to retrieve links from websites.
[1457] BeautifulSoup: To parse links from retrieved HTML.
[1458] Natural language processing engines: SpaCy, etc.
[1459] Specific example
[1460] For example, if a user types the question "Can you recommend a smartphone?", the following steps will be taken:
[1461] 1. The user enters the question "What smartphone do you recommend?" into their device.
[1462] 2. The device sends the question to the server.
[1463] 3. The server analyzes the question and extracts the main keywords "recommendation" and "smartphone".
[1464] 4. Based on the analysis results, the server uses a generation method to generate the response "The latest iPhone or Samsung Galaxy is recommended" from the artificial intelligence model.
[1465] 5. The server searches for links related to "iPhone" and "Samsung Galaxy".
[1466] 6. The server compiles the answers and links and sends them to the device in the format, "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details."
[1467] 7. The device displays this answer and link to the user.
[1468] Examples of prompts to input into a generative AI model:
[1469] Question: What smartphone do you recommend?
[1470] answer:
[1471] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1472] Step 1:
[1473] The user enters the question in natural language.
[1474] Input: The user enters the question "What smartphone do you recommend?" in a smartphone application.
[1475] Output: The entered question is sent to the terminal.
[1476] Specific action: The user enters a question into the app's interface and taps the submit button.
[1477] Step 2:
[1478] The terminal sends the entered question to the server.
[1479] Input: A natural language question entered by the user.
[1480] Output: The question content is sent to the server.
[1481] Specific operation: The terminal receives user input and forwards the question to the server as an HTTP request.
[1482] Step 3:
[1483] The server performs natural language processing on the question using parsing tools.
[1484] Input: A natural language question sent from the device.
[1485] Output: Analyzed keywords and their meanings.
[1486] Specific operation: The server uses a natural language processing engine (e.g., SpaCy) to analyze the question and extract keywords such as "recommended" and "smartphone".
[1487] Step 4:
[1488] The server generates an answer based on the analysis results.
[1489] Input: Keywords and meanings of the analyzed question.
[1490] Output: The generated answer.
[1491] Specific operation: The server inputs the analysis results into an artificial intelligence model (e.g., OpenAI's GPT-4) and generates responses such as, "We recommend the latest iPhone or Samsung Galaxy."
[1492] Step 5:
[1493] The server searches for relevant links based on the generated response.
[1494] Input: The content of the generated response.
[1495] Output: Related product links.
[1496] Specific operation: The server uses web scraping libraries (e.g., requests and BeautifulSoup) to collect product links related to "iPhone" and "Samsung Galaxy" from the internet.
[1497] Step 6:
[1498] The server outputs the answer and the link together.
[1499] Input: Generated responses and collected product links.
[1500] Output: The final response to be displayed to the user.
[1501] Specific operation: The server combines the answer text and each link into a single response and sends it back to the terminal as an HTTP response.
[1502] Step 7:
[1503] The device displays the response and link from the server.
[1504] Input: The final response sent from the server.
[1505] Output: Answers and links that users can view.
[1506] Specific action: The response received by the device is displayed on the user interface, and the message "We recommend the latest iPhone or Samsung Galaxy. Please see the link below for details." is presented.
[1507] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1508] This invention combines an information retrieval system that provides appropriate answers using generative AI and simultaneously presents related links based on a single question entered by the user, with an emotion engine that recognizes the user's emotions. This system is implemented in the following specific forms.
[1509] User
[1510] The user first inputs a question in natural language into the terminal's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" The terminal receives this input and prepares to send the question to the server. It also recognizes emotions using the user's input and voice interface.
[1511] terminal
[1512] The terminal receives user input and sends it to the server as a request. The request includes the user's question and sentiment data, and the server receives this request.
[1513] server
[1514] The server processes the following steps.
[1515] 1. Analysis of the question:
[1516] The server first processes the received question using an analysis tool and runs it through a natural language processing engine. This extracts the main keywords and meaning of the question. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[1517] 2. Recognition of emotions:
[1518] The server uses emotion recognition mechanisms to analyze the user's emotional data and recognize the user's emotional state. For example, emotions such as "joy," "sadness," and "surprise" can be identified.
[1519] 3. Generating the answer:
[1520] The server invokes a generation mechanism based on the analysis results and recognized emotions, and uses an artificial intelligence model to generate an appropriate response. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated with a tone and content that matches the user's emotions.
[1521] 4. Search for related links:
[1522] Based on the generated response, the server searches for relevant links. It uses search methods to collect the appropriate links from the internet. For example, it collects links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[1523] 5. Summary and output of results:
[1524] The server combines the generated answers and collected links to create the final response. This response is sent to the terminal for the user to view. For example, it might be provided in the format: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[1525] Specific example
[1526] For example, if a user enters the question, "What are some good tourist spots in Kyoto?", the specific actions would be as follows:
[1527] 1. The user enters the question "What are some good tourist spots in Kyoto?" into the device. The user's input also detects the emotion "fun".
[1528] 2. The device sends the question and sentiment data to the server.
[1529] 3. The server analyzes the question and extracts the main keywords "Kyoto," "tourist attractions," and "recommendations."
[1530] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[1531] 5. Based on the analysis results and sentiment data, the server uses a generation method to generate a sentiment-appropriate response from an artificial intelligence model, such as "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[1532] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[1533] 7. The server compiles the answers and links and sends them to the device in the format: "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. You're sure to have a wonderful time! Please see the link below for more details."
[1534] 8. The device displays this answer and link to the user.
[1535] This invention allows users to quickly obtain the necessary information with a single question, receive empathetic answers, and make the search process more user-friendly.
[1536] The following describes the processing flow.
[1537] Step 1:
[1538] The user enters a question. The user enters "What are some recommended ramen restaurants in Tokyo?" into the input field on the device.
[1539] Step 2:
[1540] The terminal retrieves the input. The terminal retrieves the user's input, and in addition, collects sentiment data obtained from the voice interface to construct the request.
[1541] Step 3:
[1542] The terminal sends a request to the server. The terminal establishes a connection to the server and sends a request to the server containing the user's question and sentiment data.
[1543] Step 4:
[1544] The server receives the request. The server receives the request from the terminal and prepares to analyze the question content and sentiment data.
[1545] Step 5:
[1546] The server analyzes the question. The server uses a natural language processing engine to analyze the received question and extract key keywords. For example, keywords such as "Tokyo," "recommended," and "ramen restaurant" might be extracted.
[1547] Step 6:
[1548] The server analyzes emotional data. Using emotion recognition tools, the server analyzes emotions from user input and voice data to recognize the user's emotional state. For example, emotions such as "excitement" and "anticipation" may be identified.
[1549] Step 7:
[1550] The server sends a question to a generative AI. Based on the analysis results and recognized emotions, the server sends the question to the generative AI (e.g., GPT-4) to generate an appropriate answer that is sensitive to those emotions.
[1551] Step 8:
[1552] A generative AI generates the answer. The generative AI generates an answer that matches the emotion, such as, "For ramen restaurants in Tokyo, I recommend Ichiran and Tsukemen Daioh. You're sure to have a wonderful time!"
[1553] Step 9:
[1554] The server searches for relevant links. Based on the answers obtained from the generative AI, the server searches the internet for relevant links (for example, the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo").
[1555] Step 10:
[1556] The server compiles the results. The server combines the generated answers and searched links to create the final response.
[1557] Step 11:
[1558] The server sends a response to the terminal. The response contains the generated answer and related links.
[1559] Step 12:
[1560] The terminal receives a response. The terminal receives a response from the server.
[1561] Step 13:
[1562] The device displays the results to the user. The device displays the received response and link to the user, showing the response and link: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. You're sure to have a wonderful time! Please see the link below for details."
[1563] This allows users to quickly obtain answers and relevant links to their questions, and also provides emotionally resonant responses, making the search experience more user-friendly and satisfying.
[1564] (Example 2)
[1565] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1566] Traditional information retrieval systems required users to go through multiple steps to obtain appropriate answers after entering a question, resulting in a complex user experience. Furthermore, they lacked the ability to recognize and respond to user emotions, failing to adequately enhance user satisfaction. Therefore, there is a need for a system that provides appropriate answers and related information with a single question, and even provides answers that are sensitive to the user's emotions.
[1567] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1568] In this invention, the server includes an input means, a means for receiving questions and sentiment data transmitted from a terminal, an analysis means using a natural language processing engine for analyzing the input questions, an analysis means using an sentiment recognition engine for analyzing the sentiment data, a generation means using an artificial intelligence model for generating answers based on the analyzed question content and sentiment data, a search means using a search engine for searching the internet for links related to the generated answers, and an output means for sending the answers and related links together to the terminal. As a result, the user can obtain quick and appropriate answers and related information with a single question, and furthermore, answers that are sensitive to the user's emotions can be provided.
[1569] "Input method" refers to a device or software that provides an interface for users to input questions.
[1570] "Receiving means" refers to the function used to retrieve questions and sentiment data sent from a terminal on the server side.
[1571] A "natural language processing engine" refers to software or algorithms that analyze input natural language questions and extract their main keywords and meanings.
[1572] "Analysis means" refers to methods and techniques for analyzing received data and extracting necessary information.
[1573] An "emotion recognition engine" refers to software or algorithms that analyze and identify emotional states from user input data.
[1574] "Generation means" refers to functions and technologies that use artificial intelligence models to generate answers based on analyzed question content and sentiment data.
[1575] An "artificial intelligence model" refers to a system that has been trained using technologies such as machine learning and deep learning, and is capable of generating appropriate answers to questions.
[1576] A "search engine" refers to software or services used to search the internet for and collect links related to generated answers.
[1577] "Search methods" refer to methods and techniques for collecting relevant information.
[1578] "Output means" refers to methods and technologies for compiling generated answers and related links and providing them to the user.
[1579] A "terminal" refers to a device used by a user to input questions or view information received from a server.
[1580] A "question" refers to the content entered by the user in natural language regarding the information they want to know.
[1581] "Emotional data" refers to information indicating the emotional state, extracted from user input and voice.
[1582] This invention combines an information retrieval system that provides an appropriate answer using generative artificial intelligence (AI) and simultaneously presents related links based on the user's input of a single question, with an emotion engine that recognizes the user's emotions. A specific embodiment of this system is described below.
[1583] User actions
[1584] The user first inputs a question in natural language through the device's interface. For example, they might input, "What are some recommended ramen restaurants in Tokyo?" At this time, the device acquires the user's input and uses the voice interface to recognize the user's emotions. For example, the tone of voice and facial expressions while the user is inputting may be used to detect that the user is "happy."
[1585] Terminal operation
[1586] The terminal combines the user's input question and recognized sentiment data, and sends this to the server. Here, it plays the role of sending the question content and sentiment data as a single request to the server. Specific devices used include personal computers, smartphones, and tablets.
[1587] Server Processing
[1588] The server takes several steps to process an incoming request. First, it extracts key keywords and meanings by analyzing the question using a natural language processing engine (for example, the Google Cloud Natural Language API). For example, from the question "What are some recommended ramen restaurants in Tokyo?", the keywords extracted would be "Tokyo," "recommended," and "ramen restaurant."
[1589] Next, an emotion recognition engine (for example, IBM Watson Tone Analyzer) is used to analyze the user's emotional data. For instance, emotions such as "happy" are identified from the user's voice tone and text.
[1590] Based on the analysis results, the server invokes a generation mechanism and uses an artificial intelligence model (e.g., OpenAI's GPT-3) to generate an appropriate response. This generated response is delivered in a content and tone that matches the user's emotions. For example, the response might be something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The response is generated in a way that resonates with the user's "happy" feelings.
[1591] Based on the generated response, the server uses search methods to find relevant links. Specifically, it uses a search engine (for example, the Google Search API) to collect links to relevant official websites and review sites. For example, it searches for links to the official websites and review sites of "Ichiran Tokyo" and "Tsukemen Daioh Tokyo".
[1592] Finally, the server combines the generated answers and collected links to create a response to send to the terminal. For example, it might be provided in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details."
[1593] Specific example
[1594] For example, consider a case where a user enters the question, "What are some good tourist spots in Kyoto?" The emotion "fun" is detected simultaneously with the user's input.
[1595] 1. The user enters the question "What are some good tourist spots in Kyoto?" into their device.
[1596] 2. The device sends the question and sentiment data to the server.
[1597] 3. The server analyzes the question and extracts the keywords "Kyoto," "tourist attractions," and "recommendations."
[1598] 4. The server analyzes the emotion data and recognizes that the user's emotion is "happy".
[1599] 5. Based on the analysis results and sentiment data, the server uses an artificial intelligence model to generate the response, "Recommended tourist spots in Kyoto include Kinkaku-ji Temple, Kiyomizu-dera Temple, and Arashiyama. I think you'll have a wonderful time!"
[1600] 6. The server searches for links related to the information "Kinkaku-ji Temple," "Kiyomizu-dera Temple," and "Arashiyama."
[1601] 7. The server compiles the answers and links and sends them to the device.
[1602] 8. The device displays this answer and link to the user.
[1603] Example of a prompt:
[1604] "What are some good tourist spots in Kyoto? I'm in a good mood."
[1605] This invention enables users to quickly obtain appropriate answers and related information with a single question, and further provides a search experience that is sensitive to the user's emotions.
[1606] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1607] Step 1:
[1608] The user enters a question into the terminal. Specifically, the user uses text input or voice input to enter a question into the interface, for example, "What are some recommended ramen restaurants in Tokyo?" The input is the text data of the question.
[1609] Step 2:
[1610] The device receives user input text and also acquires emotion data. In the case of voice input, it uses a speech recognition engine to convert it to text, and at the same time, an emotion recognition engine analyzes the emotion. For example, it might acquire the emotion "happy" as text data. The output consists of the converted text data and emotion data.
[1611] Step 3:
[1612] The device sends the question content and sentiment data to the server. The data sent includes the question text and request data containing sentiment data. Specifically, the request is sent to the server using the HTTP protocol.
[1613] Step 4:
[1614] The server receives the request data. The server analyzes the received question text and extracts key keywords. Specifically, it uses a natural language processing engine (e.g., Google Cloud Natural Language API) to extract keywords such as "Tokyo," "recommended," and "ramen restaurant" from the question text. The output is the extracted keywords.
[1615] Step 5:
[1616] The server analyzes emotional data to identify the user's emotions. Specifically, it uses an emotion recognition engine (e.g., IBM Watson Tone Analyzer) to analyze the emotional data. The output is the user's emotional state. For example, it might retrieve data indicating "happy."
[1617] Step 6:
[1618] The server generates a response based on the analysis results (extracted keywords and sentiment data). An artificial intelligence model (e.g., OpenAI's GPT-3) is used as the generation method. The input consists of extracted keywords and sentiment data. For example, it might generate the response text: "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen!" The output is the generated response text.
[1619] Step 7:
[1620] The server searches for relevant links based on the generated response. It uses a search engine (e.g., Google Search API) to perform a web search using words included in the generated response. For example, it collects links for "Ichiran Tokyo" and "Tsukemen Daioh Tokyo". The output is a list of relevant links.
[1621] Step 8:
[1622] The server combines the generated answers and collected links to create the final response. Specifically, it combines the answer text and the list of links into a single response data. For example, it might create a response in the format of, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Please see the link below for details." The output is the combined response data.
[1623] Step 9:
[1624] The server sends the final response data to the terminal. It sends the response data back to the terminal using a specific communication protocol. The output indicates that the transmission of the response data is complete.
[1625] Step 10:
[1626] The terminal displays response data obtained from the server to the user. Specifically, it either displays the response on the screen interface or reads it aloud. For example, it might display something like, "Recommended ramen restaurants in Tokyo include Ichiran and Tsukemen Daioh. Enjoy some delicious ramen! For more details, please refer to the link below." The output consists of the response displayed to the user and the link.
[1627] This specific processing flow allows users to quickly obtain appropriate answers and related information with a single question, and furthermore, to receive emotionally resonant responses.
[1628] (Application Example 2)
[1629] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1630] Traditional information retrieval systems provide answers to user questions, but they lack the ability to generate answers that take user emotions into consideration. Furthermore, their ability to provide related links to the generated answers is limited. As a result, users often experience low satisfaction when obtaining answers, and the search experience remains unimproved.
[1631] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes an input means for inputting a question, an analysis means for analyzing the input question and the user's emotions, a generation means for generating an answer based on the analyzed question content and emotion data, a search means for searching for links related to the generated answer, and an output means for outputting the answer and related links together. This makes it possible to generate an answer that is sensitive to the user's emotions and to provide related information quickly.
[1632] A "question" is the content that users input in natural language to find out what they want to know.
[1633] "Input means" refers to the devices or software that users use to input questions.
[1634] "Analysis means" refers to a device or program for analyzing input questions and user sentiment data.
[1635] "Generation means" refers to a device or program that generates answers based on analyzed question content and sentiment data.
[1636] "Search means" refers to a device or program for searching the internet for links related to the generated answer.
[1637] "Output means" refers to a device or program for displaying answers and related links to the user.
[1638] "Emotional data" refers to emotional information extracted from user input.
[1639] An "artificial intelligence model" is a model trained based on machine learning or deep learning, and is used as a means of generation.
[1640] A "natural language processing engine" is a program that analyzes text input in natural language and extracts meaning and keywords.
[1641] "Related links" refer to URLs of web pages or information that are relevant to the generated response.
[1642] This invention is an information retrieval system that provides an appropriate answer using generative AI and simultaneously presents related links, based on the user's input of a single question. This system incorporates an emotion engine that recognizes the user's emotions and generates answers in a tone and content that corresponds to the user's emotions.
[1643] System Configuration
[1644] This system consists of the following main components:
[1645] 1. Input method: This refers to the means by which the user inputs questions, and includes devices such as smartphones and smart glasses.
[1646] 2. Analysis means: This refers to means for analyzing the input questions and sentiment data, and corresponds to natural language processing engines (e.g., Transformers, spaCy) or sentiment recognition libraries (e.g., emote).
[1647] 3. Generation means: This refers to means for generating answers based on the analyzed question content and sentiment data, and a generative AI model (e.g., GPT-3) falls under this category.
[1648] 4. Search methods: These are methods for searching the internet for links related to the generated answers, and web scraping libraries (BeautifulSoup, Scrapy) are examples of this.
[1649] 5. Output means: This refers to a means of displaying the answer and related links to the user, and this includes the device's display and notification function.
[1650] Processing flow
[1651] As a concrete example, the "emotion-recognition store navigator" application using smart glasses operates in the following steps.
[1652] 1. User input: The user uses smart glasses to input a question in natural language. For example, they might ask, "What product would suit my current mood?"
[1653] 2. Emotion Recognition: The device receives user input and analyzes emotions using an emotion recognition library (such as emote).
[1654] 3. Request to the server: Send the analyzed question and sentiment data to the server.
[1655] 4. Question Analysis: The server uses a natural language processing engine (such as Transformers or spaCy) to analyze the question and extract key keywords.
[1656] 5. Response generation: Use a generative AI model (such as GPT-3) to generate appropriate responses based on analysis results and sentiment data.
[1657] 6. Search for related links: Based on the generated answers, use a web scraping library (such as BeautifulSoup or Scrapy) to search the internet for related links.
[1658] 7. Displaying Results: The answers and related links are compiled and displayed on the smart glasses' screen.
[1659] Examples of specific cases and prompt statements
[1660] As a concrete example, if a user feels tired and asks, "What product would be perfect for how I'm feeling right now?", the following process would occur:
[1661] Input text: "Please recommend a product that perfectly matches my current mood."
[1662] Emotion recognition results: "Fatigue" and "Stress"
[1663] Server output: "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details."
[1664] Related links: "Aroma Candle Details Page", "Massage Chair Purchase Page"
[1665] Examples of prompt messages are as follows:
[1666] "What are some relaxing products suitable for relieving fatigue?"
[1667] "What are some good relaxation items for when you're tired?"
[1668] This system allows users to not only get quick answers to their questions, but also experience friendly and personalized responses tailored to their own emotions.
[1669] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1670] Step 1: User Input
[1671] The user uses smart glasses to input a question in natural language. The input might be something like, "Tell me what product would suit my current mood." The user asks the question via voice or text input. At this time, the smart glasses interface receives the question and retrieves the data.
[1672] Step 2: Recognizing Emotions
[1673] The device receives user input and analyzes emotions using an emotion recognition library (e.g., Emote). In this step, emotion data is extracted from the input text to recognize emotional states such as "fatigue" or "stress." The input data is in text format, and the output data consists of emotion labels and emotion scores.
[1674] Step 3: Request to the server
[1675] The terminal sends the analyzed question and sentiment data to the server. Specifically, it sends the question text and sentiment data, converted to JSON format, to the server via an HTTP POST request. The input data is in JSON format, and the output data is the response from the server.
[1676] Step 4: Analyzing the Question
[1677] The server uses a natural language processing engine (e.g., Transformers, spaCy) to analyze the received question and extract key keywords. For example, from the question "Tell me a product that suits my current mood," it extracts the keywords "mood" and "product." The input data is the question text in JSON format, and the output data is the analyzed keywords.
[1678] Step 5: Generating the answer
[1679] The server generates an appropriate answer using a generative AI model (e.g., GPT-3) based on the analysis results of the question and sentiment data. In this step, the generative AI model generates an answer based on the input data (analyzed keywords and sentiment data). For example, an answer such as "Relaxing aromatherapy candles and massage chairs are recommended" might be generated. The input data consists of keywords and sentiment data, and the output data is the generated answer text.
[1680] Step 6: Search for related links
[1681] The server uses a web scraping library (e.g., BeautifulSoup, Scrapy) to search the internet for relevant links based on the generated response. For example, it collects relevant links such as "details page for aromatherapy candles" and "purchase page for massage chairs." The input data is the generated response text, and the output data is a list of relevant links.
[1682] Step 7: Displaying the results
[1683] The server compiles the answers and related links and sends them to the device. The device displays the received answers and links to the user. At this time, the smart glasses display shows a message along with a link that reads, "We recommend aromatherapy candles and massage chairs for relaxation. Please see the link below for details." The input data is the response from the server, and the output data is information that the user can visually confirm.
[1684] As described above, a series of processes, from user questioning and sentiment recognition to natural language processing, the use of generative AI models, the search for related links, and output, are performed, making it possible to provide users with useful information that resonates with their emotions.
[1685] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1686] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1687] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1688] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1689] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1690] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1691] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1692] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1693] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1694] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1695] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1696] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1697] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1698] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1699] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1700] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1701] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1702] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1703] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1704] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1705] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1706] The following is further disclosed regarding the embodiments described above.
[1707] (Claim 1)
[1708] An input method for entering questions,
[1709] An analysis means for analyzing the input questions,
[1710] A generation means that generates answers based on the analyzed question content,
[1711] A search method for finding links related to the generated answers,
[1712] An output method for outputting answers and related links together,
[1713] A system that includes this.
[1714] (Claim 2)
[1715] An analysis means that analyzes the input question using a natural language processing engine,
[1716] A search method for finding links related to the generated answer on the internet,
[1717] The system according to claim 1, including the following:
[1718] (Claim 3)
[1719] The system according to claim 1, characterized in that the generation means generates answers to questions using an artificial intelligence model.
[1720] "Example 1"
[1721] (Claim 1)
[1722] Input means and
[1723] A terminal for receiving questions entered by the user through an input means,
[1724] An analysis means for analyzing questions sent from a terminal,
[1725] A generation means that generates answers based on the analyzed question content,
[1726] A search means for searching for links related to the answers generated by the generation means,
[1727] An output means that sends the generated answers and related links to the terminal together,
[1728] A system that includes this.
[1729] (Claim 2)
[1730] The analysis means analyzes the input question using a natural language processing engine,
[1731] The search method is characterized by searching for links related to answers generated from the internet.
[1732] The system according to claim 1.
[1733] (Claim 3)
[1734] The generation means is characterized by generating answers to questions using a generation AI model.
[1735] The system according to claim 1.
[1736] "Application Example 1"
[1737] (Claim 1)
[1738] An input method for entering questions,
[1739] An analysis means for analyzing the input questions,
[1740] A generation means that generates answers based on the analyzed question content,
[1741] A search method for finding links related to the generated answers,
[1742] An output method for outputting answers and related links together,
[1743] A product link display means for displaying the corresponding product link along with the generated answer,
[1744] A system that includes this.
[1745] (Claim 2)
[1746] An analysis means that analyzes the input question using a natural language processing engine,
[1747] A search method for finding links related to the generated answer on the internet,
[1748] A product suggestion method that proposes appropriate products based on the generated responses,
[1749] The system according to claim 1, including the following:
[1750] (Claim 3)
[1751] The system according to claim 1, characterized in that the generation means generates answers to questions using an artificial intelligence model.
[1752] "Example 2 of combining an emotion engine"
[1753] (Claim 1)
[1754] An input method for entering questions,
[1755] A means of receiving questions and sentiment data sent from a device,
[1756] An analysis method using a natural language processing engine to analyze the input question,
[1757] An analysis method using an emotion recognition engine for analyzing emotion data,
[1758] A generation method using an artificial intelligence model that generates answers based on analyzed question content and sentiment data,
[1759] A search method using a search engine that searches the internet for links related to the generated answer,
[1760] An output method for sending the answers and related links to the terminal together,
[1761] A system that includes this.
[1762] (Claim 2)
[1763] An analysis means that analyzes the input question using a natural language processing engine,
[1764] An analysis method that analyzes user emotion data using an emotion recognition engine,
[1765] A search method using a search engine that searches the internet for links related to the generated answer,
[1766] The system according to claim 1, including the following:
[1767] (Claim 3)
[1768] The generation means is characterized by generating answers to questions using an artificial intelligence model, and generating answers that include expressions tailored to the user's emotions.
[1769] The system according to claim 1.
[1770] "Application example 2 when combining with an emotional engine"
[1771] (Claim 1)
[1772] An input method for entering questions,
[1773] An analytical means for analyzing the input questions and the user's emotions,
[1774] A generation means for generating answers based on analyzed question content and sentiment data,
[1775] A search method for finding links related to the generated answers,
[1776] An output method for outputting answers and related links together,
[1777] A system that includes this.
[1778] (Claim 2)
[1779] An analysis method that uses a natural language processing engine to analyze the input question and recognizes the user's emotional state based on the user's emotional data,
[1780] A search method for finding links related to the generated answer on the internet,
[1781] The system according to claim 1, including the following:
[1782] (Claim 3)
[1783] The system according to claim 1, characterized in that the generation means generates answers to questions and emotional states using an artificial intelligence model. [Explanation of symbols]
[1784] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An input method for entering questions, An analysis means for analyzing the input questions, A generation means that generates answers based on the analyzed question content, A search method for finding links related to the generated answers, An output method for outputting answers and related links together, A system that includes this.
2. An analysis means that analyzes the input question using a natural language processing engine, A search method for finding links related to the generated answer on the internet, The system according to claim 1, including the following:
3. The system according to claim 1, characterized in that the generation means generates answers to questions using an artificial intelligence model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A