System

The system addresses the challenge of text-based AI responses by integrating question analysis, response generation, and image creation to provide intuitive and efficient visual explanations.

JP2026014902APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116376
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional generative AI systems provide text-based responses that are difficult for users to intuitively understand, lacking visual aids that hinder deep understanding of complex information.

Method used

A system that integrates question analysis, response generation, image diagram creation, and data transmission to provide visually easy-to-understand responses by combining text and images, utilizing natural language processing and diagram generation engines.

Benefits of technology

Enables users to grasp complex information intuitively and efficiently through integrated text and image responses, enhancing understanding and communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014902000001_ABST
    Figure 2026014902000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: question receiving means for receiving a question from a user; question analyzing means for analyzing the received question to identify related information; response sentence generating means for generating a response sentence based on the information identified by the question analyzing means; image diagram generating means for generating an image diagram related to the response sentence; response integrating means for integrating the response sentence and the image diagram to generate an integrated response; and data transmitting means for transmitting the integrated response to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Answers from conventional generative AI are often provided only as text, making it difficult for users to intuitively understand the information. Furthermore, there is a lack of visual aids to enable users to instantly obtain the information they want, making it difficult to gain a deep understanding of a specific question. To solve these issues, a system that provides users with visually easy-to-understand information is needed. [Means for solving the problem]

[0005] In order to solve the above-mentioned problems, the present invention provides the following means: a system including a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response generation means for generating a response based on the information identified by the question analysis means, an image generation means for generating an image related to the response, a response integration means for integrating the response and the image to generate an integrated response, and a data transmission means for transmitting the integrated response to the user. This system enables the user to receive a visual image along with the response to the question, making it possible to provide information that is more intuitive and easy to understand.

[0006] The "question receiving means" is a function for receiving a question message sent by a user.

[0007] The "question analysis means" is a function for analyzing a received question message and identifying the intent of the question and related information.

[0008] The "response sentence generation means" is a function that generates a response sentence to be provided to the user based on the information identified by the question analysis means.

[0009] The "image diagram generating means" is a function for generating a visual image diagram related to a response sentence.

[0010] The "response integration means" is a function for integrating the generated response sentence and image diagram to generate one integrated response.

[0011] The "data transmission means" is a function for transmitting an integrated response to the user.

[0012] A "database" is a system for storing information in an organized manner and retrieving that information as needed.

[0013] A "knowledge base" is a source of information that systematically collects and stores knowledge related to a particular field and makes it available for response generation.

[0014] An "illustration generation engine" is a software and hardware system for generating visual diagrams and images based on text information. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The present invention is a system that provides visually easy-to-understand responses to questions entered by a user. The system is implemented by the following modules:

[0037] Question receiving module

[0038] The user submits a question to the chatbot, which is entered in text format and sent to the system in real-time or non-real-time.

[0039] Question Analysis Module

[0040] The server sends the question received from the user to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. For example, in the question "Please explain how plants photosynthesize," the words "plant," "photosynthesis," and "mechanism" are analyzed.

[0041] Response generation module

[0042] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. Specifically, it retrieves relevant information from an internal database and knowledge base and generates grammatically correct responses. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0043] Image diagram generation module

[0044] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine to generate a visual representation of the photosynthesis process, showing elements such as light energy, carbon dioxide, water, oxygen, and organic matter.

[0045] Response Integration Module

[0046] The server's response integration module integrates the response text and the image diagram to generate a single integrated response. The generated response text is followed by a related diagram, making it easy for users to understand at a glance. For example, it might look like this: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0047] Data Transmission Module

[0048] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0049] Viewing the response

[0050] The terminal displays the response text and image received from the server to the user. The response text and image are provided through the chat interface, allowing the user to immediately understand the information.

[0051] Specific examples

[0052] A user submits the following question to the chatbot: "How does photosynthesis work?"

[0053] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[0054] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0055] The image generation module generates a diagram showing the process of photosynthesis.

[0056] The response integration module integrates the response sentence and the image diagram.

[0057] A data transmission module transmits the consolidated response to the user.

[0058] The terminal displays the response text and an image to the user, who then understands the content.

[0059] In this way, intuitive and easy-to-understand responses are provided to questions entered by the user. All of the above modules work together to provide information efficiently.

[0060] The processing flow will be explained below.

[0061] Step 1:

[0062] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[0063] Step 2:

[0064] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[0065] Step 3:

[0066] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[0067] Step 4:

[0068] The server's response generation module generates a response based on the analyzed keywords. It references an internal database or knowledge base to obtain appropriate information and create a grammatically correct response. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0069] Step 5:

[0070] The server's image generation module creates an image related to the response statement. In this case, it uses a diagram generation engine to generate a visual representation of the "process of photosynthesis." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0071] Step 6:

[0072] The server's response integration means integrates the generated response sentence with the image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0073] Step 7:

[0074] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[0075] Step 8:

[0076] The terminal displays the integrated response received from the server to the user. The response text and image diagram are provided via the chat interface, allowing the user to intuitively understand the information.

[0077] This series of processes allows users to get instant, easy-to-understand answers to their questions, and the visualization of information promotes deeper understanding.

[0078] Example 1

[0079] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0080] Conventional chatbots and FAQ systems only provide text-based responses to questions entered by users, making it difficult to provide answers efficiently by including visual information. In particular, when explaining complex concepts or processes, it is difficult to understand using text alone, and visual aids are required to help users understand. By providing visual information as well, it is necessary to enable users to understand information more intuitively and quickly.

[0081] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0082] In this invention, the server includes means for receiving a question from a user, means for analyzing the received question and identifying key keywords and the intent of the question, means for generating a response sentence based on the identified keywords and phrases, means for generating a related image based on the response sentence, means for integrating the response sentence and the image to generate an integrated response, and means for sending the integrated response to the user. This makes it possible to provide a response to a question entered by a user in a format that integrates text data and visual data. This makes it easier for the user to intuitively understand complex information, and realizes smooth communication of information.

[0083] A "user" is an entity that sends a question or request to the system.

[0084] The "question receiving means" is a device or module that provides a function for receiving a question sent by a user.

[0085] The "question analysis means" is a device or module that provides a function for analyzing a received question and identifying the main keywords and intent of the question.

[0086] The "response sentence generation means" is a device or module that provides a function for generating an appropriate response sentence based on the keywords and phrases identified by the question analysis means.

[0087] The "image diagram generating means" is a device or module that provides a function for generating a related image diagram based on a response sentence.

[0088] The "response integration means" is a device or module that provides a function for integrating a response sentence and an image diagram to generate one integrated response.

[0089] A "data transmission means" is a device or module that provides the functionality for transmitting a consolidated response to a user.

[0090] A "database" is a storage system for systematically storing information and enabling efficient search and retrieval.

[0091] A "knowledge base" is an information system that stores knowledge and information in a specific domain and makes it available for response generation.

[0092] An "illustration generation engine" is a software tool for generating visual illustrations based on given data or information.

[0093] This invention is a system that provides visually easy-to-understand responses to questions entered by users. This system is composed of a server, a terminal, and a user, and is implemented by the following modules:

[0094] Question receiving method

[0095] The user enters a question into the chatbot and sends it. The question is in text format and is sent to the system in real time or non-real time. The server receives the question.

[0096] Question analysis means

[0097] The server forwards the received question to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. Specifically, it uses an NLP engine (e.g., Google NLP API or SpaCy). For example, from the question "Please explain how plants photosynthesize," the words "plants," "photosynthesis," and "mechanism" are identified.

[0098] Response sentence generation means

[0099] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. This module consults internal databases and knowledge bases (e.g., Wikipedia API, corporate knowledge bases). For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0100] Image diagram generation means

[0101] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine (e.g., D3.js, Graphviz). For example, it generates a visual diagram that shows the process of photosynthesis, illustrating the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[0102] Response Integration Measures

[0103] The server's response integration module integrates the response text and the image to generate a single integrated response, such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0104] Data transmission method

[0105] The server's data transmission module sends the consolidated response to the user using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0106] Viewing the response

[0107] The terminal displays the response text and image received from the server. The response text and image are provided through the chat interface, allowing the user to instantly understand the information.

[0108] Specific examples

[0109] User: Sends a question to the chatbot: "How does photosynthesis work?"

[0110] Server: The question analysis module identifies "photosynthesis" and "mechanism" and sends them to the response generation module.

[0111] Response sentence generation means: Generate the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0112] Image diagram generation means: Generates a diagram showing the process of photosynthesis.

[0113] Response integration method: Integrate the response sentence and the image diagram.

[0114] Data sending means: Sends the consolidated response to the user.

[0115] Terminal: Displays the response text and an image diagram, allowing the user to understand the content.

[0116] The system provides intuitive and easy-to-understand responses to questions entered by users. All of the above modules work together to provide information efficiently.

[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0118] Step 1: Receiving the question

[0119] The user enters a question in text format into the chatbot and sends it. Specifically, the user enters the question into the chat interface of a web browser or mobile app and clicks the "Send" button.

[0120] The server receives a question sent by a user, the question data being in text format.

[0121] Input: A text question submitted by the user

[0122] Output: Received question data

[0123] Step 2: Parsing the Question

[0124] The server forwards the received question to the question analysis module, which uses an NLP engine (e.g., Google NLP API or SpaCy) to analyze the question content and identify key keywords and the intent of the question.

[0125] Specifically, the keywords "photosynthesis" and "mechanism" are extracted from the question "Please tell me how photosynthesis works."

[0126] Input: Received question data

[0127] Output: Extracted keywords and question intent

[0128] Step 3: Generate a response

[0129] The server's response generation module generates responses based on the keywords and phrases identified by the question analysis module. This module consults an internal database or knowledge base (e.g., Wikipedia API, corporate knowledge base).

[0130] Specifically, the response sentence generated is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0131] Input: Extracted keywords and question intent

[0132] Output: Generated response

[0133] Step 4: Generate an image

[0134] The server's image generation module generates the relevant image based on the generated response. This module uses an image generation engine (e.g., D3.js, Graphviz).

[0135] Specifically, it generates a diagram that visually shows the process of photosynthesis and explains the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[0136] Input: Generated response sentence

[0137] Output: Generated image

[0138] Step 5: Consolidating the response

[0139] The server's response integration module integrates the response text and the image to generate a single integrated response. Specifically, the response text is followed by a related image, and the format is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0140] Input: Generated response text and image

[0141] Output: Consolidated response

[0142] Step 6: Sending a Response

[0143] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0144] Input: Consolidated response

[0145] Output: The integration response sent to the user

[0146] Step 7: View the response

[0147] The terminal displays the integrated response received from the server, and the response text and image are displayed to the user through the chat interface.

[0148] This allows the user to instantly understand the answer to the question.

[0149] Input: The integration response sent by the server

[0150] Output: Response text and image displayed to the user

[0151] (Application example 1)

[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0153] When obtaining maintenance and troubleshooting information for robots and machinery used in factories, it is difficult for on-site staff to obtain the information quickly and accurately. In particular, there is a lack of information provided in a visually easy-to-understand format, which can result in reduced work efficiency and safety. Another issue is that current technology does not have a widespread system that provides appropriate instructions for complex machine operation.

[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0155] In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generating means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, a means for transmitting information to a display device that displays the response, and a means for inputting a question by text or voice using the device and displaying the result in real time. This enables on-site staff in a factory to quickly learn complex machine operations and maintenance procedures in a visually easy-to-understand format.

[0156] The "question receiving means" is a device or software that has the function of receiving a question from a user in text or voice format.

[0157] The "question analysis means" is a device or software that uses natural language processing technology to analyze received questions and identify their content, intent, and keywords.

[0158] The "response sentence generation means" is a device or software for generating an appropriate response sentence based on the analyzed information.

[0159] The "image diagram generating means" is a device or software for automatically generating illustrations or diagrams related to a response sentence.

[0160] The "response integration means" is a device or software that has the function of integrating the generated response sentence and image diagram to generate one integrated response.

[0161] A "data transmission means" is a device or software that uses a communication protocol to transmit the generated integrated response to a user.

[0162] The "means for transmitting information to a display device" is a device or software for transmitting instructions to cause the display device to display the aggregate response.

[0163] "Means for text or voice input and real-time display of results" refers to a device or software that receives a user's text or voice input and displays answers generated based on that input in real time.

[0164] A "generative AI model" is an artificial intelligence model that uses deep learning and neural networks to generate response sentences and image diagrams.

[0165] A "prompt" is an instruction or question given to a generative AI model, and is the input information that enables the model to generate an appropriate response.

[0166] In this invention, a system is constructed that allows a user to input a question about a robot or machine equipment used in a factory and provides a response in a visually easy-to-understand format. Detailed embodiments of this system are described below.

[0167] System configuration

[0168] The system consists of the following components:

[0169] 1. Question receiving means: The user inputs a question via text or voice. This function is realized via a smartglasses or smartphone application. For example, the user inputs a question such as, "Please tell me how to maintain this robot."

[0170] 2. Question analysis: The server receives the entered question and analyzes it using natural language processing (NLP) technology. The OpenAI API is used to identify key keywords and the intent of the question.

[0171] 3. Response generation: Based on the keywords identified by the question analysis, an appropriate response is generated. This process also uses OpenAI's generative AI model to obtain relevant information.

[0172] 4. Image diagram generation means: Generate an image diagram related to the generated response sentence using a diagram generation engine. For example, automatically generate the related diagram through an API.

[0173] 5. Response integration: The generated response sentences and images are integrated to create a single integrated response, which is displayed in a format that is easy for the user to understand visually.

[0174] 6. Data transmission means: The integrated response is transmitted to the user's display device using a communication protocol (such as HTTP or WebSocket).

[0175] 7. Means for sending information to a display device: Send information to a device (head-mounted display or smart glasses) for displaying the integrated response. The response is displayed in real time, allowing the user to instantly understand the content.

[0176] Processing Details

[0177] The server analyzes the question received from the user and extracts key keywords. This analysis is performed using OpenAI's API, based on the generated prompt text. The response generation means generates a response text based on the analysis results. An illustration generation engine is used to generate an image diagram related to this generated response text. The generated response text and image diagram are integrated and sent to the user. A communication protocol (e.g., HTTP, WebSocket, etc.) is used in this process, and the results are displayed in real time on a display device (e.g., a head-mounted display, smart glasses).

[0178] Specific examples

[0179] The user speaks the following question into the HMD application: "Please tell me how to maintain this robot."

[0180] 1. The question receiving means receives a question from a user.

[0181] 2. The question analysis tool analyzes the question and identifies “robot,” “maintenance,” and “method” as the main keywords.

[0182] 3. The response generation means uses OpenAI's generative AI model to generate a "detailed response regarding the robot's maintenance procedures."

[0183] 4. The image diagram generating means generates a related diagram (for example, a parts layout diagram or a work procedure diagram) based on the generated response sentence.

[0184] 5. The response synthesis means creates an integrated response by integrating the response sentence and the image diagram.

[0185] 6. The data transmission means transmits the integrated response to the user's HMD.

[0186] 7. A means for transmitting information to a display device displays the integrated response in real time on the HMD.

[0187] Example prompt sentence:

[0188] Question analysis prompt:

[0189] Analyze the following question and extract the main keywords: How do I maintain this robot?

[0190] Prompt for generating a response:

[0191] Please provide more details on robot maintenance methods.

[0192] As described above, the present invention allows on-site staff to visually and quickly obtain information about robots and machinery equipment.

[0193] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0194] Step 1:

[0195] The user inputs a question by text or voice. This question is sent to the server through an application on smart glasses or a smartphone. If the user inputs, "Please tell me how to maintain this robot," the text or voice data is sent to the server.

[0196] Step 2:

[0197] The server analyzes the received question using a question analysis method. Specifically, it uses natural language processing (NLP) technology to understand the content and intent of the question. The input is the user's question text or voice data, and the output is key keywords such as "robot," "maintenance," and "method." This analysis uses OpenAI's API.

[0198] Step 3:

[0199] The server uses OpenAI's generative AI model to generate an appropriate response based on the analyzed keywords. The input is a prompt related to the analyzed keyword "robot maintenance method," and the output is a "detailed response regarding robot maintenance procedures." This generative AI model automatically obtains relevant information and generates a grammatically correct response.

[0200] Step 4:

[0201] Based on the generated response, the server uses a diagram generation engine to generate related diagrams. The input is the response, and the output is diagrams such as "diagrams of robot maintenance procedures" or "parts layout diagrams." The server automatically generates related diagrams by calling the appropriate external API.

[0202] Step 5:

[0203] The server integrates the generated response sentence and image to generate a single integrated response. The input is the response sentence and image, and the output is a single integrated visual response containing text and illustrations. These are integrated using a response integration tool.

[0204] Step 6:

[0205] The server transmits the integrated response to the user's display device via a data transmission means. The input is the integrated response, and the output is data displayed on the user's display device (e.g., head-mounted display, smart glasses). Information is transmitted in real time using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0206] Step 7:

[0207] The terminal uses a means for transmitting information to a display device to display the generated integrated response in real time. The input is the response data transmitted from the server, and the output is a display that the user can visually confirm. The user can check the integrated response text and image diagram and intuitively understand the necessary information.

[0208] Through the above steps, users can visually and efficiently obtain answers to their questions.

[0209] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0210] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the addition of a function to recognize the user's emotions and adjust the content of the response. This system is implemented by the following modules and emotion engine.

[0211] Question receiving module

[0212] The user sends a text question to the chatbot, for example, "How does photosynthesis work?"

[0213] Question Analysis Module

[0214] The server sends the question received from the user to the question analysis module, which uses natural language processing (NLP) techniques to analyze and identify the question content and key keywords (e.g., "photosynthesis," "mechanism").

[0215] Response generation module

[0216] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0217] Image diagram generation module

[0218] The server's image generation module creates an image related to the response statement. Using the diagram generation engine, it generates a visual representation of the "photosynthesis process." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0219] Emotion Engine

[0220] The emotion engine on the server analyzes the user's emotions based on user data such as voice tone, facial expressions, and input text. The emotion engine recognizes the user's emotions (e.g., excitement, anxiety, joy).

[0221] Response Adjustment Measures

[0222] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. If the user is anxious, the response will be adjusted to be more reassuring. Conversely, if the user is excited, calming words will be emphasized.

[0223] Response Integration Module

[0224] The server's response integration module integrates the adjusted response sentence with the generated image diagram. The diagram is placed after the response sentence and arranged in a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0225] Data Transmission Module

[0226] The server's data transmission module sends the consolidated response to the user, sending the response to the chat interface using a communication protocol (e.g. HTTP, WebSocket, etc.).

[0227] Viewing the response

[0228] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[0229] Specific examples

[0230] A user submits the following question to the chatbot: "How does photosynthesis work?"

[0231] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[0232] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0233] The image generation module generates a diagram showing the process of photosynthesis.

[0234] The emotion engine analyzes the user's emotions and recognizes that the user is in an anxious state.

[0235] A response adjustment means adjusts the response sentence to a tone that puts the user at ease.

[0236] The response integration module integrates the response sentence and the image diagram.

[0237] A data transmission module transmits the consolidated response to the user.

[0238] The terminal displays the response text and an image to the user, who then understands the content.

[0239] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand the information more deeply and comfortably.

[0240] The processing flow will be explained below.

[0241] Step 1:

[0242] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[0243] Step 2:

[0244] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[0245] Step 3:

[0246] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[0247] Step 4:

[0248] The server's emotion engine analyzes the user's emotion from the question text, using NLP techniques to identify emotions from the linguistic tone and keywords in the text (e.g., excitement, anxiety, joy, etc.).

[0249] Step 5:

[0250] The server's response generation module generates a response based on the analyzed keywords and the emotion analysis results from the emotion engine. It retrieves relevant information from the internal database and knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0251] Step 6:

[0252] The server's response adjustment means adjusts the generated response sentence based on the user's emotions. For example, if the user is anxious, the response sentence will be changed to a tone that gives a sense of security. Example: "Don't worry, I'll explain how photosynthesis works. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0253] Step 7:

[0254] The server's image generation module creates an image related to the tailored response. This module uses a diagram generation engine to generate a visual representation of the "process of photosynthesis," showing light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0255] Step 8:

[0256] The server's response integration means integrates the adjusted response sentence with the generated image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Don't worry, I'll explain the mechanism of photosynthesis. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0257] Step 9:

[0258] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[0259] Step 10:

[0260] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[0261] This process allows users to get instant, easy-to-understand answers to their questions, and by visualizing the information and adjusting responses based on emotion, users can more comfortably understand the information.

[0262] Example 2

[0263] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0264] Conventional question-answering systems generate uniform responses without considering the user's emotions, resulting in poor user satisfaction. In particular, when a user feels anxious or confused, the system may be unable to respond appropriately, resulting in a poor user experience. Furthermore, while image diagrams are generated to aid visual understanding, they are often mismatched with the text, making it difficult to intuitively understand the information. To solve these problems, a system is needed that provides visually understandable responses while taking the user's emotions into account.

[0265] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a sentiment analysis means for analyzing the user's sentiment, a response adjustment means for adjusting the tone and content of the response sentence based on the analysis result of the sentiment analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, and a data transmission means for transmitting the integrated response to the user. This makes it possible to provide a flexible response that takes into consideration the user's sentiment and a response that is visually easy to understand.

[0266] The "question receiving means" is a function that receives text questions from users to the chatbot.

[0267] The "question analysis means" is a function that analyzes received questions and identifies the question content and main keywords using natural language processing technology.

[0268] The "response sentence generation means" is a function that generates a grammatically correct response sentence by obtaining information from an internal database or knowledge base based on the information identified by the question analysis means.

[0269] The "image diagram generating means" is a function that generates visual information related to a response sentence as an illustration and expresses it in an easy-to-understand form.

[0270] The "emotion analysis means" is a function that analyzes data such as the user's voice tone, facial expression, and input text, and identifies the user's emotions.

[0271] The "response adjustment means" is a function that adjusts the tone and content of the response sentence to suit the user's emotions based on the results of the emotion analysis means.

[0272] The "response integration means" is a function that integrates the adjusted response sentence with the generated image diagram to generate an integrated response that is arranged in a visually intuitive format.

[0273] The "data transmission means" is a function for transmitting a response using a communication protocol (e.g., HTTP, WebSocket, etc.) for transmitting an integrated response to a user.

[0274] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the added function of recognizing the user's emotions and adjusting the content of the response. The configuration and operation of this system are described in detail below.

[0275] System configuration:

[0276] Hardware and Software Use:

[0277] Server: The server that is the center of data processing and response generation. It has a high-performance CPU and memory.

[0278] Terminal: A device used by a user (e.g., a PC, smartphone, or tablet) with a web browser or dedicated application installed.

[0279] Natural Language Processing Engine (NLP): Software used for text analysis (e.g., SpaCy, NLTK).

[0280] Diagram generation engine: Software for generating visual images (e.g., D3.js, Chart.js).

[0281] Database and knowledge base: Stores information for generating response sentences.

[0282] Sentiment analysis engine: Software that analyzes user emotions (e.g., Microsoft Azure Emotion API, Google Cloud Vision API).

[0283] System behavior:

[0284] Receiving questions:

[0285] Users send questions to the chatbot in text format, for example, "Please explain how photosynthesis works."

[0286] Parsing the question:

[0287] The server sends the received question to a question analysis module, which uses natural language processing (NLP) techniques to analyze the question and identify key keywords (e.g., "photosynthesis," "mechanism").

[0288] Generate a response:

[0289] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from a database or knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0290] Generate image diagrams:

[0291] The server's image generation module generates an image related to the response statement. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a visual representation of the "process of photosynthesis."

[0292] Sentiment Analysis:

[0293] The server's emotion engine analyzes the user's emotions by extracting emotions from voice tone, facial expressions (using facial recognition technology), and input text. At this stage, it may be recognized that the user is in an anxious state.

[0294] Tailoring response content:

[0295] The server adjusts the response based on the analysis results of the emotion engine. For example, if the user is feeling anxious, the response will be adjusted to something that gives a sense of security, such as "Don't worry, photosynthesis is a natural process for plants."

[0296] Response integration:

[0297] The server's response integration module integrates the adjusted response sentences with the generated image diagrams, arranging the response sentences followed by the image diagrams in a visually intuitive format.

[0298] Sending data:

[0299] The server's data transmission module sends the consolidated response to the user's device, and sends the response to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0300] View response:

[0301] The terminal displays the integrated response received from the server. The response text and image are provided to the user through the chat interface, allowing the user to intuitively understand the information.

[0302] Examples:

[0303] 1. A user submits the following question to the chatbot: "How does photosynthesis work?"

[0304] 2. The server identifies "photosynthesis" and "mechanism" using the question analysis module.

[0305] 3. The response generation module generates the response sentence, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0306] 4. The image generation module generates a diagram showing the process of photosynthesis.

[0307] 5. The emotion engine analyzes the user's emotions and recognizes when the user is in an anxious state.

[0308] 6. The response adjuster adjusts the response to a more reassuring tone.

[0309] 7. The response integration module integrates the response sentence and the image diagram.

[0310] 8. The data transmission module sends the consolidated response to the user.

[0311] 9. The terminal displays the response text and image to the user, and the user understands the content.

[0312] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand information more deeply and comfortably.

[0313] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0314] Step 1:

[0315] The user sends a question to the chatbot in text format. For example, they can enter a question like, "Please tell me how photosynthesis works." The input data is the text data that the user types into the input field. The sent text data is transferred to the server.

[0316] Step 2:

[0317] The server sends the question data received from the user to the question analysis module, which uses natural language processing (NLP) technology to analyze the question content. The input data is the question data in text format, and the output is the main keywords (e.g., "photosynthesis" and "mechanism").

[0318] Step 3:

[0319] The server's response generation module generates a response based on the keywords output from the question analysis module. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. The input data are keywords, and the output is a response (e.g., "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water").

[0320] Step 4:

[0321] The server's image generation module creates an image related to the generated response text. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a diagram that visually represents the "photosynthesis process." The input data is the response text, and the output is an image showing the photosynthesis process.

[0322] Step 5:

[0323] The emotion engine on the server analyzes the user's emotion. The input data are voice tone, facial expression (using facial recognition technology), and entered text. The emotion engine analyzes these data and identifies the emotion that the user is anxious. The output is the user's emotional state (e.g., anxious).

[0324] Step 6:

[0325] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. The input data is the analyzed emotional state and the original response. Based on these input data, the server changes the response to include reassuring elements (e.g., "Don't worry, photosynthesis is a natural process for plants"). The output is the adjusted response.

[0326] Step 7:

[0327] The server's response synthesis module synthesizes the adjusted response sentence and the generated image diagram. The input data are the adjusted response sentence and the image diagram. By synthesizing these, a visually intuitive response is generated. The output is the synthesized response.

[0328] Step 8:

[0329] The server's data transmission module sends the consolidated response to the user's device. The input data is the consolidated response, which is sent to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.). The output is the response data sent to the user's device.

[0330] Step 9:

[0331] The terminal displays the integrated response received from the server to the user. The input data is the response data sent from the server, and the output is the response text and image displayed on the chat interface. The user can intuitively understand the information through this display.

[0332] (Application example 2)

[0333] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0334] Conventional information provision systems only provide text-based responses to user input, which are often difficult to understand visually and do not allow for flexible responses that reflect the user's emotions.The present invention aims to provide more appropriate and reassuring information by providing visually easy-to-understand responses to questions entered by the user, recognizing the user's emotions, and adjusting the response content.

[0335] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, an emotion analysis means for analyzing the user's emotion, a response adjustment means for adjusting the tone and content of the response sentence based on the emotion analyzed by the emotion analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, and a display means for displaying the integrated response on the smart device. This allows the user to obtain visually easy-to-understand information and receive a flexible response suited to their individual emotional state.

[0336] The "question receiving means" is a device or function for receiving a question in text or voice format sent by a user.

[0337] The "question analysis means" is a device or function for analyzing a received question and identifying the content of the question and related main keywords.

[0338] The "response sentence generation means" is a device or function for generating a response sentence based on the analyzed keywords and providing the necessary information.

[0339] The "image diagram generating means" is a device or function for generating a visual image diagram related to a response sentence.

[0340] The "emotion analysis means" is a device or function for analyzing the user's emotions and adjusting the response content based on those emotions.

[0341] The "response adjustment means" is a device or function for adjusting the tone and content of a response sentence to match the user's emotions based on the emotion analysis results.

[0342] The "response integration means" is a device or function for integrating the response sentence and the generated image diagram to generate one integrated response.

[0343] A "data transmission means" is a communication device or function for transmitting an integrated response to a user.

[0344] The "display means" is a device or function for displaying the integrated response on the smart device.

[0345] A "smart device" is an electronic device used by a user, such as smart glasses, a smartphone, a head-mounted display, or a robot.

[0346] A system for implementing this invention provides visual and emotional responses to user questions when the user is engaged in activities such as shopping through a smart device (such as smart glasses, a smartphone, a head-mounted display, or a robot). An embodiment of this system is described in detail below.

[0347] Program Structure

[0348] The system includes a question receiving means, a question analyzing means, a response sentence generating means, an image diagram generating means, a sentiment analyzing means, a response adjusting means, a response integrating means, a data transmitting means, and a display means.

[0349] Program processing and use of hardware and software

[0350] Question receiving method

[0351] Users enter questions into smart glasses or smartphones by text or voice, using voice recognition software or a text input interface.

[0352] Question analysis means

[0353] The server analyzes the questions received from users using natural language processing (NLP) technology to identify key keywords, such as Google AI's natural language processing API.

[0354] Response sentence generation means

[0355] Based on the analyzed keywords, the server uses a generative AI model such as GPT-4 to generate a response, which is composed of information retrieved from a database or knowledge base.

[0356] Image diagram generation means

[0357] An image generation engine such as DALL-E is used to generate images related to the response, providing visual content appropriate for the response.

[0358] Emotion analysis means

[0359] To analyze user emotions, emotion analysis tools (e.g., Amazon Rekognition or Microsoft Azure's Emotion API) are used that process voice tone and facial expression data.

[0360] Response Adjustment Measures

[0361] Based on the results of the emotion analysis, the server adjusts the tone and content of the response to suit the user's emotions. If the user is anxious, the content will be adjusted to provide a sense of security.

[0362] Response Integration Measures

[0363] The response text and the generated image are integrated to generate a single integrated response, allowing the user to receive both text and visual information in a single interface.

[0364] Data transmission method

[0365] The server sends the integration response to the user's smart device using HTTP or WebSocket protocol.

[0366] Display means

[0367] The smart device displays the integrated response to the user.

[0368] Specific examples

[0369] When a user asks the smart glasses, "What are the characteristics of these shoes?", the following process is performed:

[0370] 1. Question analysis: The NLU module identifies the "shoes" and "features."

[0371] 2. Response generation: The generative AI model generates an answer such as, "These shoes are waterproof and feature a lightweight design."

[0372] 3. Image generation: DALL-E generates an image showing the waterproof function and lightweight design of the shoes.

[0373] 4. Sentiment Analysis: The emotion engine recognizes when the user is excited.

[0374] 5. Response adjustment: "These shoes are very popular. They fit you perfectly."

[0375] 6. Response Integration: Integrate the coordinated response statement with the illustration.

[0376] 7. Data transmission: sent to smart glasses and displayed to the user.

[0377] Prompt Sentence Examples

[0378] "Tell me the features of these shoes."

[0379] Analyzed keywords: shoes, features

[0380] Generated response: "These shoes are waterproof and feature a lightweight design."

[0381] Sentiment Analysis: Users are excited

[0382] Tailored response: "These shoes are very popular. They look great on you."

[0383] Final response: "These shoes are very popular. They're perfect for you. They're waterproof and lightweight. (Illustration: Waterproofing and Design)"

[0384] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0385] Step 1:

[0386] The user types or speaks a question into the smart device.

[0387] Input: Text or voice input from the user.

[0388] Output: Question text or audio file.

[0389] Specific operation: The user asks the smart glasses verbally, "Tell me the features of these shoes."

[0390] Step 2:

[0391] The device converts voice input into text data (voice recognition).

[0392] Input: User's voice input.

[0393] Output: The converted text data.

[0394] What it does: Speech recognition software converts the speech data into text, such as "What are the features of these shoes?"

[0395] Step 3:

[0396] The server receives the question and analyzes it using the question analysis means.

[0397] Input: Text data.

[0398] Output: Primary keywords (e.g., "shoes", "features").

[0399] How it works: The question receiving module receives the text and uses NLP technology (such as Google AI's natural language processing API) to identify key keywords.

[0400] Step 4:

[0401] The server generates a response sentence using a response sentence generation means.

[0402] Input: Primary keyword.

[0403] Output: A response statement (e.g., "These shoes are waterproof and feature a lightweight design").

[0404] Specific operation: The response generation module uses a generative AI model such as GPT-4 to obtain information from a database or knowledge base and generate a response.

[0405] Step 5:

[0406] The server generates an image diagram using an image diagram generating means.

[0407] Input: Keywords related to the response sentence.

[0408] Output: Image diagram (e.g., diagram showing waterproof function and design).

[0409] Specific operation: The image generation module uses an image generation engine such as DALL-E to generate a visual image related to the response sentence.

[0410] Step 6:

[0411] The server analyzes the user's emotions using an emotion analysis means.

[0412] Input: User's voice tone and facial expression data.

[0413] Output: User's emotional state (e.g., excited, anxious).

[0414] What it does: Sentiment analysis tools (such as Amazon Rekognition or Microsoft Azure's Emotion API) analyze the user's emotions and identify their emotional state.

[0415] Step 7:

[0416] The server adjusts the tone and content of the response using a response adjustment means.

[0417] Input: A response sentence and the user's emotional state.

[0418] Output: A tailored response (e.g., "These shoes are very popular. They fit you perfectly.").

[0419] Specific operation: Based on the results of emotion analysis, the tone and content of the response are changed to match the user's emotional state.

[0420] Step 8:

[0421] The server integrates the response sentence and the image diagram using a response integration means.

[0422] Input: Adjusted response sentence and image diagram.

[0423] Output: A consolidated response (e.g., a set of tailored sentences and image diagrams).

[0424] Specific action: Combine the response sentence and image into one integrated response.

[0425] Step 9:

[0426] The server transmits the integrated response to the user's smart device via the data transmission means.

[0427] Input: The consolidated response.

[0428] Output: The response sent to the user's smart device.

[0429] Specific operation: Send the integration response to the smart glasses using HTTP or WebSocket protocol.

[0430] Step 10:

[0431] The terminal displays the integrated response to the user.

[0432] Input: The consolidated response.

[0433] Output: The response and image displayed to the user.

[0434] Specific operation: The smart glasses display the tailored response sentence and image to the user.

[0435] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0436] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0437] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0438] [Second embodiment]

[0439] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0440] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0441] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0442] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0443] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0445] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0446] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0447] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0448] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0449] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0450] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0451] The present invention is a system that provides visually easy-to-understand responses to questions entered by a user. The system is implemented by the following modules:

[0452] Question receiving module

[0453] The user submits a question to the chatbot, which is entered in text format and sent to the system in real-time or non-real-time.

[0454] Question Analysis Module

[0455] The server sends the question received from the user to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. For example, in the question "Please explain how plants photosynthesize," the words "plant," "photosynthesis," and "mechanism" are analyzed.

[0456] Response generation module

[0457] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. Specifically, it retrieves relevant information from an internal database and knowledge base and generates grammatically correct responses. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0458] Image diagram generation module

[0459] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine to generate a visual representation of the photosynthesis process, showing elements such as light energy, carbon dioxide, water, oxygen, and organic matter.

[0460] Response Integration Module

[0461] The server's response integration module integrates the response text and the image diagram to generate a single integrated response. The generated response text is followed by a related diagram, making it easy for users to understand at a glance. For example, it might look like this: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0462] Data Transmission Module

[0463] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0464] Viewing the response

[0465] The terminal displays the response text and image received from the server to the user. The response text and image are provided through the chat interface, allowing the user to immediately understand the information.

[0466] Specific examples

[0467] A user submits the following question to the chatbot: "How does photosynthesis work?"

[0468] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[0469] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0470] The image generation module generates a diagram showing the process of photosynthesis.

[0471] The response integration module integrates the response sentence and the image diagram.

[0472] A data transmission module transmits the consolidated response to the user.

[0473] The terminal displays the response text and an image to the user, who then understands the content.

[0474] In this way, intuitive and easy-to-understand responses are provided to questions entered by the user. All of the above modules work together to provide information efficiently.

[0475] The processing flow will be explained below.

[0476] Step 1:

[0477] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[0478] Step 2:

[0479] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[0480] Step 3:

[0481] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[0482] Step 4:

[0483] The server's response generation module generates a response based on the analyzed keywords. It references an internal database or knowledge base to obtain appropriate information and create a grammatically correct response. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0484] Step 5:

[0485] The server's image generation module creates an image related to the response statement. In this case, it uses a diagram generation engine to generate a visual representation of the "process of photosynthesis." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0486] Step 6:

[0487] The server's response integration means integrates the generated response sentence with the image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0488] Step 7:

[0489] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[0490] Step 8:

[0491] The terminal displays the integrated response received from the server to the user. The response text and image diagram are provided via the chat interface, allowing the user to intuitively understand the information.

[0492] This series of processes allows users to get instant, easy-to-understand answers to their questions, and the visualization of information promotes deeper understanding.

[0493] Example 1

[0494] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0495] Conventional chatbots and FAQ systems only provide text-based responses to questions entered by users, making it difficult to provide answers efficiently by including visual information. In particular, when explaining complex concepts or processes, it is difficult to understand using text alone, and visual aids are required to help users understand. By providing visual information as well, it is necessary to enable users to understand information more intuitively and quickly.

[0496] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0497] In this invention, the server includes means for receiving a question from a user, means for analyzing the received question and identifying key keywords and the intent of the question, means for generating a response sentence based on the identified keywords and phrases, means for generating a related image based on the response sentence, means for integrating the response sentence and the image to generate an integrated response, and means for sending the integrated response to the user. This makes it possible to provide a response to a question entered by a user in a format that integrates text data and visual data. This makes it easier for the user to intuitively understand complex information, and realizes smooth communication of information.

[0498] A "user" is an entity that sends a question or request to the system.

[0499] The "question receiving means" is a device or module that provides a function for receiving a question sent by a user.

[0500] The "question analysis means" is a device or module that provides a function for analyzing a received question and identifying the main keywords and intent of the question.

[0501] The "response sentence generation means" is a device or module that provides a function for generating an appropriate response sentence based on the keywords and phrases identified by the question analysis means.

[0502] The "image diagram generating means" is a device or module that provides a function for generating a related image diagram based on a response sentence.

[0503] The "response integration means" is a device or module that provides a function for integrating a response sentence and an image diagram to generate one integrated response.

[0504] A "data transmission means" is a device or module that provides the functionality for transmitting a consolidated response to a user.

[0505] A "database" is a storage system for systematically storing information and enabling efficient search and retrieval.

[0506] A "knowledge base" is an information system that stores knowledge and information in a specific domain and makes it available for response generation.

[0507] An "illustration generation engine" is a software tool for generating visual illustrations based on given data or information.

[0508] This invention is a system that provides visually easy-to-understand responses to questions entered by users. This system is composed of a server, a terminal, and a user, and is implemented by the following modules:

[0509] Question receiving method

[0510] The user enters a question into the chatbot and sends it. The question is in text format and is sent to the system in real time or non-real time. The server receives the question.

[0511] Question analysis means

[0512] The server forwards the received question to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. Specifically, it uses an NLP engine (e.g., Google NLP API or SpaCy). For example, from the question "Please explain how plants photosynthesize," the words "plants," "photosynthesis," and "mechanism" are identified.

[0513] Response sentence generation means

[0514] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. This module consults internal databases and knowledge bases (e.g., Wikipedia API, corporate knowledge bases). For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0515] Image diagram generation means

[0516] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine (e.g., D3.js, Graphviz). For example, it generates a visual diagram that shows the process of photosynthesis, illustrating the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[0517] Response Integration Measures

[0518] The server's response integration module integrates the response text and the image to generate a single integrated response, such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0519] Data transmission method

[0520] The server's data transmission module sends the consolidated response to the user using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0521] Viewing the response

[0522] The terminal displays the response text and image received from the server. The response text and image are provided through the chat interface, allowing the user to instantly understand the information.

[0523] Specific examples

[0524] User: Sends a question to the chatbot: "How does photosynthesis work?"

[0525] Server: The question analysis module identifies "photosynthesis" and "mechanism" and sends them to the response generation module.

[0526] Response sentence generation means: Generate the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0527] Image diagram generation means: Generates a diagram showing the process of photosynthesis.

[0528] Response integration method: Integrate the response sentence and the image diagram.

[0529] Data sending means: Sends the consolidated response to the user.

[0530] Terminal: Displays the response text and an image diagram, allowing the user to understand the content.

[0531] The system provides intuitive and easy-to-understand responses to questions entered by users. All of the above modules work together to provide information efficiently.

[0532] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0533] Step 1: Receiving the question

[0534] The user enters a question in text format into the chatbot and sends it. Specifically, the user enters the question into the chat interface of a web browser or mobile app and clicks the "Send" button.

[0535] The server receives a question sent by a user, the question data being in text format.

[0536] Input: A text question submitted by the user

[0537] Output: Received question data

[0538] Step 2: Parsing the Question

[0539] The server forwards the received question to the question analysis module, which uses an NLP engine (e.g., Google NLP API or SpaCy) to analyze the question content and identify key keywords and the intent of the question.

[0540] Specifically, the keywords "photosynthesis" and "mechanism" are extracted from the question "Please tell me how photosynthesis works."

[0541] Input: Received question data

[0542] Output: Extracted keywords and question intent

[0543] Step 3: Generate a response

[0544] The server's response generation module generates responses based on the keywords and phrases identified by the question analysis module. This module consults an internal database or knowledge base (e.g., Wikipedia API, corporate knowledge base).

[0545] Specifically, the response sentence generated is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0546] Input: Extracted keywords and question intent

[0547] Output: Generated response

[0548] Step 4: Generate an image

[0549] The server's image generation module generates the relevant image based on the generated response. This module uses an image generation engine (e.g., D3.js, Graphviz).

[0550] Specifically, it generates a diagram that visually shows the process of photosynthesis and explains the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[0551] Input: Generated response sentence

[0552] Output: Generated image

[0553] Step 5: Consolidating the response

[0554] The server's response integration module integrates the response text and the image to generate a single integrated response. Specifically, the response text is followed by a related image, and the format is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0555] Input: Generated response text and image

[0556] Output: Consolidated response

[0557] Step 6: Sending a Response

[0558] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0559] Input: Consolidated response

[0560] Output: The integration response sent to the user

[0561] Step 7: View the response

[0562] The terminal displays the integrated response received from the server, and the response text and image are displayed to the user through the chat interface.

[0563] This allows the user to instantly understand the answer to the question.

[0564] Input: The integration response sent by the server

[0565] Output: Response text and image displayed to the user

[0566] (Application example 1)

[0567] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0568] When obtaining maintenance and troubleshooting information for robots and machinery used in factories, it is difficult for on-site staff to obtain the information quickly and accurately. In particular, there is a lack of information provided in a visually easy-to-understand format, which can result in reduced work efficiency and safety. Another issue is that current technology does not have a widespread system that provides appropriate instructions for complex machine operation.

[0569] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0570] In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generating means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, a means for transmitting information to a display device that displays the response, and a means for inputting a question by text or voice using the device and displaying the result in real time. This enables on-site staff in a factory to quickly learn complex machine operations and maintenance procedures in a visually easy-to-understand format.

[0571] The "question receiving means" is a device or software that has the function of receiving a question from a user in text or voice format.

[0572] The "question analysis means" is a device or software that uses natural language processing technology to analyze received questions and identify their content, intent, and keywords.

[0573] The "response sentence generation means" is a device or software for generating an appropriate response sentence based on the analyzed information.

[0574] The "image diagram generating means" is a device or software for automatically generating illustrations or diagrams related to a response sentence.

[0575] The "response integration means" is a device or software that has the function of integrating the generated response sentence and image diagram to generate one integrated response.

[0576] A "data transmission means" is a device or software that uses a communication protocol to transmit the generated integrated response to a user.

[0577] The "means for transmitting information to a display device" is a device or software for transmitting instructions to cause the display device to display the aggregate response.

[0578] "Means for text or voice input and real-time display of results" refers to a device or software that receives a user's text or voice input and displays answers generated based on that input in real time.

[0579] A "generative AI model" is an artificial intelligence model that uses deep learning and neural networks to generate response sentences and image diagrams.

[0580] A "prompt" is an instruction or question given to a generative AI model, and is the input information that enables the model to generate an appropriate response.

[0581] In this invention, a system is constructed that allows a user to input a question about a robot or machine equipment used in a factory and provides a response in a visually easy-to-understand format. Detailed embodiments of this system are described below.

[0582] System configuration

[0583] The system consists of the following components:

[0584] 1. Question receiving means: The user inputs a question via text or voice. This function is realized via a smartglasses or smartphone application. For example, the user inputs a question such as, "Please tell me how to maintain this robot."

[0585] 2. Question analysis: The server receives the entered question and analyzes it using natural language processing (NLP) technology. The OpenAI API is used to identify key keywords and the intent of the question.

[0586] 3. Response generation: Based on the keywords identified by the question analysis, an appropriate response is generated. This process also uses OpenAI's generative AI model to obtain relevant information.

[0587] 4. Image diagram generation means: Generate an image diagram related to the generated response sentence using a diagram generation engine. For example, automatically generate the related diagram through an API.

[0588] 5. Response integration: The generated response sentences and images are integrated to create a single integrated response, which is displayed in a format that is easy for the user to understand visually.

[0589] 6. Data transmission means: The integrated response is transmitted to the user's display device using a communication protocol (such as HTTP or WebSocket).

[0590] 7. Means for sending information to a display device: Send information to a device (head-mounted display or smart glasses) for displaying the integrated response. The response is displayed in real time, allowing the user to instantly understand the content.

[0591] Processing Details

[0592] The server analyzes the question received from the user and extracts key keywords. This analysis is performed using OpenAI's API, based on the generated prompt text. The response generation means generates a response text based on the analysis results. An illustration generation engine is used to generate an image diagram related to this generated response text. The generated response text and image diagram are integrated and sent to the user. A communication protocol (e.g., HTTP, WebSocket, etc.) is used in this process, and the results are displayed in real time on a display device (e.g., a head-mounted display, smart glasses).

[0593] Specific examples

[0594] The user speaks the following question into the HMD application: "Please tell me how to maintain this robot."

[0595] 1. The question receiving means receives a question from a user.

[0596] 2. The question analysis tool analyzes the question and identifies “robot,” “maintenance,” and “method” as the main keywords.

[0597] 3. The response generation means uses OpenAI's generative AI model to generate a "detailed response regarding the robot's maintenance procedures."

[0598] 4. The image diagram generating means generates a related diagram (for example, a parts layout diagram or a work procedure diagram) based on the generated response sentence.

[0599] 5. The response synthesis means creates an integrated response by integrating the response sentence and the image diagram.

[0600] 6. The data transmission means transmits the integrated response to the user's HMD.

[0601] 7. A means for transmitting information to a display device displays the integrated response in real time on the HMD.

[0602] Example prompt sentence:

[0603] Question analysis prompt:

[0604] Analyze the following question and extract the main keywords: How do I maintain this robot?

[0605] Prompt for generating a response:

[0606] Please provide more details on robot maintenance methods.

[0607] As described above, the present invention allows on-site staff to visually and quickly obtain information about robots and machinery equipment.

[0608] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0609] Step 1:

[0610] The user inputs a question by text or voice. This question is sent to the server through an application on smart glasses or a smartphone. If the user inputs, "Please tell me how to maintain this robot," the text or voice data is sent to the server.

[0611] Step 2:

[0612] The server analyzes the received question using a question analysis method. Specifically, it uses natural language processing (NLP) technology to understand the content and intent of the question. The input is the user's question text or voice data, and the output is key keywords such as "robot," "maintenance," and "method." This analysis uses OpenAI's API.

[0613] Step 3:

[0614] The server uses OpenAI's generative AI model to generate an appropriate response based on the analyzed keywords. The input is a prompt related to the analyzed keyword "robot maintenance method," and the output is a "detailed response regarding robot maintenance procedures." This generative AI model automatically obtains relevant information and generates a grammatically correct response.

[0615] Step 4:

[0616] Based on the generated response, the server uses a diagram generation engine to generate related diagrams. The input is the response, and the output is diagrams such as "diagrams of robot maintenance procedures" or "parts layout diagrams." The server automatically generates related diagrams by calling the appropriate external API.

[0617] Step 5:

[0618] The server integrates the generated response sentence and image to generate a single integrated response. The input is the response sentence and image, and the output is a single integrated visual response containing text and illustrations. These are integrated using a response integration tool.

[0619] Step 6:

[0620] The server transmits the integrated response to the user's display device via a data transmission means. The input is the integrated response, and the output is data displayed on the user's display device (e.g., head-mounted display, smart glasses). Information is transmitted in real time using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0621] Step 7:

[0622] The terminal uses a means for transmitting information to a display device to display the generated integrated response in real time. The input is the response data transmitted from the server, and the output is a display that the user can visually confirm. The user can check the integrated response text and image diagram and intuitively understand the necessary information.

[0623] Through the above steps, users can visually and efficiently obtain answers to their questions.

[0624] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0625] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the addition of a function to recognize the user's emotions and adjust the content of the response. This system is implemented by the following modules and emotion engine.

[0626] Question receiving module

[0627] The user sends a text question to the chatbot, for example, "How does photosynthesis work?"

[0628] Question Analysis Module

[0629] The server sends the question received from the user to the question analysis module, which uses natural language processing (NLP) techniques to analyze and identify the question content and key keywords (e.g., "photosynthesis," "mechanism").

[0630] Response generation module

[0631] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0632] Image diagram generation module

[0633] The server's image generation module creates an image related to the response statement. Using the diagram generation engine, it generates a visual representation of the "photosynthesis process." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0634] Emotion Engine

[0635] The emotion engine on the server analyzes the user's emotions based on user data such as voice tone, facial expressions, and input text. The emotion engine recognizes the user's emotions (e.g., excitement, anxiety, joy).

[0636] Response Adjustment Measures

[0637] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. If the user is anxious, the response will be adjusted to be more reassuring. Conversely, if the user is excited, calming words will be emphasized.

[0638] Response Integration Module

[0639] The server's response integration module integrates the adjusted response sentence with the generated image diagram. The diagram is placed after the response sentence and arranged in a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0640] Data Transmission Module

[0641] The server's data transmission module sends the consolidated response to the user, sending the response to the chat interface using a communication protocol (e.g. HTTP, WebSocket, etc.).

[0642] Viewing the response

[0643] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[0644] Specific examples

[0645] A user submits the following question to the chatbot: "How does photosynthesis work?"

[0646] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[0647] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0648] The image generation module generates a diagram showing the process of photosynthesis.

[0649] The emotion engine analyzes the user's emotions and recognizes that the user is in an anxious state.

[0650] A response adjustment means adjusts the response sentence to a tone that puts the user at ease.

[0651] The response integration module integrates the response sentence and the image diagram.

[0652] A data transmission module transmits the consolidated response to the user.

[0653] The terminal displays the response text and an image to the user, who then understands the content.

[0654] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand the information more deeply and comfortably.

[0655] The processing flow will be explained below.

[0656] Step 1:

[0657] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[0658] Step 2:

[0659] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[0660] Step 3:

[0661] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[0662] Step 4:

[0663] The server's emotion engine analyzes the user's emotion from the question text, using NLP techniques to identify emotions from the linguistic tone and keywords in the text (e.g., excitement, anxiety, joy, etc.).

[0664] Step 5:

[0665] The server's response generation module generates a response based on the analyzed keywords and the emotion analysis results from the emotion engine. It retrieves relevant information from the internal database and knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0666] Step 6:

[0667] The server's response adjustment means adjusts the generated response sentence based on the user's emotions. For example, if the user is anxious, the response sentence will be changed to a tone that gives a sense of security. Example: "Don't worry, I'll explain how photosynthesis works. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0668] Step 7:

[0669] The server's image generation module creates an image related to the tailored response. This module uses a diagram generation engine to generate a visual representation of the "process of photosynthesis," showing light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0670] Step 8:

[0671] The server's response integration means integrates the adjusted response sentence with the generated image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Don't worry, I'll explain the mechanism of photosynthesis. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0672] Step 9:

[0673] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[0674] Step 10:

[0675] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[0676] This process allows users to get instant, easy-to-understand answers to their questions, and by visualizing the information and adjusting responses based on emotion, users can more comfortably understand the information.

[0677] Example 2

[0678] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0679] Conventional question-answering systems generate uniform responses without considering the user's emotions, resulting in poor user satisfaction. In particular, when a user feels anxious or confused, the system may be unable to respond appropriately, resulting in a poor user experience. Furthermore, while image diagrams are generated to aid visual understanding, they are often mismatched with the text, making it difficult to intuitively understand the information. To solve these problems, a system is needed that provides visually understandable responses while taking the user's emotions into account.

[0680] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a sentiment analysis means for analyzing the user's sentiment, a response adjustment means for adjusting the tone and content of the response sentence based on the analysis result of the sentiment analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, and a data transmission means for transmitting the integrated response to the user. This makes it possible to provide a flexible response that takes into consideration the user's sentiment and a response that is visually easy to understand.

[0681] The "question receiving means" is a function that receives text questions from users to the chatbot.

[0682] The "question analysis means" is a function that analyzes received questions and identifies the question content and main keywords using natural language processing technology.

[0683] The "response sentence generation means" is a function that generates a grammatically correct response sentence by obtaining information from an internal database or knowledge base based on the information identified by the question analysis means.

[0684] The "image diagram generating means" is a function that generates visual information related to a response sentence as an illustration and expresses it in an easy-to-understand form.

[0685] The "emotion analysis means" is a function that analyzes data such as the user's voice tone, facial expression, and input text, and identifies the user's emotions.

[0686] The "response adjustment means" is a function that adjusts the tone and content of the response sentence to suit the user's emotions based on the results of the emotion analysis means.

[0687] The "response integration means" is a function that integrates the adjusted response sentence with the generated image diagram to generate an integrated response that is arranged in a visually intuitive format.

[0688] The "data transmission means" is a function for transmitting a response using a communication protocol (e.g., HTTP, WebSocket, etc.) for transmitting an integrated response to a user.

[0689] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the added function of recognizing the user's emotions and adjusting the content of the response. The configuration and operation of this system are described in detail below.

[0690] System configuration:

[0691] Hardware and Software Use:

[0692] Server: The server that is the center of data processing and response generation. It has a high-performance CPU and memory.

[0693] Terminal: A device used by a user (e.g., a PC, smartphone, or tablet) with a web browser or dedicated application installed.

[0694] Natural Language Processing Engine (NLP): Software used for text analysis (e.g., SpaCy, NLTK).

[0695] Diagram generation engine: Software for generating visual images (e.g., D3.js, Chart.js).

[0696] Database and knowledge base: Stores information for generating response sentences.

[0697] Sentiment analysis engine: Software that analyzes user emotions (e.g., Microsoft Azure Emotion API, Google Cloud Vision API).

[0698] System behavior:

[0699] Receiving questions:

[0700] Users send questions to the chatbot in text format, for example, "Please explain how photosynthesis works."

[0701] Parsing the question:

[0702] The server sends the received question to a question analysis module, which uses natural language processing (NLP) techniques to analyze the question and identify key keywords (e.g., "photosynthesis," "mechanism").

[0703] Generate a response:

[0704] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from a database or knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0705] Generate image diagrams:

[0706] The server's image generation module generates an image related to the response statement. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a visual representation of the "process of photosynthesis."

[0707] Sentiment Analysis:

[0708] The server's emotion engine analyzes the user's emotions by extracting emotions from voice tone, facial expressions (using facial recognition technology), and input text. At this stage, it may be recognized that the user is in an anxious state.

[0709] Tailoring response content:

[0710] The server adjusts the response based on the analysis results of the emotion engine. For example, if the user is feeling anxious, the response will be adjusted to something that gives a sense of security, such as "Don't worry, photosynthesis is a natural process for plants."

[0711] Response integration:

[0712] The server's response integration module integrates the adjusted response sentences with the generated image diagrams, arranging the response sentences followed by the image diagrams in a visually intuitive format.

[0713] Sending data:

[0714] The server's data transmission module sends the consolidated response to the user's device, and sends the response to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0715] View response:

[0716] The terminal displays the integrated response received from the server. The response text and image are provided to the user through the chat interface, allowing the user to intuitively understand the information.

[0717] Examples:

[0718] 1. A user submits the following question to the chatbot: "How does photosynthesis work?"

[0719] 2. The server identifies "photosynthesis" and "mechanism" using the question analysis module.

[0720] 3. The response generation module generates the response sentence, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0721] 4. The image generation module generates a diagram showing the process of photosynthesis.

[0722] 5. The emotion engine analyzes the user's emotions and recognizes when the user is in an anxious state.

[0723] 6. The response adjuster adjusts the response to a more reassuring tone.

[0724] 7. The response integration module integrates the response sentence and the image diagram.

[0725] 8. The data transmission module sends the consolidated response to the user.

[0726] 9. The terminal displays the response text and image to the user, and the user understands the content.

[0727] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand information more deeply and comfortably.

[0728] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0729] Step 1:

[0730] The user sends a question to the chatbot in text format. For example, they can enter a question like, "Please tell me how photosynthesis works." The input data is the text data that the user types into the input field. The sent text data is transferred to the server.

[0731] Step 2:

[0732] The server sends the question data received from the user to the question analysis module, which uses natural language processing (NLP) technology to analyze the question content. The input data is the question data in text format, and the output is the main keywords (e.g., "photosynthesis" and "mechanism").

[0733] Step 3:

[0734] The server's response generation module generates a response based on the keywords output from the question analysis module. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. The input data are keywords, and the output is a response (e.g., "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water").

[0735] Step 4:

[0736] The server's image generation module creates an image related to the generated response text. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a diagram that visually represents the "photosynthesis process." The input data is the response text, and the output is an image showing the photosynthesis process.

[0737] Step 5:

[0738] The emotion engine on the server analyzes the user's emotion. The input data are voice tone, facial expression (using facial recognition technology), and entered text. The emotion engine analyzes these data and identifies the emotion that the user is anxious. The output is the user's emotional state (e.g., anxious).

[0739] Step 6:

[0740] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. The input data is the analyzed emotional state and the original response. Based on these input data, the server changes the response to include reassuring elements (e.g., "Don't worry, photosynthesis is a natural process for plants"). The output is the adjusted response.

[0741] Step 7:

[0742] The server's response synthesis module synthesizes the adjusted response sentence and the generated image diagram. The input data are the adjusted response sentence and the image diagram. By synthesizing these, a visually intuitive response is generated. The output is the synthesized response.

[0743] Step 8:

[0744] The server's data transmission module sends the consolidated response to the user's device. The input data is the consolidated response, which is sent to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.). The output is the response data sent to the user's device.

[0745] Step 9:

[0746] The terminal displays the integrated response received from the server to the user. The input data is the response data sent from the server, and the output is the response text and image displayed on the chat interface. The user can intuitively understand the information through this display.

[0747] (Application example 2)

[0748] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0749] Conventional information provision systems only provide text-based responses to user input, which are often difficult to understand visually and do not allow for flexible responses that reflect the user's emotions.The present invention aims to provide more appropriate and reassuring information by providing visually easy-to-understand responses to questions entered by the user, recognizing the user's emotions, and adjusting the response content.

[0750] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, an emotion analysis means for analyzing the user's emotion, a response adjustment means for adjusting the tone and content of the response sentence based on the emotion analyzed by the emotion analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, and a display means for displaying the integrated response on the smart device. This allows the user to obtain visually easy-to-understand information and receive a flexible response suited to their individual emotional state.

[0751] The "question receiving means" is a device or function for receiving a question in text or voice format sent by a user.

[0752] The "question analysis means" is a device or function for analyzing a received question and identifying the content of the question and related main keywords.

[0753] The "response sentence generation means" is a device or function for generating a response sentence based on the analyzed keywords and providing the necessary information.

[0754] The "image diagram generating means" is a device or function for generating a visual image diagram related to a response sentence.

[0755] The "emotion analysis means" is a device or function for analyzing the user's emotions and adjusting the response content based on those emotions.

[0756] The "response adjustment means" is a device or function for adjusting the tone and content of a response sentence to match the user's emotions based on the emotion analysis results.

[0757] The "response integration means" is a device or function for integrating the response sentence and the generated image diagram to generate one integrated response.

[0758] A "data transmission means" is a communication device or function for transmitting an integrated response to a user.

[0759] The "display means" is a device or function for displaying the integrated response on the smart device.

[0760] A "smart device" is an electronic device used by a user, such as smart glasses, a smartphone, a head-mounted display, or a robot.

[0761] A system for implementing this invention provides visual and emotional responses to user questions when the user is engaged in activities such as shopping through a smart device (such as smart glasses, a smartphone, a head-mounted display, or a robot). An embodiment of this system is described in detail below.

[0762] Program Structure

[0763] The system includes a question receiving means, a question analyzing means, a response sentence generating means, an image diagram generating means, a sentiment analyzing means, a response adjusting means, a response integrating means, a data transmitting means, and a display means.

[0764] Program processing and use of hardware and software

[0765] Question receiving method

[0766] Users enter questions into smart glasses or smartphones by text or voice, using voice recognition software or a text input interface.

[0767] Question analysis means

[0768] The server analyzes the questions received from users using natural language processing (NLP) technology to identify key keywords, such as Google AI's natural language processing API.

[0769] Response sentence generation means

[0770] Based on the analyzed keywords, the server uses a generative AI model such as GPT-4 to generate a response, which is composed of information retrieved from a database or knowledge base.

[0771] Image diagram generation means

[0772] An image generation engine such as DALL-E is used to generate images related to the response, providing visual content appropriate for the response.

[0773] Emotion analysis means

[0774] To analyze user emotions, emotion analysis tools (e.g., Amazon Rekognition or Microsoft Azure's Emotion API) are used that process voice tone and facial expression data.

[0775] Response Adjustment Measures

[0776] Based on the results of the emotion analysis, the server adjusts the tone and content of the response to suit the user's emotions. If the user is anxious, the content will be adjusted to provide a sense of security.

[0777] Response Integration Measures

[0778] The response text and the generated image are integrated to generate a single integrated response, allowing the user to receive both text and visual information in a single interface.

[0779] Data transmission method

[0780] The server sends the integration response to the user's smart device using HTTP or WebSocket protocol.

[0781] Display means

[0782] The smart device displays the integrated response to the user.

[0783] Specific examples

[0784] When a user asks the smart glasses, "What are the characteristics of these shoes?", the following process is performed:

[0785] 1. Question analysis: The NLU module identifies the "shoes" and "features."

[0786] 2. Response generation: The generative AI model generates an answer such as, "These shoes are waterproof and feature a lightweight design."

[0787] 3. Image generation: DALL-E generates an image showing the waterproof function and lightweight design of the shoes.

[0788] 4. Sentiment Analysis: The emotion engine recognizes when the user is excited.

[0789] 5. Response adjustment: "These shoes are very popular. They fit you perfectly."

[0790] 6. Response Integration: Integrate the coordinated response statement with the illustration.

[0791] 7. Data transmission: sent to smart glasses and displayed to the user.

[0792] Prompt Sentence Examples

[0793] "Tell me the features of these shoes."

[0794] Analyzed keywords: shoes, features

[0795] Generated response: "These shoes are waterproof and feature a lightweight design."

[0796] Sentiment Analysis: Users are excited

[0797] Tailored response: "These shoes are very popular. They look great on you."

[0798] Final response: "These shoes are very popular. They're perfect for you. They're waterproof and lightweight. (Illustration: Waterproofing and Design)"

[0799] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0800] Step 1:

[0801] The user types or speaks a question into the smart device.

[0802] Input: Text or voice input from the user.

[0803] Output: Question text or audio file.

[0804] Specific operation: The user asks the smart glasses verbally, "Tell me the features of these shoes."

[0805] Step 2:

[0806] The device converts voice input into text data (voice recognition).

[0807] Input: User's voice input.

[0808] Output: The converted text data.

[0809] What it does: Speech recognition software converts the speech data into text, such as "What are the features of these shoes?"

[0810] Step 3:

[0811] The server receives the question and analyzes it using the question analysis means.

[0812] Input: Text data.

[0813] Output: Primary keywords (e.g., "shoes", "features").

[0814] How it works: The question receiving module receives the text and uses NLP technology (such as Google AI's natural language processing API) to identify key keywords.

[0815] Step 4:

[0816] The server generates a response sentence using a response sentence generation means.

[0817] Input: Primary keyword.

[0818] Output: A response statement (e.g., "These shoes are waterproof and feature a lightweight design").

[0819] Specific operation: The response generation module uses a generative AI model such as GPT-4 to obtain information from a database or knowledge base and generate a response.

[0820] Step 5:

[0821] The server generates an image diagram using an image diagram generating means.

[0822] Input: Keywords related to the response sentence.

[0823] Output: Image diagram (e.g., diagram showing waterproof function and design).

[0824] Specific operation: The image generation module uses an image generation engine such as DALL-E to generate a visual image related to the response sentence.

[0825] Step 6:

[0826] The server analyzes the user's emotions using an emotion analysis means.

[0827] Input: User's voice tone and facial expression data.

[0828] Output: User's emotional state (e.g., excited, anxious).

[0829] What it does: Sentiment analysis tools (such as Amazon Rekognition or Microsoft Azure's Emotion API) analyze the user's emotions and identify their emotional state.

[0830] Step 7:

[0831] The server adjusts the tone and content of the response using a response adjustment means.

[0832] Input: A response sentence and the user's emotional state.

[0833] Output: A tailored response (e.g., "These shoes are very popular. They fit you perfectly.").

[0834] Specific operation: Based on the results of emotion analysis, the tone and content of the response are changed to match the user's emotional state.

[0835] Step 8:

[0836] The server integrates the response sentence and the image diagram using a response integration means.

[0837] Input: Adjusted response sentence and image diagram.

[0838] Output: A consolidated response (e.g., a set of tailored sentences and image diagrams).

[0839] Specific action: Combine the response sentence and image into one integrated response.

[0840] Step 9:

[0841] The server transmits the integrated response to the user's smart device via the data transmission means.

[0842] Input: The consolidated response.

[0843] Output: The response sent to the user's smart device.

[0844] Specific operation: Send the integration response to the smart glasses using HTTP or WebSocket protocol.

[0845] Step 10:

[0846] The terminal displays the integrated response to the user.

[0847] Input: The consolidated response.

[0848] Output: The response and image displayed to the user.

[0849] Specific operation: The smart glasses display the tailored response sentence and image to the user.

[0850] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0851] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0852] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0853] [Third embodiment]

[0854] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0855] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0856] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0857] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0858] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0859] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0860] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0861] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0862] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0863] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0864] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0865] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0866] The present invention is a system that provides visually easy-to-understand responses to questions entered by a user. The system is implemented by the following modules:

[0867] Question receiving module

[0868] The user submits a question to the chatbot, which is entered in text format and sent to the system in real-time or non-real-time.

[0869] Question Analysis Module

[0870] The server sends the question received from the user to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. For example, in the question "Please explain how plants photosynthesize," the words "plant," "photosynthesis," and "mechanism" are analyzed.

[0871] Response generation module

[0872] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. Specifically, it retrieves relevant information from an internal database and knowledge base and generates grammatically correct responses. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0873] Image diagram generation module

[0874] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine to generate a visual representation of the photosynthesis process, showing elements such as light energy, carbon dioxide, water, oxygen, and organic matter.

[0875] Response Integration Module

[0876] The server's response integration module integrates the response text and the image diagram to generate a single integrated response. The generated response text is followed by a related diagram, making it easy for users to understand at a glance. For example, it might look like this: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0877] Data Transmission Module

[0878] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0879] Viewing the response

[0880] The terminal displays the response text and image received from the server to the user. The response text and image are provided through the chat interface, allowing the user to immediately understand the information.

[0881] Specific examples

[0882] A user submits the following question to the chatbot: "How does photosynthesis work?"

[0883] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[0884] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0885] The image generation module generates a diagram showing the process of photosynthesis.

[0886] The response integration module integrates the response sentence and the image diagram.

[0887] A data transmission module transmits the consolidated response to the user.

[0888] The terminal displays the response text and an image to the user, who then understands the content.

[0889] In this way, intuitive and easy-to-understand responses are provided to questions entered by the user. All of the above modules work together to provide information efficiently.

[0890] The processing flow will be explained below.

[0891] Step 1:

[0892] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[0893] Step 2:

[0894] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[0895] Step 3:

[0896] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[0897] Step 4:

[0898] The server's response generation module generates a response based on the analyzed keywords. It references an internal database or knowledge base to obtain appropriate information and create a grammatically correct response. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0899] Step 5:

[0900] The server's image generation module creates an image related to the response statement. In this case, it uses a diagram generation engine to generate a visual representation of the "process of photosynthesis." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[0901] Step 6:

[0902] The server's response integration means integrates the generated response sentence with the image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[0903] Step 7:

[0904] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[0905] Step 8:

[0906] The terminal displays the integrated response received from the server to the user. The response text and image diagram are provided via the chat interface, allowing the user to intuitively understand the information.

[0907] This series of processes allows users to get instant, easy-to-understand answers to their questions, and the visualization of information promotes deeper understanding.

[0908] Example 1

[0909] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0910] Conventional chatbots and FAQ systems only provide text-based responses to questions entered by users, making it difficult to provide answers efficiently by including visual information. In particular, when explaining complex concepts or processes, it is difficult to understand using text alone, and visual aids are required to help users understand. By providing visual information as well, it is necessary to enable users to understand information more intuitively and quickly.

[0911] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0912] In this invention, the server includes means for receiving a question from a user, means for analyzing the received question and identifying key keywords and the intent of the question, means for generating a response sentence based on the identified keywords and phrases, means for generating a related image based on the response sentence, means for integrating the response sentence and the image to generate an integrated response, and means for sending the integrated response to the user. This makes it possible to provide a response to a question entered by a user in a format that integrates text data and visual data. This makes it easier for the user to intuitively understand complex information, and realizes smooth communication of information.

[0913] A "user" is an entity that sends a question or request to the system.

[0914] The "question receiving means" is a device or module that provides a function for receiving a question sent by a user.

[0915] The "question analysis means" is a device or module that provides a function for analyzing a received question and identifying the main keywords and intent of the question.

[0916] The "response sentence generation means" is a device or module that provides a function for generating an appropriate response sentence based on the keywords and phrases identified by the question analysis means.

[0917] The "image diagram generating means" is a device or module that provides a function for generating a related image diagram based on a response sentence.

[0918] The "response integration means" is a device or module that provides a function for integrating a response sentence and an image diagram to generate one integrated response.

[0919] A "data transmission means" is a device or module that provides the functionality for transmitting a consolidated response to a user.

[0920] A "database" is a storage system for systematically storing information and enabling efficient search and retrieval.

[0921] A "knowledge base" is an information system that stores knowledge and information in a specific domain and makes it available for response generation.

[0922] An "illustration generation engine" is a software tool for generating visual illustrations based on given data or information.

[0923] This invention is a system that provides visually easy-to-understand responses to questions entered by users. This system is composed of a server, a terminal, and a user, and is implemented by the following modules:

[0924] Question receiving method

[0925] The user enters a question into the chatbot and sends it. The question is in text format and is sent to the system in real time or non-real time. The server receives the question.

[0926] Question analysis means

[0927] The server forwards the received question to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. Specifically, it uses an NLP engine (e.g., Google NLP API or SpaCy). For example, from the question "Please explain how plants photosynthesize," the words "plants," "photosynthesis," and "mechanism" are identified.

[0928] Response sentence generation means

[0929] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. This module consults internal databases and knowledge bases (e.g., Wikipedia API, corporate knowledge bases). For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0930] Image diagram generation means

[0931] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine (e.g., D3.js, Graphviz). For example, it generates a visual diagram that shows the process of photosynthesis, illustrating the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[0932] Response Integration Measures

[0933] The server's response integration module integrates the response text and the image to generate a single integrated response, such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0934] Data transmission method

[0935] The server's data transmission module sends the consolidated response to the user using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0936] Viewing the response

[0937] The terminal displays the response text and image received from the server. The response text and image are provided through the chat interface, allowing the user to instantly understand the information.

[0938] Specific examples

[0939] User: Sends a question to the chatbot: "How does photosynthesis work?"

[0940] Server: The question analysis module identifies "photosynthesis" and "mechanism" and sends them to the response generation module.

[0941] Response sentence generation means: Generate the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0942] Image diagram generation means: Generates a diagram showing the process of photosynthesis.

[0943] Response integration method: Integrate the response sentence and the image diagram.

[0944] Data sending means: Sends the consolidated response to the user.

[0945] Terminal: Displays the response text and an image diagram, allowing the user to understand the content.

[0946] The system provides intuitive and easy-to-understand responses to questions entered by users. All of the above modules work together to provide information efficiently.

[0947] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0948] Step 1: Receiving the question

[0949] The user enters a question in text format into the chatbot and sends it. Specifically, the user enters the question into the chat interface of a web browser or mobile app and clicks the "Send" button.

[0950] The server receives a question sent by a user, the question data being in text format.

[0951] Input: A text question submitted by the user

[0952] Output: Received question data

[0953] Step 2: Parsing the Question

[0954] The server forwards the received question to the question analysis module, which uses an NLP engine (e.g., Google NLP API or SpaCy) to analyze the question content and identify key keywords and the intent of the question.

[0955] Specifically, the keywords "photosynthesis" and "mechanism" are extracted from the question "Please tell me how photosynthesis works."

[0956] Input: Received question data

[0957] Output: Extracted keywords and question intent

[0958] Step 3: Generate a response

[0959] The server's response generation module generates responses based on the keywords and phrases identified by the question analysis module. This module consults an internal database or knowledge base (e.g., Wikipedia API, corporate knowledge base).

[0960] Specifically, the response sentence generated is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[0961] Input: Extracted keywords and question intent

[0962] Output: Generated response

[0963] Step 4: Generate an image

[0964] The server's image generation module generates the relevant image based on the generated response. This module uses an image generation engine (e.g., D3.js, Graphviz).

[0965] Specifically, it generates a diagram that visually shows the process of photosynthesis and explains the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[0966] Input: Generated response sentence

[0967] Output: Generated image

[0968] Step 5: Consolidating the response

[0969] The server's response integration module integrates the response text and the image to generate a single integrated response. Specifically, the response text is followed by a related image, and the format is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[0970] Input: Generated response text and image

[0971] Output: Consolidated response

[0972] Step 6: Sending a Response

[0973] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[0974] Input: Consolidated response

[0975] Output: The integration response sent to the user

[0976] Step 7: View the response

[0977] The terminal displays the integrated response received from the server, and the response text and image are displayed to the user through the chat interface.

[0978] This allows the user to instantly understand the answer to the question.

[0979] Input: The integration response sent by the server

[0980] Output: Response text and image displayed to the user

[0981] (Application example 1)

[0982] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0983] When obtaining maintenance and troubleshooting information for robots and machinery used in factories, it is difficult for on-site staff to obtain the information quickly and accurately. In particular, there is a lack of information provided in a visually easy-to-understand format, which can result in reduced work efficiency and safety. Another issue is that current technology does not have a widespread system that provides appropriate instructions for complex machine operation.

[0984] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0985] In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generating means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, a means for transmitting information to a display device that displays the response, and a means for inputting a question by text or voice using the device and displaying the result in real time. This enables on-site staff in a factory to quickly learn complex machine operations and maintenance procedures in a visually easy-to-understand format.

[0986] The "question receiving means" is a device or software that has the function of receiving a question from a user in text or voice format.

[0987] The "question analysis means" is a device or software that uses natural language processing technology to analyze received questions and identify their content, intent, and keywords.

[0988] The "response sentence generation means" is a device or software for generating an appropriate response sentence based on the analyzed information.

[0989] The "image diagram generating means" is a device or software for automatically generating illustrations or diagrams related to a response sentence.

[0990] The "response integration means" is a device or software that has the function of integrating the generated response sentence and image diagram to generate one integrated response.

[0991] A "data transmission means" is a device or software that uses a communication protocol to transmit the generated integrated response to a user.

[0992] The "means for transmitting information to a display device" is a device or software for transmitting instructions to cause the display device to display the aggregate response.

[0993] "Means for text or voice input and real-time display of results" refers to a device or software that receives a user's text or voice input and displays answers generated based on that input in real time.

[0994] A "generative AI model" is an artificial intelligence model that uses deep learning and neural networks to generate response sentences and image diagrams.

[0995] A "prompt" is an instruction or question given to a generative AI model, and is the input information that enables the model to generate an appropriate response.

[0996] In this invention, a system is constructed that allows a user to input a question about a robot or machine equipment used in a factory and provides a response in a visually easy-to-understand format. Detailed embodiments of this system are described below.

[0997] System configuration

[0998] The system consists of the following components:

[0999] 1. Question receiving means: The user inputs a question via text or voice. This function is realized via a smartglasses or smartphone application. For example, the user inputs a question such as, "Please tell me how to maintain this robot."

[1000] 2. Question analysis: The server receives the entered question and analyzes it using natural language processing (NLP) technology. The OpenAI API is used to identify key keywords and the intent of the question.

[1001] 3. Response generation: Based on the keywords identified by the question analysis, an appropriate response is generated. This process also uses OpenAI's generative AI model to obtain relevant information.

[1002] 4. Image diagram generation means: Generate an image diagram related to the generated response sentence using a diagram generation engine. For example, automatically generate the related diagram through an API.

[1003] 5. Response integration: The generated response sentences and images are integrated to create a single integrated response, which is displayed in a format that is easy for the user to understand visually.

[1004] 6. Data transmission means: The integrated response is transmitted to the user's display device using a communication protocol (such as HTTP or WebSocket).

[1005] 7. Means for sending information to a display device: Send information to a device (head-mounted display or smart glasses) for displaying the integrated response. The response is displayed in real time, allowing the user to instantly understand the content.

[1006] Processing Details

[1007] The server analyzes the question received from the user and extracts key keywords. This analysis is performed using OpenAI's API, based on the generated prompt text. The response generation means generates a response text based on the analysis results. An illustration generation engine is used to generate an image diagram related to this generated response text. The generated response text and image diagram are integrated and sent to the user. A communication protocol (e.g., HTTP, WebSocket, etc.) is used in this process, and the results are displayed in real time on a display device (e.g., a head-mounted display, smart glasses).

[1008] Specific examples

[1009] The user speaks the following question into the HMD application: "Please tell me how to maintain this robot."

[1010] 1. The question receiving means receives a question from a user.

[1011] 2. The question analysis tool analyzes the question and identifies “robot,” “maintenance,” and “method” as the main keywords.

[1012] 3. The response generation means uses OpenAI's generative AI model to generate a "detailed response regarding the robot's maintenance procedures."

[1013] 4. The image diagram generating means generates a related diagram (for example, a parts layout diagram or a work procedure diagram) based on the generated response sentence.

[1014] 5. The response synthesis means creates an integrated response by integrating the response sentence and the image diagram.

[1015] 6. The data transmission means transmits the integrated response to the user's HMD.

[1016] 7. A means for transmitting information to a display device displays the integrated response in real time on the HMD.

[1017] Example prompt sentence:

[1018] Question analysis prompt:

[1019] Analyze the following question and extract the main keywords: How do I maintain this robot?

[1020] Prompt for generating a response:

[1021] Please provide more details on robot maintenance methods.

[1022] As described above, the present invention allows on-site staff to visually and quickly obtain information about robots and machinery equipment.

[1023] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1024] Step 1:

[1025] The user inputs a question by text or voice. This question is sent to the server through an application on smart glasses or a smartphone. If the user inputs, "Please tell me how to maintain this robot," the text or voice data is sent to the server.

[1026] Step 2:

[1027] The server analyzes the received question using a question analysis method. Specifically, it uses natural language processing (NLP) technology to understand the content and intent of the question. The input is the user's question text or voice data, and the output is key keywords such as "robot," "maintenance," and "method." This analysis uses OpenAI's API.

[1028] Step 3:

[1029] The server uses OpenAI's generative AI model to generate an appropriate response based on the analyzed keywords. The input is a prompt related to the analyzed keyword "robot maintenance method," and the output is a "detailed response regarding robot maintenance procedures." This generative AI model automatically obtains relevant information and generates a grammatically correct response.

[1030] Step 4:

[1031] Based on the generated response, the server uses a diagram generation engine to generate related diagrams. The input is the response, and the output is diagrams such as "diagrams of robot maintenance procedures" or "parts layout diagrams." The server automatically generates related diagrams by calling the appropriate external API.

[1032] Step 5:

[1033] The server integrates the generated response sentence and image to generate a single integrated response. The input is the response sentence and image, and the output is a single integrated visual response containing text and illustrations. These are integrated using a response integration tool.

[1034] Step 6:

[1035] The server transmits the integrated response to the user's display device via a data transmission means. The input is the integrated response, and the output is data displayed on the user's display device (e.g., head-mounted display, smart glasses). Information is transmitted in real time using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1036] Step 7:

[1037] The terminal uses a means for transmitting information to a display device to display the generated integrated response in real time. The input is the response data transmitted from the server, and the output is a display that the user can visually confirm. The user can check the integrated response text and image diagram and intuitively understand the necessary information.

[1038] Through the above steps, users can visually and efficiently obtain answers to their questions.

[1039] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1040] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the addition of a function to recognize the user's emotions and adjust the content of the response. This system is implemented by the following modules and emotion engine.

[1041] Question receiving module

[1042] The user sends a text question to the chatbot, for example, "How does photosynthesis work?"

[1043] Question Analysis Module

[1044] The server sends the question received from the user to the question analysis module, which uses natural language processing (NLP) techniques to analyze and identify the question content and key keywords (e.g., "photosynthesis," "mechanism").

[1045] Response generation module

[1046] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1047] Image diagram generation module

[1048] The server's image generation module creates an image related to the response statement. Using the diagram generation engine, it generates a visual representation of the "photosynthesis process." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[1049] Emotion Engine

[1050] The emotion engine on the server analyzes the user's emotions based on user data such as voice tone, facial expressions, and input text. The emotion engine recognizes the user's emotions (e.g., excitement, anxiety, joy).

[1051] Response Adjustment Measures

[1052] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. If the user is anxious, the response will be adjusted to be more reassuring. Conversely, if the user is excited, calming words will be emphasized.

[1053] Response Integration Module

[1054] The server's response integration module integrates the adjusted response sentence with the generated image diagram. The diagram is placed after the response sentence and arranged in a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[1055] Data Transmission Module

[1056] The server's data transmission module sends the consolidated response to the user, sending the response to the chat interface using a communication protocol (e.g. HTTP, WebSocket, etc.).

[1057] Viewing the response

[1058] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[1059] Specific examples

[1060] A user submits the following question to the chatbot: "How does photosynthesis work?"

[1061] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[1062] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1063] The image generation module generates a diagram showing the process of photosynthesis.

[1064] The emotion engine analyzes the user's emotions and recognizes that the user is in an anxious state.

[1065] A response adjustment means adjusts the response sentence to a tone that puts the user at ease.

[1066] The response integration module integrates the response sentence and the image diagram.

[1067] A data transmission module transmits the consolidated response to the user.

[1068] The terminal displays the response text and an image to the user, who then understands the content.

[1069] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand the information more deeply and comfortably.

[1070] The processing flow will be explained below.

[1071] Step 1:

[1072] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[1073] Step 2:

[1074] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[1075] Step 3:

[1076] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[1077] Step 4:

[1078] The server's emotion engine analyzes the user's emotion from the question text, using NLP techniques to identify emotions from the linguistic tone and keywords in the text (e.g., excitement, anxiety, joy, etc.).

[1079] Step 5:

[1080] The server's response generation module generates a response based on the analyzed keywords and the emotion analysis results from the emotion engine. It retrieves relevant information from the internal database and knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1081] Step 6:

[1082] The server's response adjustment means adjusts the generated response sentence based on the user's emotions. For example, if the user is anxious, the response sentence will be changed to a tone that gives a sense of security. Example: "Don't worry, I'll explain how photosynthesis works. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1083] Step 7:

[1084] The server's image generation module creates an image related to the tailored response. This module uses a diagram generation engine to generate a visual representation of the "process of photosynthesis," showing light energy, carbon dioxide, water, oxygen, organic matter, etc.

[1085] Step 8:

[1086] The server's response integration means integrates the adjusted response sentence with the generated image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Don't worry, I'll explain the mechanism of photosynthesis. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[1087] Step 9:

[1088] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[1089] Step 10:

[1090] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[1091] This process allows users to get instant, easy-to-understand answers to their questions, and by visualizing the information and adjusting responses based on emotion, users can more comfortably understand the information.

[1092] Example 2

[1093] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1094] Conventional question-answering systems generate uniform responses without considering the user's emotions, resulting in poor user satisfaction. In particular, when a user feels anxious or confused, the system may be unable to respond appropriately, resulting in a poor user experience. Furthermore, while image diagrams are generated to aid visual understanding, they are often mismatched with the text, making it difficult to intuitively understand the information. To solve these problems, a system is needed that provides visually understandable responses while taking the user's emotions into account.

[1095] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a sentiment analysis means for analyzing the user's sentiment, a response adjustment means for adjusting the tone and content of the response sentence based on the analysis result of the sentiment analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, and a data transmission means for transmitting the integrated response to the user. This makes it possible to provide a flexible response that takes into consideration the user's sentiment and a response that is visually easy to understand.

[1096] The "question receiving means" is a function that receives text questions from users to the chatbot.

[1097] The "question analysis means" is a function that analyzes received questions and identifies the question content and main keywords using natural language processing technology.

[1098] The "response sentence generation means" is a function that generates a grammatically correct response sentence by obtaining information from an internal database or knowledge base based on the information identified by the question analysis means.

[1099] The "image diagram generating means" is a function that generates visual information related to a response sentence as an illustration and expresses it in an easy-to-understand form.

[1100] The "emotion analysis means" is a function that analyzes data such as the user's voice tone, facial expression, and input text, and identifies the user's emotions.

[1101] The "response adjustment means" is a function that adjusts the tone and content of the response sentence to suit the user's emotions based on the results of the emotion analysis means.

[1102] The "response integration means" is a function that integrates the adjusted response sentence with the generated image diagram to generate an integrated response that is arranged in a visually intuitive format.

[1103] The "data transmission means" is a function for transmitting a response using a communication protocol (e.g., HTTP, WebSocket, etc.) for transmitting an integrated response to a user.

[1104] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the added function of recognizing the user's emotions and adjusting the content of the response. The configuration and operation of this system are described in detail below.

[1105] System configuration:

[1106] Hardware and Software Use:

[1107] Server: The server that is the center of data processing and response generation. It has a high-performance CPU and memory.

[1108] Terminal: A device used by a user (e.g., a PC, smartphone, or tablet) with a web browser or dedicated application installed.

[1109] Natural Language Processing Engine (NLP): Software used for text analysis (e.g., SpaCy, NLTK).

[1110] Diagram generation engine: Software for generating visual images (e.g., D3.js, Chart.js).

[1111] Database and knowledge base: Stores information for generating response sentences.

[1112] Sentiment analysis engine: Software that analyzes user emotions (e.g., Microsoft Azure Emotion API, Google Cloud Vision API).

[1113] System behavior:

[1114] Receiving questions:

[1115] Users send questions to the chatbot in text format, for example, "Please explain how photosynthesis works."

[1116] Parsing the question:

[1117] The server sends the received question to a question analysis module, which uses natural language processing (NLP) techniques to analyze the question and identify key keywords (e.g., "photosynthesis," "mechanism").

[1118] Generate a response:

[1119] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from a database or knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1120] Generate image diagrams:

[1121] The server's image generation module generates an image related to the response statement. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a visual representation of the "process of photosynthesis."

[1122] Sentiment Analysis:

[1123] The server's emotion engine analyzes the user's emotions by extracting emotions from voice tone, facial expressions (using facial recognition technology), and input text. At this stage, it may be recognized that the user is in an anxious state.

[1124] Tailoring response content:

[1125] The server adjusts the response based on the analysis results of the emotion engine. For example, if the user is feeling anxious, the response will be adjusted to something that gives a sense of security, such as "Don't worry, photosynthesis is a natural process for plants."

[1126] Response integration:

[1127] The server's response integration module integrates the adjusted response sentences with the generated image diagrams, arranging the response sentences followed by the image diagrams in a visually intuitive format.

[1128] Sending data:

[1129] The server's data transmission module sends the consolidated response to the user's device, and sends the response to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1130] View response:

[1131] The terminal displays the integrated response received from the server. The response text and image are provided to the user through the chat interface, allowing the user to intuitively understand the information.

[1132] Examples:

[1133] 1. A user submits the following question to the chatbot: "How does photosynthesis work?"

[1134] 2. The server identifies "photosynthesis" and "mechanism" using the question analysis module.

[1135] 3. The response generation module generates the response sentence, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1136] 4. The image generation module generates a diagram showing the process of photosynthesis.

[1137] 5. The emotion engine analyzes the user's emotions and recognizes when the user is in an anxious state.

[1138] 6. The response adjuster adjusts the response to a more reassuring tone.

[1139] 7. The response integration module integrates the response sentence and the image diagram.

[1140] 8. The data transmission module sends the consolidated response to the user.

[1141] 9. The terminal displays the response text and image to the user, and the user understands the content.

[1142] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand information more deeply and comfortably.

[1143] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1144] Step 1:

[1145] The user sends a question to the chatbot in text format. For example, they can enter a question like, "Please tell me how photosynthesis works." The input data is the text data that the user types into the input field. The sent text data is transferred to the server.

[1146] Step 2:

[1147] The server sends the question data received from the user to the question analysis module, which uses natural language processing (NLP) technology to analyze the question content. The input data is the question data in text format, and the output is the main keywords (e.g., "photosynthesis" and "mechanism").

[1148] Step 3:

[1149] The server's response generation module generates a response based on the keywords output from the question analysis module. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. The input data are keywords, and the output is a response (e.g., "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water").

[1150] Step 4:

[1151] The server's image generation module creates an image related to the generated response text. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a diagram that visually represents the "photosynthesis process." The input data is the response text, and the output is an image showing the photosynthesis process.

[1152] Step 5:

[1153] The emotion engine on the server analyzes the user's emotion. The input data are voice tone, facial expression (using facial recognition technology), and entered text. The emotion engine analyzes these data and identifies the emotion that the user is anxious. The output is the user's emotional state (e.g., anxious).

[1154] Step 6:

[1155] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. The input data is the analyzed emotional state and the original response. Based on these input data, the server changes the response to include reassuring elements (e.g., "Don't worry, photosynthesis is a natural process for plants"). The output is the adjusted response.

[1156] Step 7:

[1157] The server's response synthesis module synthesizes the adjusted response sentence and the generated image diagram. The input data are the adjusted response sentence and the image diagram. By synthesizing these, a visually intuitive response is generated. The output is the synthesized response.

[1158] Step 8:

[1159] The server's data transmission module sends the consolidated response to the user's device. The input data is the consolidated response, which is sent to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.). The output is the response data sent to the user's device.

[1160] Step 9:

[1161] The terminal displays the integrated response received from the server to the user. The input data is the response data sent from the server, and the output is the response text and image displayed on the chat interface. The user can intuitively understand the information through this display.

[1162] (Application example 2)

[1163] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1164] Conventional information provision systems only provide text-based responses to user input, which are often difficult to understand visually and do not allow for flexible responses that reflect the user's emotions.The present invention aims to provide more appropriate and reassuring information by providing visually easy-to-understand responses to questions entered by the user, recognizing the user's emotions, and adjusting the response content.

[1165] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, an emotion analysis means for analyzing the user's emotion, a response adjustment means for adjusting the tone and content of the response sentence based on the emotion analyzed by the emotion analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, and a display means for displaying the integrated response on the smart device. This allows the user to obtain visually easy-to-understand information and receive a flexible response suited to their individual emotional state.

[1166] The "question receiving means" is a device or function for receiving a question in text or voice format sent by a user.

[1167] The "question analysis means" is a device or function for analyzing a received question and identifying the content of the question and related main keywords.

[1168] The "response sentence generation means" is a device or function for generating a response sentence based on the analyzed keywords and providing the necessary information.

[1169] The "image diagram generating means" is a device or function for generating a visual image diagram related to a response sentence.

[1170] The "emotion analysis means" is a device or function for analyzing the user's emotions and adjusting the response content based on those emotions.

[1171] The "response adjustment means" is a device or function for adjusting the tone and content of a response sentence to match the user's emotions based on the emotion analysis results.

[1172] The "response integration means" is a device or function for integrating the response sentence and the generated image diagram to generate one integrated response.

[1173] A "data transmission means" is a communication device or function for transmitting an integrated response to a user.

[1174] The "display means" is a device or function for displaying the integrated response on the smart device.

[1175] A "smart device" is an electronic device used by a user, such as smart glasses, a smartphone, a head-mounted display, or a robot.

[1176] A system for implementing this invention provides visual and emotional responses to user questions when the user is engaged in activities such as shopping through a smart device (such as smart glasses, a smartphone, a head-mounted display, or a robot). An embodiment of this system is described in detail below.

[1177] Program Structure

[1178] The system includes a question receiving means, a question analyzing means, a response sentence generating means, an image diagram generating means, a sentiment analyzing means, a response adjusting means, a response integrating means, a data transmitting means, and a display means.

[1179] Program processing and use of hardware and software

[1180] Question receiving method

[1181] Users enter questions into smart glasses or smartphones by text or voice, using voice recognition software or a text input interface.

[1182] Question analysis means

[1183] The server analyzes the questions received from users using natural language processing (NLP) technology to identify key keywords, such as Google AI's natural language processing API.

[1184] Response sentence generation means

[1185] Based on the analyzed keywords, the server uses a generative AI model such as GPT-4 to generate a response, which is composed of information retrieved from a database or knowledge base.

[1186] Image diagram generation means

[1187] An image generation engine such as DALL-E is used to generate images related to the response, providing visual content appropriate for the response.

[1188] Emotion analysis means

[1189] To analyze user emotions, emotion analysis tools (e.g., Amazon Rekognition or Microsoft Azure's Emotion API) are used that process voice tone and facial expression data.

[1190] Response Adjustment Measures

[1191] Based on the results of the emotion analysis, the server adjusts the tone and content of the response to suit the user's emotions. If the user is anxious, the content will be adjusted to provide a sense of security.

[1192] Response Integration Measures

[1193] The response text and the generated image are integrated to generate a single integrated response, allowing the user to receive both text and visual information in a single interface.

[1194] Data transmission method

[1195] The server sends the integration response to the user's smart device using HTTP or WebSocket protocol.

[1196] Display means

[1197] The smart device displays the integrated response to the user.

[1198] Specific examples

[1199] When a user asks the smart glasses, "What are the characteristics of these shoes?", the following process is performed:

[1200] 1. Question analysis: The NLU module identifies the "shoes" and "features."

[1201] 2. Response generation: The generative AI model generates an answer such as, "These shoes are waterproof and feature a lightweight design."

[1202] 3. Image generation: DALL-E generates an image showing the waterproof function and lightweight design of the shoes.

[1203] 4. Sentiment Analysis: The emotion engine recognizes when the user is excited.

[1204] 5. Response adjustment: "These shoes are very popular. They fit you perfectly."

[1205] 6. Response Integration: Integrate the coordinated response statement with the illustration.

[1206] 7. Data transmission: sent to smart glasses and displayed to the user.

[1207] Prompt Sentence Examples

[1208] "Tell me the features of these shoes."

[1209] Analyzed keywords: shoes, features

[1210] Generated response: "These shoes are waterproof and feature a lightweight design."

[1211] Sentiment Analysis: Users are excited

[1212] Tailored response: "These shoes are very popular. They look great on you."

[1213] Final response: "These shoes are very popular. They're perfect for you. They're waterproof and lightweight. (Illustration: Waterproofing and Design)"

[1214] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1215] Step 1:

[1216] The user types or speaks a question into the smart device.

[1217] Input: Text or voice input from the user.

[1218] Output: Question text or audio file.

[1219] Specific operation: The user asks the smart glasses verbally, "Tell me the features of these shoes."

[1220] Step 2:

[1221] The device converts voice input into text data (voice recognition).

[1222] Input: User's voice input.

[1223] Output: The converted text data.

[1224] What it does: Speech recognition software converts the speech data into text, such as "What are the features of these shoes?"

[1225] Step 3:

[1226] The server receives the question and analyzes it using the question analysis means.

[1227] Input: Text data.

[1228] Output: Primary keywords (e.g., "shoes", "features").

[1229] How it works: The question receiving module receives the text and uses NLP technology (such as Google AI's natural language processing API) to identify key keywords.

[1230] Step 4:

[1231] The server generates a response sentence using a response sentence generation means.

[1232] Input: Primary keyword.

[1233] Output: A response statement (e.g., "These shoes are waterproof and feature a lightweight design").

[1234] Specific operation: The response generation module uses a generative AI model such as GPT-4 to obtain information from a database or knowledge base and generate a response.

[1235] Step 5:

[1236] The server generates an image diagram using an image diagram generating means.

[1237] Input: Keywords related to the response sentence.

[1238] Output: Image diagram (e.g., diagram showing waterproof function and design).

[1239] Specific operation: The image generation module uses an image generation engine such as DALL-E to generate a visual image related to the response sentence.

[1240] Step 6:

[1241] The server analyzes the user's emotions using an emotion analysis means.

[1242] Input: User's voice tone and facial expression data.

[1243] Output: User's emotional state (e.g., excited, anxious).

[1244] What it does: Sentiment analysis tools (such as Amazon Rekognition or Microsoft Azure's Emotion API) analyze the user's emotions and identify their emotional state.

[1245] Step 7:

[1246] The server adjusts the tone and content of the response using a response adjustment means.

[1247] Input: A response sentence and the user's emotional state.

[1248] Output: A tailored response (e.g., "These shoes are very popular. They fit you perfectly.").

[1249] Specific operation: Based on the results of emotion analysis, the tone and content of the response are changed to match the user's emotional state.

[1250] Step 8:

[1251] The server integrates the response sentence and the image diagram using a response integration means.

[1252] Input: Adjusted response sentence and image diagram.

[1253] Output: A consolidated response (e.g., a set of tailored sentences and image diagrams).

[1254] Specific action: Combine the response sentence and image into one integrated response.

[1255] Step 9:

[1256] The server transmits the integrated response to the user's smart device via the data transmission means.

[1257] Input: The consolidated response.

[1258] Output: The response sent to the user's smart device.

[1259] Specific operation: Send the integration response to the smart glasses using HTTP or WebSocket protocol.

[1260] Step 10:

[1261] The terminal displays the integrated response to the user.

[1262] Input: The consolidated response.

[1263] Output: The response and image displayed to the user.

[1264] Specific operation: The smart glasses display the tailored response sentence and image to the user.

[1265] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1266] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1267] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1268] [Fourth embodiment]

[1269] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1270] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1271] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1272] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1273] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1274] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1275] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1276] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1277] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1278] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1279] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1280] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1281] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1282] The present invention is a system that provides visually easy-to-understand responses to questions entered by a user. The system is implemented by the following modules:

[1283] Question receiving module

[1284] The user submits a question to the chatbot, which is entered in text format and sent to the system in real-time or non-real-time.

[1285] Question Analysis Module

[1286] The server sends the question received from the user to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. For example, in the question "Please explain how plants photosynthesize," the words "plant," "photosynthesis," and "mechanism" are analyzed.

[1287] Response generation module

[1288] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. Specifically, it retrieves relevant information from an internal database and knowledge base and generates grammatically correct responses. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1289] Image diagram generation module

[1290] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine to generate a visual representation of the photosynthesis process, showing elements such as light energy, carbon dioxide, water, oxygen, and organic matter.

[1291] Response Integration Module

[1292] The server's response integration module integrates the response text and the image diagram to generate a single integrated response. The generated response text is followed by a related diagram, making it easy for users to understand at a glance. For example, it might look like this: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[1293] Data Transmission Module

[1294] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1295] Viewing the response

[1296] The terminal displays the response text and image received from the server to the user. The response text and image are provided through the chat interface, allowing the user to immediately understand the information.

[1297] Specific examples

[1298] A user submits the following question to the chatbot: "How does photosynthesis work?"

[1299] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[1300] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1301] The image generation module generates a diagram showing the process of photosynthesis.

[1302] The response integration module integrates the response sentence and the image diagram.

[1303] A data transmission module transmits the consolidated response to the user.

[1304] The terminal displays the response text and an image to the user, who then understands the content.

[1305] In this way, intuitive and easy-to-understand responses are provided to questions entered by the user. All of the above modules work together to provide information efficiently.

[1306] The processing flow will be explained below.

[1307] Step 1:

[1308] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[1309] Step 2:

[1310] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[1311] Step 3:

[1312] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[1313] Step 4:

[1314] The server's response generation module generates a response based on the analyzed keywords. It references an internal database or knowledge base to obtain appropriate information and create a grammatically correct response. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1315] Step 5:

[1316] The server's image generation module creates an image related to the response statement. In this case, it uses a diagram generation engine to generate a visual representation of the "process of photosynthesis." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[1317] Step 6:

[1318] The server's response integration means integrates the generated response sentence with the image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[1319] Step 7:

[1320] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[1321] Step 8:

[1322] The terminal displays the integrated response received from the server to the user. The response text and image diagram are provided via the chat interface, allowing the user to intuitively understand the information.

[1323] This series of processes allows users to get instant, easy-to-understand answers to their questions, and the visualization of information promotes deeper understanding.

[1324] Example 1

[1325] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1326] Conventional chatbots and FAQ systems only provide text-based responses to questions entered by users, making it difficult to provide answers efficiently by including visual information. In particular, when explaining complex concepts or processes, it is difficult to understand using text alone, and visual aids are required to help users understand. By providing visual information as well, it is necessary to enable users to understand information more intuitively and quickly.

[1327] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1328] In this invention, the server includes means for receiving a question from a user, means for analyzing the received question and identifying key keywords and the intent of the question, means for generating a response sentence based on the identified keywords and phrases, means for generating a related image based on the response sentence, means for integrating the response sentence and the image to generate an integrated response, and means for sending the integrated response to the user. This makes it possible to provide a response to a question entered by a user in a format that integrates text data and visual data. This makes it easier for the user to intuitively understand complex information, and realizes smooth communication of information.

[1329] A "user" is an entity that sends a question or request to the system.

[1330] The "question receiving means" is a device or module that provides a function for receiving a question sent by a user.

[1331] The "question analysis means" is a device or module that provides a function for analyzing a received question and identifying the main keywords and intent of the question.

[1332] The "response sentence generation means" is a device or module that provides a function for generating an appropriate response sentence based on the keywords and phrases identified by the question analysis means.

[1333] The "image diagram generating means" is a device or module that provides a function for generating a related image diagram based on a response sentence.

[1334] The "response integration means" is a device or module that provides a function for integrating a response sentence and an image diagram to generate one integrated response.

[1335] A "data transmission means" is a device or module that provides the functionality for transmitting a consolidated response to a user.

[1336] A "database" is a storage system for systematically storing information and enabling efficient search and retrieval.

[1337] A "knowledge base" is an information system that stores knowledge and information in a specific domain and makes it available for response generation.

[1338] An "illustration generation engine" is a software tool for generating visual illustrations based on given data or information.

[1339] This invention is a system that provides visually easy-to-understand responses to questions entered by users. This system is composed of a server, a terminal, and a user, and is implemented by the following modules:

[1340] Question receiving method

[1341] The user enters a question into the chatbot and sends it. The question is in text format and is sent to the system in real time or non-real time. The server receives the question.

[1342] Question analysis means

[1343] The server forwards the received question to the question analysis module. This module uses natural language processing (NLP) technology to analyze the question and identify key keywords and the intent of the question. Specifically, it uses an NLP engine (e.g., Google NLP API or SpaCy). For example, from the question "Please explain how plants photosynthesize," the words "plants," "photosynthesis," and "mechanism" are identified.

[1344] Response sentence generation means

[1345] The server's response generation module generates appropriate responses based on the keywords and phrases identified by the question analysis module. This module consults internal databases and knowledge bases (e.g., Wikipedia API, corporate knowledge bases). For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1346] Image diagram generation means

[1347] The server's image generation module generates a related image based on the response text. This module uses a diagram generation engine (e.g., D3.js, Graphviz). For example, it generates a visual diagram that shows the process of photosynthesis, illustrating the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[1348] Response Integration Measures

[1349] The server's response integration module integrates the response text and the image to generate a single integrated response, such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[1350] Data transmission method

[1351] The server's data transmission module sends the consolidated response to the user using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1352] Viewing the response

[1353] The terminal displays the response text and image received from the server. The response text and image are provided through the chat interface, allowing the user to instantly understand the information.

[1354] Specific examples

[1355] User: Sends a question to the chatbot: "How does photosynthesis work?"

[1356] Server: The question analysis module identifies "photosynthesis" and "mechanism" and sends them to the response generation module.

[1357] Response sentence generation means: Generate the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1358] Image diagram generation means: Generates a diagram showing the process of photosynthesis.

[1359] Response integration method: Integrate the response sentence and the image diagram.

[1360] Data sending means: Sends the consolidated response to the user.

[1361] Terminal: Displays the response text and an image diagram, allowing the user to understand the content.

[1362] The system provides intuitive and easy-to-understand responses to questions entered by users. All of the above modules work together to provide information efficiently.

[1363] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1364] Step 1: Receiving the question

[1365] The user enters a question in text format into the chatbot and sends it. Specifically, the user enters the question into the chat interface of a web browser or mobile app and clicks the "Send" button.

[1366] The server receives a question sent by a user, the question data being in text format.

[1367] Input: A text question submitted by the user

[1368] Output: Received question data

[1369] Step 2: Parsing the Question

[1370] The server forwards the received question to the question analysis module, which uses an NLP engine (e.g., Google NLP API or SpaCy) to analyze the question content and identify key keywords and the intent of the question.

[1371] Specifically, the keywords "photosynthesis" and "mechanism" are extracted from the question "Please tell me how photosynthesis works."

[1372] Input: Received question data

[1373] Output: Extracted keywords and question intent

[1374] Step 3: Generate a response

[1375] The server's response generation module generates responses based on the keywords and phrases identified by the question analysis module. This module consults an internal database or knowledge base (e.g., Wikipedia API, corporate knowledge base).

[1376] Specifically, the response sentence generated is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1377] Input: Extracted keywords and question intent

[1378] Output: Generated response

[1379] Step 4: Generate an image

[1380] The server's image generation module generates the relevant image based on the generated response. This module uses an image generation engine (e.g., D3.js, Graphviz).

[1381] Specifically, it generates a diagram that visually shows the process of photosynthesis and explains the relationship between light energy, carbon dioxide, water, oxygen, and organic matter.

[1382] Input: Generated response sentence

[1383] Output: Generated image

[1384] Step 5: Consolidating the response

[1385] The server's response integration module integrates the response text and the image to generate a single integrated response. Specifically, the response text is followed by a related image, and the format is "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Figure: Mechanism of photosynthesis)."

[1386] Input: Generated response text and image

[1387] Output: Consolidated response

[1388] Step 6: Sending a Response

[1389] The server's data transmission module sends the consolidated response to the user via the chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1390] Input: Consolidated response

[1391] Output: The integration response sent to the user

[1392] Step 7: View the response

[1393] The terminal displays the integrated response received from the server, and the response text and image are displayed to the user through the chat interface.

[1394] This allows the user to instantly understand the answer to the question.

[1395] Input: The integration response sent by the server

[1396] Output: Response text and image displayed to the user

[1397] (Application example 1)

[1398] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1399] When obtaining maintenance and troubleshooting information for robots and machinery used in factories, it is difficult for on-site staff to obtain the information quickly and accurately. In particular, there is a lack of information provided in a visually easy-to-understand format, which can result in reduced work efficiency and safety. Another issue is that current technology does not have a widespread system that provides appropriate instructions for complex machine operation.

[1400] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1401] In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generating means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, a means for transmitting information to a display device that displays the response, and a means for inputting a question by text or voice using the device and displaying the result in real time. This enables on-site staff in a factory to quickly learn complex machine operations and maintenance procedures in a visually easy-to-understand format.

[1402] The "question receiving means" is a device or software that has the function of receiving a question from a user in text or voice format.

[1403] The "question analysis means" is a device or software that uses natural language processing technology to analyze received questions and identify their content, intent, and keywords.

[1404] The "response sentence generation means" is a device or software for generating an appropriate response sentence based on the analyzed information.

[1405] The "image diagram generating means" is a device or software for automatically generating illustrations or diagrams related to a response sentence.

[1406] The "response integration means" is a device or software that has the function of integrating the generated response sentence and image diagram to generate one integrated response.

[1407] A "data transmission means" is a device or software that uses a communication protocol to transmit the generated integrated response to a user.

[1408] The "means for transmitting information to a display device" is a device or software for transmitting instructions to cause the display device to display the aggregate response.

[1409] "Means for text or voice input and real-time display of results" refers to a device or software that receives a user's text or voice input and displays answers generated based on that input in real time.

[1410] A "generative AI model" is an artificial intelligence model that uses deep learning and neural networks to generate response sentences and image diagrams.

[1411] A "prompt" is an instruction or question given to a generative AI model, and is the input information that enables the model to generate an appropriate response.

[1412] In this invention, a system is constructed that allows a user to input a question about a robot or machine equipment used in a factory and provides a response in a visually easy-to-understand format. Detailed embodiments of this system are described below.

[1413] System configuration

[1414] The system consists of the following components:

[1415] 1. Question receiving means: The user inputs a question via text or voice. This function is realized via a smartglasses or smartphone application. For example, the user inputs a question such as, "Please tell me how to maintain this robot."

[1416] 2. Question analysis: The server receives the entered question and analyzes it using natural language processing (NLP) technology. The OpenAI API is used to identify key keywords and the intent of the question.

[1417] 3. Response generation: Based on the keywords identified by the question analysis, an appropriate response is generated. This process also uses OpenAI's generative AI model to obtain relevant information.

[1418] 4. Image diagram generation means: Generate an image diagram related to the generated response sentence using a diagram generation engine. For example, automatically generate the related diagram through an API.

[1419] 5. Response integration: The generated response sentences and images are integrated to create a single integrated response, which is displayed in a format that is easy for the user to understand visually.

[1420] 6. Data transmission means: The integrated response is transmitted to the user's display device using a communication protocol (such as HTTP or WebSocket).

[1421] 7. Means for sending information to a display device: Send information to a device (head-mounted display or smart glasses) for displaying the integrated response. The response is displayed in real time, allowing the user to instantly understand the content.

[1422] Processing Details

[1423] The server analyzes the question received from the user and extracts key keywords. This analysis is performed using OpenAI's API, based on the generated prompt text. The response generation means generates a response text based on the analysis results. An illustration generation engine is used to generate an image diagram related to this generated response text. The generated response text and image diagram are integrated and sent to the user. A communication protocol (e.g., HTTP, WebSocket, etc.) is used in this process, and the results are displayed in real time on a display device (e.g., a head-mounted display, smart glasses).

[1424] Specific examples

[1425] The user speaks the following question into the HMD application: "Please tell me how to maintain this robot."

[1426] 1. The question receiving means receives a question from a user.

[1427] 2. The question analysis tool analyzes the question and identifies “robot,” “maintenance,” and “method” as the main keywords.

[1428] 3. The response generation means uses OpenAI's generative AI model to generate a "detailed response regarding the robot's maintenance procedures."

[1429] 4. The image diagram generating means generates a related diagram (for example, a parts layout diagram or a work procedure diagram) based on the generated response sentence.

[1430] 5. The response synthesis means creates an integrated response by integrating the response sentence and the image diagram.

[1431] 6. The data transmission means transmits the integrated response to the user's HMD.

[1432] 7. A means for transmitting information to a display device displays the integrated response in real time on the HMD.

[1433] Example prompt sentence:

[1434] Question analysis prompt:

[1435] Analyze the following question and extract the main keywords: How do I maintain this robot?

[1436] Prompt for generating a response:

[1437] Please provide more details on robot maintenance methods.

[1438] As described above, the present invention allows on-site staff to visually and quickly obtain information about robots and machinery equipment.

[1439] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1440] Step 1:

[1441] The user inputs a question by text or voice. This question is sent to the server through an application on smart glasses or a smartphone. If the user inputs, "Please tell me how to maintain this robot," the text or voice data is sent to the server.

[1442] Step 2:

[1443] The server analyzes the received question using a question analysis method. Specifically, it uses natural language processing (NLP) technology to understand the content and intent of the question. The input is the user's question text or voice data, and the output is key keywords such as "robot," "maintenance," and "method." This analysis uses OpenAI's API.

[1444] Step 3:

[1445] The server uses OpenAI's generative AI model to generate an appropriate response based on the analyzed keywords. The input is a prompt related to the analyzed keyword "robot maintenance method," and the output is a "detailed response regarding robot maintenance procedures." This generative AI model automatically obtains relevant information and generates a grammatically correct response.

[1446] Step 4:

[1447] Based on the generated response, the server uses a diagram generation engine to generate related diagrams. The input is the response, and the output is diagrams such as "diagrams of robot maintenance procedures" or "parts layout diagrams." The server automatically generates related diagrams by calling the appropriate external API.

[1448] Step 5:

[1449] The server integrates the generated response sentence and image to generate a single integrated response. The input is the response sentence and image, and the output is a single integrated visual response containing text and illustrations. These are integrated using a response integration tool.

[1450] Step 6:

[1451] The server transmits the integrated response to the user's display device via a data transmission means. The input is the integrated response, and the output is data displayed on the user's display device (e.g., head-mounted display, smart glasses). Information is transmitted in real time using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1452] Step 7:

[1453] The terminal uses a means for transmitting information to a display device to display the generated integrated response in real time. The input is the response data transmitted from the server, and the output is a display that the user can visually confirm. The user can check the integrated response text and image diagram and intuitively understand the necessary information.

[1454] Through the above steps, users can visually and efficiently obtain answers to their questions.

[1455] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1456] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the addition of a function to recognize the user's emotions and adjust the content of the response. This system is implemented by the following modules and emotion engine.

[1457] Question receiving module

[1458] The user sends a text question to the chatbot, for example, "How does photosynthesis work?"

[1459] Question Analysis Module

[1460] The server sends the question received from the user to the question analysis module, which uses natural language processing (NLP) techniques to analyze and identify the question content and key keywords (e.g., "photosynthesis," "mechanism").

[1461] Response generation module

[1462] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. For example, it generates the response, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1463] Image diagram generation module

[1464] The server's image generation module creates an image related to the response statement. Using the diagram generation engine, it generates a visual representation of the "photosynthesis process." The diagram shows light energy, carbon dioxide, water, oxygen, organic matter, etc.

[1465] Emotion Engine

[1466] The emotion engine on the server analyzes the user's emotions based on user data such as voice tone, facial expressions, and input text. The emotion engine recognizes the user's emotions (e.g., excitement, anxiety, joy).

[1467] Response Adjustment Measures

[1468] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. If the user is anxious, the response will be adjusted to be more reassuring. Conversely, if the user is excited, calming words will be emphasized.

[1469] Response Integration Module

[1470] The server's response integration module integrates the adjusted response sentence with the generated image diagram. The diagram is placed after the response sentence and arranged in a visually intuitive format. Example: "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[1471] Data Transmission Module

[1472] The server's data transmission module sends the consolidated response to the user, sending the response to the chat interface using a communication protocol (e.g. HTTP, WebSocket, etc.).

[1473] Viewing the response

[1474] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[1475] Specific examples

[1476] A user submits the following question to the chatbot: "How does photosynthesis work?"

[1477] The server identifies "photosynthesis" and "mechanism" using a question analysis module.

[1478] The response sentence generation module generates the response sentence "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1479] The image generation module generates a diagram showing the process of photosynthesis.

[1480] The emotion engine analyzes the user's emotions and recognizes that the user is in an anxious state.

[1481] A response adjustment means adjusts the response sentence to a tone that puts the user at ease.

[1482] The response integration module integrates the response sentence and the image diagram.

[1483] A data transmission module transmits the consolidated response to the user.

[1484] The terminal displays the response text and an image to the user, who then understands the content.

[1485] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand the information more deeply and comfortably.

[1486] The processing flow will be explained below.

[1487] Step 1:

[1488] The user types a question into the chatbot in text format and sends it. Example: "Please tell me how photosynthesis works."

[1489] Step 2:

[1490] The server receives the question sent from the user, and the question receiving means acquires the content of the question.

[1491] Step 3:

[1492] The server's question analysis module processes the received question, using natural language processing (NLP) techniques to extract and identify keywords related to the intent of the question (e.g., "photosynthesis," "mechanism").

[1493] Step 4:

[1494] The server's emotion engine analyzes the user's emotion from the question text, using NLP techniques to identify emotions from the linguistic tone and keywords in the text (e.g., excitement, anxiety, joy, etc.).

[1495] Step 5:

[1496] The server's response generation module generates a response based on the analyzed keywords and the emotion analysis results from the emotion engine. It retrieves relevant information from the internal database and knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1497] Step 6:

[1498] The server's response adjustment means adjusts the generated response sentence based on the user's emotions. For example, if the user is anxious, the response sentence will be changed to a tone that gives a sense of security. Example: "Don't worry, I'll explain how photosynthesis works. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1499] Step 7:

[1500] The server's image generation module creates an image related to the tailored response. This module uses a diagram generation engine to generate a visual representation of the "process of photosynthesis," showing light energy, carbon dioxide, water, oxygen, organic matter, etc.

[1501] Step 8:

[1502] The server's response integration means integrates the adjusted response sentence with the generated image diagram. The relevant diagram is placed after the response sentence, creating a visually intuitive format. Example: "Don't worry, I'll explain the mechanism of photosynthesis. Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water. (Diagram: Mechanism of photosynthesis)"

[1503] Step 9:

[1504] The server's data transmission means sends the consolidated response to the user, sending the response to the chat interface using a communications protocol (e.g., HTTP, WebSocket, etc.).

[1505] Step 10:

[1506] The terminal displays the integrated response received from the server to the user. The response text and image are provided through the chat interface, allowing the user to intuitively understand the information.

[1507] This process allows users to get instant, easy-to-understand answers to their questions, and by visualizing the information and adjusting responses based on emotion, users can more comfortably understand the information.

[1508] Example 2

[1509] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1510] Conventional question-answering systems generate uniform responses without considering the user's emotions, resulting in poor user satisfaction. In particular, when a user feels anxious or confused, the system may be unable to respond appropriately, resulting in a poor user experience. Furthermore, while image diagrams are generated to aid visual understanding, they are often mismatched with the text, making it difficult to intuitively understand the information. To solve these problems, a system is needed that provides visually understandable responses while taking the user's emotions into account.

[1511] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, a sentiment analysis means for analyzing the user's sentiment, a response adjustment means for adjusting the tone and content of the response sentence based on the analysis result of the sentiment analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, and a data transmission means for transmitting the integrated response to the user. This makes it possible to provide a flexible response that takes into consideration the user's sentiment and a response that is visually easy to understand.

[1512] The "question receiving means" is a function that receives text questions from users to the chatbot.

[1513] The "question analysis means" is a function that analyzes received questions and identifies the question content and main keywords using natural language processing technology.

[1514] The "response sentence generation means" is a function that generates a grammatically correct response sentence by obtaining information from an internal database or knowledge base based on the information identified by the question analysis means.

[1515] The "image diagram generating means" is a function that generates visual information related to a response sentence as an illustration and expresses it in an easy-to-understand form.

[1516] The "emotion analysis means" is a function that analyzes data such as the user's voice tone, facial expression, and input text, and identifies the user's emotions.

[1517] The "response adjustment means" is a function that adjusts the tone and content of the response sentence to suit the user's emotions based on the results of the emotion analysis means.

[1518] The "response integration means" is a function that integrates the adjusted response sentence with the generated image diagram to generate an integrated response that is arranged in a visually intuitive format.

[1519] The "data transmission means" is a function for transmitting a response using a communication protocol (e.g., HTTP, WebSocket, etc.) for transmitting an integrated response to a user.

[1520] This invention is a system that provides visually easy-to-understand responses to questions entered by a user, with the added function of recognizing the user's emotions and adjusting the content of the response. The configuration and operation of this system are described in detail below.

[1521] System configuration:

[1522] Hardware and Software Use:

[1523] Server: The server that is the center of data processing and response generation. It has a high-performance CPU and memory.

[1524] Terminal: A device used by a user (e.g., a PC, smartphone, or tablet) with a web browser or dedicated application installed.

[1525] Natural Language Processing Engine (NLP): Software used for text analysis (e.g., SpaCy, NLTK).

[1526] Diagram generation engine: Software for generating visual images (e.g., D3.js, Chart.js).

[1527] Database and knowledge base: Stores information for generating response sentences.

[1528] Sentiment analysis engine: Software that analyzes user emotions (e.g., Microsoft Azure Emotion API, Google Cloud Vision API).

[1529] System behavior:

[1530] Receiving questions:

[1531] Users send questions to the chatbot in text format, for example, "Please explain how photosynthesis works."

[1532] Parsing the question:

[1533] The server sends the received question to a question analysis module, which uses natural language processing (NLP) techniques to analyze the question and identify key keywords (e.g., "photosynthesis," "mechanism").

[1534] Generate a response:

[1535] The server's response generation module generates a response based on the analyzed keywords. It retrieves relevant information from a database or knowledge base to create a grammatically correct response. For example, it generates a response such as "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1536] Generate image diagrams:

[1537] The server's image generation module generates an image related to the response statement. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a visual representation of the "process of photosynthesis."

[1538] Sentiment Analysis:

[1539] The server's emotion engine analyzes the user's emotions by extracting emotions from voice tone, facial expressions (using facial recognition technology), and input text. At this stage, it may be recognized that the user is in an anxious state.

[1540] Tailoring response content:

[1541] The server adjusts the response based on the analysis results of the emotion engine. For example, if the user is feeling anxious, the response will be adjusted to something that gives a sense of security, such as "Don't worry, photosynthesis is a natural process for plants."

[1542] Response integration:

[1543] The server's response integration module integrates the adjusted response sentences with the generated image diagrams, arranging the response sentences followed by the image diagrams in a visually intuitive format.

[1544] Sending data:

[1545] The server's data transmission module sends the consolidated response to the user's device, and sends the response to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.).

[1546] View response:

[1547] The terminal displays the integrated response received from the server. The response text and image are provided to the user through the chat interface, allowing the user to intuitively understand the information.

[1548] Examples:

[1549] 1. A user submits the following question to the chatbot: "How does photosynthesis work?"

[1550] 2. The server identifies "photosynthesis" and "mechanism" using the question analysis module.

[1551] 3. The response generation module generates the response sentence, "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water."

[1552] 4. The image generation module generates a diagram showing the process of photosynthesis.

[1553] 5. The emotion engine analyzes the user's emotions and recognizes when the user is in an anxious state.

[1554] 6. The response adjuster adjusts the response to a more reassuring tone.

[1555] 7. The response integration module integrates the response sentence and the image diagram.

[1556] 8. The data transmission module sends the consolidated response to the user.

[1557] 9. The terminal displays the response text and image to the user, and the user understands the content.

[1558] This system allows users to not only receive appropriate responses to their questions, but also flexible responses that reflect their emotions at the time, allowing them to understand information more deeply and comfortably.

[1559] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1560] Step 1:

[1561] The user sends a question to the chatbot in text format. For example, they can enter a question like, "Please tell me how photosynthesis works." The input data is the text data that the user types into the input field. The sent text data is transferred to the server.

[1562] Step 2:

[1563] The server sends the question data received from the user to the question analysis module, which uses natural language processing (NLP) technology to analyze the question content. The input data is the question data in text format, and the output is the main keywords (e.g., "photosynthesis" and "mechanism").

[1564] Step 3:

[1565] The server's response generation module generates a response based on the keywords output from the question analysis module. It retrieves relevant information from an internal database and knowledge base to create a grammatically correct response. The input data are keywords, and the output is a response (e.g., "Photosynthesis is the process by which plants use light energy to produce oxygen and organic matter from carbon dioxide and water").

[1566] Step 4:

[1567] The server's image generation module creates an image related to the generated response text. It uses a diagram generation engine (specifically, D3.js or Chart.js) to generate a diagram that visually represents the "photosynthesis process." The input data is the response text, and the output is an image showing the photosynthesis process.

[1568] Step 5:

[1569] The emotion engine on the server analyzes the user's emotion. The input data are voice tone, facial expression (using facial recognition technology), and entered text. The emotion engine analyzes these data and identifies the emotion that the user is anxious. The output is the user's emotional state (e.g., anxious).

[1570] Step 6:

[1571] The server adjusts the tone and content of the response based on the analysis results of the emotion engine. The input data is the analyzed emotional state and the original response. Based on these input data, the server changes the response to include reassuring elements (e.g., "Don't worry, photosynthesis is a natural process for plants"). The output is the adjusted response.

[1572] Step 7:

[1573] The server's response synthesis module synthesizes the adjusted response sentence and the generated image diagram. The input data are the adjusted response sentence and the image diagram. By synthesizing these, a visually intuitive response is generated. The output is the synthesized response.

[1574] Step 8:

[1575] The server's data transmission module sends the consolidated response to the user's device. The input data is the consolidated response, which is sent to the user's chat interface using a communication protocol (e.g., HTTP, WebSocket, etc.). The output is the response data sent to the user's device.

[1576] Step 9:

[1577] The terminal displays the integrated response received from the server to the user. The input data is the response data sent from the server, and the output is the response text and image displayed on the chat interface. The user can intuitively understand the information through this display.

[1578] (Application example 2)

[1579] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1580] Conventional information provision systems only provide text-based responses to user input, which are often difficult to understand visually and do not allow for flexible responses that reflect the user's emotions.The present invention aims to provide more appropriate and reassuring information by providing visually easy-to-understand responses to questions entered by the user, recognizing the user's emotions, and adjusting the response content.

[1581] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a question receiving means for receiving a question from a user, a question analysis means for analyzing the received question and identifying related information, a response sentence generation means for generating a response sentence based on the information identified by the question analysis means, an image generation means for generating an image diagram related to the response sentence, an emotion analysis means for analyzing the user's emotion, a response adjustment means for adjusting the tone and content of the response sentence based on the emotion analyzed by the emotion analysis means, a response integration means for integrating the response sentence and the image diagram to generate an integrated response, a data transmission means for transmitting the integrated response to the user, and a display means for displaying the integrated response on the smart device. This allows the user to obtain visually easy-to-understand information and receive a flexible response suited to their individual emotional state.

[1582] The "question receiving means" is a device or function for receiving a question in text or voice format sent by a user.

[1583] The "question analysis means" is a device or function for analyzing a received question and identifying the content of the question and related main keywords.

[1584] The "response sentence generation means" is a device or function for generating a response sentence based on the analyzed keywords and providing the necessary information.

[1585] The "image diagram generating means" is a device or function for generating a visual image diagram related to a response sentence.

[1586] The "emotion analysis means" is a device or function for analyzing the user's emotions and adjusting the response content based on those emotions.

[1587] The "response adjustment means" is a device or function for adjusting the tone and content of a response sentence to match the user's emotions based on the emotion analysis results.

[1588] The "response integration means" is a device or function for integrating the response sentence and the generated image diagram to generate one integrated response.

[1589] A "data transmission means" is a communication device or function for transmitting an integrated response to a user.

[1590] The "display means" is a device or function for displaying the integrated response on the smart device.

[1591] A "smart device" is an electronic device used by a user, such as smart glasses, a smartphone, a head-mounted display, or a robot.

[1592] A system for implementing this invention provides visual and emotional responses to user questions when the user is engaged in activities such as shopping through a smart device (such as smart glasses, a smartphone, a head-mounted display, or a robot). An embodiment of this system is described in detail below.

[1593] Program Structure

[1594] The system includes a question receiving means, a question analyzing means, a response sentence generating means, an image diagram generating means, a sentiment analyzing means, a response adjusting means, a response integrating means, a data transmitting means, and a display means.

[1595] Program processing and use of hardware and software

[1596] Question receiving method

[1597] Users enter questions into smart glasses or smartphones by text or voice, using voice recognition software or a text input interface.

[1598] Question analysis means

[1599] The server analyzes the questions received from users using natural language processing (NLP) technology to identify key keywords, such as Google AI's natural language processing API.

[1600] Response sentence generation means

[1601] Based on the analyzed keywords, the server uses a generative AI model such as GPT-4 to generate a response, which is composed of information retrieved from a database or knowledge base.

[1602] Image diagram generation means

[1603] An image generation engine such as DALL-E is used to generate images related to the response, providing visual content appropriate for the response.

[1604] Emotion analysis means

[1605] To analyze user emotions, emotion analysis tools (e.g., Amazon Rekognition or Microsoft Azure's Emotion API) are used that process voice tone and facial expression data.

[1606] Response Adjustment Measures

[1607] Based on the results of the emotion analysis, the server adjusts the tone and content of the response to suit the user's emotions. If the user is anxious, the content will be adjusted to provide a sense of security.

[1608] Response Integration Measures

[1609] The response text and the generated image are integrated to generate a single integrated response, allowing the user to receive both text and visual information in a single interface.

[1610] Data transmission method

[1611] The server sends the integration response to the user's smart device using HTTP or WebSocket protocol.

[1612] Display means

[1613] The smart device displays the integrated response to the user.

[1614] Specific examples

[1615] When a user asks the smart glasses, "What are the characteristics of these shoes?", the following process is performed:

[1616] 1. Question analysis: The NLU module identifies the "shoes" and "features."

[1617] 2. Response generation: The generative AI model generates an answer such as, "These shoes are waterproof and feature a lightweight design."

[1618] 3. Image generation: DALL-E generates an image showing the waterproof function and lightweight design of the shoes.

[1619] 4. Sentiment Analysis: The emotion engine recognizes when the user is excited.

[1620] 5. Response adjustment: "These shoes are very popular. They fit you perfectly."

[1621] 6. Response Integration: Integrate the coordinated response statement with the illustration.

[1622] 7. Data transmission: sent to smart glasses and displayed to the user.

[1623] Prompt Sentence Examples

[1624] "Tell me the features of these shoes."

[1625] Analyzed keywords: shoes, features

[1626] Generated response: "These shoes are waterproof and feature a lightweight design."

[1627] Sentiment Analysis: Users are excited

[1628] Tailored response: "These shoes are very popular. They look great on you."

[1629] Final response: "These shoes are very popular. They're perfect for you. They're waterproof and lightweight. (Illustration: Waterproofing and Design)"

[1630] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1631] Step 1:

[1632] The user types or speaks a question into the smart device.

[1633] Input: Text or voice input from the user.

[1634] Output: Question text or audio file.

[1635] Specific operation: The user asks the smart glasses verbally, "Tell me the features of these shoes."

[1636] Step 2:

[1637] The device converts voice input into text data (voice recognition).

[1638] Input: User's voice input.

[1639] Output: The converted text data.

[1640] What it does: Speech recognition software converts the speech data into text, such as "What are the features of these shoes?"

[1641] Step 3:

[1642] The server receives the question and analyzes it using the question analysis means.

[1643] Input: Text data.

[1644] Output: Primary keywords (e.g., "shoes", "features").

[1645] How it works: The question receiving module receives the text and uses NLP technology (such as Google AI's natural language processing API) to identify key keywords.

[1646] Step 4:

[1647] The server generates a response sentence using a response sentence generation means.

[1648] Input: Primary keyword.

[1649] Output: A response statement (e.g., "These shoes are waterproof and feature a lightweight design").

[1650] Specific operation: The response generation module uses a generative AI model such as GPT-4 to obtain information from a database or knowledge base and generate a response.

[1651] Step 5:

[1652] The server generates an image diagram using an image diagram generating means.

[1653] Input: Keywords related to the response sentence.

[1654] Output: Image diagram (e.g., diagram showing waterproof function and design).

[1655] Specific operation: The image generation module uses an image generation engine such as DALL-E to generate a visual image related to the response sentence.

[1656] Step 6:

[1657] The server analyzes the user's emotions using an emotion analysis means.

[1658] Input: User's voice tone and facial expression data.

[1659] Output: User's emotional state (e.g., excited, anxious).

[1660] What it does: Sentiment analysis tools (such as Amazon Rekognition or Microsoft Azure's Emotion API) analyze the user's emotions and identify their emotional state.

[1661] Step 7:

[1662] The server adjusts the tone and content of the response using a response adjustment means.

[1663] Input: A response sentence and the user's emotional state.

[1664] Output: A tailored response (e.g., "These shoes are very popular. They fit you perfectly.").

[1665] Specific operation: Based on the results of emotion analysis, the tone and content of the response are changed to match the user's emotional state.

[1666] Step 8:

[1667] The server integrates the response sentence and the image diagram using a response integration means.

[1668] Input: Adjusted response sentence and image diagram.

[1669] Output: A consolidated response (e.g., a set of tailored sentences and image diagrams).

[1670] Specific action: Combine the response sentence and image into one integrated response.

[1671] Step 9:

[1672] The server transmits the integrated response to the user's smart device via the data transmission means.

[1673] Input: The consolidated response.

[1674] Output: The response sent to the user's smart device.

[1675] Specific operation: Send the integration response to the smart glasses using HTTP or WebSocket protocol.

[1676] Step 10:

[1677] The terminal displays the integrated response to the user.

[1678] Input: The consolidated response.

[1679] Output: The response and image displayed to the user.

[1680] Specific operation: The smart glasses display the tailored response sentence and image to the user.

[1681] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1682] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1683] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1684] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1685] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1686] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1687] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1688] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1689] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1690] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1691] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1692] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1693] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1694] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1695] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1696] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1697] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1698] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1699] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1700] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1701] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1702] The following is further disclosed regarding the above embodiment.

[1703] (Claim 1)

[1704] a question receiving means for receiving a question from a user;

[1705] a question analysis means for analyzing the received question and identifying relevant information;

[1706] a response sentence generation means for generating a response sentence based on the information identified by the question analysis means;

[1707] an image diagram generating means for generating an image diagram related to the response sentence;

[1708] a response integration means for integrating the response sentence and the image to generate an integrated response;

[1709] data transmission means for transmitting the integrated response to the user;

[1710] A system including:

[1711] (Claim 2)

[1712] 2. The system according to claim 1, wherein the response sentence generating means obtains information from a database or a knowledge base.

[1713] (Claim 3)

[1714] 2. The system of claim 1, wherein the image generation means uses an illustration generation engine.

[1715] "Example 1"

[1716] (Claim 1)

[1717] means for receiving a query from a user;

[1718] A means of analyzing incoming questions and identifying key keywords and question intent;

[1719] means for generating a response based on the identified keywords and phrases;

[1720] means for generating a related image diagram based on the response sentence;

[1721] a means for integrating the response sentence and the image diagram to generate an integrated response;

[1722] means for sending the integrated response to the user;

[1723] A system including:

[1724] (Claim 2)

[1725] 2. The system according to claim 1, wherein the response sentence generating means obtains information from a database or a knowledge base.

[1726] (Claim 3)

[1727] 2. The system of claim 1, wherein the image generation means uses an illustration generation engine.

[1728] "Application Example 1"

[1729] (Claim 1)

[1730] a question receiving means for receiving a question from a user;

[1731] a question analysis means for analyzing the received question and identifying relevant information;

[1732] a response sentence generation means for generating a response sentence based on the information identified by the question analysis means;

[1733] an image diagram generating means for generating an image diagram related to the response sentence;

[1734] a response integration means for integrating the response sentence and the image to generate an integrated response;

[1735] data transmission means for transmitting the integrated response to the user;

[1736] means for transmitting information to a display device that displays the response;

[1737] a means for inputting questions by text or voice using the device and displaying the results in real time;

[1738] A system including:

[1739] (Claim 2)

[1740] The system of claim 1, wherein the response sentence generation means obtains information from a database or knowledge base and generates a response sentence using a generative AI model.

[1741] (Claim 3)

[1742] 2. The system according to claim 1, wherein the image diagram generating means uses a diagram generation engine to generate an image diagram based on the prompt sentence.

[1743] "Example 2: Combining Emotion Engines"

[1744] (Claim 1)

[1745] a question receiving means for receiving a question from a user;

[1746] a question analysis means for analyzing the received question and identifying relevant information;

[1747] a response sentence generation means for generating a response sentence based on the information identified by the question analysis means;

[1748] an image diagram generating means for generating an image diagram related to the response sentence;

[1749] emotion analysis means for analyzing the emotions of a user;

[1750] a response adjustment means for adjusting the tone and content of a response sentence based on the analysis result of the emotion analysis means;

[1751] a response integration means for integrating the response sentence and the image to generate an integrated response;

[1752] data transmission means for transmitting the integrated response to the user;

[1753] A system including:

[1754] (Claim 2)

[1755] 2. The system according to claim 1, wherein the response sentence generating means obtains information from a database or a knowledge base.

[1756] (Claim 3)

[1757] 2. The system of claim 1, wherein the image generation means uses an illustration generation engine.

[1758] "Application example 2 when combining emotion engines"

[1759] (Claim 1)

[1760] a question receiving means for receiving a question from a user;

[1761] a question analysis means for analyzing the received question and identifying relevant information;

[1762] a response sentence generation means for generating a response sentence based on the information identified by the question analysis means;

[1763] an image diagram generating means for generating an image diagram related to the response sentence;

[1764] emotion analysis means for analyzing the emotions of a user;

[1765] a response adjustment means for adjusting the tone and content of a response sentence based on the emotion analyzed by the emotion analysis means;

[1766] a response integration means for integrating the response sentence and the image to generate an integrated response;

[1767] data transmission means for transmitting the integrated response to the user;

[1768] a display means for displaying the integrated response on the smart device;

[1769] A system including:

[1770] (Claim 2)

[1771] 2. The system according to claim 1, wherein the response sentence generating means obtains information from a database or a knowledge base.

[1772] (Claim 3)

[1773] 2. The system of claim 1, wherein the image generation means uses an illustration generation engine. [Explanation of symbols]

[1774] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a question receiving means for receiving a question from a user; a question analysis means for analyzing the received question and identifying relevant information; a response sentence generation means for generating a response sentence based on the information identified by the question analysis means; an image diagram generating means for generating an image diagram related to the response sentence; a response integration means for integrating the response sentence and the image to generate an integrated response; data transmission means for transmitting the integrated response to the user; A system including:

2. 2. The system according to claim 1, wherein the response sentence generating means obtains information from a database or a knowledge base.

3. 2. The system of claim 1, wherein the image generation means uses an illustration generation engine.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A