system
The system addresses the challenge of secure and efficient information retrieval within a company's internal network by using natural language processing and generative AI for document analysis and search query generation, ensuring rapid and intuitive access to necessary information.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
Existing systems struggle to efficiently utilize in-house information while ensuring security and providing quick, intuitive access to necessary information within a company's internal network, particularly in question-and-answer systems using AI.
A system equipped with document analysis and search query generation using natural language processing technology within the company's internal network, enabling secure and rapid information provision through generative artificial intelligence and access right filtering.
Enables efficient searching and secure retrieval of relevant information within the company's internal network, ensuring quick and intuitive access to necessary data while maintaining security.
Smart Images

Figure 2026070191000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When introducing a question-and-answer system that utilizes in-house generated AI, there is a problem that it is required to efficiently utilize in-house information while ensuring security without going through the cloud. Furthermore, providing an environment where users can quickly and intuitively obtain necessary information from in-house materials is also one of the problems.
Means for Solving the Problems
[0005] This invention solves the aforementioned problems by providing a system equipped with a document analysis and search query generation means using natural language processing technology within the company's internal network. This enables efficient searching of materials accumulated within the company, generates answers in natural language using generative artificial intelligence, and presents them to the user, thereby enabling secure and rapid information provision. Furthermore, security is ensured by providing a function to appropriately filter search results according to access rights.
[0006] "Natural language processing" is a technology that enables computers to understand, analyze, and generate human language (natural language).
[0007] "Document analysis" is the process of extracting text data from a document and converting its meaning and structure into a format that a computer can understand.
[0008] "Searchable format" refers to a structured data arrangement designed to efficiently find specific information.
[0009] A "search query" is a set of terms or phrases used to find specific information in a database.
[0010] "Generative artificial intelligence" is an AI technology that has the ability to generate new expressions and answers in natural language based on diverse data and information.
[0011] "Indexing" is the process of organizing the content of documents and data to make them easily accessible and searchable, thereby facilitating the retrieval of related data.
[0012] A "database" is a collection of information that allows for the systematic storage, retrieval, and management of related data.
[0013] "User interface" is a general term for the screens and operating methods that users use when interacting with a system. [Brief explanation of the drawing]
[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0015] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention relates to a question answering system that operates within an internal network environment. Users can input questions in natural language using a terminal. The terminal receives these questions and sends them to a server. At this stage, the terminal's program analyzes the context of the question, identifies important keywords and phrases, and generates a search query.
[0036] The server receives the generated search query and searches a database that has been pre-indexed using natural language processing. This database contains various internal company documents, ensuring that important information is efficiently managed. Search results are ranked based on relevance and filtered as needed, taking access rights into consideration.
[0037] After relevant information is retrieved, the server uses a generative artificial intelligence algorithm to construct a natural language response. This process utilizes text extracted from search results to generate a response that directly addresses the user's question. The generated response is further formatted and presented in a polished style before being presented to the user.
[0038] Users can view server-generated answers through their devices. The user interface presents information in an easy-to-read layout and provides supplementary information and related links as needed. Users can also ask additional questions and provide feedback to the system.
[0039] For example, if a user asks, "Tell me the summary of the latest technical report," the system extracts keywords such as "technical report" and "summary," and searches for related documents based on them. Based on the retrieved information, the server generates a specific answer such as, "The latest technical report describes current technical challenges and their solutions." In this way, this invention enhances the efficiency of information utilization within the company and supports business operations while maintaining security.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The user inputs a question in natural language through the device. The device receives the user's input in real time and prepares to convert it into the appropriate format.
[0043] Step 2:
[0044] The terminal analyzes the entered question and extracts important keywords and phrases. Using natural language processing technology, it interprets the intent of the question and generates a search query.
[0045] Step 3:
[0046] The terminal sends the generated search query to the server. The server receives this query and accesses a pre-indexed database.
[0047] Step 4:
[0048] The server searches the database based on the query it receives. The search is performed based on keyword frequency and document relevance scores, and the most relevant information is collected.
[0049] Step 5:
[0050] The server retrieves search results and filters them according to access rights. Security is ensured so that only appropriate information proceeds to the next step.
[0051] Step 6:
[0052] The server uses generative artificial intelligence to generate appropriate response sentences based on filtered information. In this step, data obtained from multiple documents is integrated and the response is constructed using natural language expressions.
[0053] Step 7:
[0054] The server sends the generated response to the terminal. The terminal displays the received response on its user interface, presenting it in a format that is easy for the user to understand.
[0055] Step 8:
[0056] The user can review the displayed answers and have options to ask additional questions or provide feedback as needed. The device then prepares to process these actions as input for the next cycle.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] Conventional information retrieval systems have struggled to provide users with the information they need quickly and accurately. In particular, there is a need to extract highly relevant information from diverse document sets and provide clear, understandable responses in natural language. Furthermore, the provision of appropriate information according to user access rights is essential, making efficient searching and result processing crucial challenges.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes information processing means for analyzing a set of documents and converting them into a searchable format; processing means for receiving a question in natural language from a user, analyzing it, and generating a search instruction; and search means for searching the analyzed document data set using the generated search instruction and extracting highly relevant information. This makes it possible to quickly and accurately extract highly relevant information from a diverse set of documents, provide natural language responses to user questions, and perform information filtering according to access rights as needed.
[0062] A "document collection" refers to various forms of documents and data stored within a company or a specific environment.
[0063] "Information processing means" refers to hardware or software functions that analyze and transform data and perform processing according to requirements.
[0064] "User" refers to a person or entity that accesses the system to search for information or ask questions.
[0065] "Natural language" refers to the forms of words and sentences that humans use on a daily basis, and is the subject of analysis by computers.
[0066] A "search instruction" refers to a query or command generated to satisfy a user's question or request.
[0067] "Search method" refers to a function for finding information from databases or document sets based on specified conditions.
[0068] "Highly relevant information" refers to information that is deemed most appropriate for providing direct or indirect answers to users' questions or requests.
[0069] A "computer-generated natural language model" refers to algorithms and programs that use artificial intelligence technology to generate natural-sounding text.
[0070] "Response" refers to the answer or information provided in response to a user's question.
[0071] "Presentation means" refers to a function for displaying or providing information or responses generated by the system to the user.
[0072] This invention relates to an information processing system aimed at the efficient analysis and provision of information from a collection of documents. The system mainly includes users, terminals, and a server.
[0073] Users can input questions in natural language using a terminal. The terminal uses natural language processing libraries such as "NLTK" and "spaCy" to understand the context of the input question and identify important keywords. Based on this, it creates a search command and sends it to the server.
[0074] The server is responsible for searching for relevant information from the data set based on the search instructions it receives. The data set is indexed by advanced search engines such as "ElasticSearch®," enabling fast and accurate information retrieval. The search results are ranked based on their relevance.
[0075] Subsequently, the server uses generative AI models such as "GPT" and "Transformers" to generate a response based on the information obtained from the search results. This response directly answers the question and is expressed in natural and easy-to-understand language.
[0076] Ultimately, the user receives a response generated through the device's user interface. The screen provides links to relevant information and supplementary explanations to aid in understanding the information. Users can also ask additional questions and submit feedback to the system.
[0077] For example, by using a prompt such as "Tell me the summary of the latest technical report," the system provides a specific response such as "The latest technical report describes current technical challenges and their solutions." In this way, the invention aims to provide users with the information they need quickly and accurately, thereby improving the efficiency of information utilization.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The user inputs a question into the terminal using natural language. The terminal analyzes the input question using natural language processing libraries such as "NLTK" or "spaCy." In the process, it extracts important keywords and phrases from the question. Based on the input question text, it performs data analysis and outputs a set of keywords. Specifically, it performs morphological analysis to understand the structure of the question text.
[0081] Step 2:
[0082] The terminal generates a search instruction using the extracted keyword set. This search instruction is then formatted as a query to be sent to the server. Based on the keyword set as input, it converts it into an appropriate search syntax to output a query statement that the server can interpret. Specifically, it constructs the query by concatenating keywords and adding logical operators.
[0083] Step 3:
[0084] The server uses search commands received from the terminal to search the data set. The server uses a search engine such as "Elasticsearch" to quickly search the indexed database and retrieve relevant information. Based on the input query, it performs a data search within the database and outputs a list of relevant information. Specifically, it performs a full-text search, calculates a relevance score, and ranks the information.
[0085] Step 4:
[0086] The server uses a generative AI model to generate natural language responses based on the list of search results. Here, models such as "GPT" and "Transformers" are used to create appropriate answers to user questions. It analyzes the input list of information, selects relevant information, and outputs natural language sentences. Specifically, it generates natural responses that are relevant to the context.
[0087] Step 5:
[0088] The server sends the generated response to the terminal, which the user visually confirms. The terminal's user interface displays the received response in an easy-to-read format, providing links to relevant information and supplementary explanations as needed. Based on the response text as input, it outputs information formatted for screen display. Specifically, it applies layout and style using HTML and CSS.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] To improve work efficiency within a factory, it is necessary to create an environment where workers can quickly and accurately obtain the information they need. However, information retrieval on-site is not easy, and providing timely information is difficult. Furthermore, for direct interaction with machinery during work, there is a need for a means of easily accessing information without requiring user intervention.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for analyzing information using natural language processing within the company's internal communication network and converting it into a searchable format; means for receiving inquiries in natural language from users, analyzing them, and generating search queries; and means for using the generated search queries to search the analyzed information repository and extract highly relevant knowledge. This makes it possible for workers to easily obtain information by voice and improve work efficiency.
[0094] An "internal communication network" is a network used for sharing information and communicating within a company or organization.
[0095] "Natural language processing" is a technology that uses computers to understand, analyze, and generate human language.
[0096] A "searchable format" is a data format that is organized and structured in a way that allows information to be efficiently searched.
[0097] "User" refers to a person or user who uses the system.
[0098] A "search query" is a set of search conditions and keywords entered to retrieve information.
[0099] An "information repository" is a database or storage device where data and information are stored and managed.
[0100] "Highly relevant knowledge" refers to information that is appropriate and important in response to a user's question during a search.
[0101] "Generative intelligence" is an artificial intelligence technology that generates new information or answers based on given data or queries.
[0102] "Speech recognition" is a technology that analyzes audio data and converts its content into text data.
[0103] "Audio output" is a technology for playing text data as audio.
[0104] The system designed to realize this application provides the ability to quickly obtain the information workers need in a factory work environment. The system accepts questions from workers in natural language using a smartphone or a voice-input capable device.
[0105] The smartphone converts the received audio into text using the Google® Speech-to-Text API. Then, it analyzes the text using spaCy as a natural language processing engine, identifying important keywords and generating a search query. This query is sent to an API server using Flask. The server uses Elasticsearch to search the company's indexed information repository and extract highly relevant knowledge.
[0106] The server uses generative AI models such as OpenAI's GPT-3 to generate natural language responses from extracted knowledge. After formatting, these responses are output as speech using the Google Text-to-Speech API and provided to workers via smartphones.
[0107] For example, if a worker asks, "Please tell me the assembly process for the new product A and B," the system will treat "product A and B" and "assembly process" as important keywords and search for and provide relevant information. An example of a prompt in this case would be: "Question: 'Please tell me the assembly process for the new product A and B.' Search keywords: 'product A and B', 'assembly process'."
[0108] Thus, this system aims to improve work efficiency by allowing factory workers to easily obtain information through voice input.
[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0110] Step 1:
[0111] The device accepts voice input from the user. The input voice is converted to text using the Google Speech-to-Text API. At this stage, the input is voice data, and the output is the corresponding text data.
[0112] Step 2:
[0113] The terminal analyzes the generated text using spaCy, a natural language processing engine. Important keywords and phrases are extracted from the text, and search queries are generated. The input here is text data, and the output is a data structure containing the search queries.
[0114] Step 3:
[0115] The server uses Elasticsearch to search the company's indexed information repository using the search query received from the terminal. It then extracts highly relevant knowledge. In this step, the input includes the search query, and the output is the retrieved relevant information.
[0116] Step 4:
[0117] The server generates a natural language response based on the extracted relevant information, using OpenAI's GPT-3, etc. A generative AI model processes this. The input is relevant information, and the output is the natural language response text.
[0118] Step 5:
[0119] The server converts the generated response text into speech using the Google Text-to-Speech API and sends it to the device. In this case, the input is natural language response text, and the output is audio data.
[0120] Step 6:
[0121] The terminal plays audio data and provides it to the user. The input is audio data received from the server, and the output is the played audio.
[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0123] This invention combines an emotion engine with a system that provides appropriate answers to natural language questions from users within an internal network. Users can input questions via a terminal. The terminal receives the input question and uses the emotion engine to analyze metadata such as the user's word choice and input speed to identify the user's emotional state.
[0124] The emotion engine uses natural language processing techniques and machine learning algorithms to analyze the emotions expressed in the user's text. The results of this emotion analysis are sent to the server, which adjusts the wording of the response according to the emotional state. For example, if the system detects that the user is confused, it can generate a response that includes a detailed and easy-to-understand explanation.
[0125] The server analyzes the question content, taking this emotional information into account, generates a search query, and searches the company's database. After retrieving the search results, it uses generative artificial intelligence to generate the optimal answer. In this answer generation process, the server selectively emphasizes the content of the answer and adjusts the tone appropriately according to the user's emotional state.
[0126] The generated answers are presented to the user through a user interface. The device displays the answers in a user-friendly layout, making them easy for the user to understand. Users can also ask additional questions on the screen and submit feedback on the system's answers. This feedback is recorded by the system and used to improve the accuracy of future responses.
[0127] For example, if a user asks the terminal, "Why isn't this software working properly?", and the emotion engine detects frustration from the user's writing style and typing speed, the server can generate a detailed response in a gentle tone, such as, "We apologize for the inconvenience. The usual reason this software doesn't work properly is XX." This invention makes it possible to empathize with the user's emotions and provide a better user experience.
[0128] The following describes the processing flow.
[0129] Step 1:
[0130] The user enters a question in natural language using a device. The device receives this input and temporarily stores it as text data.
[0131] Step 2:
[0132] The device sends the entered question to the emotion engine, which analyzes the user's emotional state based on parameters such as word choice, sentence structure, and input speed.
[0133] Step 3:
[0134] The emotion engine identifies emotions from user input. For example, it analyzes keywords and evaluates the tone of sentences to classify emotions into categories such as "joy," "anxiety," and "frustration."
[0135] Step 4:
[0136] The terminal receives the analysis results and transmits the user's emotional state to the server. The server receives the user's emotional information along with the question data.
[0137] Step 5:
[0138] The server analyzes the question and extracts keywords necessary to understand the intent of the question. Then, it generates an appropriate search query based on those keywords.
[0139] Step 6:
[0140] The system uses server-generated search queries to search indexed internal databases and gather highly relevant information. Simultaneously, it filters search results, taking access rights into consideration as needed.
[0141] Step 7:
[0142] Based on the information collected by the server, generative artificial intelligence is used to generate the optimal response. During this process, the tone and details of the response are adjusted according to the user's emotional state. For example, if frustration is detected, polite and empathetic language will be used.
[0143] Step 8:
[0144] The generated response is sent from the server to the terminal, which then displays it on the user interface. The user can receive the presented response in an easy-to-understand format.
[0145] Step 9:
[0146] Users can request further details as needed or provide feedback on the system's response. The terminal receives these requests and feedback and prepares to send them to the server for use in future processing.
[0147] (Example 2)
[0148] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0149] In modern information systems, simply providing information mechanically without considering the emotions users feel, such as frustration or confusion, does not allow for truly user-centric service. Furthermore, simple information retrieval systems have difficulty accurately grasping the user's search intent, and therefore cannot be expected to improve the user experience.
[0150] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0151] In this invention, the server includes means for receiving questions in natural language from the user, collecting metadata and performing sentiment analysis, means for adjusting the tone and content of the response based on the analyzed sentiment state, and means for searching for information using the generated search query and generating a response in natural language using generative artificial intelligence. This makes it possible to provide information in a way that takes the user's emotions into consideration and to improve the user experience.
[0152] A "user" refers to an individual or group that uses the system to obtain information or enter questions.
[0153] "Natural language" refers to the language that humans use on a daily basis, and is treated as text data that computers can understand and analyze.
[0154] "Metadata" refers to additional information accompanying user input, including data other than the text itself, such as input speed and word choice.
[0155] "Sentiment analysis" refers to the process of identifying a user's emotional state from their written text using natural language processing technology and machine learning algorithms.
[0156] "Generative artificial intelligence" refers to artificial intelligence technology that can generate natural language text similar to that produced by humans, based on a large amount of data.
[0157] A "search query" refers to a string or syntax used internally within a system to specify search conditions for retrieving information.
[0158] "Tone" refers to the overall tone and feel of a text, and is a concept that means adjusting the expression to match the user's emotions.
[0159] This system aims to provide appropriate information in response to users' natural language questions via the company's internal network. The system primarily consists of terminals, servers, an emotion analysis engine, and a generative AI model.
[0160] The user uses a device to input a question in natural language. The device collects metadata such as the input speed and word choice, in addition to the entered question. The sentiment engine uses this metadata, employing natural language processing techniques and machine learning algorithms, to analyze the user's emotional state. For this sentiment analysis, one could utilize, for example, the Python library NLTK (Natural Language Toolkit) or a dedicated sentiment analysis API.
[0161] The analysis results are sent to a server, which adjusts the tone and content of the answers to the questions based on the user's emotional state. The server analyzes the questions and generates appropriate search queries to search the internal data store. This uses common information technology techniques, such as automatically generating search queries and using SQL queries to retrieve information.
[0162] After acquiring the information, the server generates the optimal response using a generative AI model. For example, a GPT (Generative Pre-trained Transformer) model can be used as the generative AI model. Here, prompts are used to instruct the AI model. A concrete example of a prompt could be, "Generate an answer that is easy for the user to understand when they are confused."
[0163] The responses generated by the generative AI model are sent to the device and presented in a visually organized format through the user interface. This allows the user to easily understand the responses. Users can also ask additional questions and submit feedback on the provided answers. This feedback is recorded by the server and used to improve the accuracy of future responses.
[0164] This system is expected to improve the overall user experience by allowing users to efficiently obtain information in an emotionally sensitive manner.
[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0166] Step 1:
[0167] The user inputs a question in natural language through the device. The device receives the user's input as digital data. Specifically, the user inputs a question via the keyboard interface, such as "Why isn't this function working?". The device simultaneously collects metadata such as input speed and writing style.
[0168] Step 2:
[0169] The device sends the collected question data and metadata to the sentiment analysis engine. The sentiment engine uses natural language processing and machine learning algorithms to analyze the user's emotional state from the text. In this process, it performs data calculations on the word choices and punctuation frequency of the input text and outputs sentiment labels such as "confused" or "irritated."
[0170] Step 3:
[0171] The terminal sends the sentiment analysis results to the server, which receives them. The server analyzes the user's question based on the received sentiment state and input text. Based on the analysis, it determines what information is needed and generates a search query. Specifically, it converts the natural language question into an SQL query or a query using search operators.
[0172] Step 4:
[0173] The server uses the generated query to search the company's data store. At this stage, the information obtained from the query is efficiently extracted and the necessary data is aggregated. Once the search results are obtained, the dataset is prepared as input for the AI model.
[0174] Step 5:
[0175] The server inputs the acquired data into the generative AI model and applies prompt sentences according to the user's emotional state. The generative AI model generates the optimal response based on prompts such as, "Generate an easy-to-understand response when the user is confused." Based on the prompts used and the analyzed data, the AI generates a response in natural language and outputs it as text.
[0176] Step 6:
[0177] The server sends the generated response to the terminal, which displays it through a user interface. The terminal displays the response in a format that is appropriately laid out for readability, making it easy for the user to understand. The user can then enter additional questions or provide feedback. This feedback information is sent back to the server and used for subsequent processing.
[0178] (Application Example 2)
[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0180] Conventional question-answering systems have a problem in that they do not adjust their responses according to the user's emotional state, making it difficult to obtain answers that satisfy the user. Furthermore, they lack means to quickly and accurately resolve the user's dissatisfaction and confusion, thus failing to improve the user experience.
[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0182] In this invention, the server includes means for analyzing documents using language processing technology within the company's internal communication network and converting them into a searchable format; means for receiving natural language questions from users, analyzing them, and generating search instructions; and means for generating natural language answers using generative artificial intelligence based on the extracted information. This makes it possible to identify the user's emotional state and adjust the optimal answer based on that state.
[0183] An "internal communication network" is a communication network system used within a company, which allows for the sending, receiving, and sharing of information.
[0184] "Language processing technology" refers to technologies that enable computers to analyze, understand, and generate natural language.
[0185] A "document storage device" is a data storage system in which analyzed document information is stored and used for later retrieval.
[0186] A "search instruction" is a query generated based on a user's question, and its purpose is to efficiently retrieve information.
[0187] "Generative artificial intelligence" is a system that creates and provides text in natural language based on human instructions.
[0188] "Emotional state" refers to the state of mind inferred from the user's writing and actions, and is reflected in the tone and content of their responses.
[0189] A "user interface" refers to the screens and means of operation that allow a user to interact directly with a system.
[0190] The system for realizing this invention begins by receiving natural language questions from the user's terminal via the company's internal communication network. The terminal collects the grammar and metadata of the text entered by the user and sends it to an emotion engine for sentiment analysis. This emotion engine uses language processing techniques such as Python and NLTK, and machine learning frameworks such as TENSORFLOW® and PyTorch to identify the user's emotional state. The analysis results are sent to a server, where a server with generative artificial intelligence installed generates search instructions according to the emotional state and extracts relevant information by referring to the accumulated document database.
[0191] Based on this information, the server uses generative artificial intelligence to generate natural language responses, adjusting them with appropriate tone and expression based on the emotional state. This process produces the final response displayed in the user interface. The user interface, in particular, features a user-friendly design utilizing React Native and other technologies, making it easy for users to ask additional questions and provide feedback. Furthermore, the response adjustment data is accumulated as insights necessary for future improvements, thereby enhancing user satisfaction.
[0192] As a concrete example, consider a scenario in customer support on an e-commerce site where a customer asks, "I feel like these shoes are the wrong size." If the emotion engine detects a "dissatisfied" emotional state, the server will generate a reassuring response such as, "We apologize for the inconvenience caused by the incorrect sizing. We will guide you through the return or exchange process." An example of a prompt might be, "Create a sample of a gentle response when a user inquires about product dissatisfaction or problems." In this way, the invention can provide a new means of communicating more effectively with users.
[0193] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0194] Step 1:
[0195] The device receives user questions in natural language and simultaneously collects metadata such as input speed and word choice. This data is prepared for analysis by the sentiment engine. The input data consists of the user's text information and metadata, while the output is the data sent to the sentiment engine.
[0196] Step 2:
[0197] The terminal sends the collected user text information and metadata to the server. The server uses an emotion engine to analyze the received data and identify the user's emotional state. In this step, natural language processing techniques and machine learning models are applied, and the identified emotional state is output.
[0198] Step 3:
[0199] The server generates appropriate search instructions based on the user's emotional state and question content. These search instructions are created as queries to retrieve highly relevant information from the database. The input is the question content and emotional state, and the output is the search instruction query.
[0200] Step 4:
[0201] The server uses the generated search instructions to search the document storage device and extract relevant information. A database query is executed, and the extracted relevant information is output.
[0202] Step 5:
[0203] Based on the extracted information, the server uses generative artificial intelligence to generate a response in natural language. The generated response is then adjusted to an appropriate tone based on the emotional state. The input is the extracted information and emotional state, and the output is the adjusted natural language response.
[0204] Step 6:
[0205] The generated, adjusted responses are presented to the user through the device's user interface. The user can review the responses on screen and enter additional questions or feedback, leading to further improvements in the user experience. The input is the adjusted responses, and the output is the presentation to the user and the collection of user feedback.
[0206] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0207] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0208] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0209] [Second Embodiment]
[0210] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0211] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0212] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0213] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0214] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0215] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0216] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0217] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0218] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0219] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0220] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0221] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0222] This invention relates to a question answering system that operates within an internal network environment. Users can input questions in natural language using a terminal. The terminal receives these questions and sends them to a server. At this stage, the terminal's program analyzes the context of the question, identifies important keywords and phrases, and generates a search query.
[0223] The server receives the generated search query and searches a database that has been pre-indexed using natural language processing. This database contains various internal company documents, ensuring that important information is efficiently managed. Search results are ranked based on relevance and filtered as needed, taking access rights into consideration.
[0224] After relevant information is retrieved, the server uses a generative artificial intelligence algorithm to construct a natural language response. This process utilizes text extracted from search results to generate a response that directly addresses the user's question. The generated response is further formatted and presented in a polished style before being presented to the user.
[0225] Users can view server-generated answers through their devices. The user interface presents information in an easy-to-read layout and provides supplementary information and related links as needed. Users can also ask additional questions and provide feedback to the system.
[0226] For example, if a user asks, "Tell me the summary of the latest technical report," the system extracts keywords such as "technical report" and "summary," and searches for related documents based on them. Based on the retrieved information, the server generates a specific answer such as, "The latest technical report describes current technical challenges and their solutions." In this way, this invention enhances the efficiency of information utilization within the company and supports business operations while maintaining security.
[0227] The following describes the processing flow.
[0228] Step 1:
[0229] The user inputs a question in natural language through the device. The device receives the user's input in real time and prepares to convert it into the appropriate format.
[0230] Step 2:
[0231] The terminal analyzes the entered question and extracts important keywords and phrases. Using natural language processing technology, it interprets the intent of the question and generates a search query.
[0232] Step 3:
[0233] The terminal sends the generated search query to the server. The server receives this query and accesses a pre-indexed database.
[0234] Step 4:
[0235] The server searches the database based on the query it receives. The search is performed based on keyword frequency and document relevance scores, and the most relevant information is collected.
[0236] Step 5:
[0237] The server retrieves search results and filters them according to access rights. Security is ensured so that only appropriate information proceeds to the next step.
[0238] Step 6:
[0239] The server uses generative artificial intelligence to generate appropriate response sentences based on filtered information. In this step, data obtained from multiple documents is integrated and the response is constructed using natural language expressions.
[0240] Step 7:
[0241] The server sends the generated response to the terminal. The terminal displays the received response on its user interface, presenting it in a format that is easy for the user to understand.
[0242] Step 8:
[0243] The user can review the displayed answers and have options to ask additional questions or provide feedback as needed. The device then prepares to process these actions as input for the next cycle.
[0244] (Example 1)
[0245] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0246] Conventional information retrieval systems have struggled to provide users with the information they need quickly and accurately. In particular, there is a need to extract highly relevant information from diverse document sets and provide clear, understandable responses in natural language. Furthermore, the provision of appropriate information according to user access rights is essential, making efficient searching and result processing crucial challenges.
[0247] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0248] In this invention, the server includes information processing means for analyzing a set of documents and converting them into a searchable format; processing means for receiving a question in natural language from a user, analyzing it, and generating a search instruction; and search means for searching the analyzed document data set using the generated search instruction and extracting highly relevant information. This makes it possible to quickly and accurately extract highly relevant information from a diverse set of documents, provide natural language responses to user questions, and perform information filtering according to access rights as needed.
[0249] A "document collection" refers to various forms of documents and data stored within a company or a specific environment.
[0250] "Information processing means" refers to hardware or software functions that analyze and transform data and perform processing according to requirements.
[0251] "User" refers to a person or entity that accesses the system to search for information or ask questions.
[0252] "Natural language" refers to the forms of words and sentences that humans use on a daily basis, and is the subject of analysis by computers.
[0253] A "search instruction" refers to a query or command generated to satisfy a user's question or request.
[0254] "Search method" refers to a function for finding information from databases or document sets based on specified conditions.
[0255] "Highly relevant information" refers to information that is deemed most appropriate to provide direct or indirect answers to users' questions or requests.
[0256] A "computer-generated natural language model" refers to algorithms and programs that use artificial intelligence technology to generate natural-sounding text.
[0257] "Response" refers to the answer or information provided in response to a user's question.
[0258] "Presentation means" refers to a function for displaying or providing information or responses generated by the system to the user.
[0259] This invention relates to an information processing system aimed at the efficient analysis and provision of information from a collection of documents. The system mainly includes users, terminals, and a server.
[0260] Users can input questions in natural language using a terminal. The terminal uses natural language processing libraries such as "NLTK" and "spaCy" to understand the context of the input question and identify important keywords. Based on this, it creates a search command and sends it to the server.
[0261] The server is responsible for searching for relevant information from the data set based on the search instructions it receives. The data set is indexed by advanced search engines such as Elasticsearch, enabling fast and accurate information retrieval. The search results are ranked based on their relevance.
[0262] Subsequently, the server uses generative AI models such as "GPT" and "Transformers" to generate a response based on the information obtained from the search results. This response directly answers the question and is expressed in natural and easy-to-understand language.
[0263] Ultimately, the user receives a response generated through the device's user interface. The screen provides links to relevant information and supplementary explanations to aid in understanding the information. Users can also ask additional questions and submit feedback to the system.
[0264] For example, by using a prompt such as "Tell me the summary of the latest technical report," the system provides a specific response such as "The latest technical report describes current technical challenges and their solutions." In this way, the invention aims to provide users with the information they need quickly and accurately, thereby improving the efficiency of information utilization.
[0265] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0266] Step 1:
[0267] The user inputs a question into the terminal using natural language. The terminal analyzes the input question using natural language processing libraries such as "NLTK" or "spaCy." In the process, it extracts important keywords and phrases from the question. Based on the input question text, it performs data analysis and outputs a set of keywords. Specifically, it performs morphological analysis to understand the structure of the question text.
[0268] Step 2:
[0269] The terminal generates a search instruction using the extracted keyword set. This search instruction is then formatted as a query to be sent to the server. Based on the keyword set as input, it converts it into an appropriate search syntax to output a query statement that the server can interpret. Specifically, it constructs the query by concatenating keywords and adding logical operators.
[0270] Step 3:
[0271] The server uses search commands received from the terminal to search the data set. The server uses a search engine such as "Elasticsearch" to quickly search the indexed database and retrieve relevant information. Based on the input query, it performs a data search within the database and outputs a list of relevant information. Specifically, it performs a full-text search, calculates a relevance score, and ranks the information.
[0272] Step 4:
[0273] The server uses a generative AI model to generate natural language responses based on the list of search results. Here, models such as "GPT" and "Transformers" are used to create appropriate answers to user questions. It analyzes the input list of information, selects relevant information, and outputs natural language sentences. Specifically, it generates natural responses that are relevant to the context.
[0274] Step 5:
[0275] The server sends the generated response to the terminal, which the user visually confirms. The terminal's user interface displays the received response in an easy-to-read format, providing links to relevant information and supplementary explanations as needed. Based on the response text as input, it outputs information formatted for screen display. Specifically, it applies layout and style using HTML and CSS.
[0276] (Application Example 1)
[0277] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".
[0278] In order to improve the work efficiency in the factory, it is necessary to create an environment where workers can quickly and accurately obtain the information they need. However, information retrieval at the site is not easy, and there is a problem that timely information provision is difficult. In addition, in order to directly interact with the machine during work, a means for easily accessing information without using the user's hands is required.
[0279] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0280] In this invention, the server includes means for analyzing information using natural language processing within the company communication network and converting it into a searchable format, means for receiving a natural language query from a user, analyzing it, and generating a search query, and means for searching the analyzed information database using the generated search query and extracting highly relevant knowledge. As a result, it becomes possible for workers to easily obtain information by voice and improve work efficiency.
[0281] The "company communication network" is a network for sharing information and communicating within a company or organization.
[0282] "Natural language processing" is a technology for a computer to understand, analyze, and generate human language.
[0283] The "searchable format" is a data format in which information is organized and structured so that it can be efficiently searched.
[0284] The "user" refers to a human or a user who uses the system.
[0285] A "search query" is a set of search conditions and keywords entered to search for information.
[0286] An "information repository" is a database or storage where data and information are stored and managed.
[0287] "Highly relevant knowledge" is appropriate and important information for the user's question in a search.
[0288] "Generative intelligence" is an artificial intelligence technology that generates new information and answers based on given data and queries.
[0289] "Speech recognition" is a technology that analyzes speech data and converts its content into text data.
[0290] "Voice output" is a technology for playing text data as speech.
[0291] The system for realizing this application example provides a function that enables workers to quickly obtain the information they need in the working environment of a factory. The system receives questions in natural language from workers using a smartphone or a device capable of voice input.
[0292] The smartphone converts the received voice into text using the Google Speech-to-Text API. Then, it analyzes the text using spaCy as a natural language processing engine, identifies important keywords, and generates a search query. This query is sent to an API server using Flask. The server searches the in-house indexed information repository using Elasticsearch and extracts highly relevant knowledge.
[0293] The server uses generative AI models such as OpenAI's GPT-3 to generate natural language responses from extracted knowledge. After formatting, these responses are output as speech using the Google Text-to-Speech API and provided to workers via smartphones.
[0294] For example, if a worker asks, "Please tell me the assembly process for the new product A and B," the system will treat "product A and B" and "assembly process" as important keywords and search for and provide relevant information. An example of a prompt in this case would be: "Question: 'Please tell me the assembly process for the new product A and B.' Search keywords: 'product A and B', 'assembly process'."
[0295] Thus, this system aims to improve work efficiency by allowing factory workers to easily obtain information through voice input.
[0296] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0297] Step 1:
[0298] The device accepts voice input from the user. The input voice is converted to text using the Google Speech-to-Text API. At this stage, the input is voice data, and the output is the corresponding text data.
[0299] Step 2:
[0300] The terminal analyzes the generated text using spaCy, a natural language processing engine. Important keywords and phrases are extracted from the text, and search queries are generated. The input here is text data, and the output is a data structure containing the search queries.
[0301] Step 3:
[0302] The server uses the search query received from the terminal to search the in-house indexed database by utilizing Elasticsearch and extracts highly relevant knowledge. In this step, the input is the search query, and the output is the retrieved relevant information.
[0303] Step 4:
[0304] The server uses GPT-3 of OpenAI etc. to generate a response in natural language based on the extracted relevant information. The generative AI model processes this. The input is the relevant information, and the output is the response text in natural language.
[0305] Step 5:
[0306] The server uses the Google Text-to-Speech API to convert the generated response text into voice and sends it to the terminal. Here, the input is the response text in natural language, and the output is the voice data.
[0307] Step 6:
[0308] The terminal plays the voice data and provides it to the user. The input is the voice data received from the server, and the output is the played voice.
[0309] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0310] This invention combines an emotion engine with a system that provides an appropriate answer to a question in natural language from a user in an in-house network. The user can input a question through the terminal. The terminal receives the input question, analyzes it with the emotion engine based on metadata such as the user's word usage and input speed, and identifies the user's emotional state.
[0311] The emotion engine uses natural language processing techniques and machine learning algorithms to analyze the emotions expressed in the user's text. The results of this emotion analysis are sent to the server, which adjusts the wording of the response according to the emotional state. For example, if the system detects that the user is confused, it can generate a response that includes a detailed and easy-to-understand explanation.
[0312] The server analyzes the question content, taking this emotional information into account, generates a search query, and searches the company's database. After retrieving the search results, it uses generative artificial intelligence to generate the optimal answer. In this answer generation process, the server selectively emphasizes the content of the answer and adjusts the tone appropriately according to the user's emotional state.
[0313] The generated answers are presented to the user through a user interface. The device displays the answers in a user-friendly layout, making them easy for the user to understand. Users can also ask additional questions on the screen and submit feedback on the system's answers. This feedback is recorded by the system and used to improve the accuracy of future responses.
[0314] For example, if a user asks the terminal, "Why isn't this software working properly?", and the emotion engine detects frustration from the user's writing style and typing speed, the server can generate a detailed response in a gentle tone, such as, "We apologize for the inconvenience. The usual reason this software doesn't work properly is XX." This invention makes it possible to empathize with the user's emotions and provide a better user experience.
[0315] The following describes the processing flow.
[0316] Step 1:
[0317] The user enters a question in natural language using a device. The device receives this input and temporarily stores it as text data.
[0318] Step 2:
[0319] The device sends the entered question to the emotion engine, which analyzes the user's emotional state based on parameters such as word choice, sentence structure, and input speed.
[0320] Step 3:
[0321] The emotion engine identifies emotions from user input. For example, it analyzes keywords and evaluates the tone of sentences to classify emotions into categories such as "joy," "anxiety," and "frustration."
[0322] Step 4:
[0323] The terminal receives the analysis results and transmits the user's emotional state to the server. The server receives the user's emotional information along with the question data.
[0324] Step 5:
[0325] The server analyzes the question and extracts keywords necessary to understand the intent of the question. Then, it generates an appropriate search query based on those keywords.
[0326] Step 6:
[0327] The system uses server-generated search queries to search indexed internal databases and gather highly relevant information. Simultaneously, it filters search results, taking access rights into consideration as needed.
[0328] Step 7:
[0329] Based on the information collected by the server, generative artificial intelligence is used to generate the optimal response. During this process, the tone and details of the response are adjusted according to the user's emotional state. For example, if frustration is detected, polite and empathetic language will be used.
[0330] Step 8:
[0331] The generated response is sent from the server to the terminal, which then displays it on the user interface. The user can receive the presented response in an easy-to-understand format.
[0332] Step 9:
[0333] Users can request further information as needed or provide feedback on the system's response. The terminal receives these requests and feedback and prepares to send them to the server for use in future processing.
[0334] (Example 2)
[0335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0336] In modern information systems, simply providing information mechanically without considering the emotions users feel, such as frustration or confusion, does not allow for truly user-centric service. Furthermore, simple information retrieval systems have difficulty accurately grasping the user's search intent, and therefore cannot be expected to improve the user experience.
[0337] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0338] In this invention, the server includes means for receiving questions in natural language from the user, collecting metadata and performing sentiment analysis, means for adjusting the tone and content of the response based on the analyzed sentiment state, and means for searching for information using the generated search query and generating a response in natural language using generative artificial intelligence. This makes it possible to provide information in a way that takes the user's emotions into consideration and to improve the user experience.
[0339] A "user" refers to an individual or group that uses the system to obtain information or enter questions.
[0340] "Natural language" refers to the language that humans use on a daily basis, and is treated as text data that computers can understand and analyze.
[0341] "Metadata" refers to additional information accompanying user input, including data other than the text itself, such as input speed and word choice.
[0342] "Sentiment analysis" refers to the process of identifying a user's emotional state from their written text using natural language processing technology and machine learning algorithms.
[0343] "Generative artificial intelligence" refers to artificial intelligence technology that can generate natural language text similar to that produced by humans, based on a large amount of data.
[0344] A "search query" refers to a string or syntax used internally by a system to specify search conditions for retrieving information.
[0345] "Tone" refers to the overall tone and feel of a text, and is a concept that means adjusting the expression to match the user's emotions.
[0346] This system aims to provide appropriate information in response to users' natural language questions via the company's internal network. The system primarily consists of terminals, servers, an emotion analysis engine, and a generative AI model.
[0347] The user uses a device to input a question in natural language. The device collects metadata such as the input speed and word choice, in addition to the entered question. The sentiment engine uses this metadata, employing natural language processing techniques and machine learning algorithms, to analyze the user's emotional state. For this sentiment analysis, one could utilize, for example, the Python library NLTK (Natural Language Toolkit) or a dedicated sentiment analysis API.
[0348] The analysis results are sent to a server, which adjusts the tone and content of the answers to the questions based on the user's emotional state. The server analyzes the questions and generates appropriate search queries to search the internal data store. This uses common information technology techniques, such as automatically generating search queries and using SQL queries to retrieve information.
[0349] After acquiring the information, the server generates the optimal response using a generative AI model. For example, a GPT (Generative Pre-trained Transformer) model can be used as the generative AI model. Here, prompts are used to instruct the AI model. A concrete example of a prompt could be, "Generate an answer that is easy for the user to understand when they are confused."
[0350] The responses generated by the generative AI model are sent to the device and presented in a visually organized format through the user interface. This allows the user to easily understand the responses. Users can also ask additional questions and submit feedback on the provided answers. This feedback is recorded by the server and used to improve the accuracy of future responses.
[0351] This system is expected to improve the overall user experience by allowing users to efficiently obtain information in an emotionally sensitive manner.
[0352] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0353] Step 1:
[0354] The user inputs a question in natural language through the device. The device receives the user's input as digital data. Specifically, the user inputs a question via the keyboard interface, such as "Why isn't this function working?". The device simultaneously collects metadata such as input speed and writing style.
[0355] Step 2:
[0356] The device sends the collected question data and metadata to the sentiment analysis engine. The sentiment engine uses natural language processing and machine learning algorithms to analyze the user's emotional state from the text. In this process, it performs data calculations on the word choices and punctuation frequency of the input text and outputs sentiment labels such as "confused" or "irritated."
[0357] Step 3:
[0358] The terminal sends the sentiment analysis results to the server, which receives them. The server analyzes the user's question based on the received sentiment state and input text. Based on the analysis, it determines what information is needed and generates a search query. Specifically, it converts the natural language question into an SQL query or a query using search operators.
[0359] Step 4:
[0360] The server uses the generated query to search the company's data store. At this stage, the information obtained from the query is efficiently extracted and the necessary data is aggregated. Once the search results are obtained, the dataset is prepared as input for the AI model.
[0361] Step 5:
[0362] The server inputs the acquired data into the generative AI model and applies prompt sentences according to the user's emotional state. The generative AI model generates the optimal response based on prompts such as, "Generate an easy-to-understand response when the user is confused." Based on the prompts used and the analyzed data, the AI generates a response in natural language and outputs it as text.
[0363] Step 6:
[0364] The server sends the generated response to the terminal, which displays it through a user interface. The terminal displays the response in a format that is appropriately laid out for readability, making it easy for the user to understand. The user can then enter additional questions or provide feedback. This feedback information is sent back to the server and used for subsequent processing.
[0365] (Application Example 2)
[0366] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0367] Conventional question-answering systems have a problem in that they do not adjust their responses according to the user's emotional state, making it difficult to obtain answers that satisfy the user. Furthermore, they lack means to quickly and accurately resolve the user's dissatisfaction and confusion, thus failing to improve the user experience.
[0368] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0369] In this invention, the server includes means for analyzing documents using language processing technology within the company's internal communication network and converting them into a searchable format; means for receiving natural language questions from users, analyzing them, and generating search instructions; and means for generating natural language answers using generative artificial intelligence based on the extracted information. This makes it possible to identify the user's emotional state and adjust the optimal answer based on that state.
[0370] An "internal communication network" is a communication network system used within a company, which allows for the sending, receiving, and sharing of information.
[0371] "Language processing technology" refers to technologies that enable computers to analyze, understand, and generate natural language.
[0372] A "document storage device" is a data storage system in which analyzed document information is stored and used for later retrieval.
[0373] A "search instruction" is a query generated based on a user's question, and its purpose is to efficiently retrieve information.
[0374] "Generative artificial intelligence" is a system that creates and provides text in natural language based on human instructions.
[0375] "Emotional state" refers to the state of mind inferred from the user's writing and actions, and is reflected in the tone and content of their responses.
[0376] A "user interface" refers to the screens and means of operation that allow a user to interact directly with a system.
[0377] The system for realizing this invention begins by receiving natural language questions from the user's terminal via the company's internal communication network. The terminal collects the grammar and metadata of the text entered by the user and sends it to an emotion engine for sentiment analysis. This emotion engine uses language processing techniques such as Python and NLTK, and machine learning frameworks such as TensorFlow and PyTorch to identify the user's emotional state. The analysis results are sent to a server, where a server with generative artificial intelligence installed generates search instructions according to the emotional state and extracts relevant information by referring to the accumulated document database.
[0378] Based on this information, the server uses generative artificial intelligence to generate natural language responses, adjusting them with appropriate tone and expression based on the emotional state. This process produces the final response displayed in the user interface. The user interface, in particular, features a user-friendly design utilizing React Native and other technologies, making it easy for users to ask additional questions and provide feedback. Furthermore, the response adjustment data is accumulated as insights necessary for future improvements, thereby enhancing user satisfaction.
[0379] As a concrete example, consider a scenario in customer support on an e-commerce site where a customer asks, "I feel like these shoes are the wrong size." If the emotion engine detects a "dissatisfied" emotional state, the server will generate a reassuring response such as, "We apologize for the inconvenience caused by the incorrect sizing. We will guide you through the return or exchange process." An example of a prompt might be, "Create a sample of a gentle response when a user inquires about product dissatisfaction or problems." In this way, the invention can provide a new means of communicating more effectively with users.
[0380] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0381] Step 1:
[0382] The device receives user questions in natural language and simultaneously collects metadata such as input speed and word choice. This data is prepared for analysis by the sentiment engine. The input data consists of the user's text information and metadata, while the output is the data sent to the sentiment engine.
[0383] Step 2:
[0384] The terminal sends the collected user text information and metadata to the server. The server uses an emotion engine to analyze the received data and identify the user's emotional state. In this step, natural language processing techniques and machine learning models are applied, and the identified emotional state is output.
[0385] Step 3:
[0386] The server generates appropriate search instructions based on the user's emotional state and question content. These search instructions are created as queries to retrieve highly relevant information from the database. The input is the question content and emotional state, and the output is the search instruction query.
[0387] Step 4:
[0388] The server uses the generated search instructions to search the document storage device and extract relevant information. A database query is executed, and the extracted relevant information is output.
[0389] Step 5:
[0390] Based on the extracted information, the server uses generative artificial intelligence to generate a response in natural language. The generated response is then adjusted to an appropriate tone based on the emotional state. The input is the extracted information and emotional state, and the output is the adjusted natural language response.
[0391] Step 6:
[0392] The generated, adjusted responses are presented to the user through the device's user interface. The user can review the responses on screen and enter additional questions or feedback, leading to further improvements in the user experience. The input is the adjusted responses, and the output is the presentation to the user and the collection of user feedback.
[0393] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0394] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0395] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0396] [Third Embodiment]
[0397] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0398] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0399] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0400] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0401] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0403] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0404] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0405] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0406] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0407] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0408] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0409] This invention relates to a question answering system that operates within an internal network environment. Users can input questions in natural language using a terminal. The terminal receives these questions and sends them to a server. At this stage, the terminal's program analyzes the context of the question, identifies important keywords and phrases, and generates a search query.
[0410] The server receives the generated search query and searches a database that has been pre-indexed using natural language processing. This database contains various internal company documents, ensuring that important information is efficiently managed. Search results are ranked based on relevance and filtered as needed, taking access rights into consideration.
[0411] After relevant information is retrieved, the server uses a generative artificial intelligence algorithm to construct a natural language response. This process utilizes text extracted from search results to generate a response that directly addresses the user's question. The generated response is further formatted and presented in a polished style before being presented to the user.
[0412] Users can view server-generated answers through their devices. The user interface presents information in an easy-to-read layout and provides supplementary information and related links as needed. Users can also ask additional questions and provide feedback to the system.
[0413] For example, if a user asks, "Tell me the summary of the latest technical report," the system extracts keywords such as "technical report" and "summary," and searches for related documents based on them. Based on the retrieved information, the server generates a specific answer such as, "The latest technical report describes current technical challenges and their solutions." In this way, this invention enhances the efficiency of information utilization within the company and supports business operations while maintaining security.
[0414] The following describes the processing flow.
[0415] Step 1:
[0416] The user inputs a question in natural language through the device. The device receives the user's input in real time and prepares to convert it into the appropriate format.
[0417] Step 2:
[0418] The terminal analyzes the entered question and extracts important keywords and phrases. Using natural language processing technology, it interprets the intent of the question and generates a search query.
[0419] Step 3:
[0420] The terminal sends the generated search query to the server. The server receives this query and accesses a pre-indexed database.
[0421] Step 4:
[0422] The server searches the database based on the query it receives. The search is performed based on keyword frequency and document relevance scores, and the most relevant information is collected.
[0423] Step 5:
[0424] The server retrieves search results and filters them according to access rights. Security is ensured so that only appropriate information proceeds to the next step.
[0425] Step 6:
[0426] The server uses generative artificial intelligence to generate appropriate response sentences based on filtered information. In this step, data obtained from multiple documents is integrated and the response is constructed using natural language expressions.
[0427] Step 7:
[0428] The server sends the generated response to the terminal. The terminal displays the received response on its user interface, presenting it in a format that is easy for the user to understand.
[0429] Step 8:
[0430] The user can review the displayed answers and have options to ask additional questions or provide feedback as needed. The device then prepares to process these actions as input for the next cycle.
[0431] (Example 1)
[0432] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0433] Conventional information retrieval systems have struggled to provide users with the information they need quickly and accurately. In particular, there is a need to extract highly relevant information from diverse document sets and provide clear, understandable responses in natural language. Furthermore, the provision of appropriate information according to user access rights is essential, making efficient searching and result processing crucial challenges.
[0434] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0435] In this invention, the server includes information processing means for analyzing a set of documents and converting them into a searchable format; processing means for receiving a question in natural language from a user, analyzing it, and generating a search instruction; and search means for searching the analyzed document data set using the generated search instruction and extracting highly relevant information. This makes it possible to quickly and accurately extract highly relevant information from a diverse set of documents, provide natural language responses to user questions, and perform information filtering according to access rights as needed.
[0436] A "document collection" refers to various forms of documents and data stored within a company or a specific environment.
[0437] "Information processing means" refers to hardware or software functions that analyze and transform data and perform processing according to requirements.
[0438] "User" refers to a person or entity that accesses the system to search for information or ask questions.
[0439] "Natural language" refers to the forms of words and sentences that humans use on a daily basis, and is the subject of analysis by computers.
[0440] A "search instruction" refers to a query or command generated to satisfy a user's question or request.
[0441] "Search method" refers to a function for finding information from databases or document sets based on specified conditions.
[0442] "Highly relevant information" refers to information that is deemed most appropriate to provide direct or indirect answers to users' questions or requests.
[0443] A "computer-generated natural language model" refers to algorithms and programs that use artificial intelligence technology to generate natural-sounding text.
[0444] "Response" refers to the answer or information provided in response to a user's question.
[0445] "Presentation means" refers to a function for displaying or providing information or responses generated by the system to the user.
[0446] This invention relates to an information processing system aimed at the efficient analysis and provision of information from a collection of documents. The system mainly includes users, terminals, and a server.
[0447] Users can input questions in natural language using a terminal. The terminal uses natural language processing libraries such as "NLTK" and "spaCy" to understand the context of the input question and identify important keywords. Based on this, it creates a search command and sends it to the server.
[0448] The server is responsible for searching for relevant information from the data set based on the search instructions it receives. The data set is indexed by advanced search engines such as Elasticsearch, enabling fast and accurate information retrieval. The search results are ranked based on their relevance.
[0449] Subsequently, the server uses generative AI models such as "GPT" and "Transformers" to generate a response based on the information obtained from the search results. This response directly answers the question and is expressed in natural and easy-to-understand language.
[0450] Ultimately, the user receives a response generated through the device's user interface. The screen provides links to relevant information and supplementary explanations to aid in understanding the information. Users can also ask additional questions and submit feedback to the system.
[0451] For example, by using a prompt such as "Tell me the summary of the latest technical report," the system provides a specific response such as "The latest technical report describes current technical challenges and their solutions." In this way, the invention aims to provide users with the information they need quickly and accurately, thereby improving the efficiency of information utilization.
[0452] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0453] Step 1:
[0454] The user inputs a question into the terminal using natural language. The terminal analyzes the input question using natural language processing libraries such as "NLTK" or "spaCy." In the process, it extracts important keywords and phrases from the question. Based on the input question text, it performs data analysis and outputs a set of keywords. Specifically, it performs morphological analysis to understand the structure of the question text.
[0455] Step 2:
[0456] The terminal generates a search instruction using the extracted keyword set. This search instruction is then formatted as a query to be sent to the server. Based on the keyword set as input, it converts it into an appropriate search syntax to output a query statement that the server can interpret. Specifically, it constructs the query by concatenating keywords and adding logical operators.
[0457] Step 3:
[0458] The server uses search commands received from the terminal to search the data set. The server uses a search engine such as "Elasticsearch" to quickly search the indexed database and retrieve relevant information. Based on the input query, it performs a data search within the database and outputs a list of relevant information. Specifically, it performs a full-text search, calculates a relevance score, and ranks the information.
[0459] Step 4:
[0460] The server uses a generative AI model to generate natural language responses based on the list of search results. Here, models such as "GPT" and "Transformers" are used to create appropriate answers to user questions. It analyzes the input list of information, selects relevant information, and outputs natural language sentences. Specifically, it generates natural responses that are relevant to the context.
[0461] Step 5:
[0462] The server sends the generated response to the terminal, which the user visually confirms. The terminal's user interface displays the received response in an easy-to-read format, providing links to relevant information and supplementary explanations as needed. Based on the response text as input, it outputs information formatted for screen display. Specifically, it applies layout and style using HTML and CSS.
[0463] (Application Example 1)
[0464] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0465] To improve work efficiency within a factory, it is necessary to create an environment where workers can quickly and accurately obtain the information they need. However, information retrieval on-site is not easy, and providing timely information is difficult. Furthermore, for direct interaction with machinery during work, there is a need for a means of easily accessing information without requiring user intervention.
[0466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0467] In this invention, the server includes means for analyzing information using natural language processing within the company's internal communication network and converting it into a searchable format; means for receiving inquiries in natural language from users, analyzing them, and generating search queries; and means for using the generated search queries to search the analyzed information repository and extract highly relevant knowledge. This makes it possible for workers to easily obtain information by voice and improve work efficiency.
[0468] An "internal communication network" is a network used for sharing information and communicating within a company or organization.
[0469] "Natural language processing" is a technology that uses computers to understand, analyze, and generate human language.
[0470] A "searchable format" is a data format that is organized and structured in a way that allows information to be efficiently searched.
[0471] "User" refers to a person or user who uses the system.
[0472] A "search query" is a set of search conditions and keywords entered to retrieve information.
[0473] An "information repository" is a database or storage device where data and information are stored and managed.
[0474] "Highly relevant knowledge" refers to information that is appropriate and important in response to a user's question during a search.
[0475] "Generative intelligence" is an artificial intelligence technology that generates new information or answers based on given data or queries.
[0476] "Speech recognition" is a technology that analyzes speech data and converts its content into text data.
[0477] "Audio output" is a technology for playing text data as audio.
[0478] The system designed to realize this application provides the ability to quickly obtain the information workers need in a factory work environment. The system accepts questions from workers in natural language using a smartphone or a voice-input capable device.
[0479] The smartphone uses the Google Speech-to-Text API to convert received audio into text. Then, spaCy, a natural language processing engine, analyzes the text to identify important keywords and generate a search query. This query is sent to an API server using Flask. The server uses Elasticsearch to search the company's indexed information repository and extract highly relevant knowledge.
[0480] The server uses generative AI models such as OpenAI's GPT-3 to generate natural language responses from extracted knowledge. After formatting, these responses are output as speech using the Google Text-to-Speech API and provided to workers via smartphones.
[0481] For example, if a worker asks, "Please tell me the assembly process for the new product A and B," the system will treat "product A and B" and "assembly process" as important keywords and search for and provide relevant information. An example of a prompt in this case would be: "Question: 'Please tell me the assembly process for the new product A and B.' Search keywords: 'product A and B', 'assembly process'."
[0482] Thus, this system aims to improve work efficiency by allowing factory workers to easily obtain information through voice input.
[0483] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0484] Step 1:
[0485] The device accepts voice input from the user. The input voice is converted to text using the Google Speech-to-Text API. At this stage, the input is voice data, and the output is the corresponding text data.
[0486] Step 2:
[0487] The terminal analyzes the generated text using spaCy, a natural language processing engine. Important keywords and phrases are extracted from the text, and search queries are generated. The input here is text data, and the output is a data structure containing the search queries.
[0488] Step 3:
[0489] The server uses Elasticsearch to search the company's indexed information repository using the search query received from the terminal. It then extracts highly relevant knowledge. In this step, the input includes the search query, and the output is the retrieved relevant information.
[0490] Step 4:
[0491] The server generates a natural language response based on the extracted relevant information, using OpenAI's GPT-3, etc. A generative AI model processes this. The input is relevant information, and the output is the natural language response text.
[0492] Step 5:
[0493] The server converts the generated response text into speech using the Google Text-to-Speech API and sends it to the device. In this case, the input is natural language response text, and the output is audio data.
[0494] Step 6:
[0495] The terminal plays audio data and provides it to the user. The input is audio data received from the server, and the output is the played audio.
[0496] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0497] This invention combines an emotion engine with a system that provides appropriate answers to natural language questions from users within an internal network. Users can input questions via a terminal. The terminal receives the input question and uses the emotion engine to analyze metadata such as the user's word choice and input speed to identify the user's emotional state.
[0498] The emotion engine uses natural language processing techniques and machine learning algorithms to analyze the emotions expressed in the user's text. The results of this emotion analysis are sent to the server, which adjusts the wording of the response according to the emotional state. For example, if the system detects that the user is confused, it can generate a response that includes a detailed and easy-to-understand explanation.
[0499] The server analyzes the question content, taking this emotional information into account, generates a search query, and searches the company's database. After retrieving the search results, it uses generative artificial intelligence to generate the optimal answer. In this answer generation process, the server selectively emphasizes the content of the answer and adjusts the tone appropriately according to the user's emotional state.
[0500] The generated answers are presented to the user through a user interface. The device displays the answers in a user-friendly layout, making them easy for the user to understand. Users can also ask additional questions on the screen and submit feedback on the system's answers. This feedback is recorded by the system and used to improve the accuracy of future responses.
[0501] For example, if a user asks the terminal, "Why isn't this software working properly?", and the emotion engine detects frustration from the user's writing style and typing speed, the server can generate a detailed response in a gentle tone, such as, "We apologize for the inconvenience. The usual reason this software doesn't work properly is XX." This invention makes it possible to empathize with the user's emotions and provide a better user experience.
[0502] The following describes the processing flow.
[0503] Step 1:
[0504] The user enters a question in natural language using a device. The device receives this input and temporarily stores it as text data.
[0505] Step 2:
[0506] The device sends the entered question to the emotion engine, which analyzes the user's emotional state based on parameters such as word choice, sentence structure, and input speed.
[0507] Step 3:
[0508] The emotion engine identifies emotions from user input. For example, it analyzes keywords and evaluates the tone of sentences to classify emotions into categories such as "joy," "anxiety," and "frustration."
[0509] Step 4:
[0510] The terminal receives the analysis results and transmits the user's emotional state to the server. The server receives the user's emotional information along with the question data.
[0511] Step 5:
[0512] The server analyzes the question and extracts keywords necessary to understand the intent of the question. Then, it generates an appropriate search query based on those keywords.
[0513] Step 6:
[0514] The system uses server-generated search queries to search indexed internal databases and gather highly relevant information. Simultaneously, it filters search results, taking access rights into consideration as needed.
[0515] Step 7:
[0516] Based on the information collected by the server, generative artificial intelligence is used to generate the optimal response. During this process, the tone and details of the response are adjusted according to the user's emotional state. For example, if frustration is detected, polite and empathetic language will be used.
[0517] Step 8:
[0518] The generated response is sent from the server to the terminal, which then displays it on the user interface. The user can receive the presented response in an easy-to-understand format.
[0519] Step 9:
[0520] Users can request further information as needed or provide feedback on the system's response. The terminal receives these requests and feedback and prepares to send them to the server for use in future processing.
[0521] (Example 2)
[0522] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0523] In modern information systems, simply providing information mechanically without considering the emotions users feel, such as frustration or confusion, does not allow for truly user-centric service. Furthermore, simple information retrieval systems have difficulty accurately grasping the user's search intent, and therefore cannot be expected to improve the user experience.
[0524] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0525] In this invention, the server includes means for receiving questions in natural language from the user, collecting metadata and performing sentiment analysis, means for adjusting the tone and content of the response based on the analyzed sentiment state, and means for searching for information using the generated search query and generating a response in natural language using generative artificial intelligence. This makes it possible to provide information in a way that takes the user's emotions into consideration and to improve the user experience.
[0526] A "user" refers to an individual or group that uses the system to obtain information or enter questions.
[0527] "Natural language" refers to the language that humans use on a daily basis, and is treated as text data that computers can understand and analyze.
[0528] "Metadata" refers to additional information accompanying user input, including data other than the text itself, such as input speed and word choice.
[0529] "Sentiment analysis" refers to the process of identifying a user's emotional state from their written text using natural language processing technology and machine learning algorithms.
[0530] "Generative artificial intelligence" refers to artificial intelligence technology that can generate natural language text similar to that produced by humans, based on a large amount of data.
[0531] A "search query" refers to a string or syntax used internally by a system to specify search conditions for retrieving information.
[0532] "Tone" refers to the overall tone and feel of a text, and is a concept that means adjusting the expression to match the user's emotions.
[0533] This system aims to provide appropriate information in response to users' natural language questions via the company's internal network. The system primarily consists of terminals, servers, an emotion analysis engine, and a generative AI model.
[0534] The user uses a device to input a question in natural language. The device collects metadata such as the input speed and word choice, in addition to the entered question. The sentiment engine uses this metadata, employing natural language processing techniques and machine learning algorithms, to analyze the user's emotional state. For this sentiment analysis, one could utilize, for example, the Python library NLTK (Natural Language Toolkit) or a dedicated sentiment analysis API.
[0535] The analysis results are sent to a server, which adjusts the tone and content of the answers to the questions based on the user's emotional state. The server analyzes the questions and generates appropriate search queries to search the internal data store. This uses common information technology techniques, such as automatically generating search queries and using SQL queries to retrieve information.
[0536] After acquiring the information, the server generates the optimal response using a generative AI model. For example, a GPT (Generative Pre-trained Transformer) model can be used as the generative AI model. Here, prompts are used to instruct the AI model. A concrete example of a prompt could be, "Generate an answer that is easy for the user to understand when they are confused."
[0537] The responses generated by the generative AI model are sent to the device and presented in a visually organized format through the user interface. This allows the user to easily understand the responses. Users can also ask additional questions and submit feedback on the provided answers. This feedback is recorded by the server and used to improve the accuracy of future responses.
[0538] This system is expected to improve the overall user experience by allowing users to efficiently obtain information in an emotionally sensitive manner.
[0539] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0540] Step 1:
[0541] The user inputs a question in natural language through the device. The device receives the user's input as digital data. Specifically, the user inputs a question via the keyboard interface, such as "Why isn't this function working?". The device simultaneously collects metadata such as input speed and writing style.
[0542] Step 2:
[0543] The device sends the collected question data and metadata to the sentiment analysis engine. The sentiment engine uses natural language processing and machine learning algorithms to analyze the user's emotional state from the text. In this process, it performs data calculations on the word choices and punctuation frequency of the input text and outputs sentiment labels such as "confused" or "irritated."
[0544] Step 3:
[0545] The terminal sends the sentiment analysis results to the server, which receives them. The server analyzes the user's question based on the received sentiment state and input text. Based on the analysis, it determines what information is needed and generates a search query. Specifically, it converts the natural language question into an SQL query or a query using search operators.
[0546] Step 4:
[0547] The server uses the generated query to search the company's data store. At this stage, the information obtained from the query is efficiently extracted and the necessary data is aggregated. Once the search results are obtained, the dataset is prepared as input for the AI model.
[0548] Step 5:
[0549] The server inputs the acquired data into the generative AI model and applies prompt sentences according to the user's emotional state. The generative AI model generates the optimal response based on prompts such as, "Generate an easy-to-understand response when the user is confused." Based on the prompts used and the analyzed data, the AI generates a response in natural language and outputs it as text.
[0550] Step 6:
[0551] The server sends the generated response to the terminal, which displays it through a user interface. The terminal displays the response in a format that is appropriately laid out for readability, making it easy for the user to understand. The user can then enter additional questions or provide feedback. This feedback information is sent back to the server and used for subsequent processing.
[0552] (Application Example 2)
[0553] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0554] Conventional question-answering systems have a problem in that they do not adjust their responses according to the user's emotional state, making it difficult to obtain answers that satisfy the user. Furthermore, they lack means to quickly and accurately resolve the user's dissatisfaction and confusion, thus failing to improve the user experience.
[0555] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0556] In this invention, the server includes means for analyzing documents using language processing technology within the company's internal communication network and converting them into a searchable format; means for receiving natural language questions from users, analyzing them, and generating search instructions; and means for generating natural language answers using generative artificial intelligence based on the extracted information. This makes it possible to identify the user's emotional state and adjust the optimal answer based on that state.
[0557] An "internal communication network" is a communication network system used within a company, which allows for the sending, receiving, and sharing of information.
[0558] "Language processing technology" refers to technologies that enable computers to analyze, understand, and generate natural language.
[0559] A "document storage device" is a data storage system in which analyzed document information is stored and used for later retrieval.
[0560] A "search instruction" is a query generated based on a user's question, and its purpose is to efficiently retrieve information.
[0561] "Generative artificial intelligence" is a system that creates and provides text in natural language based on human instructions.
[0562] "Emotional state" refers to the state of mind inferred from the user's writing and actions, and is reflected in the tone and content of their responses.
[0563] A "user interface" refers to the screens and means of operation that allow a user to interact directly with a system.
[0564] The system for realizing this invention begins by receiving natural language questions from the user's terminal via the company's internal communication network. The terminal collects the grammar and metadata of the text entered by the user and sends it to an emotion engine for sentiment analysis. This emotion engine uses language processing techniques such as Python and NLTK, and machine learning frameworks such as TensorFlow and PyTorch to identify the user's emotional state. The analysis results are sent to a server, where a server with generative artificial intelligence installed generates search instructions according to the emotional state and extracts relevant information by referring to the accumulated document database.
[0565] Based on this information, the server uses generative artificial intelligence to generate natural language responses, adjusting them with appropriate tone and expression based on the emotional state. This process produces the final response displayed in the user interface. The user interface, in particular, features a user-friendly design utilizing React Native and other technologies, making it easy for users to ask additional questions and provide feedback. Furthermore, the response adjustment data is accumulated as insights necessary for future improvements, thereby enhancing user satisfaction.
[0566] As a concrete example, consider a scenario in customer support on an e-commerce site where a customer asks, "I feel like these shoes are the wrong size." If the emotion engine detects a "dissatisfied" emotional state, the server will generate a reassuring response such as, "We apologize for the inconvenience caused by the incorrect sizing. We will guide you through the return or exchange process." An example of a prompt might be, "Create a sample of a gentle response when a user inquires about product dissatisfaction or problems." In this way, the invention can provide a new means of communicating more effectively with users.
[0567] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0568] Step 1:
[0569] The device receives user questions in natural language and simultaneously collects metadata such as input speed and word choice. This data is prepared for analysis by the sentiment engine. The input data consists of the user's text information and metadata, while the output is the data sent to the sentiment engine.
[0570] Step 2:
[0571] The terminal sends the collected user text information and metadata to the server. The server uses an emotion engine to analyze the received data and identify the user's emotional state. In this step, natural language processing techniques and machine learning models are applied, and the identified emotional state is output.
[0572] Step 3:
[0573] The server generates appropriate search instructions based on the user's emotional state and question content. These search instructions are created as queries to retrieve highly relevant information from the database. The input is the question content and emotional state, and the output is the search instruction query.
[0574] Step 4:
[0575] The server uses the generated search instructions to search the document storage device and extract relevant information. A database query is executed, and the extracted relevant information is output.
[0576] Step 5:
[0577] Based on the extracted information, the server uses generative artificial intelligence to generate a response in natural language. The generated response is then adjusted to an appropriate tone based on the emotional state. The input is the extracted information and emotional state, and the output is the adjusted natural language response.
[0578] Step 6:
[0579] The generated, adjusted responses are presented to the user through the device's user interface. The user can review the responses on screen and enter additional questions or feedback, leading to further improvements in the user experience. The input is the adjusted responses, and the output is the presentation to the user and the collection of user feedback.
[0580] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0581] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0582] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0583] [Fourth Embodiment]
[0584] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0585] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0586] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0587] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0588] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0589] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0590] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0591] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0592] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0593] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0594] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0595] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0596] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0597] This invention relates to a question answering system that operates within an internal network environment. Users can input questions in natural language using a terminal. The terminal receives these questions and sends them to a server. At this stage, the terminal's program analyzes the context of the question, identifies important keywords and phrases, and generates a search query.
[0598] The server receives the generated search query and searches a database that has been pre-indexed using natural language processing. This database contains various internal company documents, ensuring that important information is efficiently managed. Search results are ranked based on relevance and filtered as needed, taking access rights into consideration.
[0599] After relevant information is retrieved, the server uses a generative artificial intelligence algorithm to construct a natural language response. This process utilizes text extracted from search results to generate a response that directly addresses the user's question. The generated response is further formatted and presented in a polished style before being presented to the user.
[0600] Users can view server-generated answers through their devices. The user interface presents information in an easy-to-read layout and provides supplementary information and related links as needed. Users can also ask additional questions and provide feedback to the system.
[0601] For example, if a user asks, "Tell me the summary of the latest technical report," the system extracts keywords such as "technical report" and "summary," and searches for related documents based on them. Based on the retrieved information, the server generates a specific answer such as, "The latest technical report describes current technical challenges and their solutions." In this way, this invention enhances the efficiency of information utilization within the company and supports business operations while maintaining security.
[0602] The following describes the processing flow.
[0603] Step 1:
[0604] The user inputs a question in natural language through the device. The device receives the user's input in real time and prepares to convert it into the appropriate format.
[0605] Step 2:
[0606] The terminal analyzes the entered question and extracts important keywords and phrases. Using natural language processing technology, it interprets the intent of the question and generates a search query.
[0607] Step 3:
[0608] The terminal sends the generated search query to the server. The server receives this query and accesses a pre-indexed database.
[0609] Step 4:
[0610] The server searches the database based on the query it receives. The search is performed based on keyword frequency and document relevance scores, and the most relevant information is collected.
[0611] Step 5:
[0612] The server retrieves search results and filters them according to access rights. Security is ensured so that only appropriate information proceeds to the next step.
[0613] Step 6:
[0614] The server uses generative artificial intelligence to generate appropriate response sentences based on filtered information. In this step, data obtained from multiple documents is integrated and the response is constructed using natural language expressions.
[0615] Step 7:
[0616] The server sends the generated response to the terminal. The terminal displays the received response on its user interface, presenting it in a format that is easy for the user to understand.
[0617] Step 8:
[0618] The user can review the displayed answers and have options to ask additional questions or provide feedback as needed. The device then prepares to process these actions as input for the next cycle.
[0619] (Example 1)
[0620] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0621] Conventional information retrieval systems have struggled to provide users with the information they need quickly and accurately. In particular, there is a need to extract highly relevant information from diverse document sets and provide clear, understandable responses in natural language. Furthermore, the provision of appropriate information according to user access rights is essential, making efficient searching and result processing crucial challenges.
[0622] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0623] In this invention, the server includes information processing means for analyzing a set of documents and converting them into a searchable format; processing means for receiving a question in natural language from a user, analyzing it, and generating a search instruction; and search means for searching the analyzed document data set using the generated search instruction and extracting highly relevant information. This makes it possible to quickly and accurately extract highly relevant information from a diverse set of documents, provide natural language responses to user questions, and perform information filtering according to access rights as needed.
[0624] A "document collection" refers to various forms of documents and data stored within a company or a specific environment.
[0625] "Information processing means" refers to hardware or software functions that analyze and transform data and perform processing according to requirements.
[0626] "User" refers to a person or entity that accesses the system to search for information or ask questions.
[0627] "Natural language" refers to the forms of words and sentences that humans use on a daily basis, and is the subject of analysis by computers.
[0628] A "search instruction" refers to a query or command generated to satisfy a user's question or request.
[0629] "Search method" refers to a function for finding information from databases or document sets based on specified conditions.
[0630] "Highly relevant information" refers to information that is deemed most appropriate to provide direct or indirect answers to users' questions or requests.
[0631] A "computer-generated natural language model" refers to algorithms and programs that use artificial intelligence technology to generate natural-sounding text.
[0632] "Response" refers to the answer or information provided in response to a user's question.
[0633] "Presentation means" refers to a function for displaying or providing information or responses generated by the system to the user.
[0634] This invention relates to an information processing system aimed at the efficient analysis and provision of information from a collection of documents. The system mainly includes users, terminals, and a server.
[0635] Users can input questions in natural language using a terminal. The terminal uses natural language processing libraries such as "NLTK" and "spaCy" to understand the context of the input question and identify important keywords. Based on this, it creates a search command and sends it to the server.
[0636] The server is responsible for searching for relevant information from the data set based on the search instructions it receives. The data set is indexed by advanced search engines such as Elasticsearch, enabling fast and accurate information retrieval. The search results are ranked based on their relevance.
[0637] Subsequently, the server uses generative AI models such as "GPT" and "Transformers" to generate a response based on the information obtained from the search results. This response directly answers the question and is expressed in natural and easy-to-understand language.
[0638] Ultimately, the user receives a response generated through the device's user interface. The screen provides links to relevant information and supplementary explanations to aid in understanding the information. Users can also ask additional questions and submit feedback to the system.
[0639] For example, by using a prompt such as "Tell me the summary of the latest technical report," the system provides a specific response such as "The latest technical report describes current technical challenges and their solutions." In this way, the invention aims to provide users with the information they need quickly and accurately, thereby improving the efficiency of information utilization.
[0640] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0641] Step 1:
[0642] The user inputs a question into the terminal using natural language. The terminal analyzes the input question using natural language processing libraries such as "NLTK" or "spaCy." In the process, it extracts important keywords and phrases from the question. Based on the input question text, it performs data analysis and outputs a set of keywords. Specifically, it performs morphological analysis to understand the structure of the question text.
[0643] Step 2:
[0644] The terminal generates a search instruction using the extracted keyword set. This search instruction is then formatted as a query to be sent to the server. Based on the keyword set as input, it converts it into an appropriate search syntax to output a query statement that the server can interpret. Specifically, it constructs the query by concatenating keywords and adding logical operators.
[0645] Step 3:
[0646] The server uses search commands received from the terminal to search the data set. The server uses a search engine such as "Elasticsearch" to quickly search the indexed database and retrieve relevant information. Based on the input query, it performs a data search within the database and outputs a list of relevant information. Specifically, it performs a full-text search, calculates a relevance score, and ranks the information.
[0647] Step 4:
[0648] The server uses a generative AI model to generate natural language responses based on the list of search results. Here, models such as "GPT" and "Transformers" are used to create appropriate answers to user questions. It analyzes the input list of information, selects relevant information, and outputs natural language sentences. Specifically, it generates natural responses that are relevant to the context.
[0649] Step 5:
[0650] The server sends the generated response to the terminal, which the user visually confirms. The terminal's user interface displays the received response in an easy-to-read format, providing links to relevant information and supplementary explanations as needed. Based on the response text as input, it outputs information formatted for screen display. Specifically, it applies layout and style using HTML and CSS.
[0651] (Application Example 1)
[0652] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0653] To improve work efficiency within a factory, it is necessary to create an environment where workers can quickly and accurately obtain the information they need. However, information retrieval on-site is not easy, and providing timely information is difficult. Furthermore, for direct interaction with machinery during work, there is a need for a means of easily accessing information without requiring user intervention.
[0654] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0655] In this invention, the server includes means for analyzing information using natural language processing within the company's internal communication network and converting it into a searchable format; means for receiving inquiries in natural language from users, analyzing them, and generating search queries; and means for using the generated search queries to search the analyzed information repository and extract highly relevant knowledge. This makes it possible for workers to easily obtain information by voice and improve work efficiency.
[0656] An "internal communication network" is a network used for sharing information and communicating within a company or organization.
[0657] "Natural language processing" is a technology that uses computers to understand, analyze, and generate human language.
[0658] A "searchable format" is a data format that is organized and structured in a way that allows information to be efficiently searched.
[0659] "User" refers to a person or user who uses the system.
[0660] A "search query" is a set of search conditions and keywords entered to retrieve information.
[0661] An "information repository" is a database or storage device where data and information are stored and managed.
[0662] "Highly relevant knowledge" refers to information that is appropriate and important in response to a user's question during a search.
[0663] "Generative intelligence" is an artificial intelligence technology that generates new information or answers based on given data or queries.
[0664] "Speech recognition" is a technology that analyzes speech data and converts its content into text data.
[0665] "Audio output" is a technology for playing text data as audio.
[0666] The system designed to realize this application provides the ability to quickly obtain the information workers need in a factory work environment. The system accepts questions from workers in natural language using a smartphone or a voice-input capable device.
[0667] The smartphone uses the Google Speech-to-Text API to convert received audio into text. Then, spaCy, a natural language processing engine, analyzes the text to identify important keywords and generate a search query. This query is sent to an API server using Flask. The server uses Elasticsearch to search the company's indexed information repository and extract highly relevant knowledge.
[0668] The server uses generative AI models such as OpenAI's GPT-3 to generate natural language responses from extracted knowledge. After formatting, these responses are output as speech using the Google Text-to-Speech API and provided to workers via smartphones.
[0669] For example, if a worker asks, "Please tell me the assembly process for the new product A and B," the system will treat "product A and B" and "assembly process" as important keywords and search for and provide relevant information. An example of a prompt in this case would be: "Question: 'Please tell me the assembly process for the new product A and B.' Search keywords: 'product A and B', 'assembly process'."
[0670] Thus, this system aims to improve work efficiency by allowing factory workers to easily obtain information through voice input.
[0671] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0672] Step 1:
[0673] The device accepts voice input from the user. The input voice is converted to text using the Google Speech-to-Text API. At this stage, the input is voice data, and the output is the corresponding text data.
[0674] Step 2:
[0675] The terminal analyzes the generated text using spaCy, a natural language processing engine. Important keywords and phrases are extracted from the text, and search queries are generated. The input here is text data, and the output is a data structure containing the search queries.
[0676] Step 3:
[0677] The server uses Elasticsearch to search the company's indexed information repository using the search query received from the terminal. It then extracts highly relevant knowledge. In this step, the input includes the search query, and the output is the retrieved relevant information.
[0678] Step 4:
[0679] The server generates a natural language response based on the extracted relevant information, using OpenAI's GPT-3, etc. A generative AI model processes this. The input is relevant information, and the output is the natural language response text.
[0680] Step 5:
[0681] The server converts the generated response text into speech using the Google Text-to-Speech API and sends it to the device. In this case, the input is natural language response text, and the output is audio data.
[0682] Step 6:
[0683] The terminal plays audio data and provides it to the user. The input is audio data received from the server, and the output is the played audio.
[0684] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0685] This invention combines an emotion engine with a system that provides appropriate answers to natural language questions from users within an internal network. Users can input questions via a terminal. The terminal receives the input question and uses the emotion engine to analyze metadata such as the user's word choice and input speed to identify the user's emotional state.
[0686] The emotion engine uses natural language processing techniques and machine learning algorithms to analyze the emotions expressed in the user's text. The results of this emotion analysis are sent to the server, which adjusts the wording of the response according to the emotional state. For example, if the system detects that the user is confused, it can generate a response that includes a detailed and easy-to-understand explanation.
[0687] The server analyzes the question content, taking this emotional information into account, generates a search query, and searches the company's database. After retrieving the search results, it uses generative artificial intelligence to generate the optimal answer. In this answer generation process, the server selectively emphasizes the content of the answer and adjusts the tone appropriately according to the user's emotional state.
[0688] The generated answers are presented to the user through a user interface. The device displays the answers in a user-friendly layout, making them easy for the user to understand. Users can also ask additional questions on the screen and submit feedback on the system's answers. This feedback is recorded by the system and used to improve the accuracy of future responses.
[0689] For example, if a user asks the terminal, "Why isn't this software working properly?", and the emotion engine detects frustration from the user's writing style and typing speed, the server can generate a detailed response in a gentle tone, such as, "We apologize for the inconvenience. The usual reason this software doesn't work properly is XX." This invention makes it possible to empathize with the user's emotions and provide a better user experience.
[0690] The following describes the processing flow.
[0691] Step 1:
[0692] The user enters a question in natural language using a device. The device receives this input and temporarily stores it as text data.
[0693] Step 2:
[0694] The device sends the entered question to the emotion engine, which analyzes the user's emotional state based on parameters such as word choice, sentence structure, and input speed.
[0695] Step 3:
[0696] The emotion engine identifies emotions from user input. For example, it analyzes keywords and evaluates the tone of sentences to classify emotions into categories such as "joy," "anxiety," and "frustration."
[0697] Step 4:
[0698] The terminal receives the analysis results and transmits the user's emotional state to the server. The server receives the user's emotional information along with the question data.
[0699] Step 5:
[0700] The server analyzes the question and extracts keywords necessary to understand the intent of the question. Then, it generates an appropriate search query based on those keywords.
[0701] Step 6:
[0702] The system uses server-generated search queries to search indexed internal databases and gather highly relevant information. Simultaneously, it filters search results, taking access rights into consideration as needed.
[0703] Step 7:
[0704] Based on the information collected by the server, generative artificial intelligence is used to generate the optimal response. During this process, the tone and details of the response are adjusted according to the user's emotional state. For example, if frustration is detected, polite and empathetic language will be used.
[0705] Step 8:
[0706] The generated response is sent from the server to the terminal, which then displays it on the user interface. The user can receive the presented response in an easy-to-understand format.
[0707] Step 9:
[0708] Users can request further information as needed or provide feedback on the system's response. The terminal receives these requests and feedback and prepares to send them to the server for use in future processing.
[0709] (Example 2)
[0710] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0711] In modern information systems, simply providing information mechanically without considering the emotions users feel, such as frustration or confusion, does not allow for truly user-centric service. Furthermore, simple information retrieval systems have difficulty accurately grasping the user's search intent, and therefore cannot be expected to improve the user experience.
[0712] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0713] In this invention, the server includes means for receiving questions in natural language from the user, collecting metadata and performing sentiment analysis, means for adjusting the tone and content of the response based on the analyzed sentiment state, and means for searching for information using the generated search query and generating a response in natural language using generative artificial intelligence. This makes it possible to provide information in a way that takes the user's emotions into consideration and to improve the user experience.
[0714] A "user" refers to an individual or group that uses the system to obtain information or enter questions.
[0715] "Natural language" refers to the language that humans use on a daily basis, and is treated as text data that computers can understand and analyze.
[0716] "Metadata" refers to additional information accompanying user input, including data other than the text itself, such as input speed and word choice.
[0717] "Sentiment analysis" refers to the process of identifying a user's emotional state from their written text using natural language processing technology and machine learning algorithms.
[0718] "Generative artificial intelligence" refers to artificial intelligence technology that can generate natural language text similar to that produced by humans, based on a large amount of data.
[0719] A "search query" refers to a string or syntax used internally by a system to specify search conditions for retrieving information.
[0720] "Tone" refers to the overall tone and feel of a text, and is a concept that means adjusting the expression to match the user's emotions.
[0721] This system aims to provide appropriate information in response to users' natural language questions via the company's internal network. The system primarily consists of terminals, servers, an emotion analysis engine, and a generative AI model.
[0722] The user uses a device to input a question in natural language. The device collects metadata such as the input speed and word choice, in addition to the entered question. The sentiment engine uses this metadata, employing natural language processing techniques and machine learning algorithms, to analyze the user's emotional state. For this sentiment analysis, one could utilize, for example, the Python library NLTK (Natural Language Toolkit) or a dedicated sentiment analysis API.
[0723] The analysis results are sent to a server, which adjusts the tone and content of the answers to the questions based on the user's emotional state. The server analyzes the questions and generates appropriate search queries to search the internal data store. This uses common information technology techniques, such as automatically generating search queries and using SQL queries to retrieve information.
[0724] After acquiring the information, the server generates the optimal response using a generative AI model. For example, a GPT (Generative Pre-trained Transformer) model can be used as the generative AI model. Here, prompts are used to instruct the AI model. A concrete example of a prompt could be, "Generate an answer that is easy for the user to understand when they are confused."
[0725] The responses generated by the generative AI model are sent to the device and presented in a visually organized format through the user interface. This allows the user to easily understand the responses. Users can also ask additional questions and submit feedback on the provided answers. This feedback is recorded by the server and used to improve the accuracy of future responses.
[0726] This system is expected to improve the overall user experience by allowing users to efficiently obtain information in an emotionally sensitive manner.
[0727] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0728] Step 1:
[0729] The user inputs a question in natural language through the device. The device receives the user's input as digital data. Specifically, the user inputs a question via the keyboard interface, such as "Why isn't this function working?". The device simultaneously collects metadata such as input speed and writing style.
[0730] Step 2:
[0731] The device sends the collected question data and metadata to the sentiment analysis engine. The sentiment engine uses natural language processing and machine learning algorithms to analyze the user's emotional state from the text. In this process, it performs data calculations on the word choices and punctuation frequency of the input text and outputs sentiment labels such as "confused" or "irritated."
[0732] Step 3:
[0733] The terminal sends the sentiment analysis results to the server, which receives them. The server analyzes the user's question based on the received sentiment state and input text. Based on the analysis, it determines what information is needed and generates a search query. Specifically, it converts the natural language question into an SQL query or a query using search operators.
[0734] Step 4:
[0735] The server uses the generated query to search the company's data store. At this stage, the information obtained from the query is efficiently extracted and the necessary data is aggregated. Once the search results are obtained, the dataset is prepared as input for the AI model.
[0736] Step 5:
[0737] The server inputs the acquired data into the generative AI model and applies prompt sentences according to the user's emotional state. The generative AI model generates the optimal response based on prompts such as, "Generate an easy-to-understand response when the user is confused." Based on the prompts used and the analyzed data, the AI generates a response in natural language and outputs it as text.
[0738] Step 6:
[0739] The server sends the generated response to the terminal, which displays it through a user interface. The terminal displays the response in a format that is appropriately laid out for readability, making it easy for the user to understand. The user can then enter additional questions or provide feedback. This feedback information is sent back to the server and used for subsequent processing.
[0740] (Application Example 2)
[0741] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0742] Conventional question-answering systems have a problem in that they do not adjust their responses according to the user's emotional state, making it difficult to obtain answers that satisfy the user. Furthermore, they lack means to quickly and accurately resolve the user's dissatisfaction and confusion, thus failing to improve the user experience.
[0743] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0744] In this invention, the server includes means for analyzing documents using language processing technology within the company's internal communication network and converting them into a searchable format; means for receiving natural language questions from users, analyzing them, and generating search instructions; and means for generating natural language answers using generative artificial intelligence based on the extracted information. This makes it possible to identify the user's emotional state and adjust the optimal answer based on that state.
[0745] An "internal communication network" is a communication network system used within a company, which allows for the sending, receiving, and sharing of information.
[0746] "Language processing technology" refers to technologies that enable computers to analyze, understand, and generate natural language.
[0747] A "document storage device" is a data storage system in which analyzed document information is stored and used for later retrieval.
[0748] A "search instruction" is a query generated based on a user's question, and its purpose is to efficiently retrieve information.
[0749] "Generative artificial intelligence" is a system that creates and provides text in natural language based on human instructions.
[0750] "Emotional state" refers to the state of mind inferred from the user's writing and actions, and is reflected in the tone and content of their responses.
[0751] A "user interface" refers to the screens and means of operation that allow a user to interact directly with a system.
[0752] The system for realizing this invention begins by receiving natural language questions from the user's terminal via the company's internal communication network. The terminal collects the grammar and metadata of the text entered by the user and sends it to an emotion engine for sentiment analysis. This emotion engine uses language processing techniques such as Python and NLTK, and machine learning frameworks such as TensorFlow and PyTorch to identify the user's emotional state. The analysis results are sent to a server, where a server with generative artificial intelligence installed generates search instructions according to the emotional state and extracts relevant information by referring to the accumulated document database.
[0753] Based on this information, the server uses generative artificial intelligence to generate natural language responses, adjusting them with appropriate tone and expression based on the emotional state. This process produces the final response displayed in the user interface. The user interface, in particular, features a user-friendly design utilizing React Native and other technologies, making it easy for users to ask additional questions and provide feedback. Furthermore, the response adjustment data is accumulated as insights necessary for future improvements, thereby enhancing user satisfaction.
[0754] As a concrete example, consider a scenario in customer support on an e-commerce site where a customer asks, "I feel like these shoes are the wrong size." If the emotion engine detects a "dissatisfied" emotional state, the server will generate a reassuring response such as, "We apologize for the inconvenience caused by the incorrect sizing. We will guide you through the return or exchange process." An example of a prompt might be, "Create a sample of a gentle response when a user inquires about product dissatisfaction or problems." In this way, the invention can provide a new means of communicating more effectively with users.
[0755] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0756] Step 1:
[0757] The device receives user questions in natural language and simultaneously collects metadata such as input speed and word choice. This data is prepared for analysis by the sentiment engine. The input data consists of the user's text information and metadata, while the output is the data sent to the sentiment engine.
[0758] Step 2:
[0759] The terminal sends the collected user text information and metadata to the server. The server uses an emotion engine to analyze the received data and identify the user's emotional state. In this step, natural language processing techniques and machine learning models are applied, and the identified emotional state is output.
[0760] Step 3:
[0761] The server generates appropriate search instructions based on the user's emotional state and question content. These search instructions are created as queries to retrieve highly relevant information from the database. The input is the question content and emotional state, and the output is the search instruction query.
[0762] Step 4:
[0763] The server uses the generated search instructions to search the document storage device and extract relevant information. A database query is executed, and the extracted relevant information is output.
[0764] Step 5:
[0765] Based on the extracted information, the server uses generative artificial intelligence to generate a response in natural language. The generated response is then adjusted to an appropriate tone based on the emotional state. The input is the extracted information and emotional state, and the output is the adjusted natural language response.
[0766] Step 6:
[0767] The generated, adjusted responses are presented to the user through the device's user interface. The user can review the responses on screen and enter additional questions or feedback, leading to further improvements in the user experience. The input is the adjusted responses, and the output is the presentation to the user and the collection of user feedback.
[0768] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0769] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0770] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0771] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0772] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0773] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0774] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0775] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0776] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0777] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0778] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0779] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0780] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0781] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0782] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0783] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0784] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0785] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0786] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0787] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0788] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0789] The following is further disclosed regarding the embodiments described above.
[0790] (Claim 1)
[0791] A means of analyzing documents using natural language processing within the company network and converting them into a searchable format,
[0792] A means of receiving natural language questions from users, analyzing them, and generating search queries,
[0793] A means of searching the analyzed document database using the generated search query and extracting highly relevant information,
[0794] A means of generating a natural language response using generative artificial intelligence based on extracted information,
[0795] A means of presenting the generated answer to the user,
[0796] A system that includes this.
[0797] (Claim 2)
[0798] The system according to claim 1, further comprising means for searching an indexed database and filtering the search results according to access rights.
[0799] (Claim 3)
[0800] The system according to claim 1, comprising means for displaying answers through a user interface and for receiving additional questions or feedback.
[0801] "Example 1"
[0802] (Claim 1)
[0803] Information processing means for analyzing a set of documents and converting them into a searchable format,
[0804] A processing means that receives a natural language question from a user, analyzes it, and generates a search instruction,
[0805] A search means that searches a set of document data analyzed using generated search instructions and extracts highly relevant information,
[0806] A generation means that generates a response using a computer-generated natural language model based on the extracted information,
[0807] A presentation means for presenting the generated response to the user,
[0808] An information processing system that includes this.
[0809] (Claim 2)
[0810] The information processing system according to claim 1, further comprising means for searching an indexed data set and processing the search results according to the user's access rights.
[0811] (Claim 3)
[0812] The information processing system according to claim 1, comprising means for displaying responses through the user's operation screen and for receiving additional questions and opinions from the user.
[0813] "Application Example 1"
[0814] (Claim 1)
[0815] A means of analyzing information using natural language processing within the company's internal communication network and converting it into a searchable format,
[0816] A means of receiving natural language queries from users, analyzing them, and generating search queries,
[0817] A means of using the generated search queries to explore the analyzed information repository and extract highly relevant knowledge,
[0818] A means of generating a natural language response using generative intelligence from extracted knowledge,
[0819] A means of presenting the generated response to the user,
[0820] A method for inputting questions about a machine via voice and converting them into text using speech recognition,
[0821] A means for outputting the generated response as audio,
[0822] A system that includes this.
[0823] (Claim 2)
[0824] The system according to claim 1, further comprising means for searching an indexed information repository and excluding search results according to access rights.
[0825] (Claim 3)
[0826] The system according to claim 1, comprising means for displaying responses through a user interface and for receiving additional inquiries and feedback.
[0827] "Example 2 of combining an emotion engine"
[0828] (Claim 1)
[0829] A method for receiving questions in natural language from users, collecting metadata, and performing sentiment analysis,
[0830] A means of adjusting the tone and content of responses based on the analyzed emotional state,
[0831] A method for analyzing questions while considering emotional states and generating search queries,
[0832] A method for searching for information using generated search queries and generating answers in natural language using generative artificial intelligence,
[0833] A means of presenting the generated answer to the user through a user interface,
[0834] A system that includes this.
[0835] (Claim 2)
[0836] The system according to claim 1, further comprising means for searching an indexed data store and filtering search results according to access rights, and means for providing search results based on sentiment analysis.
[0837] (Claim 3)
[0838] The system according to claim 1, comprising a means for displaying emotionally sensitive responses through a user interface and for receiving additional questions or feedback.
[0839] "Application example 2 when combining with an emotional engine"
[0840] (Claim 1)
[0841] A means for analyzing documents using language processing technology within the company's internal communication network and converting them into a searchable format,
[0842] A means of receiving natural language questions from users, analyzing them, and generating search instructions,
[0843] A means for searching the analyzed document storage device using the generated search instructions and extracting highly relevant information,
[0844] A means of generating a natural language response using generative artificial intelligence based on extracted information,
[0845] A means of adjusting the generated response to take into account the user's emotional state,
[0846] A method for identifying a user's emotional state by analyzing their writing style and typing speed,
[0847] A means of presenting the generated answer to the user,
[0848] A system that includes this.
[0849] (Claim 2)
[0850] The system according to claim 1, further comprising means for searching an indexed database and filtering the search results according to access rights.
[0851] (Claim 3)
[0852] The system according to claim 1, comprising means for displaying answers through a user interface, accepting additional questions and feedback, and adjusting responses based on emotional information. [Explanation of Symbols]
[0853] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of analyzing documents using natural language processing within the company network and converting them into a searchable format, A means of receiving natural language questions from users, analyzing them, and generating search queries, A means of searching the analyzed document database using the generated search query and extracting highly relevant information, A means of generating a natural language response using generative artificial intelligence based on extracted information, A means of presenting the generated answer to the user, A system that includes this.
2. The system according to claim 1, further comprising means for searching an indexed database and filtering search results according to access rights.
3. The system according to claim 1, comprising means for displaying answers through a user interface and for receiving additional questions and feedback.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A