Method of supporting chatbot for conversation with user
The method enhances chatbot response accuracy by using vector conversion and database searches to refine answers, addressing the limitations of existing chatbot technologies.
Patent Information
- Application Number
- JP2025048512
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-03-24
- Publication Date
- 2025-10-08
AI Technical Summary
Existing chatbot technologies lack accuracy in generating conversational responses, necessitating improved methods for enhancing response accuracy.
A computer program utilizing a machine learning model for vector conversion, database search, and language model generation to enhance response accuracy by generating and refining answers based on user questions.
The method provides improved response accuracy in chatbot conversations by leveraging vector values, database searches, and language models to generate precise and relevant answers.
Smart Images

Figure 2025149944000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for supporting a conversation between a chatbot and a user. [Background technology]
[0002] Conventionally, chatbots that automatically generate conversational sentences based on technologies such as machine learning and natural language processing have been introduced as a method of supporting conversations, including questions and answers, with users of specific services.
[0003] For example, Patent Document 1 discloses a technology in which prompts that specify instructions for automatically generating conversations are configured to include instructions describing the character and conversation samples, so that a chatbot that converses with a user can maintain a consistent dialogue as a character. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0005] However, although the technology disclosed in Patent Document 1 discloses a technology for maintaining the persona of a chatbot, there is a growing need to improve the accuracy of responses as conversation content.
[0006] Therefore, an object of the present invention is to provide a method for supporting conversations with users using a chatbot, which method realizes improved response accuracy. [Means for solving the problem]
[0007] A computer program according to one aspect of the present invention causes a computer to function as a receiving means for receiving a user question, a vector value generating means for vector converting the user question using a machine learning model for vector conversion to generate a vector value, an approximate value data acquiring means for inputting the vector value into at least one of a plurality of pre-set databases and acquiring data corresponding to an approximate value of the detected vector value, an answer sentence generating means for generating an answer using a language model based on the data corresponding to the approximate value, and a transmitting means for transmitting the answer to the user terminal. [Effects of the Invention]
[0008] According to the present invention, a method for supporting conversation with a user using a chatbot can be provided, which method realizes improved response accuracy. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a system for supporting conversations between a chatbot and a user according to a first embodiment of the present invention.
[0023] FIG. [Figure 2] FIG. 2 is a functional block diagram showing the server terminal 100 of FIG. [Figure 3] FIG. 2 is a functional block diagram showing the user terminal 200 of FIG. [Figure 4] FIG. 2 is a diagram showing an example of user data stored in the server 100. [Figure 5] FIG. 2 is a diagram showing an example of chatbot data stored in the server 100. [Figure 6] 1 is a flowchart illustrating an example of a method for assisting a chatbot in a conversation with a user according to a first embodiment of the present invention. [Figure 7] 10 is a flowchart illustrating another example of a method for assisting a conversation between a chatbot and a user according to the first embodiment of the present invention. [Figure 8] 10 is a flowchart illustrating yet another example of a method for assisting a chatbot in a conversation with a user according to the first embodiment of the present invention. [Figure 9] 10 is a flowchart illustrating yet another example of a method for assisting a chatbot in a conversation with a user according to the first embodiment of the present invention. [Figure 10] 10 is a flowchart illustrating yet another example of a method for assisting a chatbot in a conversation with a user according to the first embodiment of the present invention. [Figure 11] 10 is a flowchart illustrating yet another example of a method for assisting a chatbot in a conversation with a user according to the first embodiment of the present invention. [Figure 12] 10 is a flowchart illustrating yet another example of a method for assisting a chatbot in a conversation with a user according to the first embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments described below do not unduly limit the content of the present invention described in the claims. Furthermore, not all of the components shown in the embodiments are necessarily essential components of the present invention.
[0011] <Configuration> FIG. 1 is a block diagram showing a system for supporting conversations with users using a chatbot according to a first embodiment of the present invention. This system 1 includes a server terminal 100 that receives, for example, a question about a user's specific product or service from a user terminal or from another terminal via an API (Application Programming Interface), analyzes the question, generates a response to the question using an API server (hereinafter referred to as an LLM or API server) 300 capable of executing a large-scale language model (LLM), and executes a process of outputting the response to the question via a chatbot. The system also includes the LLM 300 and user terminals 200A and 200B managed by the user. For convenience of explanation, each terminal is described as a single terminal or a specific number of terminals, but the number of each terminal is not limited. The server terminal 100 may be managed by a service provider that provides specific products and / or services to users, or by a chatbot service provider that operates a chatbot service in cooperation with the service provider. A large-scale language model is a machine learning model for natural language processing that has parameters trained using large-scale training data and has the function of generating sentences based on input information. For example, a large-scale language model is a machine learning model based on a Transformer model with a self-attention mechanism. An example of a Transformer model is a GPT (Generative Pre-trained Transformer) model. By training a large-scale language model using a large dataset, it becomes possible to perform not only general natural language generation tasks but also highly specialized sentence generation tasks. For example, a model trained using a large dataset using the GPT-3 or GPT-4 algorithm, or a model combining these with reinforcement learning, can be applied as a large-scale language model.
[0012] The server terminal 100 and the user terminals 200A and 200B are connected to each other via a network NW1, which may include the Internet, an intranet, a wireless LAN (Local Area Network), or a WAN (Wide Area Network).
[0013] The server terminal 100 may be, for example, a general-purpose computer such as a workstation or a personal computer, or may be configured as a cloud computing system, a cluster or multicomputer consisting of multiple computers, a virtual machine constructed virtually by software, or a quantum computer.
[0014] The user terminals 200A and 200B are, for example, information processing devices such as personal computers and tablet terminals, but may also be configured as smartphones, mobile phones, PDAs, or the like.
[0015] In this embodiment, the system 1 is described as having a configuration including a server terminal 100 and user terminals 200A and 200B, with users of each terminal using their respective terminals to operate the server terminal 100, but the server terminal 100 may be configured as a standalone unit, with the server terminal itself provided with a function for each user to directly operate it. Hereinafter, where necessary, the user terminals 200A and 200B will be collectively referred to as user terminal 200.
[0016] Fig. 2 is a functional block configuration diagram of the server terminal 100 of Fig. 1. The server terminal 100 includes a communication unit 110, a storage unit 120, and a control unit .
[0017] The communication unit 110 is a communication interface for communicating with the user terminal 200, the user terminal 300, the LLM 300, etc. via the network NW1, and communication is performed according to a communication protocol such as TCP / IP (Transmission Control Protocol / Internet Protocol).
[0018] The storage unit 120 stores input data, programs for executing various control processes and functions in the control unit 130, and is composed of RAM (Random Access Memory), ROM (Read Only Memory), etc. The storage unit 120 also has a user data storage unit 121 that stores various data related to users, and a product data storage unit 122 that stores information about products registered by users. A database (not shown) that stores various data may be constructed outside the storage unit 120 or the server terminal 100.
[0019] The control unit 130 controls the overall operation of the server terminal 100 by executing a program stored in the storage unit 120, and is composed of circuits such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The functions of the control unit 130 include an information receiving unit 131 that receives information from each user terminal, and an information processing unit 132 that processes the information. The information receiving unit 131 and the information processing unit 132 are activated by a program stored in the storage unit 120 and executed by the server terminal 100, which is a computer (electronic calculator).
[0020] The information receiving unit 131 receives information from the user terminal 200 via the communication unit 110. For example, a question regarding a predetermined product and / or service is received from the user terminal 200 via a chat interface displayed on the user terminal 200.
[0021] The information processing unit 132 performs predetermined information processing, such as analyzing a question received from the user terminal 200 and generating an answer to the question. The server terminal 100 can also perform predetermined processing in cooperation with the LLM 300 or another API server connected via an API. The LLM 300 and the API server may be the same server. For example, the LLM 300 may be a server that can execute, via an API, not only a sentence generation function that generates sentences using a large-scale language model, but also a vector conversion function that vectorizes input sentences, a question splitting function that splits a question included in an input sentence to generate multiple split questions, and a function calling function that defines functions to be called based on the input sentence (for example, database connection processing or calculation processing).
[0022] The control unit 130 may also have a screen generation unit (not shown), which generates screen information to be displayed via the user interface of the user terminal 200 upon request. For example, the control unit 130 generates a user interface by using image and text data (not shown) stored in the storage unit 120 as material and arranging various images and text in predetermined areas of the user interface based on predetermined layout rules. Processing related to the image generation unit may also be executed by a GPU (Graphics Processing Unit).
[0023] Fig. 3 is a functional block diagram showing the user terminal 200 of Fig. 1. The user terminal 200 includes a communication unit 210, a display operation unit 220, a storage unit 230, and a control unit 240.
[0024] The communication unit 210 is a communication interface for communicating with the server terminal 100 via the network NW, and communication is performed according to a communication standard such as TCP / IP.
[0025] The display operation unit 220 is a user interface used by the provider to input instructions, analyze the input data according to the input data such as text, audio, and images from the control unit 240, and display the text, audio, and images as output, and is composed of a display, keyboard, and mouse when the user terminal 200 is configured as a personal computer, and is composed of a touch panel, etc. when the user terminal 200 is configured as a smartphone or tablet terminal. This display operation unit 220 is started up by a control program stored in the storage unit 230 and executed by the user terminal 200, which is a computer (electronic calculator).
[0026] The storage unit 230 stores input data, programs for executing various control processes and functions in the control unit 240, and is composed of RAM, ROM, etc. The storage unit 230 also temporarily stores the contents of communication with the server terminal 100.
[0027] The control unit 240 controls the overall operation of the user terminal 200 by executing the programs stored in the storage unit 230, and is composed of circuits such as a CPU and a GPU.
[0028] FIG. 4 is a diagram showing an example of user data stored in the server 100. As shown in FIG.
[0029] The user data 1000 shown in Fig. 4 stores various data related to a user. For convenience of explanation, Fig. 4 shows an example of a user identified by a user ID "10001", but information on multiple users can be stored. Note that the user ID is an example per user. The various data related to a user can include, for example, basic information about the user (e.g., user name, company name (personal name), date of birth (age), address, name of person in charge, department name, job title, contact information (email address, phone number), used services / products, etc.) and conversation information (e.g., conversation history via a chat interface, etc.), but is not limited to this.
[0030] 5 is a diagram showing an example of chatbot data stored in the server 100. The chatbot data is stored in the chatbot data storage unit 122 of the server 100, for example.
[0031] The chatbot data 2000 is data including various data related to the operation of the chatbot. The various data related to the operation of the chatbot can include, but are not limited to, questions, answers, conversation history data (e.g., natural language and / or vectorized data and datasets of natural language), and prompt data (e.g., conditional statements for controlling answer sentences generated by a language model). A feature of this embodiment is that, for each user (business) and / or dataset, a dataset for generating appropriate answer sentences to question sentences is recorded, and a prompt for controlling the answer sentences generated by a language model is also recorded.
[0032] <Processing flow> The flow of processing of a method for generating a response sentence by a chatbot in response to a question sentence by a user, which is executed by the system 1 of this embodiment, will be described with reference to the figures from Figure 6 onwards. Figure 6 is a flowchart showing a basic flow as an example of a method for supporting a conversation with a user by a chatbot, according to the first embodiment of the present invention.
[0033] First, in the process of step S101, the user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200. The information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0034] Next, in step S102, the information processing unit 132 performs a predetermined preprocessing process (described later) on the question, and transmits the preprocessed question data to the API server (LLM 300). The API server then vectorizes the question to generate a vector value. Note that generating a vector value by vectorizing the question is also called embedding processing, and is processing in which a vector value determined based on the relevance between texts is obtained using a trained machine learning model.
[0035] Next, in the process of step S103, the information processing unit 132 performs a search process to acquire an approximate value that approximates the vector value from a database (vector DB) connected via the API server or directly, based on the vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the memory unit 120 of the server terminal 100. Here, for data corresponding to approximate values extracted as vector search results, the information processing unit 132 can filter the vector search results by inputting the user's question, multiple search results, and a prompt message confirming whether the search results are relevant to the user's question to the API server (LLM300). For example, if 20 pieces of data corresponding to approximate values are extracted as vector search results by the LLM300, it is possible to identify three pieces of data that are highly relevant to the question and perform a process of generating an answer message based on the three pieces of data, or to perform a process of searching for approximate values that are close to the vector value based on the three pieces of data. Furthermore, to improve the accuracy of answers, the vector DB vectorizes pairs of question and answer sentences, enabling multifaceted semantic searches from both the question and answer sentences. For example, in the context of the story of Momotaro, the question "What did Grandma find?" and the answer "She found a big peach while doing laundry in the river" are vectorized. This enables multifaceted semantic searches from both the question and answer sentences for questions such as "What was Grandma doing?", "What were you doing in the river?", "When did you find the big peach?", and "Where did you do laundry?". As a result of the vector search, the extracted data may contain paired data of a question and answer sentence, or it may contain only the question or answer sentence. A prompt can be used to instruct the generation of an answer sentence so that the paired data is recognized as important information, the answer sentence alone as second most important information, and finally the question alone. In this example, a database of question sentences and a database of answer sentences may be provided, and vector searches may be performed on the question sentences and answer sentences for each. Furthermore, when vectorizing a question and answer pair, if the number of characters is large, the matching rate with the question tends to decrease, so it is also possible to summarize the answer and vectorize the summarized answer pair with the question.
[0036] Next, in the process of step S104, the information processing unit 132 transmits data corresponding to the approximate value acquired as the search result and an answer generation command (prompt sentence) to the language model (LLM300). The language model (LLM300) generates an answer sentence to the user's question sentence based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300. The server terminal 100 transmits the answer sentence to the user terminal 200. The information processing unit 132 stores the generated answer sentence in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120 as user data 1000 and / or chatbot data 2000.
[0037] FIG. 7 is a diagram illustrating an example of deepening the vector search results in the above basic flow.
[0038] First, in the process of step S201, when a user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200, the information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0039] Next, in the process of step S202, the information processing unit 132 performs a predetermined preprocessing, which will be described later, on the question, and transmits the preprocessed question data to the API server (LLM 300). Next, the API server vectorizes the question and generates a vector value.
[0040] Next, in the process of step S203, the information processing unit 132 performs a search process to acquire an approximate value that approximates the vector value from a database (vector DB) connected via the API server or directly, based on the vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the storage unit 120 of the server terminal 100.
[0041] Next, in step S204, the information processing unit 132 generates a primary answer to the user's question based on data corresponding to the approximate value acquired as the search result. In this example, the server terminal 100 can perform a further search based on the primary answer, thereby deepening the vector search. Specifically, the information processing unit 132 generates a secondary question based on the acquired primary answer data, vectorizes the secondary question, and then executes the vector search process again. Here, the information processing unit 132 transmits data (e.g., TOP 5) that is highly relevant to the question among the generated multiple primary answer data and an answer generation command (prompt sentence) to the language model (LLM300), which can then generate a final answer sentence.
[0042] Next, in the process of step S205, the information processing unit 132 generates a secondary answer to the user's question based on data corresponding to the approximate value acquired as the re-search result. As an example, the information processing unit 132 transmits data (e.g., TOP 5) that is highly relevant to the question from among the multiple secondary answer data generated based on the primary answer and an answer generation command (prompt sentence) to the language model (LLM300). The language model (LLM300) generates an answer sentence to the user's question based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300 as a final answer. The server terminal 100 transmits the answer sentence to the user terminal 200. The information processing unit 132 stores the generated answer sentence in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120 as user data 1000 and / or chatbot data 2000. Here, the vector DB may refer to a database common to the primary and secondary searches, or different databases may be referenced for the primary and secondary searches. Also, by generating the primary answer (and secondary answer) in advance and re-searching the vectors in this way, highly accurate answers can be prepared, and the final answer sentence can be generated in one go by the LLM, thereby saving tokens and speeding up processing.
[0043] For example, after acquiring a question sentence "What did Grandma find?" from the user, the information processing unit 132 searches a database of answer sentences to extract vector values similar to the vector values based on the question sentence from the user, and generates a sentence "peach" corresponding to the similar vector value as a primary answer. Furthermore, the information processing unit 132 searches a database of question sentences to extract vector values similar to the vector value based on the primary answer "peach," extracts questions "How big is the peach?" and "What does the peach sound like when it floats by?" which are sentences corresponding to the similar vector values, and sets these as secondary questions. Furthermore, the information processing unit 132 searches a database of answer sentences to extract vector values similar to the vector value based on the secondary question, and extracts sentences "A big peach" and "A peach floated by, floating down the river" which are sentences corresponding to the similar vector values, as secondary answers. Then, the user question "What did the old lady find in the river?", the primary answer "peach", the secondary answers "a big peach" and "a peach floated down the river", and a prompt sentence that generates an answer to the user question based on the user question and the obtained answer are sent to the language model (LLM300), and the language model generates a final answer sentence. In other words, by generating an answer based on the answer obtained based on the user's question sentence and the answer obtained based on a new question sentence generated based on the user's question sentence, it is possible to improve the accuracy of the answer.
[0044] Furthermore, the information processing unit 132 may generate multiple different new questions (e.g., questions that conceptualize the user's question in multiple directions) based on the user's question, obtain answers using different databases for each question, and then generate a final answer. For example, the information processing unit 132 inputs into the language model a prompt statement including the user's question ("What was born from the peaches that the grandmother picked up in the river?"), an instruction to generate a sentence that abstractly conceptualizes the user's question, and an instruction to generate a sentence that anti-conceptualizes the user's question, and generates a new question ("What was born?", "What was the grandfather doing?"). The information processing unit 132 extracts answers for each question by referring to different databases that are pre-associated with each other (for example, a vector search is performed in a default answer database for the user's question, an answer database for abstract concepts for the abstract-conceptualized question, and an answer database for anti-conceptualized question, and the corresponding answers are extracted). The information processing unit 132 then inputs a prompt sentence containing the user's question, a new question generated based on the user's question, answers corresponding to each question, information on the conditions for generating the newly generated questions (for example, information indicating that the newly generated first question is an abstract conceptualization of the user's question and that the newly generated second question is an anti-conceptual conceptualization of the user's question), and a command to generate a final answer based on each question and the corresponding answer to the language model, and obtains the generated final answer. In other words, by generating a new question from a different perspective based on the user's question and generating an answer based on answers obtained from the database corresponding to each perspective, it is possible to improve the accuracy of the answer.
[0045] Furthermore, in the above-described answer generation process, instead of a primary answer and a secondary answer, multiple answers to a question can be generated in advance, and then a final answer can be generated. As an example, based on a user's question, a search process can be performed in parallel with reference to multiple vector DBs, and multiple answers can be generated based on data corresponding to approximate values extracted from each vector DB. To determine which vector DB to reference, preprocessing can be performed by extracting features of the user's question. For example, assuming that there is a vector DB related to the story of Momotaro and a vector DB related to user reviews, a vector search can be performed on the Momotaro story DB based on the entire question, and a vector search on the user review DB based only on the feature term "Momotaro." Furthermore, for an answer generated based on data corresponding to approximate values obtained from each vector DB, vector DB answer information set for each vector DB can be input to the LLM 300 along with the generated answer, and an answer can be generated. For example, by inputting information such as "The Momotaro Story DB is a DB that explains the story of Momotaro, and the User Review DB is a DB that shows the opinions of other users" into the LLM 300, the LLM 300 can generate an answer with an explanation of the story in the first half and "Other users think that..." in the second half. If this example is applied to a system that clearly explains how to read a securities chart system, vector DBs such as 1) a call center customer service history DB, 2) a financial trading system manual DB, 3) a screen transition flow DB, 4) a flow detail and detailed screen flow DB, and 5) a dictionary DB of symbols used on each screen can be prepared. For a question, a vector search can be performed on 1) the call center customer service history DB, and based on the primary answer generated based on the search results, a second search can be performed on 2) the financial trading system manual DB, and then a second search can be performed on the vector DBs of 3), 4), and 5) in sequence to generate an answer (or, for the answer in 1), a search can be immediately performed on the vector DB in 5). This allows for the generation of a highly accurate answer.In addition, the Vector DB can also complement knowledge by combining it with the following purposes and uses: 1) a DB that provides a purpose (e.g., a past history DB (main DB)), 2) a DB that provides knowledge (e.g., a manual DB (sub-DB)), 3) a DB that provides meaning (sub-DB), and 4) a supplementary DB (e.g., a dictionary DB (sub-DB)).
[0046] For example, the information processing unit 132 may generate a plurality of new questions using a language model based on the user's question sentence, and determine a database to be used for vector search for each question from among a plurality of databases.
[0047] For example, the information processing unit 132 inputs a prompt sentence including a user's question, a command for estimating a purpose based on the user's question, and a command for estimating a use based on the user's question into a language model, thereby generating a purpose and use of the user's question. Furthermore, the information processing unit 132 obtains a corresponding answer regarding the purpose by performing a vector search on a database related to purposes for the generated purpose, and obtains a corresponding answer regarding the use by performing a vector search on a database related to uses for the generated use. Furthermore, the information processing unit 132 obtains a corresponding answer regarding knowledge and an answer regarding meaning by performing a vector search on a database related to meaning based on the user's question, the answer regarding the purpose, and the answer regarding the use. The information processing unit 132 obtains a final answer by inputting a prompt sentence including a user's question, an answer regarding the purpose (business purpose), an answer regarding the use (destination of use), an answer regarding knowledge (business content), an answer regarding the meaning (explanation of the wording), and an instruction for generating a final answer taking into account the characteristics of each answer into the language model. In other words, new questions are generated based on the user's question from different perspectives, such as purpose, use, knowledge, and meaning, and answers are generated based on answers obtained from databases corresponding to each perspective, thereby improving the accuracy of answers.
[0048] FIG. 8 is a diagram illustrating an example of executing preprocessing for a user's question in the above basic flow.
[0049] First, in the process of step S201, when a user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200, the information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0050] Next, in step S302, the information processing unit 132 performs predetermined preprocessing on the question. Specifically, the user's question is divided into multiple questions. For example, in a topic related to Momotaro, if the user's question is "What were your grandparents doing?", the information processing unit 132 divides the question into "What was your grandmother doing?" and "What was your grandfather doing?" For example, the LLM 300 (API server) determines whether to divide the question and generates the divided questions, and the information processing unit 132 acquires the generated divided questions.
[0051] Next, in the process of step S303, the plurality of question sentence data acquired by the above preprocessing is transmitted to the API server (LLM 300). Next, the API server vectorizes each question sentence to generate a vector value.
[0052] Next, in the process of step S304, the information processing unit 132 performs a search process to acquire approximate values that approximate the vector values from a database (vector DB) connected via the API server or directly, based on each vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the memory unit 120 of the server terminal 100.
[0053] For example, the information processing unit 132 may use a different database for vector search for each divided question. For example, the information processing unit 132 may be configured to divide a question into a plurality of perspectives using a prompt sentence, which includes a user's question sentence and a predetermined command to divide the user's question sentence into a plurality of perspectives, in a language model, and perform a vector search using a database predetermined for each perspective. This allows a different reference destination for each perspective, which is expected to further improve accuracy.
[0054] Next, in the process of step S305, the information processing unit 132 transmits data corresponding to the approximate value acquired as the search result and an answer generation command (prompt sentence) to the language model (LLM300). The language model (LLM300) generates an answer sentence to the user's question sentence based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300. The server terminal 100 transmits the answer sentence to the user terminal 200. The information processing unit 132 stores the generated answer sentence in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120 as user data 1000 and / or chatbot data 2000. Here, when splitting a question, it is possible to do so using the Function Calling function. For example, when a user asks, "How many meters is Tokyo Tower taller than Skytree?", a tool that determines which function should be selected from the question can split the question into two questions: "How many meters is Tokyo Tower tall? How many meters is Skytree tall?" and "How many meters is Skytree tall?". The answer "Tokyo Tower is Xm, Skytree is Ym" is then obtained from the "Height DB." Based on this answer, the "Calculation Tool" determines that "Ym - Xm is Zm," and the LLM200 can combine these answers to generate the final answer.
[0055] FIG. 9 is a diagram illustrating an example of strengthening user questions in the basic flow.
[0056] First, in the process of step S401, when a user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200, the information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0057] Next, in step S402, the information processing unit 132 performs predetermined preprocessing on the question and transmits the preprocessed question data to the API server (LLM 300). Next, the API server vectorizes the question and generates a vector value.
[0058] Next, in the process of step S403, the information processing unit 132 performs a search process to acquire an approximate value that approximates the vector value from a database (vector DB) connected via the API server or directly, based on the vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the memory unit 120 of the server terminal 100.
[0059] Next, in step S404, the information processing unit 132 acquires a similar question (related question) as data corresponding to the approximate value acquired as the search result, and transmits to the language model (LLM 300) an instruction to generate a new question (second question) based on the user question and the similar question, and an answer generation command (prompt sentence). For example, if the initial user question is "Why did you cut (the peach)?" and the search results are "1) the grandmother cut the peach open, and 2) a boy came out after cutting it open," the information processing unit 132 determines the user's intention as to what kind of question the user wants to ask, and then generates a new question, "Why did you cut the peach open?" Typically, data held by a company is data with few questions (a DB centered on answers). In such cases, strengthening the questions can realize a chatbot that can respond flexibly. Here, the information processing unit 132 can refer to the chatbot data 2000 and sort the vector search results based on the number of times the relevant question has been selected (clicked) by the user or other users, or the number of times it has been read out in the past, as related questions generated based on answers to user questions. Furthermore, as an extension of memory, past question contents are stored in the memory unit 120 as chatbot data 2000, and a prompt statement can be used to instruct the LLM 300 to generate a user question that combines the past user question with the current user question. This allows the LLM 300 to generate a new question, perform a vector search, and generate an answer. This enables a vector search based on the context of the current user question. For example, if a past question was "Tell me about XX for product A," and the current question is "Tell me about △ as well," the question can be "Tell me about △ for product A," and then an answer can be generated. Furthermore, when multiple past question or answer statements are stored to assist with user questions, since it is difficult to store long statements due to memory capacity limitations, the LLM 300 summarizes pairs of user questions and answers, It is also possible to assist with user questions based on summary memory. Alternatively, the most recent memory can be stored as a full sentence, the Xth previous memory as a summary memory, and the conversation before the summary can be stored in a separate vector DB, making the vector searchable. By storing summarized question and answer sentences, for example, when suggesting (asking) a user based on three questions (history), a separate LLM (e.g., an LLM for dividing the route, an LLM for saving data, and an LLM for asking a question) can be activated for each question, and personalized answers can be generated based on each LLM and provided to the user. In addition, to compress the conversation content (expand the memory area), the question and answer sentences can be overwritten and stored as a single sentence each time the summary is repeated.
[0060] Additionally, in the process of step S404, the information processing unit 132 preferably uses different databases for vector search of the user's question and vector search of the generated new question (second question). For example, it is preferable to obtain answers by using a simple database that collects pairs of simple questions and simple answers for vector search of the user's question and a database that collects detailed questions and detailed answers for vector search of the new question. Because user questions are often ambiguous or short, errors tend to be large in vector comparison with detailed questions. Therefore, improved accuracy can be expected by using a simple database for vector search of the user's question and a detailed database for vector search of the new question that includes various additional information. Note that simple questions have a smaller amount of text than detailed questions. Furthermore, for each of multiple new questions generated based on the user's question, a database to be used for vector search from among multiple databases may be determined for each question.
[0061] For example, the information processing unit 132 may acquire similar questions (related questions) similar to the user's question by vector searching a question sentence database, and send an instruction (prompt sentence) to the language model (LLM300) to generate multiple new questions based on the user's question and similar questions. At this time, for each generated question, a vector search is performed in a predetermined database based on the question type and attributes, and a sentence corresponding to the extracted vector value is acquired as an answer. Note that the question type and attributes may be types and attributes previously associated with each question in the question sentence database, or the language model may determine the question type and attributes for each acquired similar question. The question type is, for example, which of the 5W1H the question is about. The question attributes are, for example, information indicating the difficulty level of the question or the target of the question.
[0062] For example, the information processing unit 132 may generate a new question based on the user's question and recent memory (questions previously input by the user and answers to those questions) as assistance to the user based on summary memory. When the user's first question is "Does Momotaro have any dogs?" and the language model outputs "There are dog friends" as the answer to the first user's question, the language model stores the sentence "There are dog friends" that summarizes the user's question and answer as summary memory. When the user's second question is "What are Momotaro's other friends?", the information processing unit 132 sends a prompt sentence including the user's question, the summary memory, and an instruction to generate a new question based on the user's question and summary memory to the language model, and obtains a new question (e.g., "What are Momotaro's other friends other than dogs?") based on the user's question and summary memory. The information processing unit 132 then performs a vector search in the answer sentence database for each of the user's second question sentence and the new question sentence to obtain answers corresponding to approximate vector values, and transmits a prompt sentence including these answers and an instruction to generate a final answer based on these answers to the language model, thereby obtaining a final answer. This allows for improved answer accuracy using summary memory.
[0063] Next, in the process of step S405, the language model (LLM300) generates an answer sentence to the user's question sentence based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300. The server terminal 100 transmits the answer sentence to the user terminal 200. The information processing unit 132 stores the generated answer sentence in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120 as user data 1000 and / or chatbot data 2000.
[0064] FIG. 10 is a diagram illustrating an example of personalization using extended memory in response to a user question in the basic flow.
[0065] First, in the process of step S501, when a user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200, the information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0066] Next, in step S502, the information processing unit 132 performs predetermined preprocessing on the question and transmits the preprocessed question data to the API server (LLM 300). Next, the API server vectorizes the question and generates a vector value.
[0067] Next, in the process of step S503, the information processing unit 132 performs a search process to acquire an approximate value that approximates the vector value from a database (vector DB) connected via the API server or directly, based on the vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the storage unit 120 of the server terminal 100.
[0068] Next, in step S504, the information processing unit 132 transmits a command (prompt sentence) for generating an answer based on the user's persona to the language model (LLM 300) as data corresponding to the approximate value acquired as the search result. As part of the extended storage, the user data storage unit 121 stores user attributes (e.g., gender, age) as user data 1000 from data acquired from the user or the user's past questions. A user's question, "What cosmetics do you recommend?" can be corrected to "What cosmetics do you recommend for women in their 20s?" based on the user's attributes, thereby generating an answer that matches the user's persona. Here, user attribute data alone may not be sufficient for personalization. Therefore, in response to a user's question, not only definitive user information but also definable user persona data (e.g., "Hair type: natural curls, hairstyle: long hair, desired hairstyle: straight hair, budget: under 500 yen, concern: people messing with my natural curls") based on the user's question (e.g., "I'm bothered by people messing with my natural curls at school today. Is there a spray that can straighten long hair for around 500 yen?") can be transmitted to the LLM 300. A final answer can be generated by further instructing the LLM 300 to integrate the content of the answers generated by the LLM 300 and then to combine the content of the answers generated by the LLM 300.
[0069] Next, in step S505, the language model (LLM300) generates an answer sentence to the user's question sentence based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300. The server terminal 100 transmits the answer sentence to the user terminal 200. The information processing unit 132 stores the generated answer sentence in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120 as user data 1000 and / or chatbot data 2000.
[0070] For example, when the information processing unit 132 fills in the user's persona from the user's input according to a prepared template, the information processing unit 132 may generate a question to ask about the persona items that are not completely filled in. This allows the user's persona to be stored appropriately according to the template.
[0071] FIG. 11 is a flowchart illustrating an example of processing for improving the user experience in answer output in the basic flow.
[0072] First, in the process of step S601, when a user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200, the information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0073] Next, in step S602, the information processing unit 132 performs predetermined preprocessing on the question and transmits the preprocessed question data to the API server (LLM 300). Next, the API server vectorizes the question and generates a vector value.
[0074] Next, in the process of step S603, the information processing unit 132 performs a search process to acquire an approximate value that approximates the vector value from a database (vector DB) connected via the API server or directly, based on the vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the storage unit 120 of the server terminal 100.
[0075] Next, in step S604, the information processing unit 132 transmits data corresponding to the approximate value acquired as the search result and an answer generation command (prompt sentence) to the language model (LLM300). The language model (LLM300) generates an answer sentence to the user's question sentence based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300. Here, when the server terminal 100 acquires the answer sentence generated by the LLM300, the characters constituting the answer sentence are transmitted to the server terminal 100, for example, gradually, sentence by sentence, via streaming. If the server terminal 100 acquires the entire sentence and transmits it to the user terminal 200, it takes a long time to acquire the answer, resulting in a poor user experience. This phenomenon is particularly noticeable when the server terminal 100 and the LLM300 do not communicate directly, but rather the server terminal 100 communicates with the LLM via another application (e.g., an app such as LINE (registered trademark)) via an API.
[0076] Therefore, in the processing of step S605, the information processing unit 132 determines whether the sentences that constitute part of the answer text output from the LLM 300 are coherent sentences, for example, based on punctuation marks, number of characters, etc., and determines whether they meet the conditions for transmitting the sentences as coherent sentences to the user terminal 200.
[0077] Therefore, in the processing of step S606, if the information processing unit 132 determines in the processing of step S606 that the sentence received by the server terminal 100 from the LLM 300 is a coherent sentence and thus meets the transmission conditions, it transmits the sentence received from the LLM 300 as part of a response sentence to the user terminal 200. On the other hand, if the information processing unit 132 determines that the sentence received from the LLM 300 does not meet the transmission conditions, that is, the received sentence does not reach a predetermined number of characters or does not contain punctuation, etc., it waits to receive a response from the LLM 300 until the transmission conditions are met. Here, the check (for example, NG check) of the answer sentence to be transmitted from the server terminal 100 to the user terminal 200 can be performed for each streaming transmission, rather than being performed after the server terminal 100 receives the entire answer sentence from the LLM 300. Also, when converting an output sentence into video or audio, it is preferable to perform parallel processing for each output. For example, while a process of converting a certain output text into video or audio is being performed, a process of receiving another output text is being performed in parallel, thereby shortening the response time to the user terminal 200.
[0078] FIG. 12 is a flowchart illustrating an example of data search processing using SQL in the basic flow.
[0079] First, in the process of step S701, when a user inputs a question about a specific topic (including a question about a product and / or a service (hereinafter referred to as "service")) via a chat interface implemented in a web browser or application of the user terminal 200, the information receiving unit 131 of the control unit 130 of the server terminal 100 receives the question input by the user from the user terminal 200 via the communication unit 110. The information processing unit 132 of the control unit 130 stores the received question as user data 1000 in the user data storage unit 121 and / or chatbot data storage unit 122 of the memory unit 120.
[0080] Here, if the query contains numeric data, filtering of numeric values and the like cannot be performed in vector search, so in step S702, the information processing unit 132 causes the LLM 300 to reference an SQL server (not shown) and generate a select statement for acquiring the desired information based on the numeric values contained in the query, and then performs processing to acquire the desired information from the SQL server. If the column information is known in advance, the select statement can be generated by inputting the column information into a prompt statement, or the LLM can independently generate a prompt statement that determines which SQL to reference and how to reference it to acquire the desired information, and then generate the select statement after the LLM has initially understood the SQL data structure.
[0081] Next, in step S703, the information processing unit 132 performs predetermined preprocessing on the question and transmits the preprocessed question data to the API server (LLM 300). Next, the API server vectorizes the question and generates a vector value.
[0082] Next, in the process of step S704, the information processing unit 132 performs a search process to acquire an approximate value that approximates the vector value from a database (vector DB) connected via the API server or directly, based on the vector value generated by the API server (LLM300). Here, the vector DB may be a database connected to an external device via the API server, or may be a database managed as chatbot data 2000 in the memory unit 120 of the server terminal 100.
[0083] Next, in the process of step S705, the information processing unit 132 transmits data corresponding to the approximate value acquired as the search result and an answer generation command (prompt sentence) including an instruction to generate the select sentence to the language model (LLM300). The language model (LLM300) generates an answer sentence to the user's question sentence based on the prompt sentence, and the server terminal 100 acquires the answer sentence generated by the LLM300.
[0084] As described above, according to this embodiment, when a user question is vectorized, a vector search is performed in a database, and an answer sentence is generated using the vector search results, the accuracy of the chatbot's answer can be improved by performing preprocessing on the user question or the vector search results.
[0085] Although the embodiments of the present invention have been described above, they can be embodied in various other forms, and various omissions, substitutions, and modifications can be made. These embodiments, modifications, and omissions, substitutions, and modifications are included in the technical scope of the claims and their equivalents. [Explanation of symbols]
[0086] 1 System 100 Server terminal, 110 Communication unit, 120 Storage unit, 130 Control unit, 200 User terminal, 300 LLM, NW1 Network
Claims
1. A program executed on a computer, On the computer, obtaining a user question; A step of obtaining a vector value generated by vector-converting the user question using a machine learning model for vector conversion; inputting the vector value into at least one of a plurality of preset databases to obtain data corresponding to an approximate value of the detected vector value; generating an answer using a language model based on data corresponding to the approximation; transmitting data for displaying the answer on the user terminal; A computer program that executes the following:
2. The computer, moreover, generating new questions from the user questions; A step of vector-converting the new question using a machine learning model for vector conversion to obtain a vector value generated by the vector conversion; a step of inputting the vector value based on the user question and the vector values based on the new plurality of questions into at least one database out of a plurality of databases preset for each question, and acquiring data corresponding to an approximation of the detected vector value; The program according to claim 1, which executes the above.
3. the step of generating a plurality of new questions from the user question includes inputting the user question and a prompt sentence including an instruction to generate a new question from the user question into a language model; The program according to claim 2.
4. the step of generating a plurality of new questions from the user question includes inputting the user question and a sentence based on a question sentence from a user asked in the past before the user question into a language model; The program according to claim 3.
5. the step of generating new questions from the user question includes inputting a prompt sentence including the user question and an instruction to generate questions corresponding to a plurality of predetermined viewpoints based on the user question into a language model; The new questions are subjected to a vector search using a database corresponding to the predetermined plurality of perspectives. The program according to claim 2.
6. The computer program according to claim 1 , further comprising: generating a summary based on the question and the answer; and storing the generated summary as information indicating the user's past questions and user attributes.
7. 7. The computer program of claim 6, further comprising: retrieving, from the database, data corresponding to an approximation that approximates a vector value generated based on the question and the summary, based on the question; and inputting, into the database, a vector value corresponding to a fifth question generated based on the approximation that approximates a vector value generated based on the question and the summary.
8. The computer program of claim 1 , further comprising: determining a user's persona based on the question; and generating a new question based on the question and the persona.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A
Cited By
Dialogue generation method, device, electronic device, and storage media
JP2025094263A
Dialogue generation method, apparatus, electronic device, and storage medium
JP7911104B2