Automatic response data improvement device, method and program

The automatic response data improvement device addresses the challenge of unclear data input in automated systems by categorizing and analyzing responses, recommending optimal learning data through user evaluation and visualization tools, thereby improving response accuracy and user engagement.

JP2025146322AActive Publication Date: 2025-10-03DTOSH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024047030
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03
Estimated Expiration
2044-03-22

AI Technical Summary

Technical Problem

Existing automated response systems using large-scale language models face challenges in achieving accurate responses due to unclear data input methods, leading to insufficient answer quality and the need for user consultations on suitable learning data.

Method used

An automatic response data improvement device that includes user evaluation, answer evaluation, topic estimation, and evaluation data analysis to categorize and analyze responses, recommending optimal learning data through visualization and simulation tools.

Benefits of technology

Enhances response accuracy by providing users with actionable insights for improving learning data, increasing user motivation and convenience in enhancing response quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025146322000001_ABST
    Figure 2025146322000001_ABST
Patent Text Reader

Abstract

To provide an automatic response data improvement device, method and program which can effectively present information to be a base of determining learning data suitable for an input to a user and has high convenience.SOLUTION: A device for analyzing and presenting information to be a clue for learning data to be recommended for an input in order to improve answer accuracy of an automatic response device 2, and comprises user evaluation means 11, answer evaluation means 12, evaluation data analysis means 13, topic estimation means 14, answer evaluation display means 15, simulation 16 and topic correction means 17. The user evaluation means 11 receives user evaluation. The answer evaluation means 12 generates answer evaluation data obtained by evaluating the quality of an automatic answer. The topic estimation means 14 performs automatic language processing, and performs classification into categories by using topic modeling. The evaluation data analysis means 13 analyzes the quality of the automatic answer in each category.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for improving the accuracy of responses from a device that automatically responds to questions from users by using pre-input learning data and generation AI. [Background technology]

[0002] In recent years, systems that automatically respond to user inputs such as questions using a program have become widespread (see, for example, Patent Document 1). While these systems are sometimes used to communicate information about a company to external parties, such as customers, they are also used by companies to share information about their internal agreements and the products and services they offer within their organizations. Such automated response systems are expected to provide improved response accuracy.

[0003] In addition, services using generative AI are rapidly increasing, and large-scale language models (LLMs) are rapidly advancing as text-oriented AI models. Large-scale language models are natural language processing models built using massive amounts of text data and deep learning technology. Large-scale language models are also used in automated response systems, and companies can operate automated response systems by inputting large amounts of text data they hold into large-scale language models.

[0004] However, the data input into large-scale language models is entered manually, and it is not clear in advance what data and how much should be input. This has led to problems such as insufficient accuracy in answers due to the necessary data not being input. Therefore, currently, providers of automatic response systems provide users with consultations on suitable learning data for input. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-99968 Summary of the Invention [Problem to be solved by the invention]

[0006] In view of this situation, the present invention aims to provide a highly convenient automatic response data improvement device, method, and program that can effectively present users with information that serves as the basis for determining learning data suitable for input. [Means for solving the problem]

[0007] In order to solve the above problems, the automatic response data improvement device of the present invention is a device that improves a device that automatically responds to questions from users by using pre-input learning data and generation AI, and is equipped with: user evaluation means that accepts user evaluations regarding answer data and stores them in a database as user evaluation data; answer evaluation means that generates answer evaluation data that evaluates the quality of an automatic answer using the question data, answer data, and user evaluation data; topic estimation means that performs natural language processing on the question data and answer data, classifies them into categories using topic modeling, and stores them in the database; and evaluation data analysis means that uses the answer evaluation data and categories to analyze the quality of the automatic answer for at least each category. By evaluating the quality of the automated responses based on user evaluations and categorizing and analyzing them, it is possible to effectively analyze information that is useful for improving the automated response data. The analyzed data can be displayed visually on the display of the terminal used by the user, or output as audio, etc. The term "automatic response data" used here includes both question data and the corresponding answer data, which is intended to improve not only the response performance of the automatic response device but also the quality of the question data entered by the user. The user evaluation received by the user evaluation means may be an evaluation such as "Good" or "Bad" or a graded evaluation such as a five-point scale. Alternatively, the means may be configured to receive input of free text and analyze the text using natural language processing technology. The topic modeling used in the topic estimation means is not limited to one, but multiple topic modelings may be used to automatically extract the most suitable topic from the results of the algorithms of each topic modeling.

[0008] The automatic response device to which the automatic response data improvement device of the present invention is applied broadly includes devices that automatically respond to user questions using pre-input learning data and a generation AI. It is preferable that the automatic response device uses Retrieval-Augmented Generation (RAG) technology, but it may also be configured to perform fine tuning, for example. That is, in the case of an automatic response device that uses RAG technology, the pre-input learning data refers to the learning data pre-input into the automatic response device. Furthermore, in the case of an automatic response device that performs fine tuning, the pre-input learning data refers to the learning data that has been retrained by the generation AI. A wide range of data, such as text, images, and audio, can be input as learning data. As for generative AI, a wide range of generative AIs can be used that utilize well-known large-scale language models (LLMs) such as "BERT" (Bidirectional Encoder Representations from Transformers) and "GPT" (Generative Pre-trained Transformer).

[0009] In the automatic response data improvement device of the present invention, the evaluation data analysis means may include a learning data analysis means for analyzing learning data that is recommended for input. By including the learning data analysis means, it becomes possible to specifically present to the user data that is recommended for input as learning data.

[0010] In the automatic response data improvement device of the present invention, the evaluation data analysis means may include question data analysis means for analyzing question data that is recommended to be input. By including the question data analysis means, it becomes possible to specifically present data that is recommended to be input as question data.

[0011] The automatic response data improvement device of the present invention preferably further comprises an answer evaluation display means for visually displaying on the display at least one of the answer success rate for each category over a predetermined period and a list of questions for which the answer success rate is below a predetermined threshold. By visually displaying the answer success rate for each category over a predetermined period, the user can at a glance confirm which category of learning data they should enter, improving convenience. It may also be possible to display graphs of changes in answer success rate compared to the past, or of the answer success rate for each number of questions. By visually displaying a list of questions whose answer success rate is below a predetermined threshold, users can easily check which specific questions are likely to be answered unsuccessfully, improving convenience.

[0012] The automatic response data improvement device of the present invention preferably further comprises a simulation means including a simulation data input means for accepting input of learning data for simulation, a question content selection means for selecting from a list of question content whose answer success rate is below a predetermined threshold, and a preview means for displaying answer content related to the answer data in parallel before and after input of the simulation data. By providing the simulation means, it becomes easier to realize the effect of improving answer performance by inputting appropriate learning data for questions with a low answer success rate, in other words, questions that are likely to be answered incorrectly. This can increase the user's motivation to improve the automatic response data.

[0013] The automatic response data improvement device of the present invention preferably further includes a topic correction means for accepting corrections to the category classification made by the topic estimation means, and stores the corrected category classification in a database in association with the question data and answer data. By including the topic correction means, a user can manually correct the category classification estimated by the topic estimation means. Furthermore, by storing data related to the corrected category classification in the database, it can be used to improve the accuracy of the topic estimation means.

[0014] In the automatic response data improvement device of the present invention, the answer evaluation means may extract a question immediately before or immediately after a question based on user information and the question date and time contained in the question data, and may include a relevance determination means for determining the relevance of the question with the extracted immediately before or immediately after question based on at least one of the question content contained in the question data and the answer content contained in the answer data. The answer evaluation means may use user evaluation data for the immediately before or immediately after question that is determined to be relevant to the question in question to evaluate the answer to the question. Generally, questions asked by the same user in close time intervals are often related questions. Therefore, by using user evaluation data for questions immediately before or immediately after the question in question evaluation, more accurate evaluation is possible. Furthermore, for example, questions asked by users in the same or similar departments or jobs, even if they are not the same user, the immediately before or immediately after questions may be extracted. The relevance determination based on the question data and answer data is performed, for example, by referring to the similarity of keywords extracted using natural language processing.

[0015] In the automatic response data improvement device of the present invention, the answer evaluation means may include grouping means for grouping multiple questions in a question group based on at least one of the user information, question date and time, question content, and answer content contained in the answer data, and may evaluate the answers based on at least one of the type, number, and order of user evaluation data within the group. Generally, when a user asks a question, the question may be repeated multiple times until a sufficient understanding is achieved. Therefore, by regarding multiple questions, regardless of whether they are asked immediately before or after the question in question, as belonging to the same group and utilizing the user evaluation data within the group, more accurate evaluation is possible. When grouping users, the identity of the user and the identity and similarity of the department and work are taken into consideration. The proximity of the question date and time is taken into consideration, and the similarity of keywords extracted using natural language processing is taken into consideration for the question content and answer content. There is no particular limit to the number of questions that make up a group. Furthermore, the handling of user evaluation data within a group may be based on the evaluation trends of the user.

[0016] The automatic response data improvement method of the present invention is a method for improving a device that automatically responds to questions from users by using pre-input learning data and a generation AI, and includes a user evaluation step of accepting user evaluations regarding answer data and storing them in a database as user evaluation data, an answer evaluation step of generating answer evaluation data that evaluates the quality of an automatic response using the question data, answer data, and user evaluation data, a topic estimation step of performing natural language processing on the question data and answer data, classifying them into categories using topic modeling, and storing them in the database, and an evaluation data analysis step of analyzing the quality of the automatic response for at least each category using the answer evaluation data and categories.

[0017] The automatic response data improving program of the present invention causes a computer to execute each step of the automatic response data improving method described above. [Effects of the Invention]

[0018] The automatic response data improvement device, method, and program of the present invention have the advantage of effectively presenting users with information that serves as the basis for determining suitable learning data for input. They also have the advantage of being highly convenient and increasing users' motivation to improve response quality. [Brief explanation of the drawings]

[0019] [Figure 1] Functional block diagram of the automatic response data improvement device according to the first embodiment [Figure 2] System configuration diagram of the automatic response system [Figure 3] Schematic flow diagram of the automatic response data improvement method of Example 1 [Figure 4] User evaluation image [Figure 5] Explanation of topic inference method [Figure 6] Image of improvement of learning data [Figure 7] Response evaluation display image [Figure 8] Recommendation data input simulation image [Figure 9] Functional block diagram of an automatic response data improvement device according to a second embodiment [Figure 10] An explanatory diagram of the response evaluation means of the second embodiment [Figure 11] Functional block diagram of an automatic response data improvement device according to a third embodiment [Figure 12] An explanatory diagram of the response evaluation means of the third embodiment [Figure 13] Illustrative diagram of the evaluation method when grouped DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of the present invention will be described in detail below with reference to the drawings. Note that the scope of the present invention is not limited to the following examples and illustrated examples, and many modifications and variations are possible. [Example]

[0021] 1 shows a functional block diagram of an automatic response data improvement device of Example 1. As shown in FIG. 1, an automatic response system 100 includes an automatic response data improvement device 1, an automatic response device 2, and a generation AI (LLM) 3. The automatic response device 2 is a device that automatically responds to questions from a user (not shown) using pre-input learning data 4a and a generation AI 3, and is equipped with a database 7. As the generation AI 3, a wide range of generation AIs that utilize well-known large-scale language models (LLMs) such as "BERT" and "GPT" can be used.

[0022] The automated answering system 2 uses a search augmentation generation (RAG) technology. RAG technology improves answer accuracy by combining search of external information, such as internal data, with the generation of answer data using a large-scale language model. For example, when the automated answering system 2 is used within a company, internal data is input as training data 4a and stored in database 7. The data stored in database 7 has a vector database structure that stores data as numerical vectors (embedded data) and is managed for efficient search. When a user inputs question data 4b into the automated answering system 2, the automated answering system 2 automatically searches for data (Word, PDF, PPT, Excel, etc.) related to the input question data 4b. The input question data 4b and the automatically searched related data are input to the generation AI 3 via API communication, and the answer results output by the generation AI 3 are displayed on the display of the automated answering system 2.

[0023] The automatic response data improvement device 1 is a device that analyzes and presents information that serves as a clue to learning data that is recommended to be input in order to improve the response accuracy of the automatic response device 2, and is composed of a user evaluation means 11, an answer evaluation means 12, an evaluation data analysis means 13, a topic estimation means 14, an answer evaluation display means 15, a simulation means 16, and a topic correction means 17. The user evaluation means 11 receives user evaluations regarding the response data 4c and stores them in the database 7 as user evaluation data. The answer evaluation means 12 generates answer evaluation data that evaluates the quality of the automatic answer using the question data 4b, the answer data 4c, and the user evaluation data. The topic estimation means 14 performs natural language processing on the question data 4b and answer data 4c, classifies them into categories using topic modeling, links them to the question data 4b and answer data 4c, and stores them in the database 7. The topic correction means 17 accepts corrections to the category classification made by the topic estimation means 14, and stores the corrected category classification in the database 7, linking them to the question data 4b and answer data 4c. The evaluation data analysis means 13 uses the answer evaluation data and categories to analyze the quality of at least the automatic answers for each category, and includes a learning data analysis means 13a and a question data analysis means 13b. The learning data analysis means 13a analyzes learning data that is recommended for input. The question data analysis means 13b analyzes question data that is recommended for input.

[0024] The answer evaluation display means 15 visually displays on a display at least one of the answer success rate for each category during a predetermined period and a list of questions whose answer success rate is below a predetermined threshold. The simulation means 16 includes a simulation data input means 161, a question content selection means 162, and a preview means 163. The simulation data input means 161 accepts input of learning data 4a for simulation. The accepted learning data 4a is transmitted to the automatic answering device 2. The question content selection means 162 selects from a list question content for which the answer success rate is below a predetermined threshold. The preview means 163 accepts answer data 4c for the learning data 4a from the automatic answering device 2, and displays the answer content for the answer data 4c before and after input of the simulation data in parallel.

[0025] FIG. 2 shows a system configuration diagram of an automatic response system. As shown in FIG. 2, the automatic response system 100 is a system in which a server (20, 30) transmits and receives data to and from client terminals such as a PC 51, a smartphone 52, and a tablet terminal 53 via a network 6. The server 20 functions as an automatic response device 2 and an automatic response data improvement device 1. The server 30 functions as a generation AI 3. Learning data 4a and question data 4b can be input using a client terminal such as a PC 51. In addition, the analysis data and recommendation data presented by the automatic response data improvement device 1 are output to a client terminal such as a PC 51 and displayed on a display 51a of the PC 51. The analysis and recommendation data may be presented not only visually but also audibly, such as by voice. In this embodiment, a user uses the automatic response device 2 and the automatic response data improvement device 1 using a PC 51.

[0026] Fig. 3 shows a schematic flow diagram of the automatic response data improvement method of Example 1. As shown in Fig. 3, first, a user evaluation regarding the response data 4c is accepted and stored in the database 7 as user evaluation data (step S01: user evaluation step). FIG. 4 shows an image of a user evaluation. As shown in FIG. 4, an answer 42a generated by the generation AI 3 is displayed in response to a question 41a submitted by a user. Below the answer 42a, a "Good" evaluation button 11a and a "Bad" evaluation button 11b are displayed. After carefully examining the content of the answer 42a to the question 41a, if the user feels that the answer is good, the user clicks the "Good" evaluation button 11a. If the user feels that the answer is bad, the user clicks the "Bad" evaluation button 11b to evaluate the answer. Note that a user evaluation is not required; if neither evaluation button is clicked, the user is treated as having no evaluation. The input user evaluation data is linked to the question data 4b and the answer data 4c and stored in the database 7.

[0027] Next, as shown in Fig. 3, answer evaluation data is generated that evaluates the quality of the automatic answer using the question data, answer data, and user evaluation data (step S02: answer evaluation step). The quality of the automatic answer is evaluated by answer evaluation means 12. In this embodiment, when the user evaluation for the answer data is "Good," the automatic answer is evaluated as successful and given a value of "+1," and when the user evaluation is "Bad," the automatic answer is evaluated as unsuccessful and given a value of "-1." Furthermore, when there is no user evaluation, the automatic answer is evaluated as neither successful nor unsuccessful and given a value of "0."

[0028] Furthermore, natural language processing is performed on the question data 4b and answer data 4c, and the data is classified into categories using topic modeling, and stored in the database 7 (step S03: topic estimation step). Note that the order of steps S02 and S03 does not matter. FIG. 5 is an explanatory diagram of the topic estimation means. As shown in FIG. 5, the database 7 includes a question history database 7a and an answer history database 7b. The question history database 7a stores, for each question, the question ID, question date and time, question content, and account ID. The answer history database 7b stores the answer ID, answer content, rating, related question ID, and category. Taking a question with the question ID "q0000001" as an example, this question is stored in the question history database 7a, and "q0000001" is also stored in the answer history database 7b as a related question ID.

[0029] The topic estimation means 14 performs natural language processing on the text of the question with question ID "q0000001," "Please tell me the industry in which product X sells well," and the answer, "We cannot answer that question because we do not have that information." Specifically, first, each question is divided into words and vectorized, and the frequency of occurrence of each word is recorded (vector conversion). Then, processes such as removing words that have little meaning in the sentence and standardizing the prototype of parts of speech are performed (data preprocessing). After vector conversion and data preprocessing, the topic estimation means 14 classifies each question into a group of similar categories using topic modeling technology. Note that if a topic modeling algorithm is used to make a judgment based on only one algorithm, there is a possibility that the algorithm will be biased. Therefore, in this embodiment, a method is adopted in which the results of four topic modeling algorithms, "Latent Seman+c Index (LSI)," "Probabilis+c Latent Seman+c Indexing (PLSI)," "Latent Dirichlet Allocation (LDA)," and "Cosine Similarity," are examined to automatically extract the most comprehensively optimal topic. The data relating to the classified categories is stored in the database 7 in association with the question data 4b and the answer data 4c.

[0030] Figure 6 is an image diagram of the improvement of learning data, where (1) shows before categorization, (2) shows after categorization, and (3) shows after input of improved data. Of the questions (41b to 41e) in question group 410 shown in Figure 6(1), questions (41b, 41c) were rated "Good" by the user, and the automatic response was evaluated as successful, resulting in a "+1" rating. In contrast, questions (41d, 41e) were rated "Bad" by the user, and the automatic response was evaluated as unsuccessful, resulting in a "-1" rating. As shown in Figure 6(2), questions (41b to 41e) in question group 410 shown in Figure 6(1) were classified into categories 14a and 14b. As a result, questions (41b, 41c) classified into category 14a were highly evaluated, while questions (41b, 41c) classified into category 14b were poorly evaluated, suggesting a high need for improvement. Therefore, by improving the quality of learning data 4a related to category 14b, the evaluation can be improved, as shown in Figure 6(3).

[0031] The quality of the automatic answers for each category is analyzed using the evaluation by the answer evaluation means 12 and the category classification by the topic estimation means 14 (step S04: evaluation data analysis step). In this embodiment, the analysis results are visually displayed on the display 51a using the answer evaluation display means 15. Here, an example of the visual display using the answer evaluation display means 15 will be described. FIG. 7 shows an image diagram of the answer evaluation display. As shown in FIG. 7, on the display 51a, answer evaluation results (151a, 151b) are displayed on the left side of the screen, and diagnosis results 152 obtained by analyzing the categories are displayed on the right side of the screen. It is possible to sequentially view a larger number of answer evaluations by scrolling the answer evaluation results (151a, 151b) up and down.

[0032] Buttons (15a to 15c) are arranged above the answer evaluation results (151a, 151b), and by clicking the buttons (15a to 15c), the display method of the answer evaluation results can be switched, such as displaying by "category name," by "answer success rate," or in descending order of "number of questions." In the example shown in Fig. 7, button 15a is clicked, and display by "category name" is selected. Of the answer evaluation results (151a, 151b), answer evaluation result 151a will be used as an example for explanation. In answer evaluation result 151a, "Expense Reimbursement" is displayed as "Category 1," and information related to the "Expense Reimbursement" category is displayed. Information related to the "Expense Reimbursement" category includes the question content in the question data, such as "About guests staying at hotels during remote business trips," answer evaluations for the questions, represented by "Yes" or "No," and the "AI answer success rate," represented by a pie chart and a numerical value. The number in parentheses indicates the total number of questions related to the category.

[0033] For "Expenses Reimbursement" in Category 1 and "Product Details" in Category 2, the category estimated by the topic estimation means 14 is initially displayed, but the category name can be edited by clicking button 15d provided to the right of "Expenses Reimbursement" or "Product Details." Button 15d activates topic modification means 17. A button 15e provided below the answer evaluation results (151a, 151b) increases the number of questions displayed, and if a user wishes to view more questions related to category 1 or category 2, the user clicks on the button 15e. Button 15f located at the bottom right of response evaluation result 151b is used to change the conditions. By clicking button 15f, you can narrow down the displayed content by conditions such as the period you want to analyze or whether or not to display results that belong to multiple categories.

[0034] The evaluation of the question is displayed as a "Good" or "Bad" evaluation given by the user evaluation means 11, with the answer evaluation means 12 evaluating the automatic response as a success and giving a rating of "+1" being displayed as "◯" or evaluating the automatic response as a failure and giving a rating of "-1" being displayed as "X". Although not shown, if no user evaluation is given and the evaluation by the answer evaluation means 12 is "0", it may be displayed as "△". The question content and the evaluation of the question are stored in the database 7 and output and displayed on the display 51a. The "AI answer success rate" indicates the rate at which the generated AI 3 correctly answers the user's question, and the data stored in the database 7 is analyzed by the evaluation data analysis means 13 and displayed on the screen by the answer evaluation display means 15. As shown in Figure 7, the "AI answer success rate" is visually displayed as a pie chart or percentage, which promotes intuitive understanding by the user and makes it possible to present information in a highly convenient manner.

[0035] The diagnostic results 152 displayed on the right side of the screen show the results of the category analysis and are explained in easy-to-understand text. The displayed text is automatically generated for all sentences using the evaluation data analysis means 13, which allows for a wide variety of recommendations. However, instead of this configuration, for example, a fixed format sentence may be prepared in advance, and the category name such as "Product Details" and numerical values ​​such as "3 months", "500 items", and "50%" may be automatically input based on the analysis results. The example question 152a is an excerpt from a question that was evaluated as a failed answer. Reading the example question 152a together with the explanatory text promotes understanding of the diagnosis results and how to improve them. Clicking the button 15g transitions to a simulation screen.

[0036] FIG. 8 shows an image diagram of a simulation of inputting recommendation data. As shown in FIG. 8, on the left side of the screen of the display 51a, a data addition section 16a, a data list display section 16b, a pull-down menu 16c, and questions (41f to 41h) are displayed, and a preview display section 16d is displayed on the right side of the screen. It is possible to view more data or questions by scrolling up and down the data list display section 16b or the questions (41f to 41h). The data addition section 16a, the data list display section 16b, the pull-down menu 16c, the questions (41f to 41h), and the preview display section 16d function as the simulation means 16. A user operating the PC 51 selects a file in the data adding unit 16a or adds it by dragging and dropping, and uploads the learning data 4a for simulation to the server 20. The uploaded learning data 4a is displayed in the data list display unit 16b. In the data list display unit 16b, the user can check details, delete data, and so on. Next, click on the pull-down menu 16c to select a category. Here, the "Product Details" category is selected. A list of questions in the selected category that the generation AI3 was evaluated as having failed to answer is displayed. Here, questions (41f to 41h) are displayed. Clicking on any of the questions displays the simulation results in the preview display area 16d. In Figure 8, question 41f has been clicked.

[0037] The preview display area 16d displays the icon and ID of the user who entered the question, the question 41f, the answers (42b, 42c) generated by the generation AI 3, the suggestion prompt display area 16e, and the message box 16g. The answer 42c for the question 41f when the uploaded learning data 4a is read using RAG technology (after) and the answer 42b when the answer is given without the data (before) are displayed side by side, allowing users to see in real time how the answer changes before and after the data is input. It can be seen that the answer 42b in the before state could not be answered using RAG technology due to the lack of data. It is also thought that hallucination occurred in the second sentence. In contrast, the answer 42c in the after state accurately answers the question. In this way, the provision of the simulation means 16 allows users to see at a glance the improvement results of adding data and realize the effects of the improvement.

[0038] Furthermore, in order to increase answer accuracy, it is important not only to improve the quality of the learning data but also to improve the quality of the questions. Therefore, better prompts are automatically generated based on the history of successfully answered questions stored in the database 7, and are displayed on the proposed prompt display unit 16e. The proposed prompt display unit 16e enables the question data analysis means 13b to function. Since the proposed prompt display unit 16e automatically recommends prompts that are likely to be answered successfully, if these are registered in a prompt list that can be used by all users, users who have not previously received a good answer to the question can also use the prompt, improving convenience. When the "Enter into chat" button 16f provided in the suggestion prompt display section 16e is clicked, the prompt is automatically entered into the message box, and when the send button 16h is clicked, the prompt can be registered in the prompt list. This has the advantage that the user only needs to fill in the specific information. [Example]

[0039] Generally, when questions are asked consecutively, they are often highly related. Therefore, in this embodiment, a configuration for evaluating questions asked by the same user in close proximity in time will be described. Fig. 9 shows a functional block diagram of an automatic response data improvement device of Example 2. As shown in Fig. 9, an automatic response system 101 of Example 2 is composed of an automatic response data improvement device 1a, an automatic response device 2, and a generation AI (LLM) 3. The automatic response data improvement device 1a of Example 2 differs from the automatic response data improvement device 1 of Example 1 in that the response evaluation means 120 includes a relevance determination means 12a. The other configurations are the same as those of Example 1. The relevance determination means 12a extracts the question immediately before or immediately after the question in question based on the user information and question date and time contained in the question data 4b, and determines the relevance between the question and the extracted immediately before or immediately after question based on at least one of the question data 4b and the answer data 4c, and the user evaluation data of the immediately before or immediately after question that is confirmed to be relevant to the question in question is used to evaluate the answer to the question.

[0040] FIG. 10 is an explanatory diagram of the answer evaluation means of the second embodiment, where (1) shows the answer evaluation means 12 of the first embodiment, and (2) shows the answer evaluation means 120 of the second embodiment. In each case, nine questions are asked by the same user, starting from the top, and the user evaluation data (5a to 5i) for each answer is shown as "Good," "Bad," or "-" (no evaluation). As shown in FIG. 10(1), the answer evaluation means 12 of the first embodiment judges the user evaluation data (5a to 5i) individually. Therefore, for example, when evaluating user evaluation data 5e using the answer evaluation means 12, evaluation is performed using only the user evaluation data 5e.

[0041] In contrast, as shown in FIG. 10(2), the answer evaluation means 120 of the second embodiment extracts the question immediately before or immediately after each question corresponding to the user evaluation data (5a to 5i) and determines the relevance. Therefore, for example, when evaluating the user evaluation data 5e using the answer evaluation means 120, the answer evaluation means 120 extracts the user evaluation data 5d immediately before the user evaluation data 5e and the user evaluation data 5f immediately after the user evaluation data 5e based on the identity of the user and the proximity of the question dates and times, and then determines the relevance. The relevance determination is performed based on the question data and the answer data. Here, the relevance with the user evaluation data 5d is confirmed, but the relevance with the user evaluation data 5f is denied. In contrast, in the case of the user evaluation data 5h, the relevance with both the immediately preceding user evaluation data 5g and the immediately following user evaluation data 5i is confirmed. The user evaluation data 5d that is confirmed to be related to the user evaluation data 5e is used to evaluate the question data and answer data corresponding to the user evaluation data 5e. Furthermore, the user evaluation data (5g, 5i) that is confirmed to be related to the user evaluation data 5h is used to evaluate the question data and answer data corresponding to the user evaluation data 5h. In this way, by using the evaluation regarding the immediately preceding or following question, a more accurate evaluation is possible, and useful information can be provided to the user. [Example]

[0042] Fig. 11 shows a functional block diagram of an automatic response data improvement device of Example 3. As shown in Fig. 11, an automatic response system 102 of Example 3 is composed of an automatic response data improvement device 1b, an automatic response device 2, and a generation AI (LLM) 3. The automatic response data improvement device 1b of Example 3 differs from the automatic response data improvement device 1 of Example 1 in that the answer evaluation means 121 includes a relevance determination means 12b. The other configurations are the same as those of Example 1. The grouping means 12b groups multiple questions in a question set based on at least one of the user information contained in the question data, the question date and time, the question content, and the answer content contained in the answer data, and evaluates the answers based on at least one of the type, number, and order of user evaluation data within the group.

[0043] FIG. 12 is an explanatory diagram of the answer evaluation means of the third embodiment, where (1) shows the answer evaluation means 12 of the first embodiment and (2) shows the answer evaluation means 121 of the third embodiment. In each case, nine questions are asked by the same user in order from top to bottom, and the user evaluation data (5a to 5i) for each answer is shown as "Good", "Bad" or "-" (no evaluation). As shown in FIG. 12(1), the answer evaluation means 12 of the first embodiment judges the user evaluation data (5a to 5i) individually. In contrast, as shown in Fig. 12(2), the answer evaluation means 121 of the third embodiment groups questions corresponding to user evaluation data (5a to 5i). The criteria for grouping are the identity of the user, the proximity of the question date and time, and the relevance of the question content and the answer content. In the example shown in Fig. 12(2), the user evaluation data (5a to 5c) is determined to be group 8a, the user evaluation data (5d, 5e) is determined to be group 8b, and the user evaluation data (5f to 5i) is determined to be group 8c. For each group (8a to 8c), the answers are evaluated based on the type, number, and order of the user evaluation data. The "type" of user evaluation data refers to "Good", "Bad", or "-" (no evaluation), and the "number" refers to the number of "Good", etc. In addition to the number, a percentage may also be added to the evaluation. The "order" refers to the order of "Good", "Bad", etc.

[0044] Figure 13 is an explanatory diagram of the evaluation method when grouping, where (1) shows the case where all the questions and answers ultimately received a "Good" user rating, and (2) shows the case where all the questions and answers ultimately received a "Bad" user rating. In this example, all three questions and answers are determined to be in a group. When all the evaluations in a group are "Good," as in group 8d shown in FIG. 13(1), the answer can be evaluated as successful. Both Group 8e and Group 8f had two "Good" responses and one "Bad" response, but Group 8e rated it "Bad" on the first question, but improved to "Good" on the second question. They also rated it "Good" on the third question. This suggests that they did not receive an appropriate answer on the first question, but received an appropriate answer on the second question, and their understanding deepened by the third question. In contrast, in Group 8f, the first question was rated as "Good," but the second question was rated as "Bad." Then, the third question was rated as "Good." Because the first question was rated as "Good," but the question immediately following it rated as "Bad," it is possible that the evaluation of the first question was inaccurate. Based on the above, Group 8e can be rated higher than Group 8f.

[0045] Furthermore, comparing group 8g and group 8h, group 8g has one "Good" and two "Bad," while group 8h has one "Good" and two "-" (no rating). Here, as in Example 1, it is possible to rate "0" when there is no rating, but it is generally said that many users only rate at the end when multiple questions are asked. For example, if it can be determined from the user's rating trends over a predetermined period that they only click "Good" or "Bad," it can be determined that the "-" is likely to be the opposite rating to the last rating given, and the "-" in group 8h can be treated the same as a "Bad," and groups 8g and 8h can be rated the same. Note that, as a method for determining whether a user only clicks "Good" or "Bad" based on the user's evaluation tendency over a predetermined period, for example, the criterion is whether the ratio of "Good" or "Bad" to the user's total number of questions exceeds a predetermined threshold. Furthermore, the user's evaluation tendency may be determined based on the number of evaluations out of the total number of questions, the frequency of a specific evaluation type, etc.

[0046] Next, as in group 8i shown in FIG. 13(2), if all the evaluations in the group are "Bad", it can be evaluated that the answer was unsuccessful. Both Group 8j and Group 8k had one "Good" and two "Bad" ratings, but Group 8j was rated "Good" in the first question, but "Bad" in the second question, and again "Bad" in the third question. In contrast, in group 8k, the score was rated as "Bad" in the first question, but the score improved to "Good" in the second question, and then the score was rated as "Bad" again in the third question. For example, if there are consecutive "Bad" ratings at the end, it can be evaluated that satisfaction with the answers is low. Therefore, group 8k can be given a higher rating than loop 8j.

[0047] Furthermore, comparing group 8l and group 8m, group 8l has two "Good" and one "Bad," while group 8m has one "Bad" and two "-" (no rating). As with the ratings for groups 8g and 8h described above, if it can be determined from the user's rating trends over a predetermined period that they only click "Good" or "Bad," it can be determined that there is a high possibility that the "-" is the opposite rating to the last rating given, and the "-" for group 8m can be treated the same as "Good," making it possible to give groups 8l and 8m the same rating. However, it is generally said that if a user continues to receive no ratings from the beginning and then receives a "Bad" rating at the end, this is often because the user has not received a satisfactory answer no matter how many times they have asked. For this reason, unless a particular user's rating tendency can be identified, all "-"s in group 8m should be rated as "Bad." Based on the above, Group 8l can be rated higher than Group 8m. The evaluation method shown here is merely an example, and more accurate evaluation can be achieved by setting a variety of evaluation criteria. [Industrial Applicability]

[0048] The present invention is useful as a technology for improving response accuracy in an automatic answering device using a generation AI. [Explanation of symbols]

[0049] 1, 1a, 1b Automatic response data improvement device 2. Automated answering machine 3 Generation AI 4a Training data 4b Question data 4c Response data 5a~5i user evaluation data 6 Network 7 Database 7a Question History Database 7b Answer history database 8a~8m group 11 User evaluation methods 11a,11b Rating button 12,120,121 Response evaluation method 12a Relevance determination method 12b Grouping Methods 13. Evaluation data analysis methods 13a Training data analysis methods 13b Questionnaire data analysis methods 14 Topic estimation method 14a, 14b Category 15. Answer evaluation display means 15a~15g,16f buttons 16 Simulation Methods 16a Data addition section 16b Data list display section 16c pulldown 16d Preview display area 16e Proposal prompt display section 16g Message Box 16h Send button 17 Topic Revision Methods 20,30 servers 41a~41h Questions 42a~42c Answer 51 PC 51a Display 52 Smartphones 53 Tablet devices 100~102 Automated Answering System 151a, 151b Answer evaluation results 152 Diagnosis Results 152a Example Questions 161 Simulation data input means 162 Question content selection method 163 Preview Methods 410 Questions

Claims

1. A device for improving a device that automatically responds to questions from a user using pre-input learning data and generated AI, a user evaluation means for receiving user evaluations regarding the response data and storing the evaluations in a database as user evaluation data; an answer evaluation means for generating answer evaluation data that evaluates the quality of an automatic answer using the question data, the answer data, and the user evaluation data; a topic estimation means for performing natural language processing on the question data and the answer data, classifying the data into categories using topic modeling, and storing the results in a database; evaluation data analysis means for analyzing the quality of the automatic answers at least for each category using the answer evaluation data and the categories; An automatic response data improvement device comprising:

2. 2. The automatic response data improving device according to claim 1, wherein said evaluation data analyzing means comprises learning data analyzing means for analyzing learning data recommended for input.

3. 2. The automatic response data improving device according to claim 1, wherein the evaluation data analyzing means includes question data analyzing means for analyzing question data recommended for input.

4. The automatic response data improvement device described in claim 1 further comprises an answer evaluation display means for visually displaying on a display at least one of the answer success rate for each category during a specified period and a list of questions for which the answer success rate is below a specified threshold.

5. a simulation data input means for receiving input of learning data for simulation; a question content selection means for selecting from a list a question content for which the answer success rate is below a predetermined threshold; A preview means for displaying the contents of answers regarding the answer data in parallel before and after inputting the simulation data; 5. The automatic response data improving device according to claim 4, further comprising a simulation means having:

6. The automatic response data improvement device according to claim 1, further comprising a topic correction means for accepting corrections to category classifications made by the topic estimation means, and for storing the corrected category classifications in a database in association with the question data and the answer data.

7. The response evaluation means extracting questions immediately before or after the question based on the user information and question date and time contained in the question data; a relevance determination means for determining a relevance between the question and the extracted immediately preceding question or immediately following question based on at least one of the question content included in the question data and the answer content included in the answer data, 2. The automatic response data improving device according to claim 1, wherein the user evaluation data of the immediately preceding question or the immediately following question that is confirmed to be related to the question is used to evaluate the answer to the question.

8. The response evaluation means a grouping means for grouping a plurality of questions in a question group based on at least one of user information, question date and time, question content, and answer content included in answer data, which are included in the question data; 2. The automatic response data improving device according to claim 1, wherein the answer is evaluated based on at least one of the type, number, and order of the user evaluation data in the group.

9. A method for improving a device that automatically responds to questions from a user using pre-input learning data and a generating AI, comprising: a user evaluation step of receiving a user evaluation regarding the answer data and storing it in a database as user evaluation data; an answer evaluation step of generating answer evaluation data that evaluates the quality of the automatic answer using the question data, the answer data, and the user evaluation data; a topic estimation step of performing natural language processing on the question data and the answer data, classifying the data into categories using topic modeling, and storing the results in a database; an evaluation data analysis step of analyzing the quality of the automatic answer for at least each category using the answer evaluation data and the categories; A method for improving automatic response data, comprising:

10. 10. An automatic response data improving program that causes a computer to execute each step of the automatic response data improving method of claim 9.

Citation Information

Patent Citations

  • Program, method and system

    JP2024006940A

  • Routine Evaluation of Accuracy of a Factoid Pipeline and Staleness of Associated Training Data

    US20200372109A1

  • Information providing system, information providing method, program and data structure

    JP2016099968A