Information processing system, information processing method and program
The proposed mechanism addresses the inefficiency in estimating the purpose of user interactions in chatbot systems by optimizing the estimation process, resulting in reduced processing load and costs while maintaining accurate purpose identification.
Patent Information
- Application Number
- JP2023210971
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-26
AI Technical Summary
Existing systems that automatically respond to user input, such as chatbots, face inefficiencies in estimating the purpose of user interactions due to the need for frequent category estimations, which increases processing load and costs.
A mechanism that efficiently estimates the purpose of use in a system responding to user input by acquiring user input content, outputting instructions to an external device for purpose estimation, and specifying the purpose of use based on predetermined conditions, thereby reducing unnecessary estimation processes.
This approach allows for efficient estimation of the purpose of use in chatbot systems, reducing processing load and costs while maintaining accurate purpose identification.
Smart Images

Figure 2025095156000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system, an information processing method, and a program.
Background Art
[0002] A system that automatically responds to user input (hereinafter referred to as a chatbot) is known. In recent years, due to the text generation AI technology using large language models, the usage scenarios of chatbots have increased, and chatbots are used in a wide range of applications. With the increase and generalization of chatbot usage, there is an increasing need from chatbot operators to aggregate and analyze for what purposes users are using chatbots.
[0003] To analyze the usage purpose of users, there is a method of saving the user's conversation history as a log. However, the user's input may contain confidential information, and it often costs a lot to manage the log data. Patent Document 1 discloses a mechanism for estimating the category representing the user's speech intention by referring to an intention estimation database in which extraction words corresponding to categories representing each speech intention are registered so that the user and the chatbot can continue the conversation smoothly.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Disclosure of the Invention
Problems to be Solved by the Invention
[0005] The technique described in Patent Document 1 has a problem that it is necessary to perform category estimation for each user speech for answer generation, and the number of times of category estimation increases.
[0006] Therefore, an object of the present invention is to provide a mechanism for efficiently estimating the purpose of use of a system that answers a user's input in an interactive format.
Means for Solving the Problems
[0007] The present invention includes: a first acquisition means for acquiring input content by a user in a system that answers a user's input in an interactive format; an output means for outputting an instruction to an external device to estimate the purpose of use of the system related to a series of interactions by the user based on the first input content acquired by the acquisition means; a second acquisition means for acquiring the purpose of use estimated based on the instruction output by the output means; and a specifying means for specifying, when the purpose of use acquired by the second acquisition means satisfies a predetermined condition, the purpose of use as the purpose of use of the system by the user. The output means outputs an instruction to the external device to estimate the purpose of use of the system by the user based on second input content by the user when the purpose of use estimated by an instruction based on the first input content does not satisfy a predetermined condition.
Effects of the Invention
[0008] According to the present invention, it is possible to efficiently estimate the purpose of use of a system that answers a user's input in an interactive format.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0011] FIG. 1 is a diagram showing an example of the system configuration of a chatbot (a system that answers a user's input in an interactive form) in an embodiment of the present invention.
[0012] The chatbot 100 includes a user input reception unit 110, a response generation device 120, a category estimation device 130, a log DB 140, and a client terminal.
[0013] The user input reception unit 110 is a device for receiving a user's input to the chatbot. The user can send any input to the user input reception unit 110 through a Web browser or the like of the client terminal.
[0014] The response generation device 120 includes a response generation unit 121 and a response output unit 122.
[0015] The response generation unit 121 is a device that receives a user input from the user input reception unit 110 and generates a chatbot response thereto. This may use known techniques used in chatbots, or techniques called generation AI may also be used.
[0016] The response output unit 122 is a device for returning the response generated by the response generation unit 121 to the user. The user can receive the response from the chatbot through a Web browser or the like of the client terminal. Also, for category estimation described later, the response may be returned to the user and output to the category estimation device 130.
[0017] The category estimation device 130 consists of a category estimation unit 131 and an estimation result output unit 132.
[0018] The category estimation unit 131 is a device for estimating the category of the user's purpose of using the chatbot in a series of exchanges (hereinafter referred to as conversations) consisting of user inputs and chatbot responses. The category estimation unit 131 performs category estimation based on both the user input received from the user input reception unit 110 and the chatbot response generated by the user input and the response generation device 120. Alternatively, in addition to the new user input received from the user input reception unit 110, category estimation may be performed based on previous user inputs and previous chatbot responses within the conversation.
[0019] The category estimation unit 131 determines the category of the entire conversation based on the result of this category estimation. However, when the category of the conversation has already been estimated, there are cases where category estimation is not performed and cases where category estimation is redone. Details of the mechanism for estimating the category of the conversation and the processing for skipping or redoing the estimation will be described later with reference to FIG. 5. For category estimation, known category classification techniques may be used, or it may be solved by giving a prompt for category estimation to a general-purpose language model.
[0020] FIG. 2 is a block diagram showing an example of the hardware configuration of an information processing apparatus that can be used as the client terminal and chatbot 100 in the embodiment of the present invention.
[0021] As shown in FIG. 2, the information processing apparatus has a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 connected via a system bus 204.
[0022] The CPU 201 comprehensively controls each device and controller connected to the system bus 204.
[0023] The ROM 202 or the external memory 211 holds the BIOS (Basic Input / Output System), the OS (Operating System), which are control programs executed by the CPU 201, a computer-readable and executable program for implementing this information processing method, and various necessary data (including data tables).
[0024] The RAM 203 functions as the main memory, work area, etc. of the CPU 201. When executing a process, the CPU 201 loads a program, etc. necessary for the execution from the ROM 202 or the external memory 211 into the RAM 203, and realizes various operations by executing the loaded program.
[0025] The input controller 205 controls the input from input devices such as a keyboard 209 and a pointing device such as a mouse (not shown). When the input device is a touch panel, it is assumed that the user can give various instructions by pressing (touching with a finger, etc.) in accordance with icons, cursors, buttons, etc. displayed on the touch panel.
[0026] Further, the touch panel may be a touch panel capable of detecting the positions touched by a plurality of fingers, such as a multi-touch screen.
[0027] The video controller 206 controls the display to an external output device such as a display 210. The display includes the display of a notebook personal computer integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. Also, for a device capable of receiving the above-described touch operation, an input device is also provided.
[0028] The video controller 206 can control a video memory (VRAM) for performing display control, and can use a part of the RAM 203 as a video memory area, or can separately provide a dedicated video memory.
[0029] The memory controller 207 controls access to the external memory 211. As the external memory, an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a compact flash (registered trademark) memory connected via an adapter to a PCMCIA card slot can be used.
[0030] The communication I / F controller 208 is connected to and communicates with an external device via a network, and executes communication control processing on the network. For example, communication using TCP / IP, a telephone line such as ISDN, and communication using a 3G line of a mobile phone are possible.
[0031] In addition, the CPU 201 enables display on the display 210 by executing an outline font expansion (rasterization) process on, for example, the display information area in the RAM 203. Further, the CPU 201 enables user instructions using a mouse cursor (not shown) on the display 210.
[0032] First, the overall image of the processing of the present invention will be described with reference to FIG. 8. In this embodiment, for example, in an in-house chatbot, it is assumed that the purpose of using the chatbot by the user (employee) is estimated. The user inputs "Translate the following sentence into English: The cat and the mouse quarreled" to the input form within the site. For example, the chatbot gives the user's input to the generation AI, executes the task of translation, and receives the result. Further, in order to estimate the purpose of using the user's input, a prompt (to be described later with reference to FIG. 9) including the input content and category classification of the user is given to the generation AI, and the result of category estimation is received. That is, two requests, "task execution" and "category estimation", are sent to the generation AI. Here, considering the case of using a pay-per-use service (a mechanism in which a charge is imposed on the data volume and the number of times of interaction with a system that performs estimation processing such as generation AI), a charge will occur each time a request is sent. In particular, in the case of text generation AI, a charge often occurs for each word (token) input (or output by the generation AI). That is, if the category is estimated for each user input, a charge will occur each time such a request is made. Also, as the number of estimations increases, the processing load also increases. Therefore, in the present invention, in order to efficiently identify the purpose of using the system, in a system that answers the user's input in a dialogue format, the purpose of using the system in the dialogue is identified using the input content of the user in a series of dialogues, and when the purpose of use is identified, control is performed so as not to perform the estimation process of the purpose of use using other input contents in the series of dialogues.
[0033] Figure 3 shows an example of a category list predefined in the category estimation unit 131. The category estimation of the usage purpose executed by the category estimation unit 131 estimates which record in the category list 300 the user input matches. For example, if the user input is a sentence "Translate the following sentence: The cat and the mouse had a quarrel", the category is estimated to be record 301 "Translation". Note that for category estimation, known category classification techniques may be used, or a prompt may be given to a general-purpose language model (including generative AI) for category estimation. Details of the category estimation process will be described later with reference to Figure 9.
[0034] Figure 4 shows an example of a list of retry categories (categories that require re-estimation) predefined in the category estimation unit 131. Categories that match within the retry category list 400 are considered inappropriate as the estimation result of the user's usage purpose in the entire conversation. For example, when the user starts a conversation with a greeting to the chatbot, it is inappropriate to consider that the entire conversation is related to "greeting" and the user is using the system for the purpose of "greeting". In such a case, the conversation estimated as a category that matches the retry category list 400 is re-estimated based on the next user input. The retry category list 400 is a part of the category list 300.
[0035] Returning to Figure 1, the estimation result output unit 132 is a device that outputs the category estimated by the category estimation unit 131 to the log DB 140. Note that for the convenience of the application, etc., the estimated category may be returned to the user together with the output to the DB.
[0036] Next, with reference to the flowchart of Figure 5, the process for efficiently estimating the category of a conversation executed by the category estimation unit 131 in the embodiment of the present invention will be described.
[0037] The flowchart of FIG. 5 is a process that is executed each time a user input is received, indicating that there are times when category estimation (step S506) is executed and times when category estimation is not performed (skipped). Also, for simplicity of explanation here, it is assumed that category estimation is executed only based on new user inputs. Note that in addition to the new user input received from the user input receiving unit 110, a mechanism that performs category estimation based on previous user inputs in the conversation and previous chatbot responses may also be used.
[0038] In step S501, the user input receiving unit 110 receives a user input A.
[0039] In step S502, it is determined whether the conversation has already been category-estimated. It is assumed that the determination that it has not been estimated is for the case where the user input A is the first user input in the conversation. Here, if it is determined that it has not been estimated, category estimation is performed for the user input A (move to step S506).
[0040] In step S503, the category α of the conversation that has already been estimated is acquired.
[0041] In step S504, it is determined whether the estimated category α of the conversation that has already been estimated matches a record in the retry category list 400. If it does not match, it is regarded that the category of the conversation has already been determined as the estimated category α, and category estimation is skipped (move to step S507). If it matches, it is considered that there is a possibility of re-estimating the category of the conversation with the user input A, and move to step S505.
[0042] In step S505, it is determined whether the estimated category α of the conversation is the result of having been redone up to the maximum number of redos. If the estimation has already been redone up to the maximum number of redos, in this conversation, it is considered that the topic related to the redo category is being continued, and the category estimation is skipped (move to step S507). If the maximum number of times has not been reached, it is assumed that the category α is not an appropriate category for the user's purpose in the conversation, and the category is re - estimated based on the user input A (move to step S506). Here, the maximum number of redos can take any value, and it is also possible not to perform any redo at all or to continue re - doing until an appropriate category is estimated.
[0043] In step S506, category estimation is performed based on the user input A to obtain category β. In addition, when an estimation result other than the category list in FIG. 3 is returned, the result may be ignored (discarded) as a failed estimation, so as to devise a way that the estimation result does not contain confidential information. For example, assume a case where the category is estimated from the user's input content without specifying the category classification (not including it in the prompt). For example, when the user inputs "Please create an email for convening a general shareholders' meeting regarding the dismissal of Director A.", the generative AI may estimate the category as "Creation of an email regarding the dismissal of Director A" and output it. If the category estimation result contains words such as "dismissal of Director A" input by the user and this is confidential information, the data regarding the purpose of use will become data containing confidential information, so attention is required in handling such data. Therefore, a category abstracted so as not to contain confidential information is set in advance, and when the estimation result does not match the said category, the estimation result is discarded as a failed estimation, thereby maintaining confidentiality.
[0044] In step S507, the category of the conversation is assumed to be the estimated category. That is, if step S506 is skipped, it is category α, or if it is not skipped and estimated, it is category β.
[0045] Next, with reference to FIGS. 6 and 7, a method for efficiently estimating the category of a conversation in an embodiment of the present invention will be described.
[0046] FIG. 6 is an example of a user interface that the answer output unit 122 displays on the browser of the client terminal. The conversation screen 600 includes a conversation history 610, a user input component 620, and a conversation restart component 630 in addition to the user input component 620. The conversation restart component 630 is a component assumed to be used by the user at the turning point of the topic, and can reset the previous irrelevant conversation and start a new conversation. In particular, if the answer generation unit 121 adopts an architecture that inputs the entire conversation for one answer output, the conversation before the topic conversion point can be excluded from the input.
[0047] The conversation screen 600 shows the state after the user speaks the user input 611 "Hello" and greets the chatbot for the first time, and the chatbot returns an answer. At this time, this conversation is estimated by the category estimation unit 131 to be a conversation related to the record 302 "Greeting" in the category list 300.
[0048] The conversation screen 700 in FIG. 7 shows the state after the user continues to speak the user input 711 "Translate the following sentence: The cat and the mouse quarreled" from the state of the conversation screen 600, and the chatbot returns an answer. In the state of the conversation screen 600, the estimated category of the conversation was "Greeting", but this corresponds to the record 401 in the retry category list 400. Therefore, in the state of the conversation screen 700, the category estimation is redone, and based on the user input 711, the conversation is estimated to be the record 301 "Translation" in the category list 300. However, here, the retry upper limit number is set to 1 or more for the sake of simplicity of explanation. In actuality, the retry upper limit number can be set to any value.
[0049] In the state of the conversation screen 700, the conversation is presumed to pertain to the "Translation" category that is not within the retry category list 400. Therefore, in this conversation, the category is determined as "Translation", and category estimation is skipped regardless of any subsequent user input.
[0050] As described above, in the present invention, once the category is determined in one conversation, it is considered that the conversation of that category continues thereafter. This is highly valid when the user uses the conversation restart component 630 to start a new conversation each time the topic changes. In particular, as described above, if the entire conversation is input to the response generation unit 121, it is desirable for improving the accuracy of the response not to include previous unrelated conversations at the topic conversion point. Therefore, the validity of this concept becomes higher if the user uses it in an ideal way.
[0051] FIG. 9 is a diagram showing an example of the category estimation process of a conversation. As shown in the figure, by providing the generation AI with the category list 300 and the user input as a prompt and having it select a suitable one from the category list 300 for the user's purpose of use, the purpose of use can be estimated without including highly confidential information.
[0052] As described above, according to the present invention, in a system that responds to a user's input in an interactive format, the purpose of using the system in a series of interactions is specified using the content of the user's input in the series of interactions. When the purpose of use is specified, by controlling so as not to perform the estimation process of the purpose of use regardless of other input contents in the series of interactions, it becomes possible to efficiently specify the purpose of using the system. Further, for the estimated purpose of use, when it falls within a preset purpose-of-use category (that is, when a predetermined condition is satisfied), by specifying the estimated purpose of use as the purpose of using the system by the user, it becomes possible to handle the purpose of use as data that does not include confidential information (perform aggregation, analysis, etc.). Further, for the estimated purpose of use, when it does not fall within a category that requires re-estimation (for example, "greeting", etc.) set in advance (that is, when a predetermined condition is satisfied), the estimated purpose of use is specified as the purpose of using the system by the user, and when it falls within a category that requires re-estimation, by estimating the purpose of use using other input contents, it becomes possible to specify an appropriate purpose of use in a series of interactions.
[0053] By analyzing the user's purpose of use as described above, it can be used for the development of chatbots. For example, by analyzing frequently used applications and preparing dedicated prompts that suit the purpose of use, the convenience for the user is enhanced. Further, when analyzing the purpose of use for each department, if there are applications that are used across various departments or applications that are used predominantly in a specific department, it becomes possible to develop applications specialized for that purpose, hold study sessions for specific departments, etc.
[0054] Although the embodiments have been described above, the present invention can take an embodiment as, for example, a system, device, method, program, or recording medium. Specifically, it may be applied to a system composed of a plurality of devices, or may be applied to a device composed of a single device.
[0055] In addition, the program in the present invention is a program that can be executed by a computer according to the processing method of the flowchart shown in FIG. 3, and the storage medium of the present invention stores a program that can be executed by a computer according to the processing method of FIG. 3. Note that the program in the present invention may be a program for each processing method of each device in FIG. 3.
[0056] As described above, it goes without saying that the object of the present invention can also be achieved by supplying a recording medium recording a program for realizing the functions of the above-described embodiments to a system or apparatus, and causing a computer (or CPU or MPU) of the system or apparatus to read and execute the program stored in the recording medium.
[0057] In this case, the program itself read from the recording medium realizes the novel functions of the present invention, and the recording medium recording the program constitutes the present invention.
[0058] As the recording medium for supplying the program, for example, a flexible disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a DVD-ROM, a magnetic tape, a non-volatile memory card, a ROM, an EEPROM, a silicon disk, etc. can be used.
[0059] In addition, by executing the program read by the computer, not only the functions of the above-described embodiments are realized, but also based on the instructions of the program, an OS (operating system) or the like running on the computer performs part or all of the actual processing, and the functions of the above-described embodiments are realized by the processing. This case is also included, which goes without saying.
[0060] Furthermore, after the program read from the recording medium is written into the memory provided in the function expansion board inserted into the computer or the function expansion unit connected to the computer, based on the instructions of the program code, the CPU etc. provided in the function expansion board or the function expansion unit perform part or all of the actual processing, and it goes without saying that the case where the functions of the above-described embodiments are realized by this processing is also included.
[0061] In addition, the present invention may be applied to a system composed of a plurality of devices or to an apparatus composed of a single device. Needless to say, the present invention is also applicable when achieved by supplying a program to a system or an apparatus. In this case, by reading out the recording medium storing the program for achieving the present invention to the system or the apparatus, the system or the apparatus can enjoy the effects of the present invention.
[0062] Furthermore, by downloading and reading out the program for achieving the present invention from a server, a database, etc. on the network by means of a communication program, the system or the apparatus can enjoy the effects of the present invention. Note that all configurations combining the above-described embodiments and their modified examples are also included in the present invention.
Explanation of Signs
[0063] 100 Chatbot 110 User Input Reception Unit 120 Answer Generation Device 130 Category Estimation Device 140 Log DB
Claims
1. In a system that answers a user's input in a dialogue format, a first acquisition means for acquiring the input content by the user, an output means for outputting an instruction to cause an external device to estimate the purpose of use of the system related to a series of dialogues by the user based on the first input content acquired by the acquisition means, a second acquisition means for acquiring the purpose of use estimated based on the instruction output by the output means, a specifying means for specifying, as the purpose of use of the system by the user, the purpose of use when the purpose of use acquired by the second acquisition means satisfies a predetermined condition, comprising, when the purpose of use estimated by the instruction based on the first input content does not satisfy a predetermined condition, the output means outputs an instruction to cause the external device to estimate the purpose of use of the system by the user based on the second input content by the user. The information processing apparatus is characterized by this.
2. The information processing apparatus according to claim 1, wherein after the purpose of use related to a series of dialogues is specified by the specifying means, the output means does not output an instruction to estimate the purpose of use related to the dialogue.
3. The information processing apparatus according to claim 2, wherein the second input content is input content related to the same series of dialogues as the first input content and is input after the first input content.
4. The information processing apparatus according to claim 3, wherein the case where the predetermined condition is satisfied is a case where the purpose of use acquired by the second acquisition means is included in a preset purpose of use category.
5. The information processing apparatus according to claim 4, wherein the case where the predetermined condition is not satisfied is a case where the purpose of use acquired by the second acquisition means is not included in a preset purpose of use category.
6. The information processing apparatus according to claim 4, wherein the case where the predetermined condition is not satisfied is a case where the purpose of use acquired by the second acquisition means is a category for which re - estimation is required and is preset.
7. The information processing apparatus according to claim 1, wherein the output means outputs an instruction to cause a generative AI to estimate the purpose of use of the system by the user.
8. The output means outputs an instruction to estimate the purpose of use by giving an instruction created based on preset category information and the input content by the user to the generative AI. The information processing apparatus according to claim 7, wherein the second acquisition means acquires, as a usage purpose, a category selected from the preset categories.
9. In a system that answers a user's input in an interactive form, a first acquisition means for acquiring the input content by the user, an output means for outputting an instruction to cause an external device to estimate a usage purpose of the system related to a series of interactions by the user based on the first input content acquired by the acquisition means, a second acquisition means for acquiring a usage purpose estimated based on the instruction output by the output means, a specifying means for specifying, when the usage purpose acquired by the second acquisition means satisfies a predetermined condition, the usage purpose as the usage purpose of the system by the user, comprising The information processing system, wherein when the usage purpose estimated by the instruction based on the first input content does not satisfy a predetermined condition, the output means outputs an instruction to cause the external device to estimate the usage purpose of the system by the user based on a second input content by the user.
10. A first acquisition step in which a first acquisition means acquires input content by a user in a system that answers a user's input in an interactive form, an output step in which an output means outputs an instruction to cause an external device to estimate a usage purpose of the system related to a series of interactions by the user based on the first input content acquired by the acquisition means, a second acquisition step in which a second acquisition means acquires a usage purpose estimated based on the instruction output by the output means, a specifying step in which a specifying means specifies, when the usage purpose acquired by the second acquisition means satisfies a predetermined condition, the usage purpose as the usage purpose of the system by the user, comprising The control method of an information processing apparatus, wherein when the usage purpose estimated by the instruction based on the first input content does not satisfy a predetermined condition, the output means outputs an instruction to cause the external device to estimate the usage purpose of the system by the user based on a second input content by the user.
11. A program for causing a computer to function as each means according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text data processing device and program
JP2010250480A
Voice processing device and method, and program
JP2011033680A
Classification device, classification method, and classification program
JP2018151786A
Utterance intention determination device, utterance intention determination method, and program
JP2019114141A
Dialogue control method, dialogue control program, dialogue control device, information presentation method, and information presentation device
JP2020087352A
Cited By
Apparatus, method, and program for providing AI agent services
JP7903112B1