Information processing device, information processing method, and information processing program
The information processing device enhances dialogue efficiency by selecting conversation sequences based on user states and optimizing through reinforcement learning, addressing inefficiencies in conventional chatbot communication systems.
Patent Information
- Application Number
- JP2022198796
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-12-13
AI Technical Summary
Conventional technologies for interactive communication with chatbots are inefficient in extracting information from service users of online services, particularly in preventing users from abandoning chats and not effectively obtaining reviews.
An information processing device and method that utilizes a selection unit to choose a conversation sequence based on the service user's state from predefined patterns, and an instruction unit to execute the dialogue processing, with reinforcement learning to optimize the selection model based on user reactions.
Efficiently collects information from service users by improving the quality of dialogue and increasing the likelihood of continued interaction, while optimizing conversation sequences for favorable user responses.
Smart Images

Figure 0007807364000001 
Figure 0007807364000002 
Figure 0007807364000003
Abstract
Description
[Technical Field]
[0001] The present application relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] Conventionally, techniques for performing interactive communication in cooperation with a chatbot have been proposed. For example, Patent Document 1 proposes a technique for performing interactive communication through a digital board in cooperation with a chatbot. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Re-tabled publication No. 2020-240838 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the above-described conventional technology leaves room for improvement in terms of efficiently extracting information from service users of online services through communication via chatbots. For example, the conventional technology aims to prevent users from abandoning chats, and is not intended to efficiently obtain reviews from service users, so there is still considerable room for improvement.
[0005] The present application has been made in consideration of the above, and aims to provide an information processing device, an information processing method, and an information processing program that can efficiently collect information from service users of online services. [Means for solving the problem]
[0006] The information processing device according to the present application controls processing related to a dialogue conducted with a service user of an online service through a chatbot, and includes a selection unit and an instruction unit. The selection unit selects a conversation sequence according to the state of the service user from a plurality of conversation sequences in which conversation patterns indicating a series of conversation contents expected in the dialogue are predefined. The instruction unit instructs an external device that executes processing related to the dialogue to use the conversation sequence selected by the selection unit. [Effects of the Invention]
[0007] According to one aspect of the embodiment, it is possible to efficiently collect information from service users of online services. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram showing an overview of information processing according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an outline of a conversation sequence according to the embodiment. [Figure 3] FIG. 3 is a diagram showing an example of an instruction of a conversation sequence from the second server to the first server according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an overview of reinforcement learning according to the embodiment. [Figure 5] FIG. 5 is a block diagram illustrating an example of the configuration of the second server according to the embodiment. [Figure 6] FIG. 6 is a diagram showing an outline of conversation sequence information according to the embodiment. [Figure 7] FIG. 7 is a diagram illustrating an overview of information related to a selection model according to the embodiment. [Figure 8] FIG. 8 is a diagram showing an overview of user information according to the embodiment. [Figure 9] FIG. 9 is a flowchart illustrating an example of a processing procedure executed by the second server according to the embodiment. [Figure 10]FIG. 10 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the second server according to the embodiment or each modification. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, modes for implementing an information processing device, an information processing method, and an information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to these embodiments. Furthermore, the respective embodiments can be appropriately combined within the scope of not causing any contradiction in the processing content. Furthermore, the same components in the following embodiments will be assigned the same reference numerals, and redundant explanations will be omitted.
[0010] (Embodiment) [1. System configuration according to the embodiment] First, the configuration of an information processing system SYS having a second server 200, which is an example of an information processing device according to an embodiment, will be described with reference to Fig. 1. Fig. 1 shows an example of the configuration of the information processing system SYS according to an embodiment. As shown in Fig. 1, the information processing system SYS according to an embodiment has a user terminal 10, a first server 100, and a second server 200.
[0011] The user terminal 10, the first server 100, and the second server 200 are connected to a network such as the Internet (for example, the network N shown in FIG. 5). The user terminal 10 and the first server 100 can communicate with each other through the network. The first server 100 and the second server 200 can communicate with each other through the network. The user terminal 10 and the second server 200 may also communicate with each other through the network. The information processing system SYS shown in FIG. 1 may have more user terminals 10 than the example shown in FIG. 1.
[0012] The user terminal 10 is an information processing terminal used by a service user U who is a user of various online services operated by the administrator of the first server 100 as a platform provider. For example, the user terminal 10 can be realized by a smartphone, a desktop PC (Personal Computer), a notebook PC, a tablet terminal, a mobile phone, a PDA (Personal Digital Assistant), or the like.
[0013] In addition, the user terminal 10 has communication functions for performing wireless communication networks such as LTE (Long Term Evolution), 4G (4th Generation: fourth generation mobile communication system), and 5G (5th Generation: fifth generation mobile communication system), as well as short-range wireless communication such as Bluetooth (registered trademark) and wireless LAN (Local Area Network), and can connect to a network using these communication functions.
[0014] Furthermore, the user terminal 10 can display, for example, web content of various online services provided by the first server 100 using a web browser or an application. When the user terminal 10 receives control information for realizing information display processing from the first server 100 or the like, the user terminal 10 realizes the display processing in accordance with the control information.
[0015] The service user U can operate the user terminal 10 to browse websites of various online services displayed by a web browser and use web content displayed by the web browser. In addition, the service user U can download a dedicated application program (hereinafter referred to as a "dedicated app") for using the websites of various online services from the first server 100 and install it on the user terminal U. In this case, the service user U can use the content of various online services configured for the dedicated app by operating the dedicated app.
[0016] The first server 100 is an information processing device that provides various online services to each service user. The first server 100 is typically a server device, but may also be realized by a mainframe, a workstation, or the like. Furthermore, when the first server 100 is realized by a server device, it may be realized by a single server device, or may also be realized by a cloud system in which multiple server devices and multiple storage devices operate in cooperation with each other.
[0017] The various online services provided by the first server 100 may include internet connection, search services, SNS (Social Networking Service), e-commerce services, electronic payment services, online games, online banking services, online trading services, hotel reservation services, ticket reservation services, video distribution services, music distribution services, news distribution services, map information services, route search services, route guidance services, line information services, operation information services, weather information services, etc. The various online services may also include API (Application Programming Interface) services corresponding to various applications.
[0018] Furthermore, when providing various online services, the first server 100 creates a user account including a user ID, which is user identification information for identifying each service user (for example, service user U). The user ID included in this user account is set arbitrarily by the service user (for example, service user U) when registering to use the various online services, or is individually assigned by the first server 100. The first server 100 records a service usage history (an example of "user information"), which is a usage history of the online services, linked to the user account of each service user (for example, service user U), and manages the recorded service usage history for each service user. Furthermore, the first server 100 can distribute dedicated apps for using various online services in response to a request from the service user (for example, service user U).
[0019] Furthermore, the first server 100 executes processing related to dialogue with service users of various online services (for example, service user U) through the chatbot. The processing of the first server 100 will be described later.
[0020] The second server 200 is an information processing device that controls processing related to interactions that take place through a chatbot between a service user (for example, service user U) of various online services and the first server 100. The second server 200 is typically a server device, but may be realized by a mainframe, a workstation, or the like. When the second server 200 is realized by a server device, it may be realized by a single server device, or may be realized by a cloud system in which multiple server devices and multiple storage devices operate in cooperation with each other. The second server 200 will be described later.
[0021] [2. Overview of Information Processing According to Embodiment] An overview of information processing according to the embodiment will be described below with reference to Figures 1 to 4. In the following description, the user terminal 10 may be referred to as the service user U. In other words, the service user U can be read as the user terminal 10.
[0022] In the following description, when there is no need to distinguish between the conversation sequence SQ1-1, the conversation sequence SQ2-1, etc., they will be collectively referred to as the "conversation sequence SQ."
[0023] An overview of information processing according to an embodiment is shown in Fig. 1. As shown in Fig. 1, the first server 100 provides service content CT for an online service currently being accessed by the service user U, and also executes processing related to a dialogue with the service user U through a dialogue screen of a chatbot CB displayed together with the service content CT. When providing the service content CT, the first server 100 can acquire attribute information indicating the attributes of the service user U. Furthermore, the first server 100 can acquire a conversation history in the dialogue between the chatbot CB and the service user U, and information related to the reaction of the service user U.
[0024] First, in order to execute the above-described process related to the dialogue with the service user U, the first server 100 transmits a query for an initial conversation sequence SQ to the second server 200 (step S01). Fig. 2 shows an overview of the conversation sequence according to the embodiment.
[0025] The conversation sequence SQ according to the embodiment is information that predefines conversation patterns that indicate a series of conversational contents expected in a dialogue between the chatbot CB and a service user. The administrator of the first server 100 and the second server 200 identifies multiple conversational patterns that are extracted as necessary (essential) groups of conversational contents (utterances and responses) that will take place between the chatbot CB and the service user (for example, the service user U shown in FIG. 1). The administrator then sets each of the identified conversational patterns as a conversational sequence SQ.
[0026] For example, the conversation sequence SQ1-1 shown in FIG. 2 is composed of a series of conversational contents uttered in chronological order, including utterances X1-1, X1-2, X1-3, and utterance (question) Q1-1. The utterance (question) Q1-1 is a question posed by the chatbot CB to the service user U. For example, the utterance (question) Q1-1 is associated with answer options that allow the service user U to select an answer to the question posed by the chatbot CB. When the utterance (question) Q1-1 is displayed in the chatbot CB, the answer options are also displayed.
[0027] 2 includes an utterance X2-1, an utterance (question) Q2-1, an utterance X2-2, and an utterance (question) Q2-2 as a series of conversational contents uttered in chronological order. The utterance (question) Q2-1 and the utterance (question) Q2-2 have the same properties as the utterance (question) Q1-1 described above.
[0028] When the second server 200 receives the inquiry about the first conversation sequence from the first server 100, it selects the first conversation sequence SQ from a plurality of predefined conversation sequences SQ (step S02).Then, the second server 200 transmits an instruction about the selected first conversation sequence SQ to the first server 100 (step S03).
[0029] For example, the second server 200 may preset an initial conversation sequence common to various online services. Alternatively, the second server 200 may preset an initial conversation sequence corresponding to each online service for each online service. In this case, the second server 200 acquires, from the first server 100, an inquiry about the initial conversation sequence SQ, as well as service information for identifying the online service currently being used by the service user U who will be the other party in conversation with the chatbot CB. The second server 200 then selects, as the initial conversation sequence SQ, a conversation sequence SQ that is pre-associated with the acquired service information.
[0030] The second server 200 may also set in advance an initial conversation sequence SQ corresponding to the attributes (such as demographic attributes and psychographic attributes) of the service user who will be the other party in the conversation with the chatbot CB. In this case, the second server 200 acquires, from the first server 100, an inquiry about the initial conversation sequence SQ, as well as attribute information indicating the attributes of the service user U who will be the other party in the conversation with the chatbot CB. Then, the second server 200 selects the conversation sequence SQ associated with the acquired attribute information as the initial conversation sequence SQ.
[0031] The second server 200 may also select the final conversation sequence SQ to be used in the dialogue between the chatbot CB and the service user U in accordance with a predetermined rule.
[0032] When the first server 100 receives the instruction for the first conversation sequence SQ from the second server 200, it executes processing related to the dialogue with the service user U through the chatbot CB in accordance with the received first conversation sequence SQ (step S04). According to the example of the dialogue screen of the chatbot CB shown in Fig. 1, based on the information transmitted from the first server 100, the user terminal 10 displays information D-1 to D-3 corresponding to the utterances included in the conversation sequence SQ from top to bottom in the order set in the conversation sequence SQ.
[0033] When the first server 100 completes the dialogue based on the first conversation sequence SQ, it sends a query for the next conversation sequence SQ to the second server 200 (step S05). At this time, the first server 100 also sends to the second server 200 information for identifying the immediately preceding conversation sequence SQ and information regarding the reaction of the service user U in the dialogue with the chatbot CB.
[0034] When the second server 200 receives the inquiry about the next conversation sequence SQ from the first server 100, it selects the next conversation sequence SQ according to the state of the service user U from among a plurality of predefined conversation sequences SQ using a selection model for selecting a conversation sequence SQ to be used in a dialogue through the chatbot CB (step S06).Then, the second server 200 transmits an instruction for the selected next conversation sequence SQ to the first server 100 (step S07).
[0035] An example of an instruction of a conversation sequence SQ from the second server 200 to the first server 100 will be specifically described below with reference to Fig. 3. Fig. 3 shows an example of an instruction of a conversation sequence SQ from the second server 200 to the first server 100 according to the embodiment.
[0036] 3, the first server 100 transmits a query for the first conversation sequence SQ to the second server 200 (step S11). At this time, the first server 100 transmits attribute information (attribute UA) indicating the attribute of the service user U together with the query for the first conversation sequence SQ.
[0037] When the second server 200 receives the inquiry about the first conversation sequence SQ from the first server 100, it selects the first conversation sequence SQ1-1 and transmits the sequence number "SN101" of the selected first conversation sequence SQ1-1 to the first server 100 (step S12). In addition, the second server 200 holds attribute information indicating the attributes of the service user U received from the first server 100.
[0038] When the first server 100 completes the dialogue with the service user U using the conversation sequence SQ corresponding to the sequence number "SN101" received from the second server 200, it sends a query for the next conversation sequence SQ to the second server 200 (step S13). At this time, the first server 100 sends the query for the next conversation sequence SQ, along with the sequence number "SN101" of the previous conversation sequence SQ1-1, and information "answer R-1" indicating the answer of the service user U in the dialogue with the chatbot CB.
[0039] When the second server 200 receives the inquiry for the next conversation sequence SQ from the first server 100, it uses the selection model to select the next conversation sequence SQ2-1 according to the state of the service user U, and transmits the sequence number "SN201" of the selected next conversation sequence SQ2-1 to the first server 100 (step S14). For example, the second server 200 inputs attribute information indicating the attributes of the service user U (an example of "user information"), the sequence number "SN101" of the previous conversation sequence (an example of "conversation history"), and information "answer R-1" indicating the answer result of the service user U in the dialogue with the chatbot CB (an example of "service user's reaction") into the selection model as information indicating the state of the service user U, and then transmits the sequence number "SN201" output from the selection model to the first server 100.
[0040] The attribute information indicating the attributes of the service user U is not limited to static information such as demographic attributes and psychographic attributes, but may also include dynamic information such as location information and biometric information. In this case, the second server 200 acquires dynamic information such as location information and biometric information of the service user U each time it receives an inquiry about the next conversation sequence SQ from the first server 100, and can select the next conversation sequence SQ based on the status of the service user U updated based on the acquired dynamic information. The second server 200 can also acquire the service usage history (purchase history, reservation history, etc.) of the service user U from the first server 100 as information indicating the status of the service user, and use the acquired service usage history as input information when selecting a conversation sequence.
[0041] 1, the second server 200 sets a reward based on the response of the service user U in the dialogue conducted through the chatbot CB, and performs reinforcement learning of the selection model for each conversation sequence used in the dialogue (step S08). Figure 4 shows a schematic overview of reinforcement learning according to the embodiment.
[0042] As shown in FIG. 4, in reinforcement learning according to the embodiment, the selection model can be regarded as a reinforcement learning agent, and the dialogue between the chatbot CB and the service user U can be regarded as the reinforcement learning environment. In this case, reinforcement learning proceeds in the following procedure. First, the selection model selects a conversation sequence SQ corresponding to the state of the service user U according to a policy that is believed to produce a desired result. Here, the state of the service user U includes attribute information indicating the attributes of the service user U, a conversation sequence (conversation history) used in the dialogue between the chatbot CB and the service user U, and the answer of the service user U in the dialogue with the chatbot CB (the reaction of the service user U). The conversation sequence SQ selected by the selection model is transmitted from the second server 200 to the first server 100.
[0043] Next, when the dialogue using the conversation sequence SQ is completed in the first server 100, the conversation history is transmitted from the first server 100 to the second server 200, and the state of the service user U after the dialogue using the conversation sequence SQ (the conversation sequence SQ used in the dialogue and the service user U's response in the dialogue) is fed back to the selection model. At the same time, a reward based on the reaction of the service user U is fed back to the selection model. The selection model then reviews its policy based on the state of the service user U after the dialogue using the conversation sequence SQ and the reward based on the reaction of the service user U.
[0044] That is, the second server 200 acquires information on the answer (reaction) of the service user U in the dialogue with the chatbot CB from the first server 100 as a change that the action of selecting the conversation sequence SQ has brought to the environment of the dialogue between the chatbot CB and the service user U. Then, the second server 200 sets a reward based on the reaction (answer in the dialogue) of the service user U for the conversation sequence SQ selected by the selection model as an evaluation of the change that the action of selecting the conversation sequence SQ has brought to the environment of the dialogue between the chatbot CB and the service user U.
[0045] In this way, the second server 200 can perform reinforcement learning on a conversation sequence basis to optimize the selection of the conversation sequence SQ by the selection model, for example, by setting a reward based on the response of the service user U in the previous conversation sequence SQ, so as to maximize the reward obtained from the dialogue between the chatbot CB and the service user U.
[0046] Furthermore, during reinforcement learning, the reward set by the second server 200 for the conversation sequence SQ is set in accordance with at least the content of the conversation that took place using the immediately preceding conversation sequence SQ and the outcome of the conversation.
[0047] For example, the second server 200 may set a reward based on whether the service user U responded favorably in the dialogue of the immediately preceding (previous) conversation sequence SQ. Specifically, if the response obtained from the service user U in the dialogue of the immediately preceding conversation sequence SQ was favorable, the second server 200 will provide a positive reward (e.g., +1) for the immediately preceding conversation sequence SQ. On the other hand, if the response obtained from the service user U in the dialogue of the immediately preceding conversation sequence SQ was not favorable, the second server 200 will provide a negative reward (e.g., -1) for the immediately preceding conversation sequence SQ.
[0048] When setting a reward for the most recent conversation sequence SQ, the second server 200 may set the reward according to changes in the service user U's response in past dialogues. For example, if the responses of the service user U in the dialogue before last and the service user U immediately before (previous) were both favorable, a positive reward (e.g., +2) may be given to the most recent conversation sequence SQ; if the service user U in the dialogue before last was favorable but the response of the service user U immediately before (previous) was unfavorable, no reward may be given to the most recent conversation sequence SQ; and if the responses of the service user U in the dialogue before last and the service user U immediately before (previous) were unfavorable, a negative reward (e.g., -2) may be given to the most recent conversation sequence SQ. In this way, the second server 200 can be expected to optimize the selection model so that a conversation sequence SQ is selected according to changes in the service user U's response.
[0049] Furthermore, for example, the second server 200 may set a reward based on whether or not a predetermined conversion associated with the immediately preceding (previous) conversation sequence SQ was obtained from the service user U. Specifically, if the second server 200 is able to obtain predetermined information associated with the immediately preceding conversation sequence SQ from the service user U (for example, if the second server 200 is able to obtain information about a product desired by the service user U), it will provide a positive reward (for example, +1) for the immediately preceding conversation sequence SQ. On the other hand, if the second server 200 is unable to obtain predetermined information associated with the immediately preceding conversation sequence SQ from the service user U, it will provide a negative reward (for example, -1) for the immediately preceding conversation sequence SQ.
[0050] In this way, the second server 200 sets a reward based on the response of the service user U in the immediately preceding conversation sequence SQ and performs learning so that the selection of the conversation sequence SQ by the selection model is optimized. Furthermore, the second server 200 can perform reinforcement learning of the selection model for each attribute of the service user U. As a result, by using the selection model, the second server 200 can increase the likelihood of selecting a conversation sequence SQ that will produce a desirable result depending on the attributes and state of the service user U.
[0051] The second server 200 can use any method to perform reinforcement learning of the selection model. The second server 200 may use a value-based method such as Q-learning or SARSA, or a policy-based method such as a policy gradient method.
[0052] Returning to Figure 1, when the first server 100 receives an instruction for the next conversation sequence SQ from the second server 200, it executes processing related to the dialogue with the service user U through the chatbot CB in accordance with the received next conversation sequence SQ (step S09).
[0053] 3. Configuration of the Second Server According to the Embodiment An example of the configuration of the second server 200 according to the embodiment will be described with reference to Fig. 5. Fig. 5 shows an example of the configuration of the second server 200 according to the embodiment. As shown in Fig. 5, the second server 200 includes a communication unit 210, a storage unit 220, and a control unit 230.
[0054] (Regarding the communication unit 210) The communication unit 210 is realized by, for example, a network interface card (NIC). The communication unit 210 is connected to the network N by wire or wirelessly. The second server 200 transmits and receives information to and from other devices such as the user terminal 10 and the first server 100 via the network N.
[0055] (Regarding the storage unit 220) The storage unit 220 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. For example, the storage unit 220 has a conversation sequence storage unit 221, a selection model storage unit 222, and a user information storage unit 223.
[0056] (Conversation sequence storage unit 221) The conversation sequence storage unit 221 stores information on conversation sequences used in a dialogue with a service user (for example, service user U shown in FIG. 1) through the chatbot CB. Fig. 6 is a diagram showing an overview of information on conversation sequences according to the embodiment.
[0057] 6, the conversation sequence information stored in conversation sequence storage unit 221 has multiple items such as a "sequence number" item, a "conversation pattern" item, an "answer reception content" item, an "available service" item, etc. These items in the conversation sequence information are associated with each other.
[0058] The "sequence number" field stores an identification number that is individually assigned to each conversation sequence to identify the conversation sequence. The "conversation pattern" field stores information about the conversation pattern included in the conversation sequence. The "answer reception content" field stores content for receiving answers that is displayed in association with the utterances (questions) included in the conversation pattern. Furthermore, if the answer reception content includes multiple answers for receiving evaluations from service users, each answer is associated in advance with an attribute value indicating whether the answer is favorable or not. The "compatible services" field stores information indicating various online services to which the conversation sequence is applicable.
[0059] (Selection model storage unit 222) The selection model storage unit 222 stores information about selection models used when selecting conversation sequences. Fig. 7 is a diagram showing an overview of the information about selection models according to the embodiment.
[0060] 7, the information about the selection model stored in the selection model storage unit 222 according to the embodiment has a plurality of items such as a "model ID" item, a "corresponding attribute" item, a "model information" item, etc. These items that the information about the selection model has are associated with each other.
[0061] The "Model ID" field stores identification information for identifying the selection model. The "Corresponding Attributes" field stores information indicating the attributes of the service user (for example, service user U shown in Figure 1) corresponding to the selection model. The "Model Information" field stores information about the selection model's policy and various information that constitutes the selection model, such as various parameters.
[0062] (User information storage unit 223) The user information storage unit 223 stores information about service users who use various online services (for example, service user U shown in FIG. 1). Fig. 8 is a diagram showing an overview of user information according to the embodiment.
[0063] 8, the user information stored in the user information storage unit 223 according to the embodiment has a plurality of items such as a "user ID" item, an "attribute information" item, an "interaction history" item, etc. These items in the user information are associated with each other.
[0064] The "User ID" field stores identification information for identifying service users of various online services (for example, service user U shown in Figure 1). The "Attribute Information" field stores information about the service user's attributes, such as demographic attributes, psychographic attributes, location information, and biometric information. The "Dialogue History" field stores the dialogue history, including conversation sequences selected in dialogue with chatbot CB.
[0065] The user information storage unit 223 may store the service usage history of the service user as user information related to the service user. For example, the second server 200 (control unit 230) can acquire the service usage history of the service user from the first server 100, associate the acquired service usage history with the identification information (user ID) of the service user, and register it in the user information storage unit 223.
[0066] (Regarding the control unit 230) The control unit 230 is realized by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like executing various programs stored in a storage device inside the second server 200 using RAM as a work area. The control unit 230 is also realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0067] The control unit 230 shown in FIG. 5 has a selection unit 231, an instruction unit 232, and a learning unit 233, and these units realize or execute the functions and actions of the information processing described below. The control unit 230 may have an internal configuration divided into multiple processing units that realize or execute the functions and actions of the information processing described below. The control unit 230 is not limited to the configuration shown in FIG. 5, and may have other configurations as long as they perform the information processing described below. Functional units other than those shown in FIG. 5 may be added to the control unit 230 in response to addition of processes to be executed by the second server 200.
[0068] (Selection unit 231) The selection unit 231 selects a conversation sequence SQ according to the state of a service user of an online service (for example, service user U shown in FIG. 1) from a plurality of conversation sequences SQ in which conversation patterns indicating a series of conversation contents expected in a dialogue between the service user and the service user through a chatbot are predefined. For example, when the selection unit 231 receives an inquiry about a conversation sequence SQ from the first server 100 via the communication unit 210, the selection unit 231 selects a conversation sequence.
[0069] Furthermore, the selection unit 231 selects a conversation sequence from among the multiple conversation sequences SQ stored in the conversation sequence storage unit 221, using a selection model (for example, the selection model shown in FIG. 2) that selects a conversation sequence SQ to be used in a conversation by setting a reward based on the reaction of a service user (for example, service user U shown in FIG. 1) in a conversation conducted through the chatbot. For example, the selection unit 231 acquires a selection model associated with attribute information indicating the attributes of the service user included in the inquiry for the conversation sequence SQ acquired from the first server 100. Then, the selection unit 231 inputs the most recent conversation sequence and the reaction of service user U into the acquired selection model, thereby selecting a conversation sequence linked to the sequence number output from the selection model from among the multiple conversation sequences stored in the conversation sequence storage unit 221.
[0070] (Instruction section 232) The instruction unit 232 instructs the first server 100 (an example of an "external device") that executes processing related to the dialogue to use the conversation sequence SQ selected by the selection unit 231. For example, the instruction unit 232 transmits the sequence number of the conversation sequence to the first server 100 via the communication unit 210.
[0071] (Learning Section 233) The learning unit 233 performs reinforcement learning of a selection model that selects a conversation sequence SQ to be used in a conversation by setting a reward based on the response of a service user (for example, service user U shown in Figure 1) in a dialogue conducted through the chatbot CB.
[0072] For example, the learning unit 233 performs reinforcement learning of the selection model by setting a reward based on whether or not the service user's reaction was favorable in at least the dialogue based on the immediately preceding conversation sequence SQ.
[0073] Furthermore, for example, the learning unit 133 performs reinforcement learning of the selection model by setting a reward based on whether or not a predetermined conversion associated with the immediately preceding conversation sequence has been obtained from the service user.
[0074] 4. Processing Procedure According to the Embodiment The following describes the procedure of information processing executed by the second server 200 according to the embodiment. Fig. 9 shows an example of the procedure of processing executed by the second server 200 according to the embodiment. The procedure of processing shown in Fig. 9 is executed by the control unit 230 of the second server 200. The procedure of processing shown in Fig. 9 is repeatedly executed while the second server 200 is operating.
[0075] As shown in FIG. 9, the selection unit 231 acquires a query for a conversation sequence SQ from the first server 100 (step S101).
[0076] Furthermore, the selection unit 231 uses the selection model to select a conversation sequence according to the state of the service user (for example, the service user shown in FIG. 1) who is having a conversation with the chatbot CB (step S102).
[0077] Furthermore, the instruction unit 232 instructs the first server 100 about the conversation sequence selected by the selection unit 231 (step S103).
[0078] Furthermore, the learning unit 233 sets a reward based on the reaction of the service user in the dialogue based on the conversation sequence, executes reinforcement learning of the selection model (step S104), and ends the processing procedure shown in FIG.
[0079] [5. Modifications] The information processing device, the information processing method, and the information processing program according to the present application may be implemented in various different forms other than the above-described embodiment. Modifications of the above-described embodiment will be described below.
[0080] (5-1. About conversation sequences) The conversation sequences according to the above embodiments may be set for each predetermined topic, such as a search sequence, a today's mood sequence, or a campaign sequence. For example, a search sequence may have a conversation pattern such as "What kind of book are you looking for? → Please select a genre → ...". For example, a today's mood sequence may have a conversation pattern such as "How are you feeling today? → ...". For example, a campaign sequence may have a conversation pattern such as "We're running a great campaign today → ...".
[0081] Furthermore, the first server 100 may execute control such as not displaying a predetermined conversation more than N times (N is a natural number).
[0082] (5-2. Learning the selection model) In the above embodiment, the learning of the selection model executed in the second server 200 may be executed for each attribute of the service user. That is, the second server 200 sets a common selection model for each service user with the same attribute and executes reinforcement learning. In this case, the second server 200 may collect the reaction (content of the response in the dialogue) of each service user for each conversation sequence at a predetermined timing and execute reinforcement learning based on the collected reaction.
[0083] (5-3. About Chatbots) In the above embodiment, the first server 100 may display a virtual character image corresponding to the chatbot CB on the dialogue screen of the chatbot CB. At this time, the first server 100 may change the facial expression of the character image depending on the content of the response of the service user who is the dialogue partner. Furthermore, the first server 100 may change the appearance of the character image depending on the attributes of the service user, etc.
[0084] (6. Effects) The second server 200 according to the embodiment is an information processing device that controls processing related to a dialogue conducted with a service user of an online service through a chatbot, and includes a selection unit 231 and an instruction unit 232. The selection unit 231 selects a conversation sequence according to the state of the service user from a plurality of conversation sequences in which conversation patterns indicating the contents of a series of conversations expected in the dialogue are predefined. The instruction unit 232 instructs the first server 100, which executes processing related to the dialogue, to use the conversation sequence selected by the selection unit 231.
[0085] For this reason, the second server 200, which is an example of an information processing device according to the embodiment, can efficiently collect information from service users of online services. For example, the second server 200 according to the embodiment can be expected to have an effect of improving the quality of the user experience in the dialogue by conducting a dialogue with the service user in units of predefined conversation sequences. That is, the second server 200 according to the embodiment can realize a more natural conversation with the service user U through the chatbot CB. As a result, it is possible to increase the possibility that the dialogue with the chatbot CB will continue, and to efficiently collect information from the service user through the dialogue.
[0086] The second server 200 further includes a learning unit 233 that performs reinforcement learning of a selection model for selecting a conversation sequence to be used in a conversation by setting a reward for the conversation sequence based on the reaction of the service user in the conversation conducted through the chatbot. The selection unit 231 selects a conversation sequence using the selection model.
[0087] Furthermore, the selection unit 231 uses the selection model to select a conversation sequence based on user information about the service user, a conversation history indicating the contents of the most recent conversation, and the service user's reaction in the conversation.
[0088] Therefore, the second server can produce a natural conversation for each service user according to the reaction of the service user in the conversation.
[0089] The user information also includes attribute information indicating the attributes of the service user and the service usage history of the online service.
[0090] Therefore, the second server can produce a natural conversation that matches the attributes of the service user and the usage situation of the service.
[0091] In this way, the second server 200 executes reinforcement learning of the selection model on a conversation sequence basis, thereby preventing degradation of the quality of the user experience in dialogues using conversation sequences due to the randomness of learning inherent in reinforcement learning. Furthermore, by using conversation sequences to perform learning on a conversation sequence basis, the second server 200 can reduce the amount of learning compared to normal learning using reinforcement learning for the chatbot CB, thereby improving system efficiency.
[0092] Furthermore, the learning unit 233 performs reinforcement learning of the selection model by setting a reward based on whether or not the service user's reaction was favorable in at least the dialogue in the immediately preceding conversation sequence.
[0093] Therefore, the second server 200 can optimize the selection of a conversation sequence by the selection model so that the service user's reaction is favorable in a dialogue using the conversation sequence.
[0094] Furthermore, the learning unit 233 performs reinforcement learning of the selection model by setting a reward based on whether or not a predetermined conversion associated with the immediately preceding conversation sequence has been obtained from the service user.
[0095] Therefore, the second server 200 can optimize the selection of conversation sequences by the selection model so that desirable results are obtained from the service user through dialogue using the conversation sequences.
[0096] [7. Hardware Configuration] The second server 200 according to the above-described embodiment and each of the modifications is realized, for example, by a computer 1000 having a configuration as shown in Fig. 10. Fig. 10 is a hardware configuration diagram showing an example of a computer that realizes the functions of the second server 200 according to the embodiment and each of the modifications.
[0097] The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which an arithmetic unit 1030, a primary storage device 1040, a secondary storage device 1050, an output IF (Interface) 1060, an input IF 1070, and a network IF 1080 are connected by a bus 1090.
[0098] The arithmetic device 1030 operates based on programs stored in the primary storage device 1040 and secondary storage device 1050, programs read from the input device 1020, and the like, and executes various processes. The primary storage device 1040 is a memory device, such as a RAM, that temporarily stores data used by the arithmetic device 1030 for various calculations. The secondary storage device 1050 is a storage device in which data used by the arithmetic device 1030 for various calculations and various databases are registered, and is realized by a ROM (Read Only Memory), HDD, flash memory, or the like.
[0099] The output IF 1060 is an interface for transmitting information to be output to an output device 1010 that outputs various types of information, such as a monitor or a printer, and is realized by a connector conforming to a standard such as USB (Universal Serial Bus), DVI (Digital Visual Interface), or HDMI (High Definition Multimedia Interface), etc. The input IF 1070 is an interface for receiving information from various input devices 1020, such as a mouse, keyboard, and scanner, and is realized by, for example, USB.
[0100] The input device 1020 may be a device that reads information from an optical recording medium such as a CD (Compact Disc), a DVD (Digital Versatile Disc), or a PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory. The input device 1020 may also be an external storage medium such as a USB memory.
[0101] The network IF 1080 receives data from other devices via the network N and sends it to the arithmetic device 1030, and also transmits data generated by the arithmetic device 1030 to other devices via the network N.
[0102] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output IF 1060 and the input IF 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0103] For example, when the computer 1000 functions as the second server 200 according to the embodiment, the arithmetic device 1030 of the computer 1000 executes a program (for example, an information processing program) loaded onto the primary storage device 1040, thereby realizing the same functions as the control unit 230. That is, the arithmetic device 1030 cooperates with the program (for example, an information processing program) loaded onto the primary storage device 1040 to realize the processing by the second server 200 according to the embodiment.
[0104] [8. Other] Of the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.
[0105] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. For example, the selection unit 231 and the instruction unit 232 of the control unit 230 of the second server 200 may be functionally integrated. Furthermore, for example, the first server 100 and the second server 200 in the information processing system SYS may be a single information processing device that is functionally and physically integrated.
[0106] Furthermore, the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0107] The above describes in detail the embodiments of the present application based on several drawings, but these are merely examples, and the present invention can be implemented in other forms that include the embodiments described in the Disclosure of the Invention section and that have been modified and improved in various ways based on the knowledge of those skilled in the art.
[0108] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, a control section can be read as control means or a control circuit. [Explanation of symbols]
[0109] N Network SYS Information Processing System 10 User terminal 100 First Server 200 Second Server 210 Communications Department 220 Storage section 221 Conversation sequence memory unit 222 Selection Model Memory 223 User information storage unit 230 Control Unit 231 Selection Section 232 Instruction section 233 Learning Department
Claims
1. An information processing device that controls processing related to a dialogue conducted between a service user of an online service and a chatbot, a selection unit that selects a conversation sequence from a plurality of conversation sequences in which conversation patterns indicating a series of conversation contents expected in the dialogue are defined in advance; an instruction unit that instructs an external device that executes processing related to the dialogue to use the conversation sequence selected by the selection unit; and The selection unit When receiving an inquiry about an initial conversation sequence from the external device, select, from the plurality of conversation sequences, an initial conversation sequence common to the online services, an initial conversation sequence preset for each online service, or an initial conversation sequence corresponding to an attribute of the service user who will be the other party in the conversation; Each time an inquiry for a next conversation sequence is received, dynamic information including location information and biometric information of the service user is acquired, and the next conversation sequence is selected based on the state of the service user updated based on the acquired dynamic information.
1. An information processing device comprising:
2. a learning unit that executes reinforcement learning of a selection model that selects a conversation sequence to be used in the conversation by setting a reward for the conversation sequence based on a response of the service user in the conversation conducted through the chatbot, for the conversation sequence, in units of the conversation sequence; and The selection unit A conversation sequence is selected using the selection model associated with attribute information indicating attributes of the service user included in the inquiry acquired from the external device.
2. The information processing apparatus according to claim 1, wherein:
3. The selection unit A conversation sequence linked to a sequence number output from the selection model is selected by inputting a conversation sequence immediately before the next conversation sequence and the service user's reaction in the dialogue to the selection model.
3. The information processing apparatus according to claim 2, wherein:
4. The learning unit The reward is set based on whether the service user responded favorably to the dialogue in at least the immediately preceding conversation sequence, thereby performing reinforcement learning of the selection model.
3. The information processing apparatus according to claim 2, wherein:
5. The learning unit When setting a reward for the immediately preceding conversation sequence, the reward is set according to a change in the reaction of the service user in the previous conversation.
5. The information processing apparatus according to claim 4,
6. The learning unit The reward is set based on whether or not a predetermined conversion associated with the immediately preceding conversation sequence has been obtained from the service user, thereby performing reinforcement learning of the selection model.
3. The information processing apparatus according to claim 2, wherein:
7. An information processing method for controlling processing related to a dialogue conducted between a service user of an online service and a chatbot, comprising: a selection step of selecting a conversation sequence from a plurality of conversation sequences in which conversation patterns indicating a series of conversation contents expected in the dialogue are defined in advance; an instruction step of instructing an external device that executes processing related to the dialogue to use the conversation sequence selected in the selection step; Including, The selection step includes: When receiving an inquiry about an initial conversation sequence from the external device, select, from the plurality of conversation sequences, an initial conversation sequence common to the online services, an initial conversation sequence preset for each online service, or an initial conversation sequence corresponding to an attribute of the service user who will be the other party in the conversation; Each time an inquiry for a next conversation sequence is received, dynamic information including location information and biometric information of the service user is acquired, and the next conversation sequence is selected based on the state of the service user updated based on the acquired dynamic information.
1. An information processing method comprising:
8. A computer that controls the processing of interactions between users of an online service and the chatbot, a selection step of selecting a conversation sequence from a plurality of conversation sequences in which conversation patterns indicating a series of conversation contents expected in the dialogue are defined in advance; an instruction step of instructing an external device that executes processing related to the dialogue to use the conversation sequence selected by the selection step; Execute The selection procedure comprises: When receiving an inquiry about an initial conversation sequence from the external device, select, from the plurality of conversation sequences, an initial conversation sequence common to the online services, an initial conversation sequence preset for each online service, or an initial conversation sequence corresponding to an attribute of the service user who will be the other party in the conversation; Each time an inquiry for a next conversation sequence is received, dynamic information including location information and biometric information of the service user is acquired, and the next conversation sequence is selected based on the state of the service user updated based on the acquired dynamic information. An information processing program characterized by:
Citation Information
Patent Citations
Conversation providing device, conversation providing method, and program
JP2018181018A
Interaction control device, learning device, interaction control method, learning method, control program, and recording medium
JP2019045978A
Information processing device, information processing method, and program
JP2020027527A
Method for constructing ai chatbot obtained by combining class classification and regression classification
JP2021056941A
Information processing program, information processing method, and information processing system
JP2022113675A