Information processing device, information processing method, and information processing program

The information processing apparatus enhances one-on-one meetings by using a chatbot to guide users through career-related dialogues, addressing the lack of goal-oriented support in existing systems by generating and mediating utterances to facilitate career advancement and growth.

JP7837256B2Active Publication Date: 2026-03-30유겐가이샤티아이에스
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-20
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing chatbots lack the ability to intervene in conversations between users to support the appropriate progress towards achieving specific goals, such as career advancement and growth, often failing to provide meaningful support in one-on-one meetings.

Method used

An information processing apparatus and method that utilizes a chatbot to facilitate dialogues between users by generating utterance content based on user responses, guiding them through predefined scenarios to achieve career-related goals, and mediating interactions to ensure appropriate progression.

Benefits of technology

Enhances the effectiveness of one-on-one meetings by helping users discover their interests, strengths, and set goals, thereby supporting career development and growth through guided dialogues and intelligent intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837256000001
    Figure 0007837256000001
  • Figure 0007837256000002
    Figure 0007837256000002
  • Figure 0007837256000003
    Figure 0007837256000003
Patent Text Reader

Abstract

To provide a service capable of supporting appropriate progress of discussions between users to achieve their goals.SOLUTION: An information processing apparatus according to one embodiment of the present disclosure comprises: a control unit for executing a first interaction which is an interaction with a first user based on utterance candidates set for a scenario for guiding the first user to a specific goal; and a generation unit for generating utterance content that is output to a second user according to the scenario at the current stage of the scenario, based on the first user's response information in the first interaction, wherein the utterance content mediates a second interaction which is an interaction between the first user and the second user who has a predetermined relationship with the first user.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] There is known a chatbot that realizes a dialogue with a user based on scenario information that realizes a dialogue according to input information.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the above prior art, there is room for improvement in providing a service that can support the appropriate progress of the conversation between users toward achieving the goal.

[0005] For example, the above prior art only discloses the function of a chatbot that realizes a dialogue with a user along a scenario, and there is no concept that the chatbot intervenes in the conversation between users in the first place.

[0006] Therefore, in the above prior art, it is not always possible to provide a service that can support the appropriate progress of the conversation between users toward achieving the goal.

[0007] The present disclosure has been made in view of the above, and proposes an information processing apparatus, an information processing method, and an information processing program that can provide a service that can support the appropriate progress of the conversation between users toward achieving the goal.

Means for Solving the Problems

[0008] To solve the above problems, an information processing device according to the present disclosure comprises: a control unit that executes a first dialogue, which is a conversation with the first user, based on utterance candidates set for a scenario to guide the first user to a specific goal; a generation unit that generates utterance content to be output to the second user according to the current stage of the scenario, based on the response information of the first user in the first dialogue, a second user having a predetermined relationship with the first user, and utterance content that mediates the second dialogue, which is a conversation with the first user. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a diagram showing the configuration of the system according to the embodiment. [Figure 2] Figure 2 shows an overview of the server device's operation. [Figure 3] Figure 3 is an example of dialogue control processing by a server device, shown in Figure (1). [Figure 4] Figure 4 is an example of dialogue control processing by a server device, as shown in Figure (2). [Figure 5] Figure 5 shows an example of the configuration of a server device according to this embodiment. [Figure 6] Figure 6 shows an example of a scenario information storage unit according to the embodiment. [Figure 7] Figure 7 shows an example of a participant information storage unit according to the embodiment. [Figure 8] Figure 8 shows an example of a dialogue history storage unit according to the embodiment. [Figure 9] Figure 9 is a flowchart showing the steps of the interaction control process performed by the server device. [Figure 10] Figure 10 is a block diagram showing an example of the hardware configuration of an information processing device according to an embodiment. [Modes for carrying out the invention]

[0010] Embodiments of this disclosure will be described in detail below with reference to the drawings. Note that these embodiments do not limit the information processing apparatus, information processing method, and information processing program related to this disclosure. Furthermore, in the following embodiments, the same parts are denoted by the same reference numerals to avoid redundant explanations.

[0011] [Embodiment] [1. Introduction] Statistics show that the top concerns for working adults today include "I want to do work I can be passionate about, but I don't know what I want to do," "I don't have anyone to talk to," and "My current job isn't a good fit." Therefore, an increasing number of companies are actively creating opportunities for discussion to support career development.

[0012] For example, one-on-one meetings between a supervisor and a subordinate (often called "1on1s") are sometimes conducted with the aim of improving the subordinate's productivity, building trust between supervisors and subordinates, and increasing engagement. However, in many companies, 1on1s currently implemented only serve as a "place for managing work progress," and it cannot be said that they are fulfilling their original role as a "place for supporting career advancement and growth."

[0013] As an example, employee satisfaction was particularly low for evaluation items such as "It allows us to deepen our understanding of each other's personalities and strengths and weaknesses," "It serves as a place to align our views on future career aspirations," "It serves as a place to share company strategies and tactics," "It serves as a place to be encouraged to take on challenges," and "It serves as a place to acquire new technologies and skills." This indicates that these items need to be strengthened in particular during 1on1 meetings.

[0014] Therefore, in one aspect of the present disclosure, it is an object to support employees to maximize their abilities by discovering "what they want to do" for their own growth and making use of it in their work. The information processing apparatus according to the proposed technology of the present disclosure (hereinafter sometimes referred to as "the information processing apparatus according to the embodiment") supports, for example, a conversation between a supervisor and a subordinate (or between personnel and an employee) by intervening as a chatbot so that this conversation becomes a suitable place for career advancement and growth support.

[0015] In this embodiment, as an example of the conversation, a meeting held one-on-one between a supervisor and a subordinate, that is, a 1on1, will be taken as an example for explanation. However, the scene to which the information processing apparatus according to the embodiment is applied is not limited to 1on1. For example, the information processing apparatus according to the embodiment can also provide a service that supports a conversation among three people, namely a subordinate, the supervisor directly above the subordinate, and a personnel staff member. It is applicable not only to conversations held for a specific purpose such as career advancement and growth support, but also to any conversation (for example, daily conversations among colleagues, etc.).

[0016] [2. System Configuration] From here, the configuration of the system according to the embodiment will be described using FIG. 1. FIG. 1 is a diagram showing the configuration of the system according to the embodiment. In FIG. 1, as an example of the system according to the embodiment, System 1 is shown.

[0017] As shown in FIG. 1, System 1 may be configured to include a first user device 10, a second user device 20, and a server device 100. Also, in System 1, the first user device 10 (first user device), the second user device 20 (second user device), and the server device 100 may be communicably connected by wire or wirelessly via a network N.

[0018] The server device 100 is an example of an information processing device according to an embodiment, and has functions of interacting with users and mediating interactions between users. For this reason, the server device 100 provides a so-called chatbot service. The function of the chatbot may be controlled by an information processing program according to an embodiment.

[0019] For example, according to the control of the information processing program, the server device 100 outputs a response message to each user device (the first user device 10, the second user device 20) for input information indicating a question or the like input by the user of the first user device 10 (the first user) or the user of the second user device 20 (the second user).

[0020] Here, in the present embodiment, the first user is the supported side in 1on1 (for example, a subordinate), and the second user is the supporting side in 1on1 (for example, a supervisor). In such a case, the server device 100 provides a career support service (hereinafter referred to as "service SA") as the above chatbot service.

[0021] Service SA may be provided, for example, via an application AP (app AP) corresponding to the server device 100. As a result, the first user can interact with the chatbot (server device 100) and the second user by accessing the server device 100 via the first user device 10 in which the app AP is introduced. Also, the second user can interact with the chatbot (server device 100) and the first user by accessing the server device 100 via the second user device 20 in which the app AP is introduced. Details of the interaction performed by these three parties (the first user, the second user, the chatbot) will be described with reference to FIG. 2.

[0022] Furthermore, Figure 1 shows an example of how Service SA is used on an organizational (e.g., company) basis. Figure 1 shows an arbitrary organizational Tx that uses Service SA, and the server device 100 may be configured to provide appropriate support to each organizational Tx by managing information (e.g., interaction history information) separately for each organizational Tx.

[0023] Specifically, the server device 100 is characterized by facilitating basic conversations in accordance with pre-prepared scenarios to guide the first user toward career advancement, while appropriately generating utterances to be used in the conversation based on the first user's responses, etc. The server device 100 may use machine learning technology to generate the utterances.

[0024] For example, the server device 100 may input the current response information from the first user into an AI model trained based on the history information of the dialogue with the first user (including the first user's response information), thereby generating utterances that respond to the first user's current response. Furthermore, the generated utterances include content that prompts the second user to provide support to help the first user overcome their situation (for example, an insufficient answer to a question, or difficulty in answering a question). Thus, the server device 100 also has the function of estimating the first user's intentions from the first user's response information.

[0025] The first user device 10 is an information processing terminal used by the first user described above, and examples include smartphones, tablet devices, notebook PCs, desktop PCs, mobile phones, PDAs, etc. The second user device 20 is an information processing terminal used by the second user described above, and similarly may be a smartphone, tablet device, notebook PC, desktop PC, mobile phone, PDA, etc. Furthermore, the first user device 10 and the second user device 20 may be AI speakers or wearable devices.

[0026] The first user device 10 and the second user device 20 are equipped with application APs that can access the server device 100.

[0027] Here, the second user may include a member of the human resources department of organization Tx, and the human resources member can operate the second user device 20 and pre-register each employee's attribute information (e.g., name, age, department, etc.) in the application AP. As a result, the server device 100 can recognize the personnel structure within organization Tx (e.g., a tree structure showing relationships between departments), and can generate an appropriate chat room based on the recognition result.

[0028] As a typical example, the server device 100 may, in response to a request from a second user, create a chat room that only one first user, designated by the second user, can enter. Alternatively, the server device 100 can also create a dedicated chat room where the first user can converse only with a chatbot. Furthermore, the server device 100 can create chat rooms on a departmental (or team) basis that only individuals belonging to a specific department (or team) can enter.

[0029] [3. Overview of operation by the server device] Next, we will explain the overview of the operation of the server device 100 using Figure 2. Figure 2 is a diagram illustrating the overview of the operation of the server device 100. In Figure 2, we see a scene in which the first user U11 and the second user U12 engage in a conversation in the chat room RM1 for career advancement consultation, which was set up by the second user U12, who belongs to organization T1.

[0030] Furthermore, as shown in the example in Figure 2, the first user U11 and the second user U12 are in a subordinate / superior relationship, with the first user U11's name being "A" and the second user U12's name being "B". In the following explanation of Figure 2, the first user U11 will be referred to as "subordinate A" and the second user U12 as "superior B".

[0031] First, in the example shown in Figure 2, the server device 100 operates the chatbot BT in the chat room RM1 according to the information processing program according to the embodiment. In other words, the chatbot BT is an automated conversation tool that facilitates a dialogue about career advancement between subordinate A and superior B in the chat room RM1.

[0032] Here, the server device 100 may also function as a learning device that performs learning processing related to the AI ​​model. The AI ​​model is generated by machine learning performed by the server device 100 and incorporated into an information processing program that controls the functions of chatbot BT. For this reason, chatbot BT can also be called an AI chatbot. Furthermore, the following becomes possible under the control of the server device 100.

[0033] For example, subordinate A can view screen G10, which displays chat room RM1, via the first user device 10. Therefore, information entered by subordinate A into the input fields on screen G10 is transmitted via server device 100 and displayed on screen G10 as the content of subordinate A's speech.

[0034] Similarly, supervisor B can view screen G20, which displays chat room RM1, via the second user device 20. Information entered by supervisor B into the input fields on screen G20 is transmitted via server device 100 and displayed on screen G20 as supervisor B's spoken content.

[0035] As described above, the participants in chat room RM1 are subordinate A, superior B, and chatbot BT. A dialogue takes place through the statements of these three parties, and this dialogue can be classified into three types, as shown in detail in Figure 2. Specifically, one dialogue in chat room RM1 includes a first dialogue DL1, which is a dialogue between chatbot BT and subordinate A, and a second dialogue DL2, which is a dialogue between subordinate A and superior B. Furthermore, one dialogue in chat room RM1 includes utterances MUx, which play a mediating role in ensuring that the second dialogue DL2 progresses appropriately as a forum for career advancement and growth support.

[0036] The utterance content MUx is information that the server device 100 outputs as the utterance content of the chatbot BT, and is generated based on the response information of subordinate A in the first dialogue DL1. The server device 100 basically implements the first dialogue DL1 according to a pre-prepared scenario to guide the employees of organization T1 towards career advancement, but instead of relying on the scenario, it proceeds with a single dialogue in chat room RM1 by generating this utterance content MUx which is optimal for subordinate A's response in the first dialogue DL1.

[0037] As shown in the example in Figure 2, the server device 100 generates a speech content MUx that prompts supervisor B to provide support for subordinate A's situation, based on subordinate A's response information. For example, the server device 100 generates a speech content MUx configured to prompt supervisor B to provide support so that subordinate A's situation can be resolved (e.g., the answer to chatbot BT's question is insufficient, or subordinate A is having trouble answering chatbot BT's question). The server device 100 then outputs the speech content MUx into the chat room RM1 as a statement from chatbot BT to supervisor B. Thus, as shown in Figure 2, the speech content MUx can be understood as part of the dialogue between chatbot BT and supervisor B.

[0038] Next, using Figure 2, we will explain a scenario for guiding someone towards career advancement. The scenario used by the server device 100 is structured in stages to guide the first user (subordinate A in the example in Figure 2), who is the one being supported, towards their future goals.

[0039] The scenario group SNG shown in Figure 2 consists of six scenario pieces: Scenario SN1, Scenario SN2, Scenario SN3, Scenario SN4, Scenario SN5, and Scenario SN6. The dialogue is defined to proceed in this order, step by step. Each scenario may also be associated with a candidate utterance that can be used as the content of the chatbot BT's speech. This point will be explained in the scenario information storage unit 121 in Figure 6.

[0040] As shown in the example in Figure 2, scenario SN1 includes potential utterances that elicit "areas of interest." For example, by including this step, users may be able to find clues about "what they want to do" by referring to their "areas of interest." Furthermore, the server device 100 may generate reference information that leads to "discovering what they want to do" based on the user's response information to the utterances actually presented and the AI ​​model, and may further present this reference information to the user.

[0041] Scenario SN2 includes potential utterances that elicit information about the user's "strengths." For example, by including this step, the user may be able to find clues about "what they want to do" by referring to their "strengths." Furthermore, the server device 100 may generate reference information that leads to "discovering what they want to do" based on the user's response information to the utterances actually presented and the AI ​​model, and may further present this reference information to the user.

[0042] Scenario SN3 includes utterance suggestions that guide the user toward "discovering what they want to do." In this step, users can build upon the "areas of interest" and "strengths" they considered in the previous step, making it easier for them to "discover what they want to do."

[0043] Scenario SN4 includes utterance candidates that guide the user to "hypothesize / verify their vision." This step allows the user to formulate a hypothesis about the future they want to become. Furthermore, the server device 100 may generate reference information to assist in "hypothesizing / verifying their vision" based on the user's response information to the utterances actually presented and the AI ​​model, and may further present this reference information to the user.

[0044] Scenario SN5 includes utterance candidates that encourage "short-term or long-term goal setting." In this step, the server device 100 may generate user-specific goals based on the user's response information to the utterances actually presented and the AI ​​model, and may further present these goals to the user. The server device 100 may also provide real-time reminders of the goals.

[0045] Here, the user can also consider an action plan that suits them by conducting a gap analysis against their goals. The server device 100 may generate reference information to assist in considering an action plan based on the user's response information to the actual utterances presented and the AI ​​model, and may further present this reference information to the user.

[0046] Scenario SN6 includes utterance candidates that prompt the user to reflect on the information obtained in the previous steps (scenarios SN1-SN5) as a result of "action / practice." This step allows the user to gain a trigger to actually take action. For example, the server device 100 may generate utterances that inspire the user based on the AI ​​model and further present these utterances to the user.

[0047] Now, we have described the scenarios used by the server device 100, but for example, the number of scenarios included in the scenario group SNG, the content of the scenarios, the configuration requirements (candidate utterances) according to the content of the scenarios, and the order of the scenarios are not limited to the example in Figure 2. However, it is desirable that the series of scenarios be set based on favorable results obtained from previous career support. For example, the scenario group SNG may be set independently for each organization Tx, or it may be set by the service provider that provides the service SA.

[0048] [4. Examples of dialogue control based on a scenario] Next, using Figures 3 and 4, we will explain an example of dialogue control processing that follows the scenario shown in Figure 2. In Figures 3 and 4, we continue to use the example from Figure 2 and show a scene in which a dialogue takes place between subordinate A, superior B, and chatbot BT. In the examples of Figures 3 and 4, it is assumed that the server device 100 already has an AI model used to generate the utterance content MUx.

[0049] [4-1. Example of Dialogue Control (1)] Figure 3 is Figure (1) illustrating an example of dialogue control processing by the server device 100. Figure 3 shows a scene in which the server device 100 controls the dialogue according to scenario SN1 in Figure 2. Figure 3(a) shows screen G10 displayed on the first user device 10 on the subordinate A's side as a dialogue screen showing chat room RM1. On the other hand, Figure 3(b) shows screen G20 displayed on the second user device 20 on the superior B's side as a dialogue screen showing chat room RM1.

[0050] For example, in scenario SN1, the output of potential utterances such as "What are your areas of interest?" and "Could you also write about your field of work?" may be defined in a tree structure according to the expected utterances for the user. In the example in Figure 3, the server device 100 acquires "What are your areas of interest?" from the potential utterances included in scenario SN1 as the target utterance M11 and displays it in the chat room RM1.

[0051] Subordinate A operates the first user device 10 and inputs "cooking, soccer..." as the response M12 to the spoken content M11.

[0052] The server device 100 may determine whether there is an utterance candidate among the utterance candidates corresponding to scenario SN1 that is suitable for response content M12, and if there is, it may prioritize presenting that utterance candidate. On the other hand, if there is no utterance candidate among the utterance candidates corresponding to scenario SN1 that is suitable for response content M12, the server device 100 may generate an utterance for response content M12 based on the situation of subordinate A estimated from response content M12 and the AI ​​model.

[0053] In the example in Figure 3, the server device 100, recognizing that "Can you also write about your field of work?" exists as a suitable utterance candidate for response content M12, acquires "Can you also write about your field of work?" as the utterance content M13 to be used and displays it in chat room RM1.

[0054] Subordinate A operates the first user device 10 and inputs "marketing..." as response content M14 to utterance content M13.

[0055] The server device 100 may determine whether there is an utterance candidate among the utterance candidates corresponding to scenario SN1 that is suitable for response content M14, and if there is, it may prioritize presenting that utterance candidate. On the other hand, if there is no utterance candidate among the utterance candidates corresponding to scenario SN1 that is suitable for response content M14, the server device 100 may generate an utterance for response content M14 based on the situation (intention) of subordinate A estimated from response content M14 and the AI ​​model.

[0056] In the example in Figure 3, the server device 100, upon realizing that there are no suitable utterance candidates for response content M14, acquires situational information indicating that subordinate A's situation, "The answer to the question is insufficient," is estimated from response content M14. The server device 100 then inputs the situational information into the AI ​​model and generates utterance content M15 from the output of the AI ​​model. According to the example in Figure 3, the server device 100 generates utterance content M15, "What do you think of my boss, Mr. B?"

[0057] In this example, the AI ​​model is trained to output an utterance such as "What do you think of Mr. / Ms. XX?" to encourage the second user to provide support in order to resolve the situation if the estimated situation for the first user is "insufficient answer to the question." Furthermore, in order to obtain such an AI model, the server device 100 may train the AI ​​model using the response information of the person being supported, which is included in the dialogue history information, and the response information from this history information that plays an intermediary role, as training data.

[0058] The dialogue history information referred to here may be the history of conversations between subordinate A, superior B, and chatbot BT, or it may be the history of other 1-on-1 meetings conducted via chatbot BT. As another example, the dialogue history information may be the history of conversations conducted only by humans, without including chatbot BT.

[0059] Returning to the explanation of Figure 3, utterance M15 plays a role in mediating the second dialogue DL2 to ensure it proceeds appropriately as a setting for career advancement and growth support, and corresponds to an example of utterance MUx explained in Figure 2.

[0060] Supervisor B can gain insights in response to the utterance M15 mediated by the chatbot BT. Specifically, supervisor B operates the second user device 20 and inputs "A, what are you going to cook?" as the response M16 to the utterance M15.

[0061] [4-2. Example of Dialogue Control (2)] Figure 4 is Figure (2) showing an example of dialogue control processing by the server device 100. Figure 4 shows a scene in which the server device 100 controls the dialogue according to scenario SN3 in Figure 2. Figure 4(a) shows screen G10 displayed on the first user device 10 on the subordinate A's side as a dialogue screen showing chat room RM1. On the other hand, Figure 4(b) shows screen G20 displayed on the second user device 20 on the superior B's side as a dialogue screen showing chat room RM1.

[0062] For example, in scenario SN3, the tree structure may define which of the utterance candidates to output, such as "Think about what you really want to do," or "Think about the area of ​​expertise you mentioned earlier," is expected to be used for the user. In the example in Figure 4, the server device 100 acquires "Think about what you really want to do" from the utterance candidates included in scenario SN3 as the target utterance M21 and displays it in the chat room RM1.

[0063] Subordinate A operates the first user device 10 and inputs "Hmm, that's difficult..." as the response M22 to the spoken content M21.

[0064] The server device 100 may determine whether there is an utterance candidate among the utterance candidates corresponding to scenario SN3 that is suitable for response content M22, and if there is, it may prioritize presenting that utterance candidate. On the other hand, if there is no utterance candidate among the utterance candidates corresponding to scenario SN3 that is suitable for response content M22, the server device 100 may generate an utterance for response content M22 based on the situation of subordinate A estimated from response content M22 and the AI ​​model.

[0065] In the example shown in Figure 4, the server device 100, recognizing that "Please consider this in conjunction with the area of ​​expertise you mentioned earlier" exists as a suitable utterance candidate for response content M22, acquires "Please consider this in conjunction with the area of ​​expertise you mentioned earlier" as the utterance content M23 to be used and displays it in the chat room RM1.

[0066] Subordinate A operates the first user device 10 and inputs "I can only think of things unrelated to work" as the response M24 to the spoken content M23.

[0067] The server device 100 may determine whether there is an utterance candidate among the utterance candidates corresponding to scenario SN3 that is suitable for response content M24, and if there is, it may prioritize presenting that utterance candidate. On the other hand, if there is no utterance candidate among the utterance candidates corresponding to scenario SN3 that is suitable for response content M24, the server device 100 may generate an utterance for response content M24 based on the situation (intention) of subordinate A estimated from response content M24 and the AI ​​model.

[0068] In the example in Figure 4, the server device 100, upon finding no suitable utterance candidate for response content M24, acquires situational information indicating that subordinate A's situation, "having trouble answering the question," is estimated from response content M24. The server device 100 then inputs the situational information into the AI ​​model and generates utterance content M25 from the AI ​​model's output. According to the example in Figure 4, the server device 100 generates utterance content M25, "It's okay if it's not relevant! I'd also appreciate support from my superior, B."

[0069] In this example, the AI ​​model can be said to have learned to output an utterance such as "Please help Mr. / Ms. XX on the other end" when the estimated situation for the first user is "having trouble answering a question," prompting the second user to provide support to resolve the situation. Furthermore, in order to obtain such an AI model, the server device 100 may train the AI ​​model using the response information of the person being supported, which is included in the dialogue history information, and the response information from this history information that plays an intermediary role, as training data.

[0070] The dialogue history information referred to here may be the history of conversations between subordinate A, superior B, and chatbot BT, or it may be the history of other 1-on-1 meetings conducted via chatbot BT. As another example, the dialogue history information may be the history of conversations conducted only by humans, without including chatbot BT.

[0071] Returning to the explanation of Figure 4, utterance M25 plays a role in mediating the second dialogue DL2 to ensure it proceeds appropriately as a setting for career advancement and growth support, and corresponds to an example of utterance MUx explained in Figure 2.

[0072] Supervisor B can gain insights from the utterance M25 mediated by the chatbot BT. Specifically, supervisor B operates the second user device 20 and inputs "It seems there are people in other departments aiming for XX." as the response M26 to the utterance M25.

[0073] [4-3. Summary] As shown in the examples in Figures 3 and 4, the server device 100 reads the situation of subordinate A and the atmosphere of the dialogue between subordinate A and superior B (second dialogue DL2), and can be said to be skillfully mediating the relationship between the two. As a result, a world view is realized in which superior B skillfully supports subordinate A as the second dialogue DL2 progresses, making it easier for subordinate A to express what they are thinking. Furthermore, if subordinate A makes appropriate statements, the dialogue will proceed according to the scenario, allowing them to find out "what they should do."

[0074] [5. Configuration of the Information Processing Device] An information processing apparatus according to the embodiment will be described using Figure 5. Figure 5 is a diagram showing an example configuration of a server device 100 according to the embodiment. As shown in Figure 5, the server device 100 has a communication unit 110, a storage unit 120, and a control unit 130.

[0075] (Regarding Communications Unit 110) The communication unit 110 is implemented, for example, by a NIC (Network Interface Card). For example, the communication unit 110 is connected to the network N by wire or wireless connection and transmits and receives information between, for example, the first user device 10 and the second user device 20.

[0076] (Regarding memory unit 120) The memory unit 120 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by a storage device such as a hard disk or optical disc. The memory unit 120 includes a scenario information storage unit 121, a participant information storage unit 122, a dialogue history storage unit 123, and a learning data storage unit 124.

[0077] (Regarding the scenario information storage unit 121) The scenario information storage unit 121 stores scenario information used for dialogue control processing. The scenario information may be structured in stages to guide the person being supported toward a future goal.

[0078] Here, Figure 6 shows an example of the scenario information storage unit 121 according to the embodiment. In the example in Figure 6, the scenario information storage unit 121 has items such as "scenario type," "scenario ID," "scenario content," and "utterance candidate."

[0079] "Scenario Type" is information indicating the type of scenario, and as shown in Figure 6, there are types such as "Career Support". For example, a scenario of type "Career Support" is used when controlling the dialogue for career support. "Scenario ID" is information that identifies the scenario.

[0080] "Scenario Content" is information that outlines the content of the scenario indicated by the "Scenario ID". "Potential Utterances" indicate what should be output as utterances by the chatbot BT in order to elicit information from or have the user consider the "Scenario Content".

[0081] Figure 6 shows an example where scenario ID "SN1" is associated with scenario content "Areas of Interest" and "What are your areas of interest?". In the example in Figure 6, scenario SN1 is defined as the 1st step. This example shows that in the first step of the first dialogue DL1, the dialogue is defined to proceed with a topic that elicits "areas of interest" according to scenario SN1.

[0082] Figure 6 also shows an example where scenario ID "SN2" is associated with scenario content "What you're good at" and "Please choose what you're good at from the options." Here, scenario SN2 is defined as the 2nd step. This example shows that in the second step after the dialogue has progressed in scenario SN1, the dialogue is defined to proceed with a topic that elicits "what you're good at," in accordance with scenario SN2.

[0083] Figure 6 also shows an example where scenario ID "SN3" corresponds to the scenario content "Discovering what you want to do" and "Think about what you really want to do." Here, scenario SN3 is defined as the 3rd step. This example shows that in the third step, after the dialogue has progressed in scenario SN2, the dialogue is defined to proceed in accordance with scenario SN3 with the aim of helping the participant "discover what they want to do."

[0084] Figure 6 also shows an example where Scenario ID "SN4" corresponds to the scenario content "Hypothesizing / Verifying a Vision" and "What do you want to be like in 3 years?". Here, Scenario SN4 is defined as the 4th step. This example shows that in the 4th step, after the dialogue has progressed in Scenario SN3, the dialogue is defined to proceed with the purpose of helping participants "formulate a hypothesis about what they want to be like in the future" in accordance with Scenario SN4.

[0085] Figure 6 also shows an example where Scenario ID "SN5" corresponds to the scenario content "Short-term / Long-term goal setting" and "Let's think about what to do tomorrow." Here, Scenario SN5 is defined as the 5th step. This example shows that in the 5th step, after the dialogue has progressed in Scenario SN4, the dialogue is defined to proceed with the purpose of helping to "set short-term or long-term goals" in accordance with Scenario SN5.

[0086] Figure 6 also shows an example where Scenario ID "SN6", Scenario content "Action Practice", and "Were we able to achieve the goal of XX?" are associated. Here, Scenario SN6 is defined as the 6th step. This example shows that in the 6th step, after the dialogue has progressed in Scenario SN5, the dialogue is defined to proceed in accordance with Scenario SN6 with the purpose of assisting in "reflection".

[0087] (Regarding participant information storage unit 122) The participant information storage unit 122 stores information indicating the participants in each dialogue (including the first dialogue DL1 and the second dialogue DL2) involving the chatbot BT. Figure 7 shows an example of the participant information storage unit 122 according to this embodiment. In the example in Figure 7, the participant information storage unit 122 has items such as "organization ID", "dialogue ID", "user ID", "department", "name", and "relationship".

[0088] "Organization ID" is information that identifies the organization Tx using the service SA. "Dialogue ID" is information that identifies a dialogue that took place in a single chat room and involved the chatbot BT. "User ID" is information that identifies the user who participated in the dialogue identified by the "Dialogue ID".

[0089] "Department" is information indicating the department to which the user identified by the "User ID" belongs. "Name" is information indicating the name of the user identified by the "User ID". "Relationship" is information indicating the relationship that exists between the users indicated by each "User ID" associated with the "Dialogue ID".

[0090] Figure 7 shows an example of how the organization ID "T1", dialogue ID "LG1", and user IDs "U11" and "U12" are associated. This example corresponds to the example in Figure 2, and shows that in dialogue LG1, which takes place in chat room RM1 for career advancement consultation set up by the second user U12 belonging to organization T1, both the first user U11 and the second user U12 participate.

[0091] (Regarding the dialogue history storage unit 123) The dialogue history storage unit 123 stores the history of each dialogue (including the first dialogue DL1 and the second dialogue DL2) involving the chatbot BT. Figure 8 shows an example of the dialogue history storage unit 123 according to this embodiment. In the example in Figure 8, the dialogue history storage unit 123 has items such as "Dialogue ID", "Dialogue Type", "Previous Utterance History", and "Later Utterance History".

[0092] The "Dialogue ID" is information that identifies a conversation that took place in a single chat room, including a conversation involving the chatbot BT. The "Dialogue ID" referred to here is the same as the "Dialogue ID" in Figure 7.

[0093] The "Dialogue Type" is information that indicates which participants are having a dialogue with whom, among those who participated in the dialogue identified by the "Dialogue ID". Figure 8 shows an example where the dialogue ID "LG1" is associated with the dialogue type "bot to U11". In this example, dialogue LG1 in chat room RM1 includes a dialogue between chatbot BT and the first user U11 (i.e., the first dialogue DL1), and chatbot BT is defined as the "former" of the dialogue, and the first user U11 is defined as the "latter" of the dialogue.

[0094] Figure 8 also shows an example where the dialogue ID "LG1" is associated with the dialogue type "U12toU11". In this example, dialogue LG1 in chat room RM1 includes a dialogue between a second user U12 and a first user U11 (i.e., the second dialogue DL2), and the second user U12 is defined as the "former" of the dialogue, and the first user U11 is defined as the "latter" of the dialogue.

[0095] Furthermore, Figure 8 shows that the dialogue ID "LG1" is associated with the dialogue type "bottoU12," and it exists between the dialogue types "bottoU11" and "U12toU11." In this example, dialogue LG1 in chat room RM1 includes a dialogue between chatbot BT and a second user U12, with chatbot BT defined as the "former" of the dialogue and the second user U12 as the "latter." According to this example, the dialogue of dialogue type "bottoU12" includes utterance content MUx, which is structured to prompt the second user U12 for support.

[0096] The "Former Participant Utterance History" is the utterance history of the participant defined as "Former Participant" in the "Dialogue Type," and includes the "Content of Utterance" by this participant and the "Date and Time" when the "Content of Utterance" was entered.

[0097] The "Latter Utterance History" is the utterance history of a participant defined as "Latter" in the "Dialogue Type," and includes the "Content of Utterance" and the "Date and Time" when the "Content of Utterance" was entered.

[0098] (Regarding the learning data storage unit 124) The learning data storage unit 124 stores learning data used to train the AI ​​model. For example, the learning data storage unit 124 may store the history information stored in the dialogue history storage unit 123 as learning data. Alternatively, the learning data storage unit 124 may store the history information of dialogues conducted only by humans, without including the chatbot BT, as learning data.

[0099] Here, if the server device 100 estimates from the response information of the first user that the first user has encountered a specific situation, it may, for example, conduct a questionnaire regarding mediation with the second user after the conversation has ended. This point will be explained using examples such as Figure 2.

[0100] For example, the server device 100 may conduct a survey such as, "["During a conversation in Room RM1, Person A encountered a difficult situation. In this case, should we ask Person B for a response?", and "If we ask for a response, what kind of content would you like to receive? Please select from the following options."]". In such an example, the server device 100 may store as training data a pair of feedback from the second user (Person B) to the survey and situational information indicating the estimated situation of the first user (Person A).

[0101] (Regarding the control unit 130) Returning to Figure 5, the control unit 130 is implemented by a CPU (Central Processing Unit) or MPU (Micro Processing Unit), etc., which executes various programs (for example, the information processing program according to this embodiment) stored in the storage device inside the server device 100 using RAM as the working area. Alternatively, the control unit 130 can be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array).

[0102] As shown in Figure 5, the control unit 130 includes an input receiving unit 131, an interaction control unit 132, a generation unit 133, and a learning unit 134, and realizes or executes the information processing functions and operations described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Figure 5, and other configurations are also possible as long as they perform the information processing described later. Also, the connection relationships of the various processing units in the control unit 130 are not limited to the connection relationships shown in Figure 5, and other connection relationships are also possible.

[0103] (Regarding the input reception unit 131) The input receiving unit 131 receives input of utterance content. For example, the input receiving unit 131 receives input of utterance content displayed on the screen corresponding to the chat room from each participant participating in the chat room generated by the dialogue control unit 132. For example, the input receiving unit 131 receives input of utterance content from both the first user, who is the person receiving support according to the scenario, and the second user, who is a party involved in the dialogue with the first user.

[0104] As shown in the example in Figure 2, the input receiving unit 131 receives input of speech content from the first user U11 and the second user U12, who are both in the chat room RM1.

[0105] (Regarding the dialogue control unit 132) The dialogue control unit 132 performs dialogue control processing. For example, the dialogue control unit 132 generates a chat room in response to a user's request. Then, the dialogue control unit 132 controls the dialogue in each chat room. For example, the dialogue control unit 132 controls a user's entry into a chat room based on the user information corresponding to that chat room. This processing may be performed, for example, based on information stored in the participant information storage unit 122.

[0106] Furthermore, when the dialogue control unit 132 receives input of speech content from a user who has entered the chat room, it displays the speech content on the screen corresponding to the chat room.

[0107] Furthermore, the dialogue control unit 132 controls the content of the chatbot BT's utterances in the chat room so that the dialogue progresses according to the scenario. For example, the dialogue control unit 132 executes the first dialogue DL1, which is a dialogue between the chatbot BT and the first user, based on the utterance candidates set for a scenario to guide the first user to a specific goal. Specifically, the dialogue control unit 132 selects the most appropriate utterance to present at that moment from the utterance candidates and displays the selected utterance as the content of the chatbot BT's utterance on the screen corresponding to the chat room.

[0108] Furthermore, as explained above, in the chat room, a second dialogue DL2 takes place between the first user and the second user in parallel with the first dialogue DL1. For this reason, for example, if the dialogue control unit 132 estimates from the first user's response information that the first user is in a specific situation or has a specific intention, it may control the generation unit 133 to generate utterances that correspond to the estimation results. For example, if the dialogue control unit 132 estimates that the first user is in a specific situation or has a specific intention, it may control the generation unit 133 to generate utterances that support the situation or intention indicated by the estimation results for the second user.

[0109] The dialogue control unit 132 may also perform a process to estimate the status of the first user based on the first user's response information. For example, the dialogue control unit 132 may estimate the status of the first user based on the content of the first user's utterance as the first user's response information.

[0110] For example, the dialogue control unit 132 may calculate the degree of appropriateness of the first user's utterances in relation to the utterances presented by the chatbot BT, and estimate the first user's situation from the calculated degree of appropriateness.

[0111] As another example, the dialogue control unit 132 may estimate the first user's situation based on the combination of keywords included in the first user's utterance, the number of keywords, and the correctness of the keywords as answers. Alternatively, the dialogue control unit 132 may estimate the first user's situation based on the time elapsed from when the chatbot BT presented the utterance until the first user inputted the utterance (i.e., the time the first user was silent) as the first user's response information.

[0112] As yet another example, the dialogue control unit 132 may use a predictive model that predicts changes in the first user's situation, taking the first user's response information as input, to estimate the first user's situation.

[0113] (Regarding the generation unit 133) The generation unit 133 generates utterances that mediate the second dialogue DL2, which is a dialogue between the first user and a second user who has a predetermined relationship with the first user, based on the first user's response information in the first dialogue DL1, and which are output to the second user according to the scenario at the current stage. For example, the generation unit 133 may generate utterances that encourage the second user to support the situation of the first user, which is estimated from the first user's response information.

[0114] (Regarding Learning Section 134) The learning unit 134 generates an AI model that the generation unit 133 uses to generate speech content. For example, the learning unit 134 uses dialogue history information as training data to train the model, thereby generating an AI model that outputs information related to speech content.

[0115] For example, the learning unit 134 may generate an AI model using the history information of conversations conducted according to a scenario as training data. Alternatively, the learning unit 134 may generate an AI model using the history information of conversations conducted only by humans, without including the chatbot BT, as training data. More specifically, the learning unit 134 may generate an AI model by training the model using pairs of response information from the person being supported and response information playing an intermediary role, which are included in the history information of the conversation, as training data.

[0116] Furthermore, as described above, if the server device 100 estimates from the response information of the first user that the first user has fallen into a specific situation, it may conduct a questionnaire regarding mediation with the second user. In this case, the server device 100 may generate an AI model by training a model using the second user's feedback on the questionnaire and situational information indicating the estimated situation of the first user as training data.

[0117] Now, let's return to the explanation of the generation unit 133. The generation unit 133 generates utterances to be presented as chatbot BT based on the first user's response information and the AI ​​model. For example, the generation unit 133 generates utterances based on the first user's situation estimated from the first user's response information and the AI ​​model. Note that the first user's situation here also includes the concept of the intent of the utterance estimated from the first user's response information.

[0118] As another example, if the first user does not input any utterance after a predetermined time has elapsed since the chatbot BT presented the utterance content, the dialogue control unit 132 may infer that the first user is unsure what to say. In such a case, the generation unit 133 may generate, for example, an utterance that advises on input or an utterance that prompts input, and control the system so that the generated utterance content is presented to the first user. Alternatively, the generation unit 133 may generate an utterance that prompts the second user to provide support so that the first user can resolve their confusion, and control the system so that the generated utterance content is presented to the second user.

[0119] [6. Processing Procedure] Next, the operation procedure of the server device 100 will be explained using Figure 9. Figure 9 is a flowchart showing the procedure of the dialogue control processing performed by the server device 100. In the example of Figure 9, it is assumed that the AI ​​model has already been generated by the learning process. Also, in Figure 9, as with Figure 2, the dialogue control processing procedure will be explained using a scenario in which the first user U11 and the second user U12 engage in dialogue in the chat room RM1.

[0120] First, the dialogue control unit 132 identifies the progress of the scenario based on the history information of the first dialogue DL1, which is a conversation between the chatbot BT and the first user U11 (step S901).

[0121] Next, the dialogue control unit 132 selects an utterance to be used to start the dialogue from among the utterance candidates that correspond to the identified scenario (step S902).

[0122] Then, the dialogue control unit 132 displays the selected utterance as a statement from chatbot BT on the screen showing chat room RM1 (step S903).

[0123] In this state, the input receiving unit 131 determines whether or not it has received input of speech content from the user (step S904). If the input receiving unit 131 has not received input of speech content from the user (step S904; No), it waits until it receives input of speech content from the user.

[0124] On the other hand, if the dialogue control unit 132 receives input of utterance content from a user (step S904; Yes), it identifies the position of the user who entered the utterance content (step S905). For example, if the user who entered the utterance content is the first user U11 (subordinate A), the dialogue control unit 132 can refer to the participant information storage unit 122 and identify that the first user U11 is the person being supported. Also, if the user who entered the utterance content is the second user U12 (supervisor B), the dialogue control unit 132 can refer to the participant information storage unit 122 and identify that the second user U12 is the supporter.

[0125] The generation unit 133 determines whether the user who input the utterance is a person receiving support (step S906).

[0126] If the user who input the utterance is not a person being supported, i.e., a supporter (step S906; No), the generation unit 133 generates utterance content corresponding to the response from this user (second user U12). This utterance content is displayed as a statement from chatbot BT on the screen showing chat room RM1 (step S907), and the process returns to step S904.

[0127] On the other hand, if the user who input the utterance is in the position of being supported (step S906; Yes), the dialogue control unit 132 determines whether or not there is an utterance candidate among the utterance candidates corresponding to the current scenario that is suitable for the response content of this user (first user U11) (step S908).

[0128] If the dialogue control unit 132 finds a suitable utterance candidate (step S908; Yes), it displays this utterance candidate as a statement from chatbot BT on the screen showing chat room RM1 (step S909).

[0129] On the other hand, if there are no suitable utterance candidates for the response content (step S908; No), the dialogue control unit 132 estimates the situation of the user who entered the utterance content (step S910). For example, the dialogue control unit 132 may estimate the situation from the utterance content itself, or it may estimate the user's situation based on the time elapsed from when the utterance content from the chatbot BT was displayed until the user entered the utterance content.

[0130] The generation unit 133 generates utterances in response to the response content based on situation information indicating the user's situation and the AI ​​model (step S911). Here, according to the AI ​​model, if the user's situation is a specific negative situation (for example, an insufficient answer to a question, difficulty in answering a question, etc.), utterances may be output that instruct the user in the role of a supporter (second user U12) to support this situation. If such utterances are output, the generation unit 133 may process the utterances output by the AI ​​model based on the user in the role of a supporter and the current scenario.

[0131] For example, if the current scenario is scenario SN3, the generation unit 133 can generate "Mr. B, please give advice to Mr. A so that he can discover what he wants to do."

[0132] The dialogue control unit 132 then displays the generated utterance as a statement from chatbot BT on the screen indicating chat room RM1 (step S912). The process then returns to step S904.

[0133] [7. Other Embodiments] From here, other embodiments of the server device 100 will be described. The server device 100 may be implemented in various different forms other than those described above. Therefore, other embodiments of the server device 100 will be described below.

[0134] [7-1. Suppression of bot statements based on feedback (1)] In the above embodiment, the server device 100 demonstrated an example of controlling dialogue in a single chat room in which a first user, a second user, and chatbot BT participate. However, the server device 100 can also generate a dedicated chat room in which the first user can converse only with chatbot BT. Such a chat room may be provided under a name such as "Anything Goes Consultation Room with the Bot," and in effect, it is controlled so that only one first user can enter.

[0135] Using an example like Figure 2, the server device 100, for example, in response to a request from subordinate A, generates a consultation room RM11 (an example of a dedicated chat room) where subordinate A can have a one-on-one conversation with the chatbot BT, and controls it so that only subordinate A can enter. In consultation room RM11, no pre-defined speech content is output, and subordinate A can express various worries and complaints about work and daily life.

[0136] In this case, the server device 100 may also use the history information of conversations that took place in consultation room RM11 as training data to train the model. If this is the case, when a conversation takes place in chat room RM1, the server device 100 may present the chatbot BT with utterances generated based on private matters or complaints about superior B that it heard from subordinate A in consultation room RM11. In other words, even though superior B is also participating in chat room RM1, the server device 100 may have the chatbot BT utter matters that subordinate A would not want superior B to know.

[0137] Thus, the first user has information that they would be troubled if spoken to the chatbot BT in a chat room in which the second user is also participating. Therefore, the server device 100 may accept feedback from the first user regarding the content of speech that they would like not to be output in future chat rooms (especially a single chat room in which the first user, the second user, and the chatbot BT are participating).

[0138] For example, in a chat room in which the first user, the second user, and chatbot BT are participating, the server device 100 may receive feedback from the first user regarding utterances presented by chatbot BT that the user wishes not to be output in future chat rooms. For example, the server device 100 may conduct a survey asking the first user which utterances from the list presented by chatbot BT the user does not wish to be output in the future. In this case, the first user can specify chatbot BT utterances that include, for example, information that the user does not want others to know.

[0139] Furthermore, the server device 100 may receive feedback from the first user in a private chat room where only the first user participates, regarding the utterances the first user has entered that they wish not to be output in future chat rooms. For example, the server device 100 may conduct a survey of the first user to determine which utterances from a list of utterances entered by the first user they wish not to be included in the chatbot BT's utterances. In this case, the first user can specify, for example, utterances that include information they do not want others to know.

[0140] Thus, if the generation unit 133 receives feedback from the first user regarding utterances that the user wishes not to output in future chat rooms (utterances presented as chatbot BT that the user wishes not to output in future chat rooms / utterances entered by the first user that the user wishes not to output in future chat rooms), the generation unit 133 controls the utterances based on this feedback. For example, if the utterances generated by the generation unit 133 based on the AI ​​model contain keywords specified in the feedback, the generation unit 133 may process them to exclude these keywords before presenting them.

[0141] As a result, the server device 100 can increase user satisfaction with conversations involving the chatbot BT.

[0142] [7-2. Suppression of bot statements based on feedback (2)] Furthermore, the learning unit 134 may train its model on the tendency of utterances that many first users wish not to output, based on feedback accumulated so far. For example, the learning unit 134 may train its model using pairs of attributes of the first user and feedback from first users with these attributes as training data. In this case, the generation unit 133 may estimate the utterances that the first user wishes not to output in the current dialogue based on the attributes of the first user and the trained AI model, and generate utterances according to the estimation results.

[0143] As a result, the server device 100 can increase user satisfaction with conversations involving the chatbot BT.

[0144] [8. Hardware Configuration] Next, an example of the hardware configuration of the information processing device (server device 100) according to the embodiment will be described. Figure 10 is a block diagram showing an example of the hardware configuration of the information processing device according to the embodiment. Referring to Figure 10, the information processing device includes, for example, a processor 801, a ROM 802, a RAM 803, a host bus 804, a bridge 805, an external bus 806, an interface 807, an input device 808, an output device 809, a storage device 810, a drive 811, a connection port 812, and a communication device 813. Note that the hardware configuration shown here is just an example, and some of the components may be omitted. Furthermore, it may also include components other than those shown here.

[0145] (Processor 801) The processor 801 functions, for example, as an arithmetic processing unit or a control unit, and controls the overall operation or part of the operation of each component based on various programs recorded in the ROM 802, RAM 803, storage 810, or removable recording medium 901.

[0146] (ROM802, RAM803) ROM 802 is a means of storing programs loaded into the processor 801 and data used for calculations. RAM 803 temporarily or permanently stores, for example, programs loaded into the processor 801 and various parameters that change as needed when executing those programs.

[0147] (Host bus 804, bridge 805, external bus 806, interface 807) The processor 801, ROM 802, and RAM 803 are interconnected, for example, via a host bus 804 capable of high-speed data transmission. On the other hand, the host bus 804 is connected to an external bus 806, which has a relatively low data transmission speed, via a bridge 805. The external bus 806 is also connected to various components via an interface 807.

[0148] (Input device 808) Input devices 808 may include, for example, a mouse, keyboard, touch panel, buttons, switches, and levers. Furthermore, a remote controller (hereinafter referred to as a remote control) capable of transmitting control signals using infrared or other radio waves may also be used as an input device 808. Additionally, input devices 808 may include audio input devices such as microphones.

[0149] (Output device 809) The output device 809 is a device capable of visually or audibly notifying the user of acquired information, such as a display device like a CRT (Cathode Ray Tube), LCD, or organic EL; an audio output device like a speaker or headphones; a printer, mobile phone, or facsimile. Furthermore, the output device 809 according to this embodiment includes various vibration devices capable of outputting tactile stimuli. The output device 809 may also include an AI speaker or a wearable device.

[0150] (Storage 810) Storage 810 is a device for storing various types of data. Examples of storage devices that can be used for storage 810 include magnetic storage devices such as hard disk drives (HDDs), semiconductor storage devices, optical storage devices, or magneto-optical storage devices.

[0151] (Drive 811) The drive 811 is a device that reads information recorded on a removable recording medium 901, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, or writes information to the removable recording medium 901.

[0152] (Connection port 812) Connection port 812 is a port for connecting external devices 902, such as a USB (Universal Serial Bus) port, IEEE1394 port, SCSI (Small Computer System Interface), RS-232C port, or optical audio terminal.

[0153] (Communication device 813) The communication device 813 is a communication device for connecting to a network, and may include, for example, a communication card for wired or wireless LAN, Bluetooth®, or WUSB (Wireless USB), a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various types of communication.

[0154] (Removable recording medium 901) The removable recording medium 901 may be, for example, DVD media, Blu-ray® media, HD DVD media, or various semiconductor storage media. Of course, the removable recording medium 901 may also be, for example, an IC card equipped with a contactless IC chip, or an electronic device.

[0155] (External connection device 902) External connected devices 902 include, for example, a printer, a portable music player, a digital camera, a digital video camera, or an IC recorder.

[0156] In the case where the information processing device according to this embodiment is a server device 100, the storage unit 120 is implemented by ROM 802, RAM 803, and storage 810. Furthermore, the control unit 130 implemented by the processor 801 reads and executes the control programs (for example, the information processing program according to this embodiment) that implement the input receiving unit 131, the dialogue control unit 132, the generation unit 133, and the learning unit 134 from ROM 802, RAM 803, etc.

[0157] [9. Other] Of the processes described above as being performed automatically, all or part of them may be performed manually. Furthermore, all or part of the processes described as being performed manually may be performed automatically using known methods. In addition, the processing procedures, specific names, and various data and parameters shown in the above documents and drawings may be changed at will unless otherwise specified. For example, the various information shown in each drawing is not limited to the information illustrated.

[0158] Furthermore, each component of the illustrated device is a functional concept and does not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown. Moreover, each component may be configured by functionally or physically distributing and integrating all or part of it in any unit, depending on various loads and usage conditions. In addition, the processes described above may be combined and executed as appropriate, within a non-contradictory range.

[0159] Although embodiments of the present application have been described in detail above with reference to several drawings, these are illustrative examples, and the present invention can be implemented in various other forms with modifications and improvements based on the knowledge of those skilled in the art, starting with the embodiments described in the disclosure section of the invention. [Explanation of symbols]

[0160] 1 System 10. First user device 20 Second user device 100 Server Devices 120 Storage section 121 Scenario Information Storage Unit 122 Participant information storage unit 123 Dialogue History Memory Unit 124 Learning Data Storage Unit 130 Control Unit 131 Input Reception Section 132 Dialogue Control Unit 133 Generation part 134 Learning Department

Claims

1. A control unit that performs a first dialogue, which is a conversation with the first user, based on utterance candidates set for a scenario to guide the first user to a specific goal, Based on the response information of the first user in the first dialogue, a generation unit generates utterances that mediate the second dialogue, which is a dialogue between the first user and a second user having a predetermined relationship with the first user, and which are output to the second user according to the scenario at the current stage among the scenarios. An information processing device equipped with the following features.

2. The control unit executes a chatbot that realizes the first dialogue using utterance candidates from the scenarios that correspond to the scenario corresponding to the progress of the first dialogue. The information processing apparatus according to feature 1.

3. The system further includes a learning unit that learns a model that outputs information about the utterance content based on the history of the dialogue, The generation unit generates the utterance content based on the first user's response information and the model. The information processing apparatus according to feature 1.

4. The learning unit learns the model using learning data that includes situational information indicating the user's situation estimated from the history information of the dialogue. The generation unit generates the utterance content based on the first user's situation estimated from the first user's response information and the model. The information processing apparatus according to claim 3.

5. The generation unit generates speech content that allows the second user to support the situation of the first user, which is estimated from the first user's response information. The information processing apparatus according to feature 1.

6. If the generation unit receives feedback from the first user regarding utterances that the user wishes not to be output in the dialogue room where the second dialogue takes place, it controls the utterances based on the feedback. The information processing apparatus according to feature 1.

7. Based on the trends in utterance content indicated by the feedback, the generation unit estimates the utterance content that the first user wishes not to be output in the current dialogue, and generates utterance content according to the estimation result. The information processing apparatus according to feature 6.

8. An information processing method performed by an information processing device, A control step of executing a first dialogue, which is a conversation with the first user, based on utterance candidates set for a scenario to guide the first user to a specific goal, Based on the response information of the first user in the first dialogue, a generation step generates utterance content that mediates a second dialogue, which is a dialogue between the first user and a second user having a predetermined relationship with the first user, and which is output to the second user according to the scenario at the current stage among the scenarios. An information processing method characterized by including

9. A control procedure for executing a first dialogue, which is a conversation with the first user, based on utterance candidates set for a scenario to guide the first user to a specific goal, A generation procedure for generating utterances that mediate a second dialogue between the first user and a second user having a predetermined relationship with the first user, based on the response information of the first user in the first dialogue, wherein the utterances are output to the second user according to the scenario at the current stage among the scenarios. An information processing program that causes an information processing device to execute.

Citation Information

Patent Citations

  • Conversation collection device, conversation collection system, and conversation collection method

    JP2019074865A

  • Information processing system, method and program

    JP2020052439A

  • Information processing apparatus, information processing system, information processing method, and program

    JP2021093139A

  • Provision program, provision device, provision method, and provision system

    JP2021117845A

  • Information processing system, information processing device, information processing method, and recording medium

    WO2017163515A1