Information processing system, information processing method, and program
The information processing system effectively evaluates user interactions using machine learning models to control dialogue devices, addressing the lack of effective user interaction control methods by generating and assessing user performance in scenarios like customer service.
Patent Information
- Application Number
- JP2025171304
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-29
- Filing Date
- 2025-10-09
- Publication Date
- 2026-02-04
AI Technical Summary
Existing technologies lack effective methods for controlling devices that interact with users, particularly in evaluating user interactions based on dialogue data using machine learning models.
An information processing system that includes a dialogue device, a control device, and a terminal device, utilizing machine learning models to generate and evaluate user interactions, with evaluation results generated by different evaluators in separate locations.
Enables effective evaluation of user interactions through dialogue systems, allowing for precise assessment of user performance in various scenarios, such as customer service, using machine learning models to control dialogue devices and generate evaluation results.
Smart Images

Figure 2026017555000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing system, an information processing method, and a program. [Background technology]
[0002] There are known techniques for interacting with people, such as a technique for acquiring answers from a job seeker to predetermined questions, evaluating the job seeker's interview results numerically based on the answers and evaluation criteria, and displaying the evaluation results as information on the interview results. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-31089 Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides techniques for controlling devices that interact with a user. [Means for solving the problem]
[0005] An information processing system according to one aspect of the present disclosure includes at least one memory and at least one processor, wherein the at least one processor acquires a first system prompt from a plurality of candidate system prompts based on at least first information, and inputs the at least first system prompt into a machine learning model to execute a dialogue between an interlocutor and a dialogue device, stores dialogue data including at least information related to the interlocutor's utterance in the at least one memory, and generates one or more evaluation results related to the interlocutor using the dialogue data stored in the at least one memory, wherein at least some of the evaluation results included in the one or more evaluation results are evaluation results acquired by a terminal device of a first evaluator, and at least some of the evaluation results included in the one or more evaluation results are evaluation results acquired by a terminal device of a second evaluator. The first information includes at least one of the following: industry, occupation, attributes of the interlocutor, attributes of the first evaluator, attributes of the second evaluator, evaluation criteria of the first evaluator, evaluation criteria of the second evaluator, progress of the dialogue, or information about the dialogue between the interlocutor and the dialogue device; each of the multiple system prompt candidates can be used to obtain dialogue data used to generate evaluation results obtained by each of the multiple evaluator's terminal devices; the first system prompt includes at least information indicating the relationship between the interlocutor and the dialogue device; the evaluation results obtained by the first evaluator's terminal device and the evaluation results obtained by the second evaluator's terminal device are generated using the same dialogue data stored in at least one memory; and the dialogue device, the first evaluator's terminal device, and the second evaluator's terminal device are each installed in different locations. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 is a block diagram showing an example of the overall configuration of a dialogue system according to the first embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of a functional configuration of the control device according to the first embodiment. [Figure 3] FIG. 3 is a processing flow showing an example of the interaction method according to the first embodiment. [Figure 4]FIG. 4 is a block diagram showing an example of the overall configuration of the dialogue system according to the second embodiment. [Figure 5] FIG. 5 is a block diagram illustrating an example of a functional configuration of a management device according to the second embodiment. [Figure 6] FIG. 6 is a processing flow showing an example of the interaction method according to the second embodiment. [Figure 7] FIG. 7 is a block diagram showing an example of a dialogue system according to an application example. [Figure 8] FIG. 8 is a block diagram showing an example of the hardware configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0007] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0008] [First embodiment] The first embodiment of the present disclosure is an example of an information processing system that interacts with a user. Hereinafter, the information processing system according to this embodiment will be referred to as a "dialogue system." The dialogue system may execute a dialogue with a user by controlling a dialogue device that is the user's dialogue partner.
[0009] The dialogue system has a function of outputting an evaluation result regarding a user based on dialogue data that records a dialogue between the user and the dialogue device. As an example, the dialogue system may dialogue with the user using the dialogue device according to a specific scenario, and output an evaluation result that evaluates the user according to predetermined evaluation criteria based on the dialogue data that records the dialogue between the user and the dialogue device. As an example, the dialogue scenario may be a scenario in which the user is a retail store clerk, the dialogue device is a customer, and the user performs a task of responding to the customer's requests. Using this type of scenario, for example, the user's customer service skills can be evaluated.
[0010] This embodiment provides a technology for controlling a device that interacts with a user. In this embodiment, information about a first utterance uttered by the user is acquired, information about a second utterance is generated based on the information about the first utterance, input information, and a machine learning model, and the utterance of the interaction device is controlled based on the information about the second utterance, where the input information includes information indicating a relationship between the user and the interaction device.
[0011] In one aspect, this embodiment generates information about the utterance of the dialogue device based on a machine learning model, thereby enabling control of the device that interacts with the user. In another aspect, this embodiment enables information for evaluating the user to be obtained through the dialogue between the user and the dialogue device.
[0012] <Overall configuration of the dialogue system> The overall configuration of the dialogue system according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the overall configuration of the dialogue system according to the first embodiment.
[0013] 1, the dialogue system 1000 includes a dialogue device 10, a control device 20, a generation device 30, and a terminal device 40. The dialogue device 10, the control device 20, the generation device 30, and the terminal device 40 may be connected to each other via a communication network so as to be able to communicate data with each other. The communication network may be, for example, a network such as a LAN (Local Area Network), a WAN (Wide Area Network), a VPN (Virtual Private Network), or the Internet.
[0014] The dialogue device 10 is an example of a device that dialogues with a conversation partner S, who is an example of a user of the dialogue system 1000. The dialogue device 10 may accept an utterance by the conversation partner S. The dialogue device 10 may present the utterance by the dialogue device 10 to the conversation partner S. The dialogue device 10 may dialogue with the conversation partner S by repeatedly accepting an utterance by the conversation partner S and presenting the utterance by the dialogue device 10. Hereinafter, an utterance by the conversation partner S will also be referred to as a "user utterance." An utterance by the dialogue device 10 will also be referred to as a "system utterance."
[0015] The dialogue device 10 may be an information processing terminal such as a personal computer, a smartphone, or a tablet terminal. As an example, the dialogue device 10 may display a CG (Computer Graphics) character on a display device and control the CG character to execute a dialogue with the interlocutor S. As an example, the CG character may be a virtual human also known as an avatar. The CG character is not limited to a human being, but may be a real creature, a fictional creature, a robot, an anthropomorphized object, or the like. In this case, the interlocutor S will be interacting with the virtual CG character displayed on the display device.
[0016] The dialogue device 10 may be a robot. The dialogue device 10 may execute a dialogue with the interlocutor S by operating the robot. As an example, the robot may be a humanoid that imitates a human. The robot is not limited to a human, and may be a real creature, a fictional creature, an anthropomorphized object, or the like. In this case, the interlocutor S will be having a dialogue with a robot that has a physical entity.
[0017] The dialogue device 10 may generate information about the user utterance and transmit it to the control device 20. The dialogue device 10 may receive information about the system utterance from the control device 20 and present the system utterance to the interlocutor S. Hereinafter, information about the user utterance will also be referred to as "user utterance information." Information about the system utterance will also be referred to as "system utterance information."
[0018] The user utterance information may include at least one of an audio signal obtained by collecting the user utterance, text data indicating the content of the user utterance, or a video signal or image signal obtained by capturing an image of the face or body of the interlocutor S. The dialogue device 10 may generate an audio signal by collecting the user utterance using a microphone. The dialogue device 10 may generate a video signal or image signal by capturing an image of the face or body of the interlocutor S using a camera. The dialogue device 10 may generate text data indicating the content of the utterance by performing voice recognition on the audio signal obtained by collecting the user utterance.
[0019] The system utterance information may include at least one of a voice signal synthesized from the system utterance, text data indicating the content of the system utterance, a video signal or image signal representing a CG character, or a control signal for controlling the expression of the dialogue device 10. The dialogue device 10 may present the system utterance to the interlocutor S by outputting a voice signal indicating the system utterance using a speaker. The dialogue device 10 may generate a voice signal indicating the system utterance by voice synthesizing text data indicating the content of the system utterance. The dialogue device 10 may present the system utterance information to the interlocutor S by displaying a video or image representing a CG character on a display device. If the dialogue device 10 is a robot, the system utterance information may be presented to the interlocutor S by moving the robot's face or body.
[0020] The control device 20 is an example of an information processing device such as a personal computer, a workstation, or a server that controls the dialogue between the interlocutor S and the dialogue device 10. The control device 20 may acquire user utterance information from the dialogue device 10. The control device 20 may generate system utterance information based on the user utterance information. The control device 20 may control the dialogue device 10 to utter a system utterance based on the system utterance information.
[0021] The control device 20 may store dialogue data D. The dialogue data D may be electronic data recording a dialogue between the interlocutor S and the dialogue device 10. The dialogue data D may include at least a portion of user utterance information and system utterance information. The dialogue data D may be stored for each dialogue between the interlocutor S and the dialogue device 10. The dialogue data D may include identification information capable of identifying the dialogue. An example of the identification information capable of identifying the dialogue may be time information indicating the time when the dialogue was executed. The time information may, for example, be the time when the dialogue started or the time when the dialogue ended.
[0022] The control device 20 may include an evaluation model M1. The evaluation model M1 is an example of a machine learning model trained to output one or more evaluation values based on the dialogue data D. The evaluation model M1 may take at least a portion of the dialogue data D as input and output one or more evaluation values. The evaluation model M1 may take one or more feature amounts extracted from the dialogue data D as input and output one or more evaluation values.
[0023] The evaluation model M1 may be, for example, a neural network, a decision tree, a random forest, a gradient boosting decision tree, a support vector machine (SVM), a large language model (LLM), a foundation model, etc. The gradient boosting decision tree may be, for example, a light gradient boosting machine (LightGBM) or an eXtreme gradient boosting (XGBoost).
[0024] The generating device 30 is an example of an information processing device such as a personal computer, a workstation, or a server that executes a predetermined task based on a machine learning model. The generating device 30 may execute a task of generating a system utterance in response to a generation request from the control device 20. The generating device 30 may transmit a generation result of the system utterance to the control device 20.
[0025] The generating device 30 may include a generative model M2. The generative model M2 is an example of a machine learning model for executing a predetermined task. The generative model M2 may be a machine learning model capable of generating various types of data such as text, audio, images, and videos. The generative model M2 may be, for example, a neural network, a large-scale language model, a generative model, or a base model. The generative model M2 may be multimodal.
[0026] The generative model M2 may be realized by a single machine learning model, or may be realized by multiple machine learning models working together, or may be composed of multiple machine learning models according to the tasks to be performed.
[0027] The terminal device 40 is an example of an information processing terminal such as a personal computer, smartphone, or tablet terminal operated by an evaluator E, who is an example of a user of the dialogue system 1000. The evaluator E is an entity that evaluates the interlocutor S. For example, the evaluator E may be a person in charge of evaluating the interlocutor S, who is a job seeker, to determine whether or not to hire him / her. For example, the evaluator E may be a person in charge of personnel evaluation of the interlocutor S, who is an employee. For example, the evaluator E may be a training instructor. For example, the evaluator E may be a test grader.
[0028] The terminal device 40 may acquire an evaluation result regarding the interlocutor S from the control device 20. The terminal device 40 may present the evaluation result regarding the interlocutor S to the evaluator E. For example, the terminal device 40 may display the evaluation result on a display device of the terminal device 40. For example, the terminal device 40 may output a voice synthesized from the evaluation result from a speaker of the terminal device 40.
[0029] Presenting information to a user of the dialogue system 1000 may include a processor performing at least a portion of the processing required to display the information on a display device. The display device may be provided in the same device as the processor or in a different device from the processor. The display device may be multiple display devices.
[0030] The overall configuration of the dialogue system 1000 shown in FIG. 1 is an example, and various system configuration examples are possible depending on the application and purpose. The dialogue system 1000 may be configured with one or more devices. Each device included in the dialogue system 1000 may be a system configured with multiple devices. Each function included in the dialogue system 1000 may be realized by any device that constitutes the system. Each component included in the dialogue system 1000 may be included in any device that constitutes the system.
[0031] The dialogue system 1000 may include a plurality of one or more of the dialogue device 10, the control device 20, the generation device 30, and the terminal device 40. The control device 20 or the generation device 30 may be realized by a plurality of computers, or may be realized as a cloud computing service. The control device 20 and the generation device 30 may be realized by an integrated standalone computer. The division of devices such as the dialogue device 10, the control device 20, the generation device 30, and the terminal device 40 shown in FIG. 1 is an example.
[0032] The dialogue device 10 may be any device that dialogues with the interlocutor S, and may be any of the following, as a non-limiting example: The dialogue device 10 may be one or more information processing devices. The dialogue device 10 may be an information processing system including multiple information processing devices. The dialogue device 10 may include a terminal device having a display device. The dialogue device 10 may be a CG character such as a virtual human controlled by one or more hardware. The dialogue device 10 may be realized by another device included in the dialogue system 1000. The dialogue device 10 may be realized by an information processing device or information processing system external to the dialogue system 1000. Note that "external" means not included in the dialogue system 1000.
[0033] As an example, the dialogue system 1000 may be configured with one or more server devices and one or more client devices. The one or more server devices may have one or more of the functions of the control device 20 and the generation device 30. The one or more client devices may have one or more of the functions of the dialogue device 10 and the terminal device 40. The server device may be realized as a system including multiple information processing devices. The server device may be realized as a cloud computing service.
[0034] As another example, the dialogue system 1000 may be configured with a single information processing device. The single information processing device may include the functions of the dialogue device 10, the control device 20, the generation device 30, and the terminal device 40.
[0035] <Controller functional configuration> The functional configuration of the control device 20 will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the functional configuration of the control device according to the first embodiment.
[0036] 2, the control device 20 includes a scenario storage unit 101, a dialogue storage unit 102, a model storage unit 103, an evaluation result storage unit 104, an utterance acquisition unit 110, an utterance generation unit 120, an utterance presentation unit 130, a dialogue evaluation unit 140, and a result output unit 150. The control device 20 functions as the scenario storage unit 101, the dialogue storage unit 102, the model storage unit 103, the evaluation result storage unit 104, the utterance acquisition unit 110, the utterance generation unit 120, the utterance presentation unit 130, the dialogue evaluation unit 140, and the result output unit 150 by executing a control program installed in advance.
[0037] (Scenario memory section) The scenario storage unit 101 stores scenario information. The scenario information is information indicating a dialogue scenario. The scenario information may include information indicating a dialogue situation. The information indicating a dialogue situation may include, for example, the relationship between the interlocutor S and the dialogue device 10, prerequisites for the dialogue, constraints on utterances, etc. The relationship between the interlocutor S and the dialogue device 10 may include the relationship between the interlocutor S and a CG character controlled by the dialogue device 10 or a robot as the dialogue device 10. The relationship between the interlocutor S and the dialogue device 10 may include the role of the interlocutor S and the role of the dialogue device 10.
[0038] The relationship between the interlocutor S and the dialogue device 10 may include at least one or a combination of at least two of the following relationships: a relationship in which the interlocutor S provides a service or a product to the dialogue device 10, a relationship in which the interlocutor S instructs the dialogue device 10, a relationship in which the dialogue device 10 pays a reward to the interlocutor S, and a relationship in which the interlocutor S provides information to the dialogue device 10. Specifically, the relationship between the interlocutor S and the dialogue device 10 may include at least one of the following relationships: a relationship between a retail store clerk and a customer, a relationship between a boss and a subordinate, a relationship between a teacher and a student, a relationship between a lecturer and an audience, a relationship between a doctor and a patient, or a relationship between a requested person and a client. The requested person may include, for example, an expert such as a lawyer or a consultant.
[0039] As an example, the scenario information may be a scenario regarding a dialogue between a store clerk and a customer in a retail store or a restaurant. In this case, the role of the interlocutor S may be any of "customer," "guest," "customer," "guest," "visitor," etc. The role of the dialogue device 10 may be any of "store clerk," "staff," "store employee," etc.
[0040] The scenario information may include one or more system prompts and one or more pieces of guidance information. The system prompts may be information used to cause the generation model M2 to generate system utterances in accordance with the scenario. The guidance information may be information used to explain the scenario to the interlocutor S when starting a dialogue. The guidance information may be associated with the corresponding system prompts. The guidance information may include information indicating the relationship between the interlocutor S and the dialogue device 10.
[0041] The system prompt may include information indicating the situation of the dialogue. The system prompt may be generated in advance based on a dialogue scenario. The system prompt may include information indicating the relationship between the interlocutor S and the dialogue device 10. The system prompt may include preconditions for the dialogue. The system prompt may include constraints on utterances.
[0042] When the system prompt includes information indicating the relationship between the interlocutor S and the dialogue device 10, the system prompt may include at least information regarding the role of the dialogue device 10, and may not explicitly include information regarding the role of the interlocutor S. The system prompt may include information that indirectly indicates the role of the interlocutor S. The role of the dialogue device 10 may include at least one of a customer, a subordinate, a student, an audience member, a patient, or a client.
[0043] The system prompt may be, for example, the following string: "You are a male customer at a clothing or miscellaneous goods store. You are purchasing clothes while being asked by the store clerk. You haven't decided what clothes you want to buy. Your preferences are simple designs, blue cotton, everyday use, and a budget of 50,000 yen. Please give out information in small amounts. You can mention that you are looking for clothes. Please output only the customer's utterances in 50 characters or less."
[0044] Here, "customer" and "store clerk" are examples of information indicating the relationship between the dialogue device 10 and the interlocutor S. Note that "customer" is the role of the dialogue device 10, and "store clerk" is the role of the interlocutor S. "Please output only the utterance of the customer role in 50 characters or less" is an example of a constraint on the utterance. The other parts are examples of preconditions in the dialogue.
[0045] As another example, the system prompt may be the following character string: The following character string is a non-limiting example of a system prompt that includes information that indirectly indicates the role of interlocutor S. "You are a male customer at a clothing or accessories store. You are looking to buy some clothes. You haven't decided what clothes you want to buy yet. Your preferences are simple designs, blue cotton, everyday use, and a budget of 50,000 yen. Please provide the information in small amounts. You can mention that you are looking for clothes. Please output only the customer's utterances in 50 characters or less."
[0046] Here, "customer" is an example of information relating to the role of the dialogue device 10. Also, "customer" is an example of information indirectly indicating the role of the interlocutor S. By indicating that the role of the dialogue device 10 is "customer," it is indirectly indicated that the role of the interlocutor S is "store clerk."
[0047] The scenario information may include a plurality of system prompts. For example, the plurality of system prompts may be system prompts corresponding to the degree of progress of the dialogue. For example, the degree of progress of the dialogue may be the number of turns in the dialogue or the elapsed time since the start of the dialogue.
[0048] The guidance information may have content related to the system prompt. The guidance information may include information indicating the relationship between the interlocutor S and the dialogue device 10. The information indicating the relationship between the interlocutor S and the dialogue device 10 may have the same content as the system prompt. The guidance information may have less information than the system prompt. The guidance information may omit at least some of the prerequisites regarding the dialogue device 10 from the system prompt. The guidance information may include information instructing the start of a dialogue. The information instructing the start of a dialogue may be located at the end of the guidance information.
[0049] The guidance information may be, for example, the following character string: "You are a salesperson at a clothing store. I am standing here as a customer, so please talk to me as if you were a salesperson. Here you go."
[0050] Here, "store clerk" and "customer" are examples of information indicating the relationship between the interlocutor S and the dialogue device 10. Note that "customer" is the role of the dialogue device 10, and "store clerk" is the role of the interlocutor S. "Here you go" is an example of information instructing the start of a dialogue.
[0051] The scenario storage unit 101 may store multiple pieces of scenario information. For example, the multiple pieces of scenario information may include scenario information relating to scenarios according to industry or occupation. The multiple pieces of scenario information may be stored in the scenario storage unit 101 in association with identification information indicating the scenario. As a method for associating the identification information with the scenario information, for example, the identification information and the scenario information may be stored as a set, or information that allows one piece of information to be acquired from the other piece of information may be stored, or one piece of information and the identification information of the other piece of information may be stored as a set.
[0052] The scenario information may be created by evaluator E. The scenario information may be created based on a template created in advance. As an example, the scenario information may be created by changing proper nouns (e.g., store name, product name, service name, etc.) in the template. The scenario storage unit 101 may store scenario information selected by evaluator E from multiple pieces of scenario information created in advance.
[0053] (Dialogue memory section) The dialogue storage unit 102 stores dialogue data D. The dialogue data D may be stored for each dialogue between the interlocutor S and the dialogue device 10. The dialogue data D may include user utterance information and system utterance information acquired by the control device 20 from the start to the end of the dialogue between the interlocutor S and the dialogue device 10. The user utterance information included in the dialogue data D may be acquired by the utterance acquisition unit 110. The system utterance information included in the dialogue data D may be generated by the utterance generation unit 120.
[0054] As an example, the dialogue data D may include at least one of time information, attribute information of the interlocutor S, voice information of the interlocutor S, image information of the interlocutor S, voice information of the dialogue device 10, image information of the dialogue device 10, control information of the dialogue device 10, recorded data of the entire dialogue, and log information of the dialogue system 1000.
[0055] The attribute information of the interlocutor S may include a profile of the interlocutor S. For example, the profile may include information contained in a resume or a curriculum vitae. The attribute information of the interlocutor S may be set in advance or may be input by the interlocutor S at the start of the dialogue.
[0056] The voice information of interlocutor S may include an audio signal obtained by capturing the speech of interlocutor S. The voice information of interlocutor S may include text data obtained by recognizing the audio signal. The image information of interlocutor S may include a still image or video of at least a part of the face or body of interlocutor S. The image information of interlocutor S may include information recognized from an image of interlocutor S. As an example, the image information of interlocutor S may include an expression label or posture information. An expression label is a label indicating a class into which a person's facial expressions are classified. Posture information is information indicating the positions of a person's feature points. For example, posture may be estimated by obtaining the positions of each joint using posture estimation technology.
[0057] The voice information of the dialogue device 10 may include a voice signal synthesized from system utterances. The voice information of the dialogue device 10 may include text data used to synthesize the voice signal. The image information of the dialogue device 10 may include a still image or video including at least a part of the face or body of a CG character. The image information of the dialogue device 10 may include information used to generate the image of the CG character. As an example, the image information of the dialogue device 10 may include facial expression labels or posture information. If the dialogue device 10 is a robot, the control information of the dialogue device 10 may include body control information including the facial expressions, gestures, etc. of the robot.
[0058] The recorded data of the entire dialogue may be video data obtained by synthesizing the audio and video signals of the interlocutor S and the audio and video signals of the dialogue device 10. The recorded data of the entire dialogue may be synthesized by the dialogue device 10 or the control device 20.
[0059] The log information of the dialogue system 1000 may be history information that records the results of processing performed by the dialogue system 1000 in a dialogue between the interlocutor S and the dialogue device 10. The log information of the dialogue system 1000 may include at least one of the processing results of the dialogue device 10, the processing results of the control device 20, and the processing results of the generation device 30. As an example, the log information of the dialogue system 1000 may include the recognition results of a user utterance, input information to the generative model M2, output information from the generative model M2, the content of the system utterance, etc.
[0060] Each piece of information included in the dialogue data D may be stored in association with time information. As a method for associating each piece of information with the time information, for example, each piece of information and the time information may be stored as a set, information that allows one piece of information to be acquired from the other piece of information may be stored, or one piece of information and identification information for the other piece of information may be stored as a set.
[0061] (Model storage section) The model storage unit 103 stores an evaluation model M1. The evaluation model M1 may be stored in advance in the model storage unit 103. The evaluation model M1 may be generated by the control device 20 by learning teacher data prepared in advance. The evaluation model M1 may be generated by an information processing device or information processing system other than the control device 20, and stored in the model storage unit 103.
[0062] A plurality of evaluation models M1 may be stored in the model storage unit 103. The plurality of evaluation models M1 may include evaluation models constructed for each evaluation value. The evaluation value may include an index for evaluating at least one of facial expression, body movements (for example, head, hands, arms, posture, etc.), voice, smoothness of dialogue, and dialogue content of the interlocutor S.
[0063] The evaluation model M1 may include an evaluation model that evaluates the facial expression of the interlocutor S. The evaluation model that evaluates the facial expression may include a machine learning model (e.g., a neural network) that receives image information of the face of the interlocutor S as input and outputs an evaluation value related to the facial expression of the interlocutor S. The evaluation model M1 that evaluates the facial expression may include a machine learning model (e.g., a decision tree, a random forest, a support vector machine, a neural network, etc.) that receives feature amounts extracted from image information such as facial feature points, facial landmarks, and FAUs (Facial Action Units) as input and outputs an evaluation value related to the facial expression of the interlocutor S.
[0064] The evaluation model M1 may include an evaluation model that evaluates the body movement of the interlocutor S. The evaluation model that evaluates the body movement may include a machine learning model that receives as input a video signal capturing an image of the interlocutor S and outputs an evaluation value related to the body movement of the interlocutor S (for example, a base model, a posture estimation model, a behavior recognition model, a gesture recognition model, or a model that outputs an evaluation value from features obtained by these models).
[0065] The evaluation model M1 may include an evaluation model for evaluating the voice of interlocutor S. The evaluation model for evaluating the voice may include a machine learning model (e.g., a neural network) that receives as input an audio signal obtained by recording the speech of interlocutor S and outputs an evaluation value related to non-verbal information of the speech of interlocutor S. The non-verbal information may include, for example, tone, intonation, speech rate, voice pitch, clarity, speech duration, the number of fillers, etc.
[0066] The evaluation model M1 may include an evaluation model that evaluates the smoothness of a dialogue. The evaluation model that evaluates the smoothness of a dialogue may include a machine learning model that receives the voice signal of the entire dialogue as input and outputs an evaluation value related to the smoothness of a dialogue (for example, a base model, a speech recognition model, a voice activity detection model, a speaking rate estimation model, a filler detection model, or a model that learns evaluation values using features obtained from these models). The evaluation value related to the smoothness of a dialogue may include, for example, the length of the dialogue exchange, the timing of utterances by the interlocutor S, the proportion of speech by the interlocutor S, and periods of silence. At least a portion of the evaluation value related to the smoothness of a dialogue may be analyzed using a rule base.
[0067] The evaluation model M1 may include an evaluation model that evaluates the content of the dialogue. The evaluation model M1 that evaluates the content of the dialogue may include a machine learning model that receives as input at least a portion of text data indicating the content of the utterance and outputs an evaluation value related to the content of the dialogue (for example, a large-scale language model that has been given a system prompt instructing the user to evaluate the content of the dialogue).
[0068] The evaluation model M1 may be constructed based on the evaluation criteria of the evaluator E. The evaluation model M1 may be constructed for each evaluation criterion of a plurality of evaluators E. The evaluation model M1 may be constructed based on recorded data of a dialogue to which a correct evaluation value is assigned based on the evaluation criteria of the evaluator E.
[0069] (Evaluation result storage unit) The evaluation result storage unit 104 stores evaluation results for the interlocutor S. The evaluation results for the interlocutor S may be generated by the dialogue evaluation unit 140. The evaluation result storage unit 104 may store multiple evaluation results for the interlocutor S. The multiple evaluation results may be generated based on dialogue data D for multiple different dialogues. The multiple dialogues may differ in at least one of the time, scenario, evaluation criteria, evaluation value, etc. The evaluation result storage unit 104 may store trends or statistical information of the evaluation values indicated in the multiple evaluation results. The statistical information may be, for example, maximum, minimum, average, etc.
[0070] An external evaluation value may be stored in the evaluation result storage unit 104. The external evaluation value is an evaluation value obtained by evaluating the interlocutor S based on information other than the dialogue data D. The external evaluation value may be generated by an information processing device or information processing system external to the dialogue system 1000 and stored in the evaluation result storage unit 104. As an example, the external evaluation value may include the results of an aptitude test of the interlocutor S, information indicating language ability, information indicating work experience, information indicating expertise in a particular field, etc.
[0071] (Speech acquisition unit) The utterance acquisition unit 110 acquires user utterance information. The utterance acquisition unit 110 may generate the user utterance information based on an audio signal obtained by collecting a user utterance. The utterance acquisition unit 110 may acquire the audio signal obtained by collecting a user utterance from the interactive device 10.
[0072] The utterance acquiring unit 110 may generate user utterance information based on a video signal obtained by capturing an image of the interlocutor S. The utterance acquiring unit 110 may acquire the video signal obtained by capturing an image of the interlocutor S from the dialogue device 10.
[0073] The utterance acquisition unit 110 may generate user utterance information based on text data indicating the content of the user utterance. The utterance acquisition unit 110 may generate text data indicating the content of the user utterance by performing voice recognition on a voice signal obtained by collecting the user utterance.
[0074] The utterance acquisition unit 110 may determine whether the video signal acquired from the dialogue device 10 is a video signal in which the interlocutor S is appropriately imaged. The utterance acquisition unit 110 may determine whether the face of the interlocutor S is imaged in the video signal. The utterance acquisition unit 110 may determine whether the size of the face of the interlocutor S imaged in the video signal is within a predetermined range. The utterance acquisition unit 110 may detect the face of the interlocutor S imaged in the video signal using a face detector. The face detector may be constructed using a machine learning model, for example.
[0075] When calculating an evaluation value related to the body movement of interlocutor S, utterance acquisition unit 110 may determine whether or not the body of interlocutor S is captured in the video signal. Utterance acquisition unit 110 may determine whether or not the size of the body of interlocutor S captured in the video signal is within a predetermined range. Utterance acquisition unit 110 may use a body detector to detect the body of interlocutor S captured in the video signal. For example, the body detector may be constructed using a machine learning model.
[0076] When the interlocutor S is not properly imaged in the video signal, the utterance acquisition unit 110 may notify the interlocutor S of this fact. As an example, the utterance acquisition unit 110 may transmit a notification to the interaction device 10 indicating that the interlocutor S is not properly imaged. The interaction device 10 may display a message indicating that the interlocutor S is not properly imaged on a display device. The interaction device 10 may also utter an utterance indicating that the interlocutor S is not properly imaged. When the interlocutor S is not properly imaged in the video signal, the utterance acquisition unit 110 may stop the interaction with the interlocutor S.
[0077] The utterance acquisition unit 110 may store the acquired user utterance information in the dialogue storage unit 102. The utterance acquisition unit 110 may store the user utterance information in real time, or may store the user utterance information collectively in a predetermined processing unit. The predetermined processing unit may be, for example, the amount of data of the user utterance information, the number of turns in the dialogue, the elapsed time of the dialogue, etc.
[0078] The utterance acquisition unit 110 may include the user utterance information in the dialogue data D stored in the dialogue storage unit 102. The utterance acquisition unit 110 may store the user utterance information in the dialogue storage unit 102 in association with time information included in the dialogue data D.
[0079] (Speech generation unit) The utterance generation unit 120 generates system utterance information. The utterance generation unit 120 may generate the system utterance information based on user utterance information acquired by the utterance acquisition unit 110. The utterance generation unit 120 may generate the system utterance information based on the generative model M2. The utterance generation unit 120 may generate the system utterance information based on user utterance information, input information, and the generative model M2.
[0080] Note that generating other information based on certain information and a machine learning model may include performing one or more of the following processes: · Inputting one piece of information into a machine learning model to generate other information. · Information generated based on other information is fed into a machine learning model to generate other information. - Using information extracted from other information, a generation process is performed using a machine learning model to generate other information.
[0081] The input information may be generated based on scenario information. The scenario information may be set in advance. The scenario information may be set based on predetermined criteria. The scenario information may be set based on instructions from evaluator E.
[0082] The utterance generation unit 120 may generate input information for the generative model M2 based on the user utterance information and transmit the generated information to the generation device 30. Hereinafter, the input information for the generative model M2 will also be referred to as "generative model input information."
[0083] The generation model input information may include at least system prompts and dialogue history information, which are examples of input information. The utterance generation unit 120 may read out system prompts from scenario information stored in the scenario storage unit 101. The utterance generation unit 120 may select system prompts to read out depending on the progress of the dialogue. The utterance generation unit 120 may select system prompts to read out depending on the industry or occupation. The utterance generation unit 120 may read out multiple system prompts and combine them. As an example, the utterance generation unit 120 may read out a system prompt according to the industry and a system prompt according to the occupation, and generate a system prompt that combines them.
[0084] The utterance generation unit 120 may select a system prompt based on a machine learning model. The machine learning model may be the generative model M2 or another machine learning model. The other machine learning model may be, for example, a large-scale language model or a base model. The utterance generation unit 120 may input information about the dialogue between the interlocutor S and the dialogue device 10 and multiple system prompts into the machine learning model to obtain information for selecting one or more system prompts. The information about the dialogue may be dialogue history information. The utterance generation unit 120 may also generate a system prompt based on the machine learning model.
[0085] The dialogue history information may be information that chronologically links the utterance content of the interlocutor S and the utterance content of the dialogue device 10. The dialogue history information may include the utterance content from the start of the dialogue to the immediately previous utterance. The dialogue history information may include a predetermined number of utterance contents going back in time from the immediately previous utterance. Each utterance content may be associated with information indicating the speaker (the interlocutor S or the dialogue device 10). The information indicating the speaker may be added before or after the utterance content. The information indicating the speaker may be added so that it can be identified by a predetermined symbol. The boundary between each utterance content may be indicated by a predetermined symbol. The dialogue history information may include information indicating a speaker for whom an utterance is to be generated. The information indicating a speaker for whom an utterance is to be generated does not have to be associated with the utterance content.
[0086] The dialogue history information may be, for example, the following character string: · "Staff: Welcome \nCustomer:"
[0087] However, "Welcome" is an example of the content of the immediately preceding utterance. "Store clerk" and "Customer" are examples of information indicating the speaker. ":" is an example of a symbol separating the speaker from the content of their utterance. "¥n" is an example of a symbol indicating the boundary of the content of an utterance (in this case, the end of the content of the immediately preceding utterance). In the above example, since no utterance content is associated with "Customer", the utterance of the customer (i.e., the dialogue device 10) is the target for generation. In the above example, the dialogue history information includes only the content of the immediately preceding utterance, but if one or more turns have taken place before that, the content of the utterance in the previous turn before "Store clerk" may also be included.
[0088] The dialogue history information is not limited to text information, and may include voice information or image information, or may include text information and at least one of voice information and image information. Specifically, the dialogue history information may include information in which voice information of the interlocutor S, image information of the interlocutor S, voice information of the dialogue device 10, and image information of the dialogue device 10 are linked in chronological order.
[0089] The utterance generation unit 120 may generate the generative model input information by embedding the system prompt and the dialogue history information in a predetermined template. The utterance generation unit 120 may generate the generative model input information by processing the template, the system prompt, and the dialogue history information based on a predetermined rule. The utterance generation unit 120 may generate the generative model input information by inputting the template, the system prompt, and the dialogue history information into a machine learning model. The machine learning model may be the generative model M2 or another machine learning model. The other machine learning model may be, for example, a large-scale language model or a base model.
[0090] A template for generating generative model input information may include one or more placeholders for embedding system prompts or dialogue history information. The template may include instruction information that instructs the generation of an utterance. The instruction information may include information that instructs the output format of the output information. The output format of the output information may include, for example, at least one of text data, audio signals, video signals, and control signals. The template may include constraint information regarding the output information. The constraint information may include the length of the utterance, content that may be spoken, content that may be spoken if asked, content that must not be spoken even if asked, speaking style, etc. The instruction information and constraint information may be predetermined fixed sentences.
[0091] The template may be optimized by any prompt tuning method to obtain good output information. For example, the template may be optimized by searching for an optimal template using a genetic algorithm or the like based on a benchmark that evaluates prompts generated using the template.
[0092] The generative model input information may include text data, image data, or audio data. The text data may be, for example, a natural language sentence called a prompt. The image data may be, for example, a still image or a video. The text data may be text data obtained by speech recognition of audio data or video. The text data may be text obtained by character recognition of image data. The image data may include, for example, an image of a user. The audio data may include, for example, audio spoken by a user. The audio data may be audio data obtained by speech synthesis of text data.
[0093] The utterance generation unit 120 may receive, from the generation device 30, output information generated by the generation device 30 inputting the generation model input information to the generation model M2. The generation model input information may be input to the generation model M2 as a system prompt and dialogue history information separately. The generation model input information may be input to the generation model M2 as a combination of the system prompt and dialogue history information.
[0094] The utterance generation unit 120 may generate system utterance information based on the output information of the generative model M2. If the output information of the generative model M2 includes an audio signal indicating a system utterance or a video signal representing a CG character, the utterance generation unit 120 may use at least one of the audio signal and the video signal as the system utterance information. If the output information of the generative model M2 includes text data indicating the content of the utterance, the utterance generation unit 120 may synthesize an audio signal based on the text data. If the output information of the generative model M2 includes information for controlling the expression of the dialogue device 10 (for example, a facial expression label, information on the shape of the mouth when speaking, or posture information), the utterance generation unit 120 may synthesize a video signal representing a CG character based on the information.
[0095] The utterance generation unit 120 may generate system utterance information based on a portion of the output information of the generative model M2. When the output information of the generative model M2 is equal to or longer than a predetermined data length or a randomly determined data length, the utterance generation unit 120 may shorten the output information of the generative model M2 so that it is equal to or shorter than the predetermined data length. When the generative model M2 outputs output information in a stream, the utterance generation unit 120 may cause the generative model M2 to stop subsequent output when the data length is reached.
[0096] As an example, the utterance generation unit 120 may acquire N sentences in order from the beginning of the output information of the generative model M2. N is a positive integer and may be determined randomly. As an example, the utterance generation unit 120 may acquire N' sentences in order from the beginning of the output information of the generative model M2 so that the maximum number of sentences is M or less. M and N' are positive integers. The boundaries of sentences may be identified by predetermined symbols. As an example, periods (such as "." or ".") or symbols (such as "?" or "!") may be used as delimiters.
[0097] The utterance generation unit 120 may store the generated system utterance information in the dialogue storage unit 102. The utterance generation unit 120 may store the system utterance information in real time, or may store the system utterance information collectively in a predetermined processing unit. The predetermined processing unit may be, for example, the amount of data in the system utterance information, the number of turns in the dialogue, the elapsed time of the dialogue, etc.
[0098] The utterance generation unit 120 may include the system utterance information in the dialogue data D stored in the dialogue storage unit 102. The utterance acquisition unit 110 may store the system utterance information in the dialogue storage unit 102 in association with time information included in the dialogue data D.
[0099] (Speech presentation unit) The utterance presenting unit 130 controls the utterances of the dialogue device 10. The utterance presenting unit 130 may transmit a voice signal into which the system utterance is synthesized to the dialogue device 10. The dialogue device 10 may output a voice indicating the system utterance from a speaker based on the voice signal received from the control device 20.
[0100] The utterance presenter 130 may transmit text data indicating the content of the utterance to the dialogue device 10. The dialogue device 10 may generate a voice signal by synthesizing the text data received from the control device 20, and output the voice signal from a speaker.
[0101] The utterance presenting unit 130 may transmit a video signal representing the CG character to the interactive device 10. Based on the video signal received from the control device 20, the interactive device 10 may display a video representing the CG character on a display device.
[0102] The utterance presenting unit 130 may transmit information used for synthesizing the video signal (for example, a facial expression label, information on the shape of the mouth when speaking, or posture information) to the dialogue device 10. The dialogue device 10 may synthesize the video signal based on the information received from the control device 20, and display a video representing the CG character on the display device.
[0103] The utterance presenting unit 130 may transmit a control signal for controlling the expression of the CG character to the dialogue device 10. The dialogue device 10 may control the expression of the CG character based on the control signal received from the control device 20.
[0104] The utterance presentation unit 130 may control the dialogue device 10 to utter a system utterance that explains the dialogue situation (hereinafter also referred to as an "explanatory utterance"). The utterance presentation unit 130 may cause the dialogue device 10 to utter an explanatory utterance based on guidance information stored in the scenario storage unit 101. The utterance presentation unit 130 may transmit the guidance information stored in the scenario storage unit 101 to the dialogue device 10. The utterance presentation unit 130 may transmit a voice signal synthesized with the guidance information to the dialogue device 10.
[0105] The utterance presentation unit 130 may control the dialogue device 10 to utter a system utterance in response to a user utterance. The utterance presentation unit 130 may cause the dialogue device 10 to utter a system utterance in response to a user utterance, based on the system utterance information generated by the utterance generation unit 120. The utterance presentation unit 130 may transmit at least one of text data, an audio signal, a video signal, or a control signal included in the system utterance information to the dialogue device 10.
[0106] (Dialogue Evaluation Section) The dialogue evaluation unit 140 generates an evaluation result for the interlocutor S. The dialogue evaluation unit 140 may generate the evaluation result for the interlocutor S based on dialogue data D read out from the dialogue storage unit 102. The dialogue evaluation unit 140 may generate the evaluation result for the interlocutor S based on the evaluation model M1 read out from the model storage unit 103. The dialogue evaluation unit 140 may generate input information for the evaluation model M1 based on the dialogue data D, and input the input information for the evaluation model M1 to the evaluation model M1. The input information for the evaluation model M1 may include features extracted from the dialogue data D. Hereinafter, the input information for the evaluation model M1 will also be referred to as "evaluation model input information."
[0107] The evaluation result for interlocutor S may include one or more evaluation values. The dialogue evaluation unit 140 may select an evaluation value to be included in the evaluation result according to the evaluation criteria of evaluator E. The dialogue evaluation unit 140 may generate the evaluation result for interlocutor S based on an evaluation model M1 according to the evaluation criteria of evaluator E. The dialogue evaluation unit 140 may extract, from the dialogue data D, feature quantities to be included in the evaluation model input information according to the evaluation value to be included in the evaluation result.
[0108] When multiple evaluation models M1 are stored in the model storage unit 103, the dialogue evaluation unit 140 may generate an evaluation result for the interlocutor S based on each of the multiple evaluation models M1. The evaluation result for the interlocutor S may include multiple evaluation results based on each of the multiple evaluation models M1. The evaluation result for the interlocutor S may include one evaluation result including multiple evaluation values output by each of the multiple evaluation models M1.
[0109] The evaluation result for interlocutor S may include a reliability indicating the accuracy of the evaluation value. The reliability of the evaluation value may be output by evaluation model M1 together with the evaluation value. As an example, if evaluation model M1 is a classification model, evaluation model M1 may calculate a probability for each class to be classified, set the class with the highest probability as the evaluation value, and output the probability for that class as the reliability of the evaluation value.
[0110] As a non-limiting example, the dialogue evaluation unit 140 may generate an evaluation result for the interlocutor S by inputting evaluation model input information to the generative model M2. The evaluation model M1 may be generated based on the generative model M2. As an example, the evaluation model M1 may be generated by transfer learning the generative model M2. As an example, the evaluation model M1 may be generated by fine-tuning the generative model M2.
[0111] The dialogue evaluation unit 140 may generate an evaluation result for the interlocutor S at any timing. The dialogue evaluation unit 140 may generate the evaluation result by batch processing or by ad-hoc processing. As an example, the dialogue evaluation unit 140 may generate an evaluation result for each piece of accumulated dialogue data D at predetermined time intervals. The dialogue evaluation unit 140 may generate an evaluation result for the dialogue data D related to the dialogue when the dialogue ends.
[0112] The dialogue evaluation unit 140 may store the evaluation result for the interlocutor S in the evaluation result storage unit 104. The dialogue evaluation unit 140 may store the evaluation result in the evaluation result storage unit 104 in association with identification information that can identify the dialogue. As an example, the dialogue evaluation unit 140 may store the evaluation result in the evaluation result storage unit 104 in association with time information.
[0113] (Result output section) The result output unit 150 outputs the evaluation result regarding the interlocutor S. The result output unit 150 may read out the evaluation result regarding the interlocutor S stored in the evaluation result storage unit 104.
[0114] The result output unit 150 may include the external evaluation value for the interlocutor S in the evaluation result for the interlocutor S and output the result. The result output unit 150 may read out the external evaluation value for the interlocutor S stored in the evaluation result storage unit 104 together with the evaluation result for the interlocutor S, and include the external evaluation value in the evaluation result for the interlocutor S.
[0115] The result output unit 150 may include the dialogue data D used to evaluate the interlocutor S in the evaluation result for the interlocutor S. As an example, the result output unit 150 may include the recorded data included in the dialogue data D or the feature extracted from the recorded data in the evaluation result for the interlocutor S. The feature extracted from the recorded data may include the feature input to the evaluation model M1.
[0116] The result output unit 150 may transmit the evaluation result regarding the interlocutor S to the terminal device 40. The result output unit 150 may transmit the evaluation result regarding the interlocutor S to the terminal device 40 in response to an output request from the terminal device 40. When there are multiple evaluation results regarding the interlocutor S, the result output unit 150 may transmit to the terminal device 40 the transition of the evaluation value or statistical information indicated in the evaluation result. The terminal device 40 may present the evaluation result regarding the interlocutor S to the evaluator E.
[0117] The result output unit 150 may transmit the evaluation result regarding the interlocutor S to the dialogue device 10. The dialogue device 10 may present the evaluation result regarding the interlocutor S to the interlocutor S. The result output unit 150 may transmit the evaluation result regarding the interlocutor S to an external information processing device or information processing system. The external information processing device or information processing system may present the evaluation result regarding the interlocutor S to a user other than the interlocutor S. The user other than the interlocutor S may include an evaluator E.
[0118] For example, the result output unit 150 may transmit screen data displaying the evaluation results to the terminal device 40. For example, the screen data may be electronic data written in HTML (Hyper Text Markup Language) or the like. The result output unit 150 may transmit screen data displaying the evaluation results to the interactive device 10.
[0119] For example, the result output unit 150 may send the evaluation result to the account of the evaluator E or the interlocutor S via email, a messaging service, or the like. The result output unit 150 may send link information to a screen displaying the evaluation result to the account of the evaluator E or the interlocutor S via email, a messaging service, or the like.
[0120] The functional configuration of the control device 20 shown in Fig. 2 is an example, and it goes without saying that there are various examples of functional configurations depending on the application and purpose. The division of the storage units, such as the scenario storage unit 101, the dialogue storage unit 102, the model storage unit 103, and the evaluation result storage unit 104 shown in Fig. 2, is an example. The division of the processing units, such as the utterance acquisition unit 110, the utterance generation unit 120, the utterance presentation unit 130, the dialogue evaluation unit 140, and the result output unit 150 shown in Fig. 2, is an example.
[0121] For example, at least two of the scenario storage unit 101, the dialogue storage unit 102, the model storage unit 103, and the evaluation result storage unit 104 may be integrated into one storage unit. Also, for example, at least one of the dialogue storage unit 102, the model storage unit 103, and the evaluation result storage unit 104 may be divided into multiple storage units.
[0122] For example, at least two of the utterance acquisition unit 110, the utterance generation unit 120, the utterance presentation unit 130, the dialogue evaluation unit 140, and the result output unit 150 may be integrated into one processing unit. Also, for example, at least one of the utterance acquisition unit 110, the utterance generation unit 120, the utterance presentation unit 130, the dialogue evaluation unit 140, and the result output unit 150 may be divided into multiple processing units.
[0123] <Dialogue method flow> The dialogue method executed by the dialogue system 1000 will be described with reference to Fig. 3. Fig. 3 is a processing flow showing an example of the dialogue method according to the first embodiment. In the following, an example will be described in which a CG character is used to dialogue with a conversation partner S who operates a personal computer, which is an example of the dialogue device 10.
[0124] In step S101, the interlocutor S performs an operation to start a dialogue with the dialogue device 10. As an example, the interlocutor S may perform an operation to select a control device 20 with which to conduct a dialogue from a list of control devices 20 displayed on a display device of the dialogue device 10. When the interlocutor S performs an operation to select the control device 20, the dialogue device 10 may execute a process to connect to the control device 20.
[0125] The interlocutor S may test the microphone or camera of the dialogue device 10. The interlocutor S may turn on the microphone switch of the dialogue device 10 and check the microphone volume level. At this time, the dialogue device 10 and the control device 20 may cooperate to instruct the interlocutor S to read a specific sentence, the interlocutor S reads the specified sentence, perform voice recognition on the read sentence, and determine that the microphone is working properly if the recognized text is the same as or similar to the specified sentence. The interlocutor S may turn on the camera switch of the dialogue device 10 and check the camera image displayed on the display device. At this time, the dialogue device 10 may use a face or body detector to determine whether the interlocutor S is properly captured in the camera image. The dialogue device 10 may control not to start the dialogue until the interlocutor S is properly captured in the camera image. The interlocutor S being properly captured in the camera image may include, for example, whether the face is larger than a predetermined size, whether the face is facing forward, and whether the body is captured within a predetermined range.
[0126] In step S102, the utterance presentation unit 130 of the control device 20 reads out guidance information stored in the scenario storage unit 101. The utterance presentation unit 130 performs voice synthesis on the guidance information. As a result, a voice signal indicating an explanatory utterance is generated. The utterance presentation unit 130 transmits the voice signal indicating the explanatory utterance to the dialogue device 10. Note that the utterance presentation unit 130 may transmit the guidance information together with the voice signal or independently to the dialogue device 10.
[0127] In step S103, the dialogue device 10 receives a voice signal indicating an explanatory utterance from the control device 20. The dialogue device 10 outputs an explanatory utterance from a speaker based on the voice signal indicating the explanatory utterance. When the dialogue device 10 receives guidance information from the control device 20, the dialogue device 10 may display the guidance information on a display device. Furthermore, when the dialogue device 10 receives guidance information from the control device 20, the dialogue device 10 may output a voice signal synthesized with the guidance information from the speaker.
[0128] In step S104, the control device 20 starts recording the dialogue between the interlocutor S and the dialogue device 10. Specifically, the control device 20 requests the utterance acquisition unit 110 to store user utterance information in the dialogue storage unit 102. The control device 20 also requests the utterance generation unit 120 to store system utterance information in the dialogue storage unit 102. When the utterance acquisition unit 110 acquires the user utterance information, it enables a function to store the user utterance information in the dialogue storage unit 102. When the utterance generation unit 120 generates system utterance information, it enables a function to store the system utterance information in the dialogue storage unit 102.
[0129] In step S105, the interactive device 10 displays the CG character on the display device. The interactive device 10 transmits an audio signal picked up by the microphone and a video signal picked up by the camera to the control device 20. Thereafter, the interactive device 10 continues transmitting the audio signal and the video signal until the interaction ends.
[0130] The interlocutor S makes an utterance in response to the explanatory utterance output in step S103. The dialogue device 10 accepts the user utterance from the interlocutor S. The audio signal transmitted by the dialogue device 10 to the control device 20 includes the voice of the interlocutor S uttering the user utterance. Furthermore, the video signal transmitted by the dialogue device 10 to the control device 20 includes an image capturing the facial expression or movement of the interlocutor S when he or she utters the user utterance.
[0131] In step S106, the utterance acquisition unit 110 of the control device 20 receives an audio signal of a user utterance from the dialogue device 10. The utterance acquisition unit 110 also receives a video signal of an image of the interlocutor S from the dialogue device 10.
[0132] The speech acquisition unit 110 performs speech recognition on the audio signal obtained by collecting the user's utterance. As a result, text data indicating the content of the user's utterance is generated. The speech acquisition unit 110 sends the text data indicating the content of the user's utterance to the utterance generation unit 120.
[0133] The utterance acquisition unit 110 generates user utterance information based on an audio signal obtained by capturing a user utterance, a video signal obtained by capturing an image of the interlocutor S, and text data indicating the content of the user utterance. The utterance acquisition unit 110 stores the user utterance information in the dialogue storage unit 102.
[0134] In step S107, the utterance generation unit 120 of the control device 20 receives text data indicating the content of the user utterance from the utterance acquisition unit 110. The utterance generation unit 120 reads out a system prompt stored in the scenario storage unit 101. The utterance generation unit 120 reads out dialogue data D from the dialogue storage unit 102. The utterance generation unit 120 generates dialogue history information based on the text data indicating the content of the user utterance and the dialogue data D read out from the dialogue storage unit 102.
[0135] The utterance generation unit 120 generates generative model input information based on the system prompt and the dialogue history information. The utterance generation unit 120 may generate the generative model input information by embedding the system prompt and the dialogue history information in a predetermined template. The utterance generation unit 120 transmits the generative model input information to the generation device 30.
[0136] The generation device 30 receives generative model input information from the control device 20. The generation device 30 inputs the received generative model input information to the generative model M2. The generative model M2 executes a predetermined task based on the generative model input information and outputs output information generated as a result. The output information from the generative model M2 may include at least text data indicating the content of the system utterance. The generation device 30 transmits the output information output from the generative model M2 to the control device 20.
[0137] In step S108, the utterance generation unit 120 of the control device 20 receives output information from the generation device 30. The utterance generation unit 120 generates system utterance information based on at least a part of the received output information. The utterance generation unit 120 may perform voice synthesis on text data indicating the content of the system utterance. This generates a voice signal indicating the system utterance. The voice signal indicating the system utterance may be included in the output information from the generative model M2.
[0138] The utterance generation unit 120 may generate a video signal representing a CG character. The utterance generation unit 120 may synthesize a video signal representing a CG character based on information included in the output information from the generative model M2. The video signal representing the CG character may be included in the output information from the generative model M2. In other words, the video signal representing the CG character may be directly generated by the generative model M2.
[0139] The utterance generation unit 120 generates system utterance information based on an audio signal indicating the system utterance, a video signal representing the CG character, and text data indicating the content of the system utterance. The utterance acquisition unit 110 stores the system utterance information in the dialogue storage unit 102. The utterance generation unit 120 transmits the system utterance information to the dialogue device 10.
[0140] In step S109, the dialogue device 10 receives the system utterance information from the control device 20. The dialogue device 10 utters a system utterance based on the system utterance information. The dialogue device 10 may output a sound indicating the system utterance from a speaker. The dialogue device 10 may display an image representing a CG character on a display device. If the dialogue device 10 is a robot, it may control the robot based on the system utterance information. Note that the dialogue device 10 may switch off the microphone while uttering the system utterance. In other words, the dialogue device 10 may pick up the utterance of the interlocutor S with the microphone until the system utterance is uttered.
[0141] In step S110, the control device 20 determines whether a termination condition for terminating the dialogue has been met. For example, the termination condition may be that the number of dialogue turns or the elapsed time since the start of the dialogue exceeds a predetermined threshold. In FIG. 3, the termination determination process is performed only immediately after step S108. However, if the determination is based on the elapsed time, the termination determination process may be performed at any timing between immediately before step S106 and immediately after step S108. In this case, the termination determination may not be performed until a predetermined time has elapsed after the start and end of each utterance by the interlocutor S. This prevents the dialogue from ending abruptly and allows the dialogue to end within a time period close to the predetermined time.
[0142] At any timing between immediately before step S106 and immediately after step S108, a face or body detector may be used to check whether a face or body is captured on camera. If detection is not possible for a certain period of time, the dialogue is terminated and the process proceeds to step S111. Before the dialogue ends, the dialogue device 10 may display a text message or output a voice message to the interlocutor S that a face or body cannot be detected.
[0143] The control device 20 may determine whether to terminate the dialogue based on the content of the dialogue. As an example, the control device 20 may input dialogue history information into a machine learning model and determine whether to terminate the dialogue based on output information from the machine learning model. The machine learning model that determines whether to terminate the dialogue may be the generative model M2, or another large-scale language model or a base model.
[0144] If it is determined that the termination condition is satisfied (YES), the control device 20 proceeds to step S111. If it is determined that the termination condition is satisfied, the control device 20 may transmit a voice signal indicating a predetermined utterance to the dialogue device 10. The predetermined utterance may be an utterance indicating that the dialogue has ended. The dialogue device 10 may receive a voice signal indicating the predetermined utterance from the control device 20, and output the predetermined utterance from a speaker.
[0145] On the other hand, if it is determined that the termination condition is not satisfied (NO), the control device 20 returns the process to step S106. When returning the process to step S106, the utterance acquisition unit 110 of the control device 20 acquires user utterance information related to the next user utterance. Thereafter, the control device 20 executes the processes from step S106 to step S110 again based on the new user utterance information. In this way, the dialogue system 1000 repeatedly executes the acceptance of a user utterance and the presentation of a system utterance until it is determined in step S110 that the termination condition is satisfied. Note that the utterance generation unit 120 may change the system prompt used to generate the system utterance for each repetition.
[0146] In step S111, the dialogue evaluation unit 140 of the control device 20 reads out dialogue data D relating to a dialogue with the interlocutor S from the dialogue storage unit 102. The dialogue evaluation unit 140 also reads out an evaluation model M1 corresponding to the evaluator E from the model storage unit 103.
[0147] The dialogue evaluation unit 140 generates evaluation model input information based on the read dialogue data D. The dialogue evaluation unit 140 inputs the evaluation model input information to the evaluation model M1. The evaluation model M1 calculates an evaluation value of 1 or more based on the evaluation model input information. As a result, an evaluation result for the interlocutor S is generated. The dialogue evaluation unit 140 stores the evaluation result for the interlocutor S in the evaluation result storage unit 104.
[0148] In step S112, the evaluator E performs an operation to display the evaluation results on the terminal device 40. As an example, the evaluator E may perform an operation to open the evaluation results of the interlocutor S on the evaluation screen displayed on the display device of the terminal device 40. The terminal device 40 transmits an output request for the evaluation results regarding the interlocutor S to the control device 20.
[0149] The result output unit 150 of the control device 20 receives a request to output the evaluation result from the terminal device 40. The result output unit 150 reads out the evaluation result regarding the interlocutor S from the evaluation result storage unit 104. The result output unit 150 generates screen data for displaying the evaluation result regarding the interlocutor S. The result output unit 150 transmits the screen data to the terminal device 40.
[0150] In step S113, the terminal device 40 receives screen data from the control device 20. The terminal device 40 displays the evaluation results for the interlocutor S on the display device based on the screen data. The evaluator E may evaluate the interlocutor S by referring to the evaluation results for the interlocutor S displayed on the display device of the terminal device 40. As an example, the evaluator E may refer to the evaluation results for the interlocutor S to decide whether to hire the interlocutor S. As an example, the evaluator E, who is a training instructor, may refer to the evaluation results for the interlocutor S to judge the quality of the training. As an example, the evaluator E, who is a test grader, may refer to the evaluation results for the interlocutor S to determine the score for the interlocutor S and determine whether the test is passed or failed.
[0151] The terminal device 40 may display on a display device the recorded data included in the evaluation result regarding the interlocutor S. The evaluator E may compare the recorded data included in the evaluation result with the evaluation value and determine whether the evaluation result regarding the interlocutor S is appropriate.
[0152] The terminal device 40 may display a screen for correcting the evaluation result regarding the interlocutor S on the display device. The evaluator E may correct the evaluation result regarding the interlocutor S on the terminal device 40. The terminal device 40 may transmit the evaluation result corrected by the evaluator E to the control device 20. The control device 20 may update the evaluation result regarding the interlocutor S stored in the evaluation result storage unit 104 based on the evaluation result received from the terminal device 40. The control device 20 may additionally train the evaluation model M1 based on the updated evaluation result.
[0153] [Modification 1 of the First Embodiment] The dialogue system 1000 may store an evaluation model M1 and a system prompt in association with each other. For example, the evaluation model M1 may be stored in the model storage unit 103 in association with identification information that identifies the system prompt. The dialogue evaluation unit 140 may read the evaluation model M1 associated with the system prompt from the model storage unit 103 and calculate one or more evaluation values based on the evaluation model M1.
[0154] The dialogue system 1000 may store the evaluation model M1 and scenario information in association with each other. As an example, scenario information including identification information for identifying the evaluation model M1 may be stored in the scenario storage unit 101. The utterance generation unit 120 may read a system prompt corresponding to the scenario from the scenario storage unit 101 and include it in the input information to the generation model M2. The dialogue evaluation unit 140 may read an evaluation model M1 corresponding to the scenario from the model storage unit 103 and calculate one or more evaluation values based on the evaluation model M1. The dialogue system 1000 allows the selection of a system prompt and an evaluation model M1 simply by selecting a scenario.
[0155] [Modification 2 of the First Embodiment] The control device 20 may store information used to control the dialogue device 10. For example, the information used to control the dialogue device 10 may include text data used to synthesize a voice signal indicating system utterance, the start time and end time of the synthesized voice signal, labels used to control facial expressions or body movements (e.g., laughing, waving, nodding, etc.) of the dialogue device 10, and the start time and end time of a video indicating the movement based on the label. The time may be Japan Standard Time (JST) or Coordinated Universal Time (UTC), or the elapsed time from the time when recording of the dialogue started.
[0156] The information used to control the dialogue device 10 may be stored in a storage device of the dialogue device 10. The information used to control the dialogue device 10 may be stored in a storage device of the generation device 30. The information used to control the dialogue device 10 may be stored in an information processing device or storage device other than the dialogue device 10, the control device 20, or the generation device 30.
[0157] The information used to control the dialogue device 10 may be used together with the audio signal or video signal to calculate the evaluation value. As an example, by using the information used to synthesize the audio signal, the speech section of the dialogue device 10 can be identified from the recorded data of the entire dialogue based on the start time and end time of the playback of the audio signal. Furthermore, by using the information used to synthesize the audio signal, the speech content of the dialogue device 10 can be obtained as text without using voice recognition. As an example, by using information on the facial expression or body movement of the dialogue device 10, the content of the interlocutor S's response to the facial expression or body movement, the facial expression, the tone of voice in which the response was made, etc. can be input as additional information to the evaluation model M1.
[0158] [Modification 3 of the First Embodiment] The utterance generator 120 may generate system utterances using a technique called Retrieval Augmented Generation (RAG), which includes reference information obtained by searching external data sources in prompts to improve the results of machine learning models.
[0159] The control device 20 may store reference information including background information related to the dialogue device 10 in advance. The utterance generation unit 120 may generate a query based on the user utterance information and search for reference information based on the query. The utterance generation unit 120 may generate a system prompt based on the searched reference information. The utterance generation unit 120 may include the generated system prompt in input information to the generative model M2 and transmit the input information to the generative model M2 to the generation device 30.
[0160] As an example, a scenario will be described in which a dialogue is executed in which the interlocutor S is a doctor and the dialogue device 10 is a patient. The reference information includes text describing how the symptoms occurred, the current symptoms, medical history, chronic illnesses, medication history, vital signs, facial expressions, and, if there is an injury, the details of the injury. The utterance generation unit 120 generates a query based on the question from the interlocutor S, and acquires reference information related to the question from the interlocutor S based on the query. The utterance generation unit 120 generates a system prompt including the acquired reference information, and generates a system utterance that responds to the question from the interlocutor S based on the generation model M2. This allows the dialogue device 10 to appropriately respond to the question from the interlocutor S.
[0161] For simplicity of explanation, the reference information has been described assuming that it is text information, but the reference information is not limited to text information and may include audio information, image information, etc.
[0162] [Second embodiment] In the first embodiment, a configuration has been described in which an evaluation result for interlocutor S is generated based on a pre-constructed evaluation model M1. In the second embodiment, a configuration will be described in which an evaluation model M1 according to the evaluation criteria of evaluator E is constructed, and an evaluation result for interlocutor S is generated based on the evaluation model M1 according to the evaluation criteria of evaluator E.
[0163] The following describes the dialogue system 1000 according to the second embodiment, focusing on the differences from the first embodiment.
[0164] <Overall configuration of the dialogue system> The overall configuration of the dialogue system according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing an example of the overall configuration of the dialogue system according to the second embodiment.
[0165] 4, the dialogue system 1000 includes a dialogue device 10, a control device 20, a generation device 30, a terminal device 40, and a management device 50. That is, the dialogue system 1000 according to this embodiment differs from the dialogue system 1000 according to the first embodiment (see FIG. 1) in that it further includes a management device 50.
[0166] The management device 50 is an example of an information processing device such as a personal computer, a workstation, or a server that manages the evaluation model M1. The management device 50 may generate the evaluation model M1. The management device 50 may generate training data used to generate the evaluation model M1. The management device 50 may select the evaluation model M1 to be used for evaluating the interlocutor S. The management device 50 may transmit the evaluation model M1 to be used for evaluating the interlocutor S to the control device 20.
[0167] The management device 50 may store recorded data V of the conversation. The recorded data V may be electronic data recording a conversation between a first and a second person. The recorded data V may include an audio signal capturing speech from the first or second person. The recorded data V may also include a video signal capturing an image of the first or second person.
[0168] The first interlocutor may be a real person or the dialogue device 10. The second interlocutor may be a real person or the dialogue device 10. In other words, the recording data V may be recording data recording a dialogue between real people, recording data recording a dialogue between a real person and the dialogue device 10, or recording data recording a dialogue between the dialogue devices 10.
[0169] The management device 50 may include an evaluation model group M3. The evaluation model group M3 may include a plurality of evaluation models M1. The evaluation model group M3 may include a plurality of evaluation models M1 with different evaluation criteria. The evaluation model M1 may be generated based on the video recording data V to which a correct label according to the evaluation criteria has been assigned. The correct label may be a correct value of the evaluation value.
[0170] The evaluation criteria may be set for each evaluator E. The evaluation criteria may be set based on the attributes of the evaluator E. As an example, the evaluation criteria may be set for each industry of the evaluator E. The evaluation criteria may be set based on the attributes of the interlocutor S. As an example, the evaluation criteria may be set for each occupation of the interlocutor S. As an example, the evaluation criteria may be set for each scenario. This is because if there is a scenario in which it is better to respond with a smile and a scenario in which it is better to respond with a blank expression, the evaluation criteria for each scenario would be different.
[0171] Evaluation according to specific evaluation criteria may be performed as follows. First, one or more feature quantities are extracted based on the recorded dialogue data V. The one or more feature quantities may be feature quantities independent of the evaluation criteria. Next, evaluation according to the evaluation criteria of evaluator E is performed based on the one or more extracted feature quantities. At this time, the evaluation value may be calculated by linearly transforming the one or more feature quantities. Alternatively, the evaluation value may be calculated by adding up the feature quantities after linear transformation. Alternatively, the evaluation value may be obtained based on one or more machine learning models.
[0172] The overall configuration of the dialogue system 1000 shown in FIG. 4 is an example, and various system configuration examples are possible depending on the application and purpose. As an example, one or more of the dialogue device 10, the control device 20, the generation device 30, the terminal device 40, and the management device 50 may be included in the dialogue system 1000. The management device 50 may be realized by multiple computers or may be realized as a cloud computing service. The management device 50 may be realized by a standalone computer integrated with the control device 20 or the generation device 30.
[0173] The dialogue system 1000 may include a plurality of control devices 20 installed for each evaluator E. The dialogue system 1000 may include a single control device 20 used by a plurality of evaluators E. The dialogue system 1000 may include a plurality of management devices 50 installed for each evaluator E. The dialogue system 1000 may include a single control device 20 used by a plurality of evaluators E.
[0174] <Functional configuration of management device> The functional configuration of the management device 50 will be described with reference to Fig. 5. Fig. 5 is a block diagram showing an example of the functional configuration of the management device according to the second embodiment.
[0175] 5, the management device 50 includes a data storage unit 501, a model storage unit 502, a label assignment unit 510, a feature extraction unit 520, a model learning unit 530, a model evaluation unit 540, and a model selection unit 550. The management device 50 functions as the data storage unit 501, the model storage unit 502, the label assignment unit 510, the feature extraction unit 520, the model learning unit 530, the model evaluation unit 540, and the model selection unit 550 by executing a management program installed in advance.
[0176] (Data storage unit) The data storage unit 501 stores recorded data V. The recorded data V may be generated by the control device 20. The recorded data V may be recorded using the interactive device 10. The recorded data V may be generated by another information processing device or information processing system.
[0177] The recorded data V may include recorded data included in the dialogue data D stored in the control device 20. As an example, the recorded data V may be recorded data extracted from the dialogue data D stored in the control device 20 during a preset data collection period.
[0178] The recorded data V may include recorded data collected from unspecified users. For example, recorded data of conversations may be collected by a method such as crowdsourcing. The recorded data V may be recorded for any purpose. In other words, the recorded data V does not have to be recorded for the purpose of constructing the evaluation model M1.
[0179] The recording data V may be recording data of a conversation in which one of the interlocutors is an interlocutor S designated by an evaluator E. The evaluator E may designate an interlocutor S with a fixed evaluation. The evaluator E may designate an interlocutor S with a high evaluation. The evaluator E may designate an interlocutor S with a low evaluation.
[0180] (Model storage section) The model storage unit 502 stores an evaluation model group M3. The evaluation model group M3 may include multiple evaluation models M1 with different evaluation criteria. The evaluation model M1 included in the evaluation model group M3 may be stored in advance in the model storage unit 502. The evaluation model M1 included in the evaluation model group M3 may be acquired from the terminal device 40. The evaluation model M1 acquired from the terminal device 40 may be generated by the evaluator E. The evaluation model M1 included in the evaluation model group M3 may be generated by the model learning unit 530 and stored in the model storage unit 502.
[0181] The evaluation model M1 included in the evaluation model group M3 is the same as in the first embodiment. Therefore, the evaluation model M1 may include an evaluation model that evaluates the facial expression of the interlocutor S. The evaluation model M1 may include an evaluation model that evaluates the body movement of the interlocutor S. The evaluation model M1 may include an evaluation model that evaluates the voice of the interlocutor S. The evaluation model M1 may include an evaluation model that evaluates the smoothness of the dialogue. The evaluation model M1 may include an evaluation model that evaluates the content of the dialogue.
[0182] (Label assignment section) The label assignment unit 510 assigns a correct label to the video recording data V. The label assignment unit 510 may assign a correct label to the video recording data V stored in the data storage unit 501. The label assignment unit 510 may assign a correct label to the video recording data V in accordance with evaluation criteria. The label assignment unit 510 may assign a correct label designated by an evaluator E to the video recording data V.
[0183] The correct label is information indicating the correct value of the evaluation value included in the evaluation result. The correct label may include, for example, a numerical value (e.g., 1 to 10), a character (e.g., alphabet), or a symbol (e.g., ◯, ×, △).
[0184] The correct answer label may include an evaluation value that evaluates the audio included in the video recording data V. The correct answer label may include an evaluation value that evaluates the video included in the video recording data V. The correct answer label may include an evaluation value that evaluates the content of the dialogue included in the video recording data V.
[0185] The label assignment unit 510 may correct the correct label assigned to the recorded data related to the interlocutor S designated by the evaluator E. The label assignment unit 510 may correct the correct label assigned to the recorded data related to the interlocutor S who has received a high evaluation from the evaluator E to a correct label with a high evaluation. The label assignment unit 510 may correct the correct label assigned to the recorded data related to the interlocutor S who has received a low evaluation from the evaluator E to a correct label with a low evaluation.
[0186] An interlocutor S who is highly rated by evaluator E is considered to be an interlocutor who can have a conversation in accordance with evaluator E's evaluation criteria. On the other hand, an interlocutor S who is low rated by evaluator E is considered to be an interlocutor who has a conversation that does not comply with evaluator E's evaluation criteria. By assigning a correct answer label according to evaluator E's evaluation to the video recording data regarding interlocutor S specified by evaluator E, an evaluation result that conforms to evaluator E's evaluation criteria can be generated.
[0187] The labeling unit 510 may assign a correct label based on a machine learning model. The machine learning model may be a trained evaluation model M1, a generative model M2, or another machine learning model. The other machine learning model may be, for example, a large-scale language model or a base model.
[0188] The label assignment unit 510 may correct the correct label assigned based on the machine learning model. The label assignment unit 510 may present the correct label assigned based on the machine learning model to the evaluator E and correct the label to a correct label designated by the evaluator E.
[0189] (Feature extraction section) The feature extraction unit 520 extracts a feature vector from the recording data V. The feature extraction unit 520 may extract a feature vector to be input to the evaluation model M1 included in the evaluation model group M3. The feature extraction unit 520 may extract a feature vector including a feature amount specified by the evaluator E.
[0190] The feature vector includes the same feature amounts as in the first embodiment. Therefore, the feature vector may include feature amounts extracted from speech information of interlocutor S. The feature amounts extracted from speech information may include, for example, tone, intonation, speaking rate, voice pitch, clarity, speaking time, number of fillers, etc.
[0191] The feature vector may include feature amounts extracted from image information of the interlocutor S. The feature amounts extracted from the image information may include, for example, the degree of smiling, the degree of expressionlessness, the degree of facial expression expressing emotions such as anger or sadness, the direction of gaze, etc.
[0192] The feature vector may include features extracted from the dialogue content, such as the number of positive words, the number of negative words, and the proportion of polite language.
[0193] (Model Learning Department) The model learning unit 530 learns the evaluation model M1 to be included in the evaluation model group M3. The model learning unit 530 may learn the evaluation model M1 based on training data in which the video recording data V is assigned a correct answer label.
[0194] Specifically, the model learning unit 530 inputs the feature vector extracted from the video recording data V by the feature extraction unit 520 into the evaluation model M1. The evaluation model M1 predicts an evaluation value of 1 or more based on the input feature vector and outputs the predicted evaluation value. The model learning unit 530 calculates difference information between the predicted value output from the evaluation model M1 and the correct label assigned to the video recording data V. The model learning unit 530 updates the parameters of the evaluation model M1 based on the difference information between the predicted value and the correct label. The model learning unit 530 repeats the prediction of the evaluation value and the update of the parameters until a predetermined convergence condition is satisfied. The convergence condition may be that the number of parameter updates is equal to or greater than a threshold, or that the amount of parameter update is less than a threshold, or that a numerical value calculated from the difference information is less than a threshold.
[0195] The model learning unit 530 may learn an evaluation model M1 according to an evaluation criterion. The model learning unit 530 may learn, for each of a plurality of evaluation criteria, an evaluation model M1 that outputs an evaluation value according to the evaluation criterion. The model learning unit 530 may learn the evaluation model M1 according to the evaluation criterion based on training data to which a correct answer label according to the evaluation criterion has been assigned.
[0196] The model training unit 530 may additionally train the pre-trained evaluation model M1. The additional training may include, for example, transfer learning or fine tuning. The model training unit 530 may additionally train the evaluation model M1 based on training data in which a correct answer label is assigned to the video recording data V. For example, the model training unit 530 may additionally train the evaluation model M1, which has been pre-trained using a general-purpose evaluation criterion, based on training data to which a correct answer label is assigned according to the evaluation criterion of the evaluator E.
[0197] The model learning unit 530 may construct the evaluation model M1 by combining multiple feature amounts extracted from the recorded data. As an example, the model learning unit 530 may construct the evaluation model M1 by converting multiple feature amounts into a single evaluation value using a weighted linear sum. The model learning unit 530 may construct the evaluation model M1 based on the combination of feature amounts performed by the evaluator E. In the case of a weighted linear sum, the evaluator E may construct the evaluation model M1 by manually manipulating the weights using the terminal device 40.
[0198] (Model Evaluation Department) The model evaluation unit 540 evaluates the evaluation model M1 learned by the model learning unit 530. The model evaluation unit 540 may obtain a predicted value of the evaluation value by the evaluation model M1 using recorded data prepared in advance, and present the predicted value to the evaluator E. The recorded data prepared in advance may be recorded data that has not been used in learning the evaluation model M1.
[0199] The evaluator E may evaluate the performance of the evaluation model M1 based on the predicted value of the evaluation value presented by the model evaluation unit 540. As an example, the evaluator E may determine whether the predicted value of the evaluation value is an expected result. The evaluator E may input the evaluation result of the evaluation model M1 to the terminal device 40.
[0200] The model evaluation unit 540 may acquire the evaluation result of the evaluation model M1. The model evaluation unit 540 may acquire the evaluation result of the evaluation model M1 inputted to the terminal device 40 by the evaluator E. The model evaluation unit 540 may store the evaluation result of the evaluation model M1 in the model storage unit 502 in association with the evaluation model M1.
[0201] The evaluator E may use the evaluation result of the evaluation model M1 to select the evaluation model M1 to be used for evaluating the interlocutor S. The evaluation result of the evaluation model M1 may be shared among multiple evaluators E. That is, the evaluation result of the evaluation model M1 input by a first evaluator E may be referenced by a second evaluator E.
[0202] (Model selection section) The model selection unit 550 selects an evaluation model M1 to be used for evaluating the interlocutor S. The model selection unit 550 may select the evaluation model M1 to be used for evaluating the interlocutor S from among a plurality of evaluation models M1 included in the evaluation model group M3. The model selection unit 550 may select a plurality of evaluation models M1 to be used for evaluating the interlocutor S.
[0203] The model selection unit 550 may select one or more evaluation models M1 specified by the evaluator E. The model selection unit 550 may select multiple evaluation models M1 specified by the evaluator E. The evaluator E may select an evaluation model M1 according to scenario information used for evaluating the interlocutor S. The evaluator E may select one or more evaluation models M1 based on the evaluation result by the model evaluation unit 540.
[0204] As an example, when evaluating the customer service skills of interlocutor S, evaluator E may select evaluation model M1 constructed with evaluation criteria for customer service. When evaluating a product explanation by interlocutor S, evaluator E may select evaluation model M1 constructed with evaluation criteria for product explanation. When evaluating a technical explanation by interlocutor S, evaluator E may select evaluation model M1 constructed with evaluation criteria for technical explanation.
[0205] The model selection unit 550 may transmit the selected evaluation model M1 to the control device 20. The model selection unit 550 may transmit the evaluation model M1 to the control device 20 in response to an instruction from the evaluator E. The control device 20 may store the evaluation model M1 received from the management device 50 in the model storage unit 103.
[0206] The model selection unit 550 may select the evaluation model M1 associated with the scenario information. When scenario information is selected by the control device 20, the model selection unit 550 may select the evaluation model M1 associated with the scenario information and transmit the evaluation model M1 to the control device 20.
[0207] The functional configuration of the management device 50 shown in Fig. 5 is one example, and it goes without saying that there are various examples of functional configurations depending on the application and purpose. The division of the storage unit, such as the data storage unit 501 and model storage unit 502 shown in Fig. 5, is one example. The division of the processing unit, such as the label assignment unit 510, feature extraction unit 520, model learning unit 530, model evaluation unit 540, and model selection unit 550 shown in Fig. 5, is one example.
[0208] For example, the data storage unit 501 and the model storage unit 502 may be integrated into one storage unit, or the data storage unit 501 or the model storage unit 502 may be divided into multiple storage units.
[0209] For example, at least two of the labeling unit 510, the feature extraction unit 520, the model learning unit 530, the model evaluation unit 540, and the model selection unit 550 may be integrated into one processing unit. Also, for example, at least one of the utterance acquisition unit 110, the utterance generation unit 120, the utterance presentation unit 130, the dialogue evaluation unit 140, and the result output unit 150 may be divided into multiple processing units.
[0210] <Dialogue method flow> The dialogue method executed by the dialogue system 1000 will be described with reference to Fig. 6. Fig. 6 is a processing flow showing an example of the dialogue method according to the second embodiment.
[0211] In step S201, the terminal device 40 acquires the recorded data V from the management device 50. The terminal device 40 displays the recorded data V on a display device. The evaluator E refers to the recorded data V displayed on the display device and inputs an evaluation value for the recorded data V. The evaluator E may repeatedly evaluate multiple recorded data V.
[0212] The terminal device 40 transmits the evaluation value input by the evaluator E to the management device 50. The terminal device 40 may transmit the evaluation value to the management device 50 every time an evaluation value for the recorded data V is input. The terminal device 40 may also transmit evaluation values for multiple recorded data V to the management device 50 all at once.
[0213] In step S202, the label assignment unit 510 of the management device 50 receives the evaluation value input by the evaluator E from the terminal device 40. The label assignment unit 510 assigns a correct label to the video recording data V based on the evaluation value input by the evaluator E. This generates training data in which the correct label is assigned to the video recording data V. The label assignment unit 510 sends the training data to the model learning unit 530.
[0214] In step S203, the feature extraction unit 520 of the management device 50 extracts a feature vector from the recorded data V. The feature extraction unit 520 may extract a feature vector from the recorded data V to which the correct label was assigned in step S202. The feature extraction unit 520 sends the feature vector to the model learning unit 530.
[0215] In step S204, the model learning unit 530 of the management device 50 receives the training data from the label assignment unit 510. The model learning unit 530 also receives the feature vector from the feature extraction unit 520.
[0216] The model learning unit 530 learns the evaluation model M1 based on the training data and the feature vector. The model learning unit 530 stores the trained evaluation model M1 in the model storage unit 502. As a result, the trained evaluation model M1 is included in the evaluation model group M3.
[0217] In step S205, the evaluator E performs an operation to evaluate the evaluation model M1 on the terminal device 40. The terminal device 40 transmits an evaluation request to the management device 50 in response to the operation by the evaluator E.
[0218] The model evaluation unit 540 of the management device 50 evaluates the evaluation model M1 in response to an evaluation request from the terminal device 40. Specifically, the model evaluation unit 540 extracts a feature vector from pre-prepared recorded data and inputs it to the evaluation model M1. The evaluation model M1 predicts one or more evaluation values based on the input feature vector and outputs the predicted evaluation values. The model evaluation unit 540 transmits the predicted evaluation values to the terminal device 40. The terminal device 40 displays the predicted evaluation values received from the management device 50 on a display device.
[0219] The evaluator E may evaluate the performance of the evaluation model M1 by referring to the predicted value of the evaluation value displayed on the display device. The evaluator E may input the evaluation result of the evaluation model M1 to the terminal device 40. When the evaluation result of the evaluation model M1 is input by the evaluator E, the terminal device 40 transmits the input evaluation result of the evaluation model M1 to the management device 50.
[0220] In step S206, the model evaluation unit 540 of the management device 50 receives the evaluation result of the evaluation model M1 from the terminal device 40. The model evaluation unit 540 stores the evaluation result of the evaluation model M1 in the model storage unit 502 in association with the evaluation model M1.
[0221] In step S207, the evaluator E performs an operation to select the evaluation model M1 on the terminal device 40. The terminal device 40 transmits a selection request to the management device 50 in response to the operation by the evaluator E.
[0222] In response to a selection request from the terminal device 40, the model selection unit 550 of the management device 50 transmits a list of the evaluation models M1 included in the evaluation model group M3 and the evaluation results of each evaluation model M1 to the terminal device 40. The terminal device 40 displays the list of evaluation models M1 and the evaluation results received from the management device 50 on a display device.
[0223] The evaluator E may select the evaluation model M1 to be used for evaluating the interlocutor S from a list of evaluation models M1 displayed on the display device. When the evaluation model M1 is selected by the evaluator E, the terminal device 40 transmits the selection result of the evaluation model M1 to the management device 50. The selection result includes information indicating the evaluation model M1 selected by the evaluator E.
[0224] In step S208, the model selection unit 550 of the management device 50 receives the selection result of the evaluation model M1 from the terminal device 40. Based on the selection result of the evaluation model M1, the model selection unit 550 selects one or more evaluation models M1 to be used for evaluating the interlocutor S from among the evaluation models M1 included in the evaluation model group M3. The model selection unit 550 transmits the selected one or more evaluation models M1 to the control device 20.
[0225] In step S209, the control device 20 receives one or more evaluation models M1 from the management device 50. The control device 20 stores the received one or more evaluation models M1 in the model storage unit 103.
[0226] In step S210, the interlocutor S executes a dialogue with the dialogue device 10. The control device 20 controls the dialogue between the interlocutor S and the dialogue device 10, and stores dialogue data D recording the dialogue between the interlocutor S and the dialogue device 10 in the dialogue storage unit 102. The dialogue system 1000 may execute, as step S210, the processes from step S101 to step S110 of the dialogue method according to the first embodiment (see FIG. 3).
[0227] In step S211, the dialogue evaluation unit 140 of the control device 20 reads out dialogue data D regarding the dialogue with the interlocutor S from the dialogue storage unit 102. The dialogue evaluation unit 140 also reads out an evaluation model M1 corresponding to the evaluator E from the model storage unit 103.
[0228] The dialogue evaluation unit 140 generates evaluation model input information based on the read dialogue data D. The dialogue evaluation unit 140 inputs the evaluation model input information to the evaluation model M1. The evaluation model M1 calculates an evaluation value of 1 or more based on the evaluation model input information. As a result, an evaluation result for the interlocutor S is generated. The dialogue evaluation unit 140 stores the evaluation result for the interlocutor S in the evaluation result storage unit 104.
[0229] In step S212, the evaluator E performs an operation to display the evaluation results on the terminal device 40. As an example, the evaluator E may perform an operation to open the evaluation results of the interlocutor S on the evaluation screen displayed on the display device of the terminal device 40. The terminal device 40 transmits an output request for the evaluation results regarding the interlocutor S to the control device 20.
[0230] The result output unit 150 of the control device 20 receives a request to output the evaluation result from the terminal device 40. The result output unit 150 reads out the evaluation result regarding the interlocutor S from the evaluation result storage unit 104. The result output unit 150 generates screen data for displaying the evaluation result regarding the interlocutor S. The result output unit 150 transmits the screen data to the terminal device 40.
[0231] In step S213, the terminal device 40 receives screen data from the control device 20. The terminal device 40 displays the evaluation results for the interlocutor S on the display device based on the screen data. The evaluator E may evaluate the interlocutor S by referring to the evaluation results for the interlocutor S displayed on the display device of the terminal device 40. As an example, the evaluator E may refer to the evaluation results for the interlocutor S to decide whether to hire the interlocutor S. As an example, the evaluator E, who is a training instructor, may refer to the evaluation results for the interlocutor S to judge the quality of the training. As an example, the evaluator E, who is a test grader, may refer to the evaluation results for the interlocutor S to determine the score for the interlocutor S and judge whether the test is passed or failed.
[0232] [Modification 1 of the second embodiment] The model storage unit 103 of the control device 20 may store a plurality of evaluation models M1 corresponding to a plurality of evaluation criteria. The model storage unit 103 of the control device 20 may store all of the evaluation models M1 included in the evaluation model group M3 stored in the model storage unit 502 of the management device 50. The plurality of evaluation criteria may include a general-purpose evaluation criterion that is independent of the evaluator E. The plurality of evaluation criteria may include a plurality of general-purpose evaluation criteria.
[0233] The dialogue evaluation unit 140 of the control device 20 may generate evaluation results according to each of multiple evaluation criteria for one piece of dialogue data D. The dialogue evaluation unit 140 of the control device 20 may use all of the evaluation models M1 stored in the model storage unit 103 to generate evaluation results according to all of the evaluation criteria for one piece of dialogue data D.
[0234] When the control device 20 generates evaluation results according to multiple evaluation criteria, the evaluator E may only be able to refer to the evaluation results based on his / her own evaluation criteria. When the control device 20 generates evaluation results according to multiple evaluation criteria, the evaluator E may be able to refer to the evaluation results according to multiple evaluation criteria. In other words, the evaluator E may be able to refer to the evaluation results based on the evaluation criteria of other evaluators E. Whether the evaluator E can refer to the evaluation results based on the evaluation criteria of other evaluators E may be set by the usage plan of the evaluator E. The evaluation criteria that the evaluator E can refer to may be selected based on the usage conditions, input information, etc. of the evaluator E or the other evaluators E.
[0235] The scenario information stored by the control device 20 may be associated with the evaluator E. The scenario information may be generated by the evaluator E. The scenario information may be associated with the evaluator E who generated the scenario information. When generating an evaluation result according to the evaluation criteria of the evaluator E, the dialogue evaluation unit 140 of the control device 20 may generate an evaluation result regarding the interlocutor S using only the dialogue data D related to the dialogue executed using the scenario information associated with the evaluator E.
[0236] [Application example] The dialogue system 1000 according to the second embodiment can realize a service that provides evaluation results regarding a conversation partner S to a plurality of evaluators E. In this application example, a configuration will be described in which a service provider collects dialogue data D regarding a dialogue between the conversation partner S and the dialogue device 10, and provides each of the plurality of evaluators E with an evaluation result according to the evaluation criteria of the evaluator E.
[0237] The overall configuration of a dialogue system according to this application example will be described with reference to Fig. 7. Fig. 7 is a block diagram showing an example of the overall configuration of a dialogue system according to this application example. Note that in this application example, an example of a service that provides evaluation results regarding interlocutor S to two evaluators E (E-1, E-2) will be described, but the number of evaluators E is not limited and may be three or more.
[0238] As shown in FIG. 7, the dialogue system 1000 includes a dialogue device 10, a control device 20, two evaluation devices 25 (25-1, 25-2), a generation device 30, two terminal devices 40 (40-1, 40-2), and a management device 50.
[0239] The control device 20 and the management device 50 are installed by a service provider. As an example, the control device 20 and the management device 50 may be installed in a facility managed by the service provider, or may be installed in a remote location such as a data center. The generation device 30 may be installed by the service provider, or may be installed by an entity other than the service provider.
[0240] The dialogue device 10 is installed for each interlocutor S. The dialogue device 10 may be installed by the interlocutor S or by a service provider. As an example, the dialogue device 10 may be installed in a space where the interlocutor S is present.
[0241] The evaluation device 25 and the terminal device 40 are installed for each evaluator E. The evaluation device 25 and the terminal device 40 may be installed by the evaluator E or by the service provider. As an example, the evaluation device 25 may be installed in a facility managed by the service provider, or in a remote location such as a data center, or may be installed in a space where the evaluator E is present. The terminal device 40 may be installed in a space where the evaluator E is present.
[0242] The multiple evaluators E may be entities with different evaluation criteria. For example, evaluator E-1 may be a first company or its employee, and evaluator E-2 may be a second company or its employee. For example, evaluator E-1 may be a company in a first industry or its employee, and evaluator E-2 may be a company in a second industry or its employee.
[0243] The management device 50 stores an evaluation model group M3. The evaluation model group M3 includes an evaluation model M1-1 and an evaluation model M1-2. The evaluation model M1-1 is constructed based on training data to which a correct answer label has been assigned by an evaluator E-1 in accordance with the evaluation criteria of the evaluator E-1. The evaluation model M1-2 is constructed based on training data to which a correct answer label has been assigned by an evaluator E-2 in accordance with the evaluation criteria of the evaluator E-2.
[0244] In this application example, an example in which an evaluation model M1 is constructed for each evaluator E has been shown, but one evaluation model M1 may be constructed for multiple evaluators E. As an example, the same evaluation model M1 may be constructed for evaluators E with the same evaluation criteria. As an example, the evaluation model M1 may be constructed so as to receive identification information identifying the evaluator E as input and output an evaluation value in accordance with the evaluation criteria of the evaluator E identified by the input identification information.
[0245] The management device 50 transmits the evaluation model M1-1 to the evaluation device 25-1. The evaluation device 25-1 stores the evaluation model M1-1 received from the management device 50. The management device 50 transmits the evaluation model M1-2 to the evaluation device 25-2. The evaluation device 25-2 stores the evaluation model M1-2 received from the management device 50.
[0246] The dialogue device 10 dialogues with a party S. The control device 20 controls the dialogue between the party S and the dialogue device 10, and stores dialogue data D that records the dialogue between the party S and the dialogue device 10. The control device 20 transmits the dialogue data D to each of the evaluation devices 25-1 and 25-2. The evaluation devices 25-1 and 25-2 each store the dialogue data D.
[0247] The evaluation device 25-1 generates an evaluation result for the interlocutor S based on the evaluation model M1-1. The evaluation device 25-1 transmits the evaluation result for the interlocutor S to the terminal device 40-1. The terminal device 40-1 presents the evaluation result for the interlocutor S to the evaluator E-1. The evaluator E-1 can obtain the evaluation result in accordance with the evaluation criteria of the evaluator E-1.
[0248] The evaluation device 25-2 generates an evaluation result for the interlocutor S based on the evaluation model M1-2. The evaluation device 25-2 transmits the evaluation result for the interlocutor S to the terminal device 40-2. The terminal device 40-2 presents the evaluation result for the interlocutor S to the evaluator E-2. The evaluator E-2 can obtain the evaluation result in accordance with the evaluation criteria of the evaluator E-2.
[0249] The evaluators E-1 and E-2 may create their own evaluation criteria by customizing a pre-created evaluation criteria template. The evaluation criteria template may be provided by a service provider or another evaluator E. The evaluators E-1 and E-2 may select evaluation criteria to use in evaluating the interlocutor S from a plurality of pre-created evaluation criteria. The service provider may propose evaluation criteria based on the attributes of the evaluators E-1 and E-2 or the attributes of the interlocutor S. As an example, the attributes of the evaluators E-1 and E-2 may be industry. As an example, the attributes of the interlocutor S may be occupation.
[0250] The control device 20 may switch the scenario information based on the evaluation criteria of the evaluators E-1 and E-2. The control device 20 may switch the scenario information according to the attributes of the evaluators E-1 and E-2 or the attributes of the interlocutor S.
[0251] To configure the dialogue system 1000 according to this application example, the control device 20 may include the scenario storage unit 101, dialogue storage unit 102, utterance acquisition unit 110, utterance generation unit 120, and utterance presentation unit 130 that are included in the control device 20 according to the first embodiment. Furthermore, the evaluation device 25 may include the dialogue storage unit 102, model storage unit 103, evaluation result storage unit 104, dialogue evaluation unit 140, and result output unit 150 that are included in the control device 20 according to the first embodiment.
[0252] [Other embodiments] In the above-described embodiments, modifications, and application examples, examples have been described in which a dialogue using voice information and image information is executed between the interlocutor S and the dialogue device 10. The above-described embodiments, modifications, and application examples are not limited to a dialogue using voice information and image information, and may be applied to a dialogue using text information, for example. A dialogue using text information may include, for example, text chat, email, messaging services, etc.
[0253] <Summary> As is clear from the above description, the dialogue system 1000 according to an embodiment of the present disclosure acquires information about a first utterance uttered by a user, generates information about a second utterance based on the information about the first utterance, input information, and a machine learning model, and controls the utterance of the dialogue device based on the information about the second utterance. The input information includes information indicating the relationship between the user and the dialogue device.
[0254] The relationship between the user and the interactive device may include at least one of a relationship in which the user provides a service or product to the interactive device, a relationship in which the user instructs the interactive device, a relationship in which the interactive device pays a reward to the user, or a relationship in which a user who has information A provides information A to an interactive device that does not have information A. The relationship between the user and the interactive device may include at least one of a relationship between a store clerk and a customer, a relationship between a boss and a subordinate, a relationship between a teacher and a student, a relationship between a lecturer and an audience, a relationship between a doctor and a patient, or a relationship between a requested person and a requester.
[0255] The information indicating the relationship between the user and the interactive device may include at least information regarding a role of the interactive device. The role of the interactive device may include at least one of a customer, a subordinate, a student, an audience member, a patient, or a client. The information indicating the relationship between the user and the interactive device may include information regarding a role of the user and a role of the interactive device.
[0256] The dialogue system 1000 may generate an evaluation result for the user based on at least a part of the information about the first utterance. The dialogue system 1000 may generate second input information for a second machine learning model based on dialogue data including information about the first utterance and information about the second utterance, and input the second input information to the second machine learning model to generate the evaluation result.
[0257] The second input information may include one or more features extracted from the dialogue data, and the features may include information on at least one of a facial expression of the user, a body movement of the user, a voice of the user, a smoothness of dialogue between the user and the dialogue device, or a content of the dialogue.
[0258] The dialogue system 1000 may acquire an evaluation value for the dialogue data by an evaluator, and generate a second machine learning model based on the training data to which the evaluation value is assigned. The dialogue system 1000 may acquire an evaluation value for each of a plurality of evaluation criteria, and generate a second machine learning model for each evaluation criterion.
[0259] The dialogue system 1000 may select a second machine learning model to be used for generating the evaluation result from the second machine learning models for each evaluation criterion. The dialogue system 1000 may select a second machine learning model for the evaluation criterion corresponding to the evaluator.
[0260] The dialogue system 1000 may generate multiple evaluation results for one piece of dialogue data based on one or more second machine learning models. The dialogue system 1000 may output the evaluation results for the user and information used for the evaluation of the user. The dialogue system 1000 may present at least a portion of the information about the first utterance to the user or another user.
[0261] The dialogue system 1000 may generate input information based on preset scenario information, or may set the scenario information based on predetermined criteria, or may set the scenario information based on instructions from an evaluator.
[0262] The dialogue system 1000 may repeatedly acquire information about utterances uttered by a user, generate information about the utterances of the dialogue device based on the information about the utterances of the user, input information, and a machine learning model, and control the utterances of the dialogue device based on the information about the utterances of the dialogue device, until a predetermined condition is satisfied. The dialogue system 1000 may generate an evaluation result for the user based on dialogue data including information about multiple utterances of the user and information about multiple utterances of the dialogue device. The dialogue system 1000 may generate the evaluation result based on the dialogue data and a second machine learning model.
[0263] As a result, according to one embodiment of the present disclosure, information regarding the utterance of the dialogue device is generated based on a machine learning model, making it possible to control a device that interacts with a user. In one aspect, according to one embodiment, information for evaluating a user can be obtained through a dialogue between the user and the dialogue device. In one aspect, according to one embodiment, responses to the utterance of the user are made based on a machine learning model, allowing the user to interact with the dialogue device in a natural flow, and allowing the evaluator to appropriately evaluate the user. In another aspect, according to one embodiment, an appropriate evaluation result according to the evaluation criteria set by the evaluator can be obtained.
[0264] [Hardware configuration of information processing device] Some or all of the devices (interaction device 10, control device 20, evaluation device 25, generation device 30, terminal device 40, and management device 50) in the above-described embodiments may be configured as hardware, or may be configured as software (program) information processing executed by a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or the like. In the case of software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or a USB (Universal Serial Bus) memory, and the software information processing may be executed by loading the software into a computer. The software may also be downloaded via a communication network. Furthermore, all or part of the software processing may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array), thereby executing the software information processing by hardware.
[0265] The storage medium that stores the software may be a removable medium such as an optical disk, or a fixed medium such as a hard disk, memory, etc. The storage medium may be provided inside the computer (main storage device, auxiliary storage device, etc.) or outside the computer.
[0266] 8 is a block diagram showing an example of the hardware configuration of each device (the dialogue device 10, the control device 20, the evaluation device 25, the generation device 30, the terminal device 40, and the management device 50) in the above-described embodiment. Each device may be realized as a computer 7 including, for example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, which are connected via a bus 76.
[0267] Although the computer 7 in FIG. 8 includes one of each component, it may also include multiple of the same component. Also, while FIG. 8 shows one computer 7, the software may be installed on multiple computers, and each of the multiple computers may execute the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates with the other computers via a network interface 74 or the like to execute the processing. That is, each device in the above-described embodiment (the dialogue device 10, the control device 20, the evaluation device 25, the generation device 30, the terminal device 40, and the management device 50) may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize its functions. Furthermore, the system may be configured such that information transmitted from a terminal is processed by one or more computers provided on a cloud, and the processing results are transmitted to the terminal.
[0268] The various calculations of each device (interaction device 10, control device 20, evaluation device 25, generation device 30, terminal device 40, and management device 50) in the above-described embodiments may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processing cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. of the present disclosure may be realized by at least one of a processor and a storage device provided on a cloud that can communicate with computer 7 via a network. In this way, each device in the above-described embodiments may be implemented in the form of parallel computing using one or more computers.
[0269] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that at least controls a computer or performs calculations. The processor 71 may be a general-purpose processor, a dedicated processing circuit designed to perform a specific calculation, or a semiconductor device that includes both a general-purpose processor and a dedicated processing circuit. The processor 71 may also include an optical circuit or a calculation function based on quantum computing.
[0270] The processor 71 may perform arithmetic processing based on data or software input from each device or the like configured inside the computer 7, and may output the calculation results or control signals to each device or the like. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.
[0271] Each device in the above-described embodiment (the dialogue device 10, the control device 20, the evaluation device 25, the generation device 30, the terminal device 40, and the management device 50) may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.
[0272] The main memory device 72 may store instructions executed by the processor 71, various data, etc., and information stored in the main memory device 72 may be read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory devices for saving various data, etc. in each device (the dialogue device 10, the control device 20, the evaluation device 25, the generation device 30, the terminal device 40, and the management device 50) in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by an internal memory built into the processor 71. For example, each memory unit in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73.
[0273] When each device in the above-described embodiment (the interactive device 10, the control device 20, the evaluation device 25, the generation device 30, the terminal device 40, and the management device 50) is configured with at least one storage device (memory) and at least one processor connected (coupled) to this at least one storage device, at least one processor may be connected to one storage device. Also, at least one storage device may be connected to one processor. Also, a configuration in which at least one processor among multiple processors is connected to at least one storage device among multiple storage devices may be included. Also, this configuration may be realized by storage devices and processors included in multiple computers. Furthermore, a configuration in which a storage device is integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache) may be included.
[0274] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be an appropriate interface, such as one that conforms to an existing communication standard. Information may be exchanged with an external device 9A connected via the communication network 8 via the network interface 74. The communication network 8 may be any one of a WAN (Wide Area Network), a LAN (Local Area Network), a PAN (Personal Area Network), etc., or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.
[0275] The device interface 75 is an interface such as a USB that directly connects to the external device 9B.
[0276] The external device 9A is a device connected to the computer 7 via a network. The external device 9B is a device connected directly to the computer 7.
[0277] For example, the external device 9A or the external device 9B may be an input device. The input device may be a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 7. Alternatively, the external device 9A or the external device 9B may be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0278] Furthermore, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or a speaker that outputs sound or the like. Alternatively, the output device may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0279] Furthermore, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.
[0280] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of each device (the dialogue device 10, the control device 20, the evaluation device 25, the generation device 30, the terminal device 40, and the management device 50) in the above-described embodiments. In other words, the computer 7 may transmit some or all of the processing results to the external device 9A or the external device 9B, or may receive some or all of the processing results from the external device 9A or the external device 9B.
[0281] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.
[0282] In this specification (including claims), when expressions such as "using data as input / based on / according to / in response to data" (including similar expressions) are used, unless otherwise specified, this includes cases where the data itself is used, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of data, etc.) is used. Furthermore, when a statement is made that a result is obtained "using data as input / based on / according to / in response to data" (including similar expressions), this includes cases where the result is obtained based solely on the data, or where the result is influenced by other data, factors, conditions, and / or states other than the data itself, unless otherwise specified. Furthermore, when a statement is made that "data is output" (including similar expressions), this includes cases where the data itself is used as output, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of various data, etc.) is used as output, unless otherwise specified.
[0283] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that encompass any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.
[0284] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor, dedicated arithmetic circuit, etc., it is sufficient that the circuit structure, etc. of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.
[0285] Whenever words implying containing or possessing (e.g., "comprising / including," "having," etc.) are used in this specification (including the claims), they are intended to be open-ended terms that include the inclusion or possession of things other than the object designated by the object of the term. When the object of such words implying containing or possessing does not specify a quantity or suggests a singular number (e.g., expressions using the articles "a" or "an"), the expression should be construed as not being limited to a specific number.
[0286] In this specification (including the claims), even if expressions such as "one or more" and "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.
[0287] In this specification, when a particular advantage / result is described as being obtained with respect to a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained with one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or circumstances, and that the effect is not necessarily obtained with the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or circumstances are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.
[0288] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" (including similar expressions) is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may include electronic circuits, devices including electronic circuits, etc.
[0289] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices may store only a portion of the data, or may store the entire data. Also, a configuration in which only some of the multiple storage devices store data may be included.
[0290] In this specification (including the claims), terms such as "first," "second," etc. are used merely as a way of distinguishing between two or more elements, and are not necessarily intended to impose technical meanings such as temporal aspect, spatial aspect, sequence, quantity, etc. Thus, for example, a reference to a first element and a second element does not necessarily mean that only two elements can be employed therein, that the first element must precede the second element, that the first element must be present in order for the second element to be present, etc.
[0291] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents. For example, when numerical values or formulas are used in the above-described embodiments, they are shown for illustrative purposes and do not limit the scope of the present disclosure. Furthermore, the order of each operation shown in the embodiments is also illustrative and does not limit the scope of the present disclosure.
[0292] The disclosed technology may take the following forms as described below.
[0293] (Appendix 1) at least one memory; at least one processor; The at least one processor: Obtain information about the first utterance spoken by the user; generating information about a second utterance based on information about the first utterance, input information, and a machine learning model; controlling the utterance of the dialogue device based on information about the second utterance; the input information includes information indicating a relationship between the user and the interactive device; Dialogue system.
[0294] (Appendix 2) The relationship includes at least one of a relationship in which the user provides a service or product to the interactive device, a relationship in which the user instructs the interactive device, a relationship in which the interactive device pays a reward to the user, or a relationship in which the user who has information A provides information A to the interactive device that does not have the information A. Dialogue system according to appendix 1.
[0295] (Appendix 3) The relationship includes at least one of a relationship between a salesperson and a customer, a relationship between a superior and a subordinate, a relationship between a teacher and a student, a relationship between a lecturer and an audience, a relationship between a doctor and a patient, or a relationship between a client and a client; Dialogue system according to appendix 2.
[0296] (Appendix 4) the information indicating the relationship between the user and the dialogue device includes at least information regarding the role of the dialogue device; 4. A dialogue system according to any one of appendices 1 to 3.
[0297] (Appendix 5) The role of the interactive device includes at least one of a customer, a subordinate, a student, an audience member, a patient, or a client; Dialogue system according to appendix 4.
[0298] (Appendix 6) The at least one processor: generating an evaluation result regarding the user based at least in part on information regarding the first utterance; 6. A dialogue system according to any one of appendices 1 to 5.
[0299] (Appendix 7) The at least one processor: generating second input information for a second machine learning model based on dialogue data including information about the first utterance and information about the second utterance; generating the evaluation result by inputting the second input information into the second machine learning model; Dialogue system according to appendix 6.
[0300] (Appendix 8) the second input information includes one or more feature amounts extracted from the dialogue data; Dialogue system according to claim 7.
[0301] (Appendix 9) the feature amount includes information regarding at least one of a facial expression of the user, a body movement of the user, a voice of the user, a smoothness of the dialogue between the user and the dialogue device, or a content of the dialogue; 10. The dialogue system of claim 8.
[0302] (Appendix 10) The at least one processor: obtaining an evaluation value by an evaluator for the dialogue data; generating the second machine learning model based on training data in which the evaluation values are assigned to the dialogue data; 10. A dialogue system according to any one of appendices 6 to 9.
[0303] (Appendix 11) The at least one processor: obtaining the evaluation values for each of a plurality of evaluation criteria; generating the second machine learning model for each of the evaluation criteria; 11. The dialogue system of claim 10.
[0304] (Appendix 12) The at least one processor: selecting the second machine learning model to be used for generating the evaluation result from the second machine learning models for each evaluation criterion; 12. A dialogue system according to any one of appendices 7 to 11.
[0305] (Appendix 13) The at least one processor: selecting the second machine learning model of the evaluation criterion corresponding to the rater; 13. The dialogue system of claim 12.
[0306] (Appendix 14) The at least one processor: generating a plurality of evaluation results for one of the dialogue data based on one or more of the second machine learning models; 14. A dialogue system according to any one of appendices 7 to 13.
[0307] (Appendix 15) The at least one processor: outputting the evaluation results and information used for the user's evaluation; 15. A dialogue system according to any one of appendices 6 to 14.
[0308] (Appendix 16) The at least one processor: presenting at least a portion of information related to the first utterance to the user or another user; 16. A dialogue system according to any one of appendices 1 to 15.
[0309] (Appendix 17) The at least one processor: generating the input information based on preset scenario information; 17. A dialogue system according to any one of appendices 1 to 16.
[0310] (Appendix 18) The at least one processor: setting the scenario information based on predetermined criteria; 18. The dialogue system of claim 17.
[0311] (Appendix 19) The at least one processor: setting the scenario information based on instructions from an evaluator; 19. The dialogue system of claim 18.
[0312] (Appendix 20) The at least one processor: obtaining information about utterances made by the user; generating information about the utterance of the dialogue device based on information about the utterance of the user, the input information, and the machine learning model; controlling the utterance of the dialogue device based on information about the utterance of the dialogue device; is executed repeatedly until a predetermined condition is met. 20. A dialogue system according to any one of appendices 1 to 19.
[0313] (Appendix 21) The at least one processor: generating an evaluation result of the user based on dialogue data including information on the plurality of utterances of the user and information on the plurality of utterances of the dialogue device; 21. The dialogue system of claim 20.
[0314] (Appendix 22) The at least one processor: generating the evaluation result based on the dialogue data and a second machine learning model; 22. The dialogue system of claim 21.
[0315] (Appendix 23) At least one processor Obtain information about the first utterance spoken by the user; generating information about a second utterance based on information about the first utterance, input information, and a machine learning model; controlling the utterance of the dialogue device based on information about the second utterance; the input information includes information indicating a relationship between the user and the interactive device; How to interact.
[0316] (Appendix 24) At least one processor has Obtain information about the first utterance spoken by the user; generating information about a second utterance based on information about the first utterance, input information, and a machine learning model; controlling the utterance of the dialogue device based on information about the second utterance; the input information includes information indicating a relationship between the user and the interactive device; A program for executing a process.
[0317] This application claims priority to Provisional Application No. 63 / 564,537, filed with the U.S. Patent and Trademark Office on March 13, 2024, Provisional Application No. 63 / 648,341, filed with the U.S. Patent and Trademark Office on May 16, 2024, and Japanese Patent Application No. 2024-122898, filed with the Japan Patent Office on July 29, 2024, the entire contents of which are incorporated herein by reference. [Explanation of symbols]
[0318] 10: Interactive device 20: Control device 25: Evaluation device 30:Generation device 40: Terminal device 50: Management device 101: Scenario memory section 102: Dialogue memory unit 103: Model memory unit 104: Evaluation result storage unit 110: Speech acquisition unit 120: Speech generation unit 130: Speech presentation unit 140: Dialogue evaluation section 150: Result output section 501: Data storage unit 502: Model storage unit 510: Label assignment unit 520: Feature extraction unit 530: Model Learning Department 540: Model evaluation unit 550: Model selection section 1000: Dialogue System D: Interaction data M1: Evaluation model M2: Generative model
Claims
1. at least one memory; at least one processor; The at least one processor obtaining a first system prompt from a plurality of candidate system prompts based on at least the first information; executing a dialogue between the interlocutor and the dialogue device by inputting at least the first system prompt into a machine learning model; storing dialogue data including information on at least the utterances of the interlocutors in the at least one memory; generating one or more assessment results regarding the interlocutors using the interaction data stored in the at least one memory; at least some of the evaluation results included in the one or more evaluation results are evaluation results acquired by a terminal device of a first evaluator, at least some of the evaluation results included in the one or more evaluation results are evaluation results acquired by a terminal device of a second evaluator, the first information includes at least one of an industry, an occupation, an attribute of the interlocutor, an attribute of the first evaluator, an attribute of the second evaluator, an evaluation criterion of the first evaluator, an evaluation criterion of the second evaluator, a degree of progress of the dialogue, or information about the dialogue between the interlocutor and the dialogue device; each of the plurality of system prompt candidates can be used to acquire dialogue data used to generate an evaluation result acquired by each of the terminal devices of the plurality of evaluators; the first system prompt includes at least information indicating a relationship between the interlocutor and the interaction device; the evaluation result acquired by the terminal device of the first evaluator and the evaluation result acquired by the terminal device of the second evaluator are generated using the same dialogue data stored in the at least one memory, the dialogue device, the terminal device of the first evaluator, and the terminal device of the second evaluator are installed in different locations. Information processing system.
2. The at least one processor transmitting the evaluation result according to the evaluation criteria of the first evaluator to a terminal device of the first evaluator; transmitting the evaluation result according to the evaluation criteria of the second evaluator to a terminal device of the second evaluator; The information processing system according to claim 1 .
3. the evaluation criteria of the first evaluator and the evaluation criteria of the second evaluator are each created by customizing an evaluation criteria template; The information processing system according to claim 2 .
4. the evaluation criteria of the first evaluator and the evaluation criteria of the second evaluator are each selected from a plurality of evaluation criteria created in advance; The information processing system according to claim 2 .
5. The at least one processor proposing evaluation criteria for the first evaluator based on at least one of the attributes of the first evaluator and the attributes of the interlocutor; suggesting evaluation criteria for the second evaluator based on at least one of the attributes of the second evaluator and the attributes of the interlocutor; The information processing system according to claim 2 .
6. The at least one processor transmitting the same interaction data to the terminal device of the first evaluator and the terminal device of the second evaluator; The information processing system according to claim 1 .
7. The at least one processor Transmitting an external evaluation result to the terminal device of the first evaluator and the terminal device of the second evaluator; The external evaluation result includes at least one of the results of an aptitude test of the interlocutor, information indicating language ability, information indicating work experience, or information indicating expertise. The information processing system according to claim 1 .
8. The at least one processor determining an evaluation criterion for generating the one or more evaluation results based on at least one of the attributes of the interlocutor, the attributes of the first evaluator, or the attributes of the second evaluator; generating the one or more evaluation results based on the determined evaluation criteria; The information processing system according to claim 1 .
9. The at least one processor controlling the evaluation result acquired by the terminal device of the first evaluator based on the usage plan of the first evaluator; controlling the evaluation result acquired by the terminal device of the second evaluator based on the usage plan of the second evaluator; The information processing system according to claim 1 .
10. The interlocutor is a job seeker, the first evaluator is a first company; The second evaluator is a second company different from the first company. The information processing system according to claim 1 .
11. The at least one processor generating the one or more evaluation results based on one or more second machine learning models; The information processing system according to claim 1 .
12. The at least one processor generating inputs to the one or more second machine learning models based at least in part on the interaction data; inputting the input information into the one or more second machine learning models to generate the one or more evaluation results; The information processing system according to claim 11.
13. the input information includes one or more feature amounts extracted from at least a portion of the dialogue data; The one or more feature amounts include at least information regarding any one of a facial expression of the interlocutor, a body movement of the interlocutor, a voice of the interlocutor, a smoothness of the dialogue between the interlocutor and the dialogue device, or a content of the dialogue; The information processing system according to claim 12.
14. the input information includes one or more feature amounts extracted from at least a portion of the dialogue data; The one or more features include at least one of a feature extracted from voice information of the interlocutor included in the dialogue data, a feature extracted from image information of the interlocutor included in the dialogue data, or a feature extracted from dialogue content included in the dialogue data. The information processing system according to claim 12.
15. The at least one processor generating the second machine learning model using training data including other dialogue data and evaluation values by the first evaluator for the other dialogue data; generating the evaluation result to be acquired by the terminal device of the first evaluator using the generated second machine learning model; The information processing system according to claim 11.
16. The at least one processor selecting the one or more second machine learning models from a plurality of second machine learning models based on an instruction from the first evaluator; generating the evaluation result to be acquired by the terminal device of the first evaluator using the one or more selected second machine learning models; The information processing system according to claim 11.
17. The at least one processor selecting the one or more second machine learning models from a plurality of second machine learning models based on evaluation criteria of the first evaluator; generating the evaluation result to be acquired by the terminal device of the first evaluator using the one or more selected second machine learning models; The information processing system according to claim 11.
18. The second machine learning model is the same as the machine learning model or a model generated based on the machine learning model. The information processing system according to claim 11.
19. The attributes of the interlocutor include at least one of the resume or the curriculum vitae of the interlocutor; The information processing system according to claim 1 .
20. the at least one processor causes the interaction device to output information indicating the relationship before the interaction is performed; The information processing system according to claim 1 .
21. the information indicating the relationship output by the dialogue device has a smaller amount of information than the information indicating the relationship input to the machine learning model; 21. The information processing system according to claim 20.
22. the interactive device outputs a voice signal synthesized with information indicating the relationship.
21. The information processing system according to claim 20.
23. The relationship includes at least one of a relationship in which the interlocutor provides a service or product to the dialogue device, a relationship in which the interlocutor instructs the dialogue device, a relationship in which the dialogue device pays a reward to the interlocutor, or a relationship in which the interlocutor who has information A provides the information A to the dialogue device that does not have the information A. The information processing system according to claim 1 .
24. The relationship includes at least one of the following: a relationship in which the interlocutor is a store clerk and the interaction device is a customer; a relationship in which the interlocutor is a boss and the interaction device is a subordinate; a relationship in which the interlocutor is a teacher and the interaction device is a student; a relationship in which the interlocutor is a lecturer and the interaction device is an audience; a relationship in which the interlocutor is a doctor and the interaction device is a patient; or a relationship in which the interlocutor is a requested party and the interaction device is a requester. The information processing system according to claim 1 .
25. the information indicating the relationship includes at least information regarding the role of the interactive device; The information processing system according to claim 1 .
26. the role of the interactive device includes at least one of a customer, a subordinate, a student, an audience member, a patient, or a client; 26. The information processing system according to claim 25.
27. The information indicating the relationship includes information regarding a role of the interlocutor and a role of the interaction device. The information processing system according to claim 1 .
28. The information indicating the relationship includes information indirectly indicating the role of the interlocutor. The information processing system according to claim 1 .
29. The information indicating the relationship input to the machine learning model includes information obtained by performing a predetermined process on the information indicating the relationship. The information processing system according to claim 1 .
30. the predetermined process includes inputting information indicating the relationship into the machine learning model or another machine learning model; 30. The information processing system according to claim 29.
31. the at least one processor executes the interaction by inputting at least the first system prompt and interaction history information into the machine learning model. The information processing system according to claim 1 .
32. the first system prompt includes at least a constraint on an utterance of the interactive device; The information processing system according to claim 1 .
33. the at least one processor generates the first system prompt based on two or more system prompts. The information processing system according to claim 1 .
34. the at least one processor executes the dialogue using a second system prompt different from the first system prompt before the dialogue between the dialogue person and the dialogue device is terminated. The information processing system according to claim 1 .
35. the at least one processor uses one or more second machine learning models associated with the first system prompt to generate the one or more assessment results. The information processing system according to claim 1 .
36. the evaluation result acquired by the terminal device of the first evaluator is generated by an evaluation terminal of the first evaluator, The evaluation result acquired by the terminal device of the second evaluator is generated by the evaluation terminal of the second evaluator. The information processing system according to claim 1 .
37. The interactive device further comprises:
37. An information processing system according to any one of claims 1 to 36.
38. The system further includes a terminal device of the first evaluator and a terminal device of the second evaluator.
37. An information processing system according to any one of claims 1 to 36.
39. The interactive device, a terminal device of the first evaluator, and a terminal device of the second evaluator are further included.
37. An information processing system according to any one of claims 1 to 36.
40. Using the information processing system according to any one of claims 1 to 36, generating the one or more evaluation results regarding the interlocutor; Information processing methods.
41. causing at least one processor to perform the information processing method of claim 40; program.
Citation Information
Patent Citations
Interview support system
JP2023031089A