Information processing device, information processing method, and program
The information processing system stabilizes interview evaluations by using multiple large-scale language models to identify and resolve discrepancies in evaluation criteria, enhancing consistency and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- KDDI CORP
- Filing Date
- 2025-09-02
- Publication Date
- 2026-05-15
AI Technical Summary
The evaluation of interviewees using large language models is unstable due to variations in interpreting ambiguous evaluation criteria, leading to inconsistent results.
An information processing system utilizing multiple large-scale language models to evaluate interview dialogue content, identifying differences in evaluation results, and employing a second model to determine the cause of these discrepancies, thereby stabilizing the evaluation process.
The system stabilizes the evaluation by identifying and addressing the causes of inconsistencies across different language models, improving evaluation criteria and ensuring consistent assessment outcomes.
Smart Images

Figure 0007860319000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program for assisting an interviewer in conducting an interview with an interviewee.
Background Art
[0002] Patent Document 1 describes a system that analyzes the meaning of speech content in voice data during an interview and provides advice generated based on the analysis results to the interviewer.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is conceivable to input evaluation criteria and the dialogue content in an interview into a large language model to evaluate whether the interviewee meets the evaluation criteria by the large language model. In this case, if there are ambiguous parts in the evaluation criteria, there will be variations in the interpretation of the evaluation criteria by the large language model, so the evaluation of the interviewee by the large language model may become unstable.
[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to stabilize the evaluation by the large language model when evaluating the dialogue content in an interview by the large language model.
Means for Solving the Problems
[0006] An information processing device according to a first aspect of the present invention includes: a first acquisition unit that acquires the content of a dialogue in an interview conducted by an interviewer with an interviewee; a second acquisition unit that acquires the evaluation results of one or more evaluation criteria output by each of the multiple first large-scale language models by inputting a first instruction sentence that causes each of the multiple first large-scale language models to evaluate the interviewee using one or more evaluation criteria based on the dialogue content; an identification unit that identifies differences among the multiple evaluation results output by the multiple first large-scale language models; and an output unit that outputs the cause output by the second large-scale language model by inputting a second instruction sentence that causes the second large-scale language model to identify the cause of the difference in the one or more evaluation criteria.
[0007] The output unit may output, as the cause, the portion of the one or more evaluation criteria that the multiple first large-scale language models interpret differently from each other.
[0008] The output unit may output, as the cause, the conditions used by the multiple first large-scale language models to determine different evaluation metrics for the dialogue content, among the one or more evaluation criteria.
[0009] The second instruction statement includes a statement that causes the second large-scale language model to identify the cause and the improvement measures for one or more evaluation criteria to resolve the cause, and the output unit may output the cause and the improvement measures output by the second large-scale language model.
[0010] The aforementioned improvement measures may include proposed changes to the conditions for determining the evaluation indicators among the one or more evaluation criteria.
[0011] The aforementioned improvement measures may include proposed changes to the questions asked to the interviewee among the one or more evaluation criteria.
[0012] The first acquisition unit acquires the content of the dialogue in each of the multiple interviews, the second acquisition unit acquires the evaluation result of the content of the dialogue in each of the multiple interviews, and the output unit may output the cause by associating it with the evaluation criterion among the one or more evaluation criteria that has a difference in the evaluation result in at least one of the multiple interviews.
[0013] The identification unit may identify the differences output by the third large-scale language model by inputting a third instruction sentence to the third large-scale language model for identifying the differences among the multiple evaluation results output by the multiple first large-scale language models.
[0014] The first acquisition unit may acquire the dialogue content output by the fourth large-scale language model by inputting a fourth instruction sentence to the fourth large-scale language model for generating the dialogue content based on data from past interviews or virtual interviews.
[0015] A second aspect of the present invention is an information processing method which includes the steps of: obtaining the content of a dialogue in an interview conducted by an interviewer with an interviewee, which is executed by a processor; obtaining the evaluation results for each of the one or more evaluation criteria output by each of the multiple first large-scale language models by inputting a first instruction sentence to cause each of the multiple first large-scale language models to evaluate the interviewee using one or more evaluation criteria based on the dialogue content; identifying differences among the multiple evaluation results output by the multiple first large-scale language models; and outputting the cause output by the second large-scale language model by inputting a second instruction sentence to cause the second large-scale language model to identify the cause of the difference in the one or more evaluation criteria.
[0016] A program according to a third aspect of the present invention causes a processor to perform the following steps: acquire the content of a dialogue in an interview conducted by an interviewer with an interviewee; input a first instruction to cause a plurality of first large-scale language models of different types to evaluate the interviewee using one or more evaluation criteria based on the dialogue content, thereby acquiring the evaluation results for each of the one or more evaluation criteria output by each of the plurality of first large-scale language models; identify the differences between the plurality of evaluation results output by the plurality of first large-scale language models; and input a second instruction to cause a second large-scale language model to identify the cause of the difference in the one or more evaluation criteria, thereby outputting the cause output by the second large-scale language model. [Effects of the Invention]
[0017] According to the present invention, when the content of a dialogue in an interview is evaluated by a large-scale language model, the evaluation by the large-scale language model can be stabilized. [Brief explanation of the drawing]
[0018] [Figure 1] This is a schematic diagram of the information processing system according to the embodiment. [Figure 2] This is a block diagram of an information processing system according to an embodiment. [Figure 3] This is a diagram illustrating exemplary evaluation criteria. [Figure 4] This is a schematic diagram illustrating the process by which the evaluation unit has the first large-scale language model evaluate the content of the dialogue. [Figure 5] This is a schematic diagram illustrating the process by which the specific unit identifies the cause of the discrepancies in the second large-scale language model. [Figure 6] This is a schematic diagram of an information terminal that displays the evaluation criteria and the causes of any discrepancies. [Figure 7] This diagram shows a flowchart illustrating an example of an information processing method performed by an information processing device. [Modes for carrying out the invention]
[0019] [Overview of Information Processing System S] FIG. 1 is a schematic diagram of an information processing system S according to the present embodiment. The information processing system S includes an information processing apparatus 1 and an information terminal 2. The information processing system S may include other devices such as servers and terminals.
[0020] The information processing apparatus 1 is a computer that processes information regarding an interview conducted by an interviewer with an interviewee. The interview is, for example, an employment interview conducted by an organization to recruit members, a personnel interview conducted within an organization to evaluate members, a selection interview conducted by a school to select admitted students, etc. The interviewer is a human who conducts an interview and interacts with the interviewee to evaluate the interviewee. The interviewee is a human who receives an interview and interacts with the interviewer to receive an evaluation.
[0021] The interview is conducted face-to-face in a space where the interviewer and the interviewee are present, or remotely via a network on the information terminal 2 used by each of the interviewer and the interviewee. The information processing apparatus 1 causes the information terminal 2 to display information regarding the interview generated based on the dialogue content in the interview. The dialogue content in the interview is obtained from the voice uttered in the actually conducted interview, or is generated by a large language model.
[0022] The information terminal 2 is a computer used by the interviewer. Also, the information terminal 2 may be a computer used by the interviewee. The information terminal 2 is, for example, a smartphone, a tablet terminal, or a personal computer.
[0023] Information terminal 2 has an operation unit such as a touch panel or keyboard for receiving operations, a display unit such as a liquid crystal display for displaying information, and a sound collection unit such as a microphone for acquiring voices spoken during the interview. Information terminal 2 is pre-associated with a user by setting identification information (Identifier: ID) for identifying the user who will use information terminal 2. Information terminal 2 acquires voices spoken during the interview and transmits voice data indicating the acquired voices to information processing device 1 via the network.
[0024] The information processing system S may include a sound collection device different from the information terminal 2, either in place of or in addition to the sound collection unit of the information terminal 2. In this case, the sound collection device has a sound collection unit such as a microphone for acquiring the sound emitted during the interview, and transmits the acquired sound data to the information processing device 1 via the network.
[0025] The following describes the general process performed by the information processing system S according to this embodiment. The information processing device 1 acquires the content of the dialogue in an interview conducted by an interviewer with an interviewee. For example, the information processing device 1 receives audio data transmitted by an information terminal 2 or a sound collection device and performs known speech recognition processing on the received audio data to acquire the dialogue content showing the statements made by the interviewer and the interviewee. Alternatively, the information processing device 1 may acquire the dialogue content generated by a large-scale language model by inputting data from past interviews or virtual interviews into a large-scale language model.
[0026] The information processing device 1 inputs a first instruction sentence to multiple first large-scale language models of different types, which instructs them to evaluate the interviewee using one or more evaluation criteria based on the content of the dialogue. The evaluation criteria are standards for evaluating the interviewee based on the content of the dialogue between the interviewer and the interviewee. For example, the evaluation criteria are information that associates questions asked by the interviewer to the interviewee, the interviewee's answers to those questions, and evaluation indicators (scores, etc.) given to those answers. The information processing device 1 obtains the evaluation results for one or more evaluation criteria output by the multiple first large-scale language models.
[0027] The information processing device 1 identifies differences in multiple evaluation results output by multiple first large-scale language models. For example, the information processing device 1 identifies differences if at least some of the multiple evaluation indicators (scores, etc.) output by multiple first large-scale language models for a single evaluation criterion are different.
[0028] The information processing device 1 inputs a second instruction to the second large-scale language model to identify the cause of the difference. The second large-scale language model is a large-scale language model that is one of several first large-scale language models, or a large-scale language model of a different type from several first large-scale language models. The cause of the difference is, for example, a part of one or more evaluation criteria that several first large-scale language models interpreted differently from each other. The information processing device 1 outputs the cause of the difference output by the second large-scale language model.
[0029] In this way, the information processing system S has multiple first large-scale language models evaluate the content of the interview dialogue using predefined evaluation criteria, identifies differences in the evaluation results from the multiple first large-scale language models, and has a second large-scale language model identify the cause of these differences. This allows the information processing system S to improve the evaluation criteria, which vary depending on the type of large-scale language model, when having large-scale language models evaluate the content of the interview dialogue, and to stabilize the evaluation by the large-scale language models.
[0030] [Configuration of Information Processing System S] Figure 2 is a block diagram of the information processing system S according to this embodiment. In Figure 2, the arrows indicate the main data flows, and there may be other data flows besides those shown in Figure 2. In Figure 2, each block represents a functional unit configuration, not a hardware (device) unit configuration. Therefore, the blocks shown in Figure 2 may be implemented in a single device, or they may be implemented separately in multiple devices. Data exchange between blocks may be performed via any means, such as a data bus, network, or portable storage medium.
[0031] The information processing device 1 includes a communication unit 11, a storage unit 12, and a control unit 13. The information processing device 1 may be configured by two or more physically separate devices connected by wired or wireless connections. Alternatively, the information processing device 1 may be configured as a cloud, which is a collection of computer resources.
[0032] The communication unit 11 has a communication controller for sending and receiving data to and from the information terminal 2 via the network. The communication unit 11 notifies the control unit 13 of the data received from the information terminal 2 via the network. The communication unit 11 also transmits data output from the control unit 13 to the information terminal 2 via the network.
[0033] The storage unit 12 is a storage medium including ROM (Read Only Memory), RAM (Random Access Memory), a hard disk drive, an SSD (Solid State Drive), etc. The storage unit 12 pre-stores programs to be executed by the control unit 13. The storage unit 12 may be located outside the information processing device 1, in which case data may be exchanged with the control unit 13 via a network.
[0034] The control unit 13 includes a first acquisition unit 131, an evaluation unit 132, a second acquisition unit 133, a specification unit 134, and an output unit 135. The control unit 13 is a processor such as a CPU (Central Processing Unit), and functions as the first acquisition unit 131, evaluation unit 132, second acquisition unit 133, specification unit 134, and output unit 135 by executing a program stored in the storage unit 12.
[0035] Information terminal 2 may perform at least a portion of the processing performed by the information processing device 1 according to this embodiment. In this case, the processor of information terminal 2 functions as at least a portion of the first acquisition unit 131, evaluation unit 132, second acquisition unit 133, identification unit 134, and output unit 135.
[0036] The following describes in detail the processes performed by the information processing system S. In the information processing device 1, the first acquisition unit 131 acquires one or more evaluation criteria used to evaluate the content of the conversation in the interview.
[0037] Figure 3 shows an exemplary evaluation criterion. The evaluation criterion is information that associates, for example, a question asked by the interviewer to the interviewee (basic question in Figure 3), the interviewee's answer to that question (example answer in Figure 3), and an evaluation index (score in Figure 3) given to that answer. The question is a string of characters representing a question asked by the interviewee. The answer is a string of characters that the interviewee is expected to answer to the question. Both the question and the answer strings may be sentences or single words.
[0038] The evaluation index is a numerical value that corresponds to the interviewee's evaluation being high, for example, and a numerical value that corresponds to the interviewee's evaluation being low. In the example in Figure 3, the score of the evaluation index when the interviewee's evaluation is highest is "3", and the score of the evaluation index when the interviewee's evaluation is lowest is "0".
[0039] The evaluation indicator may be any other indicator that represents the evaluation of the interviewee. For example, the evaluation indicator may be a numerical value where a higher evaluation of the interviewee corresponds to a lower value, and a lower evaluation corresponds to a higher value. The evaluation indicator may also be a character corresponding to the evaluation, such as "acceptable" or "unacceptable," or a symbol corresponding to the evaluation, such as "○" or "×."
[0040] The memory unit 12 has pre-stored information indicating one or more evaluation criteria used to evaluate the content of the interview dialogue. One or more evaluation criteria may be specified by the administrator of the information processing device 1, or they may be generated by the large-scale language model described later. The first acquisition unit 131 reads the information indicating one or more evaluation criteria from the memory unit 12 and acquires the one or more evaluation criteria indicated by the read information.
[0041] The first acquisition unit 131 acquires the content of the dialogue that the interviewer has with the interviewee during the interview. The first acquisition unit 131 acquires the content of the dialogue from the audio spoken during the actual interview, or acquires the content of the dialogue generated by a large-scale language model. The first acquisition unit 131 may also acquire the content of the dialogue in each of multiple interviews.
[0042] When the first acquisition unit 131 acquires the content of a conversation from the audio, the information terminal 2 or sound collection device uses the sound collection unit to acquire the voices spoken by the interviewer and the person being evaluated. The information terminal 2 or sound collection device transmits audio data indicating the acquired voices to the information processing device 1. The information terminal 2 or sound collection device transmits the audio data sequentially during the interview, or transmits the audio data all at once after the interview has finished.
[0043] When the interviewer and the interviewee conduct a remote meeting via a network, the information terminal 2 used by the interviewer and the information terminal 2 used by the interviewee may each transmit audio data to the information processing device 1.
[0044] The first acquisition unit 131 acquires audio data transmitted by the information terminal 2 or the sound collection device. The first acquisition unit 131 performs known speech recognition processing on the acquired audio data to obtain the dialogue content that shows what the interviewer and the interviewee said.
[0045] The first acquisition unit 131 may acquire dialogue content from a single audio data generated by combining multiple audio data when it acquires multiple audio data from multiple information terminals 2. Alternatively, the first acquisition unit 131 may acquire dialogue content from each of the multiple audio data and combine the acquired dialogue content.
[0046] The first acquisition unit 131 may, without performing speech recognition processing itself, have an external device different from the information processing device 1 perform speech recognition processing. In this case, the first acquisition unit 131 transmits voice data to the external device, receives the dialogue content generated by the external device performing speech recognition processing on the transmitted voice data from the external device, and stores the received dialogue content in the storage unit 12.
[0047] When the first acquisition unit 131 acquires dialogue content generated by a large-scale language model, the first acquisition unit 131 acquires, for example, data from past interviews or virtual interviews. Data from past interviews is, for example, at least a portion of the dialogue content from past interviews stored in the memory unit 12 or a storage device on the network. Data from virtual interviews is, for example, the dialogue content of a virtual interview (such as a model interview) stored in the memory unit 12 or a storage device on the network.
[0048] The first acquisition unit 131 determines a dialogue content generation instruction statement for causing a large-scale language model to generate dialogue content based on data from past interviews or virtual interviews. The dialogue content generation instruction statement includes, for example, a prompt (command) that, when input into the large-scale language model, causes the large-scale language model to perform a predetermined process.
[0049] The first acquisition unit 131 acquires a large-scale language model that has been pre-stored in the storage unit 12. The large-scale language model used by the first acquisition unit 131 is a large-scale language model that is one of the multiple first large-scale language models described later, or a large-scale language model of a different type from the multiple first large-scale language models described later. The first acquisition unit 131 inputs the determined dialogue content generation instruction sentences into the acquired large-scale language model.
[0050] Furthermore, the first acquisition unit 131 may cause a large-scale language model, which is executed on an external device different from the information processing device 1, to generate the dialogue content. In this case, the first acquisition unit 131 inputs the dialogue content generation instruction to the large-scale language model executed on the external device by transmitting the dialogue content generation instruction to the external device.
[0051] The first acquisition unit 131 acquires the dialogue content output by the large-scale language model by inputting an instruction sentence for generating dialogue content into the large-scale language model.
[0052] The evaluation unit 132 performs an evaluation of each of the one or more evaluation criteria acquired by the first acquisition unit 131, based on the dialogue content acquired by the first acquisition unit 131. The evaluation unit 132 causes the first large-scale language model to evaluate the dialogue content.
[0053] Figure 4 is a schematic diagram illustrating the process by which the evaluation unit 132 causes the first large-scale language model to evaluate the content of the dialogue. Based on the content of the dialogue, the evaluation unit 132 determines an evaluation instruction sentence (first instruction sentence) for causing the first large-scale language model to evaluate the interviewee using one or more evaluation criteria. The evaluation instruction sentence includes, for example, a prompt (command) that, when input to the first large-scale language model, causes the first large-scale language model to execute a predetermined process.
[0054] The evaluation unit 132 retrieves, for example, a template for determining the evaluation instruction statement, which is pre-stored in the memory unit 12. The template is, for example, a template that includes a fixed part that does not change within the evaluation instruction statement and a variable part that changes within the evaluation instruction statement. The fixed part includes, for example, a string that does not change according to the evaluation criteria and the content of the dialogue (such as a tag used to give instructions to the machine learning model).
[0055] The evaluation unit 132 determines the evaluation instruction statement by, for example, applying (inserting) one or more evaluation criteria acquired by the first acquisition unit 131 and the dialogue content acquired by the first acquisition unit 131 into the variable portion of the template.
[0056] If the format of the evaluation instruction sentences that can be accepted differs depending on the first large-scale language model, the evaluation unit 132 may determine different evaluation instruction sentences depending on the type of each of the multiple first large-scale language models. In this case, the evaluation unit 132 determines the evaluation instruction sentences using, for example, templates that are pre-associated with each of the multiple first large-scale language models.
[0057] The evaluation unit 132 retrieves multiple first large-scale language models of different types that have been pre-stored in the memory unit 12. Each of the multiple first large-scale language models is pre-generated by machine learning a large amount of string data using known deep learning processes such as DNN (Deep Neural Network). Each of the multiple first large-scale language models is configured to output a string corresponding to the input string when a string is input. The evaluation unit 132 inputs the determined evaluation instruction sentence to each of the multiple first large-scale language models.
[0058] The evaluation unit 132 may have each of the multiple first large-scale language models, which are executed on an external device different from the information processing device 1, evaluate the interviewee. In this case, the evaluation unit 132 inputs evaluation instructions to each of the multiple first large-scale language models executed on the external device by transmitting evaluation instructions to the external device.
[0059] The second acquisition unit 133 acquires the evaluation results for one or more evaluation criteria output by the first large-scale language model by having the evaluation unit 132 input evaluation instruction sentences to each of the multiple first large-scale language models.
[0060] The evaluation results include, for example, evaluation indicators (scores, etc.) for one or more evaluation criteria. The second acquisition unit 133 acquires the evaluation results output by each of the multiple first large-scale language models executed in the information processing device 1 or an external device. If the first acquisition unit 131 acquires the dialogue content from multiple interviews, the second acquisition unit 133 may acquire the evaluation results of the dialogue content from each of the multiple interviews.
[0061] The identification unit 134 identifies the differences between the multiple evaluation results obtained by the second acquisition unit 133 and the multiple evaluation results output by the multiple first large-scale language models. The differences are, for example, parts in which the evaluation indicators (scores, etc.) shown by the evaluation results for a single evaluation criterion differ among the multiple evaluation results.
[0062] The identification unit 134, for example, compares multiple evaluation results output by multiple first large-scale language models and identifies the pair of evaluation criterion and evaluation criterion as a difference if the evaluation indices output by at least some of the multiple first large-scale language models for a single evaluation criterion are different from each other. The identification unit 134 may also identify the pair of evaluation criterion and evaluation criterion as a difference if the difference in the evaluation indices output by at least some of the multiple first large-scale language models for a single evaluation criterion is greater than or equal to a predetermined threshold value.
[0063] Furthermore, the identification unit 134 may use a large-scale language model to identify differences among multiple evaluation results. In this case, the identification unit 134 determines a difference identification instruction statement to identify the differences among multiple evaluation results. The difference identification instruction statement includes, for example, a prompt (command) that, when input into the large-scale language model, causes the large-scale language model to execute a predetermined process.
[0064] The identification unit 134 retrieves a large-scale language model that has been pre-stored in the storage unit 12. The large-scale language model used by the identification unit 134 is either a large-scale language model that is one of a plurality of first large-scale language models, or a large-scale language model of a different type from the plurality of first large-scale language models. The identification unit 134 inputs the determined difference identification instruction sentence to the retrieved large-scale language model.
[0065] Furthermore, the identification unit 134 may cause a large-scale language model executed on an external device different from the information processing device 1 to identify the differences. In this case, the identification unit 134 inputs the difference identification instruction to the large-scale language model executed on the external device by transmitting the difference identification instruction to the external device.
[0066] The identification unit 134 identifies the differences output by the large-scale language model by inputting instruction statements for identifying differences into the large-scale language model.
[0067] The identification unit 134 causes the second large-scale language model to identify the cause of the identified discrepancy. The cause of the discrepancy is a part of one or more evaluation criteria that multiple first large-scale language models interpreted differently from one another.
[0068] The discrepancies may stem from, for example, at least some of the questions asked of the interviewee as outlined in the evaluation criteria. If the wording or intent of the questions is ambiguous, the interpretation of the questions may differ depending on the type of large-scale language model used.
[0069] The discrepancies may stem from, for example, the relationship between the answers to questions and the evaluation metrics (scores, etc.) assigned to those answers within the evaluation criteria. If it is unclear which evaluation metric an interviewee's answer corresponds to, the evaluation metrics assigned to the interviewee may differ depending on the type of large-scale language model used.
[0070] Figure 5 is a schematic diagram illustrating the process by which the identification unit 134 causes the second large-scale language model to identify the cause of the discrepancy. The identification unit 134 determines a cause identification instruction statement (second instruction statement) to cause the second large-scale language model to identify the cause of the discrepancy in the multiple evaluation results output by the multiple first large-scale language models. The cause identification instruction statement includes, for example, a prompt (command) that, when input to the second large-scale language model, causes the second large-scale language model to execute a predetermined process.
[0071] The identification unit 134 retrieves, for example, a template for determining a cause identification instruction statement, which is pre-stored in the storage unit 12. The template is, for example, a template that includes a fixed part that does not change within the cause identification instruction statement and a variable part that changes within the cause identification instruction statement. The fixed part includes, for example, a string that does not change according to the evaluation criteria and differences (such as a tag used to instruct the machine learning model).
[0072] The identification unit 134 determines the cause identification instruction statement by, for example, applying (inserting) one or more evaluation criteria acquired by the first acquisition unit 131 and the differences identified by the identification unit 134 to the variable portion of the template.
[0073] The cause identification instruction statement may include, in addition to a statement that causes the second large-scale language model to identify the cause of the discrepancy in multiple evaluation results, a statement that causes the second large-scale language model to identify one or more improvement measures for evaluation criteria to resolve the cause. The improvement measures may include, for example, proposed changes to the questions asked to the interviewee among one or more evaluation criteria. The proposed changes to the questions may include, for example, the changes to the questions indicated by the evaluation criterion that caused the discrepancy.
[0074] The improvement measures may include, for example, proposed changes to the conditions for determining evaluation indicators (scores, etc.) among one or more evaluation criteria. Proposed changes to the conditions for determining evaluation indicators may include, for example, changes to the association between the interviewee's answer to a question and the evaluation indicator assigned to that answer in the evaluation criterion that caused the discrepancy.
[0075] The identification unit 134 acquires a second large-scale language model that has been pre-stored in the storage unit 12. The second large-scale language model used by the identification unit 134 is a large-scale language model that is one of a plurality of first large-scale language models, or a large-scale language model of a different type from the plurality of first large-scale language models. The identification unit 134 inputs the determined cause identification instruction sentence to the acquired second large-scale language model.
[0076] The identification unit 134 may have a second large-scale language model, which is executed on an external device different from the information processing device 1, identify the cause. In this case, the evaluation unit 132 inputs a cause identification instruction to the second large-scale language model executed on the external device by transmitting a cause identification instruction to the external device.
[0077] The identification unit 134 identifies the cause and corrective measures output by the second large-scale language model by inputting a cause identification instruction sentence to the second large-scale language model.
[0078] The output unit 135 outputs the cause of the discrepancies in multiple evaluation results in one or more evaluation criteria identified by the identification unit 134. For example, the output unit 135 outputs the portion of one or more evaluation criteria in which multiple first large-scale language models interpreted each other differently, as the cause of the discrepancies.
[0079] When the first acquisition unit 131 acquires the content of the conversation in each of the multiple interviews, the output unit 135 may output the cause of the discrepancy by associating it with the evaluation criterion in which there is a difference in the evaluation result in at least one of the multiple interviews from among one or more evaluation criteria.
[0080] The output unit 135 transmits display information to the information terminal 2 associated with the interviewee, for example, one or more evaluation criteria acquired by the first acquisition unit 131 and the cause identified by the identification unit 134. The information terminal 2 displays the evaluation criteria and the cause of the discrepancy on its display unit according to the display information transmitted by the information processing device 1.
[0081] Figure 6 is a schematic diagram of the information terminal 2 displaying evaluation criteria and the causes of discrepancies. The output unit 135 displays the causes of discrepancies on the information terminal 2 by, for example, differentiating the display of the part of one or more evaluation criteria that is the cause of the discrepancy (the color in Figure 6) from the display of the other parts. As a result, the information processing system S can visualize the evaluation criteria that cause the evaluation results to vary depending on the type of large-scale language model, and support the interviewer in improving the evaluation criteria.
[0082] Furthermore, the output unit 135 outputs improvement measures for one or more evaluation criteria identified by the identification unit 134 to resolve the cause of the discrepancy. The output unit 135 transmits, for example, display information to the information terminal 2 associated with the interviewee, for displaying the one or more evaluation criteria acquired by the first acquisition unit 131 and the improvement measures identified by the identification unit 134. The information terminal 2 displays the evaluation criteria and improvement measures on its display unit according to the display information transmitted by the information processing device 1.
[0083] The output unit 135 displays, for example, a proposed change to the question, which is an improvement measure identified by the identification unit 134, on the information terminal 2, associating it with the question that caused the discrepancy among one or more evaluation criteria. In this way, the information processing system S can present improvement measures for questions asked to the interviewee that cause the evaluation results to vary depending on the type of large-scale language model, and support the interviewer in improving the evaluation criteria.
[0084] The output unit 135 displays, for example, on the information terminal 2 a proposed change to the conditions, which is an improvement measure identified by the identification unit 134, in association with the conditions (example answer in Figure 6) used by multiple first large-scale language models to determine different evaluation indicators from one or more evaluation criteria. This allows the information processing system S to present improvement measures for the conditions used to determine the evaluation indicators that cause the evaluation results to vary depending on the type of large-scale language model, thereby supporting the interviewer in improving the evaluation criteria.
[0085] [Information Processing Method Flowchart] Figure 7 is a flowchart showing an exemplary information processing method performed by the information processing device 1 according to this embodiment. The first acquisition unit 131 acquires one or more evaluation criteria used to evaluate the content of the conversation in the interview (S11).
[0086] The first acquisition unit 131 acquires the content of the dialogue that the interviewer has with the interviewee during the interview (S12). The first acquisition unit 131 acquires the content of the dialogue from the voice spoken during the actual interview, or acquires the content of the dialogue generated by a large-scale language model.
[0087] The evaluation unit 132 determines an evaluation instruction sentence (first instruction sentence) to cause the first large-scale language model to evaluate the interviewee using one or more evaluation criteria based on the content of the dialogue. The evaluation unit 132 retrieves multiple first large-scale language models of different types that are pre-stored in the storage unit 12. The evaluation unit 132 inputs the determined evaluation instruction sentence to each of the multiple first large-scale language models (S13). Alternatively, the evaluation unit 132 may input the evaluation instruction sentence to each of the multiple first large-scale language models executed by an external device by transmitting the evaluation instruction sentence to the external device.
[0088] The second acquisition unit 133 acquires the evaluation results for one or more evaluation criteria output by the first large-scale language models, as the evaluation unit 132 inputs evaluation instruction sentences to each of the first large-scale language models. The identification unit 134 identifies the differences among the multiple evaluation results acquired by the second acquisition unit 133, which are multiple evaluation results output by the multiple first large-scale language models (S14).
[0089] The identification unit 134 determines a cause identification instruction (second instruction) to cause the second large-scale language model to identify the cause of the discrepancies that arose among the multiple evaluation results output by the multiple first large-scale language models. In addition to a sentence that causes the second large-scale language model to identify the cause of the discrepancies that arose among the multiple evaluation results, the cause identification instruction may also include a sentence that causes the second large-scale language model to identify one or more improvement measures for evaluation criteria to resolve the said cause.
[0090] The identification unit 134 acquires a second large-scale language model that has been pre-stored in the storage unit 12. The identification unit 134 inputs the determined cause identification instruction statement into the acquired second large-scale language model (S15). Alternatively, the evaluation unit 132 may input the cause identification instruction statement into the second large-scale language model executed by the external device by transmitting the cause identification instruction statement to the external device. By inputting the cause identification instruction statement into the second large-scale language model, the identification unit 134 identifies the cause and corrective measures output by the second large-scale language model.
[0091] The output unit 135 outputs the cause of the discrepancy in multiple evaluation results in one or more evaluation criteria identified by the identification unit 134, and the improvement measures for one or more evaluation criteria to resolve the cause of the discrepancy (S16).
[0092] [Effects of the Embodiment] According to the information processing system S of this embodiment, the information processing device 1 has multiple first large-scale language models evaluate the content of the interview using predefined evaluation criteria, identifies differences in the evaluation results from the multiple first large-scale language models, and has a second large-scale language model identify the cause of the differences. As a result, when the information processing system S has large-scale language models evaluate the content of the interview, it can improve the evaluation criteria that cause the evaluation results to fluctuate depending on the type of large-scale language model, and stabilize the evaluation by the large-scale language models.
[0093] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), "Build resilient infrastructure, promote inclusive and sustainable industrialization and foster innovation."
[0094] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of its gist. For example, all or part of the apparatus can be configured by functionally or physically distributing and integrating in any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combinations are combined with the effects of the original embodiments. [Explanation of Symbols]
[0095] S Information Processing System 1. Information Processing Device 11 Communications Department 12 Storage section 13 Control Unit 131 First acquisition part 132 Evaluation Department 133 Second Acquisition Department 134 Specific part 135 Output section 2. Information terminals
Claims
1. A first acquisition unit that acquires the content of the dialogue between the interviewer and the interviewee during the interview, A second acquisition unit acquires the evaluation results of each of the one or more evaluation criteria output by each of the multiple first large-scale language models, by inputting a first instruction sentence that causes each of the multiple first large-scale language models to evaluate the interviewee using one or more evaluation criteria based on the dialogue content. A unit for identifying differences among the multiple evaluation results output by the multiple first large-scale language models, An output unit outputs the cause output by the second large-scale language model, by inputting a second instruction sentence to cause the second large-scale language model to identify the cause of the difference in the one or more evaluation criteria, which indicates the part of the one or more evaluation criteria that the multiple first large-scale language models interpreted differently from one another. It has, Each of the one or more evaluation criteria is information that associates the questions the interviewer asks the interviewee, the interviewee's answers to those questions, and the evaluation indicators assigned to those answers. Information processing device.
2. A first acquisition unit that acquires the content of the dialogue in an interview conducted by an interviewer with an interviewee, A second acquisition unit acquires the evaluation results of each of the one or more evaluation criteria output by each of the multiple first large-scale language models, by inputting a first instruction sentence that causes each of the multiple first large-scale language models to evaluate the interviewee using one or more evaluation criteria based on the dialogue content. A unit for identifying differences among the multiple evaluation results output by the multiple first large-scale language models, An output unit that outputs the cause output by the second large-scale language model by inputting a second instruction sentence to cause the second large-scale language model to identify the cause of the difference in the one or more evaluation criteria, which is the condition that the multiple first large-scale language models used to determine different evaluation indicators for the dialogue content among the one or more evaluation criteria, and It has, Each of the one or more evaluation criteria is information that associates the questions the interviewer asks the interviewee, the interviewee's answers to those questions, and the evaluation indicators assigned to those answers. Information processing device.
3. The second instruction includes a sentence that causes the second large-scale language model to identify the cause and the improvement measures for one or more evaluation criteria to resolve the cause, The output unit outputs the cause and the improvement measures output by the second large-scale language model. The information processing apparatus according to claim 1 or 2.
4. The aforementioned improvement measures include a proposed change to the conditions for determining the evaluation indicators among the one or more evaluation criteria, The information processing apparatus according to claim 3.
5. The aforementioned improvement measures include proposed changes to the questions asked to the interviewee among the one or more evaluation criteria, The information processing apparatus according to claim 3.
6. The first acquisition unit acquires the content of the conversation in each of the multiple interviews, The second acquisition unit acquires the evaluation results of the dialogue content in each of the plurality of interviews, The output unit outputs the cause in association with the evaluation criterion among the one or more evaluation criteria in which the evaluation result in at least one of the multiple interviews shows a difference. The information processing apparatus according to claim 1 or 2.
7. The identification unit identifies the differences output by the third large-scale language model by inputting a third instruction sentence to the third large-scale language model for identifying the differences among the multiple evaluation results output by the multiple first large-scale language models. The information processing apparatus according to claim 1 or 2.
8. The first acquisition unit acquires the dialogue content output by the fourth large-scale language model by inputting a fourth instruction sentence to the fourth large-scale language model for generating the dialogue content based on data from past interviews or virtual interviews. The information processing apparatus according to claim 1 or 2.
9. The processor executes The steps include obtaining the content of the conversation between the interviewer and the interviewee during the interview, The steps include: inputting a first instruction sentence to multiple first large-scale language models of different types to cause them to evaluate the interviewee using one or more evaluation criteria based on the dialogue content, thereby obtaining the evaluation results for each of the one or more evaluation criteria output by each of the multiple first large-scale language models; The steps include identifying the differences between the multiple evaluation results output by the multiple first large-scale language models, A step of inputting a second instruction to cause the second large-scale language model to identify the cause of the difference in the one or more evaluation criteria, which is the part of the one or more evaluation criteria that the multiple first large-scale language models interpreted differently from each other, thereby outputting the cause output by the second large-scale language model; It has, Each of the one or more evaluation criteria is information that associates the questions the interviewer asks the interviewee, the interviewee's answers to those questions, and the evaluation indicators assigned to those answers. Information processing methods.
10. The processor executes: The steps include obtaining the content of the conversation between the interviewer and the interviewee during the interview, The steps include: inputting a first instruction sentence to multiple first large-scale language models of different types to cause them to evaluate the interviewee using one or more evaluation criteria based on the dialogue content, thereby obtaining the evaluation results for each of the one or more evaluation criteria output by each of the multiple first large-scale language models; The steps include identifying the differences between the multiple evaluation results output by the multiple first large-scale language models, A step of inputting a second instruction sentence to cause the second large-scale language model to identify the cause of the difference in the one or more evaluation criteria, which is the condition that the multiple first large-scale language models used to determine different evaluation indicators for the dialogue content, thereby outputting the cause output by the second large-scale language model; It has, Each of the one or more evaluation criteria is information that associates the questions the interviewer asks the interviewee, the interviewee's answers to those questions, and the evaluation indicators assigned to those answers. Information processing methods.
11. A program that causes a processor to execute the information processing method described in claim 9 or 10.